Companion to the essay
Anatomy of the AI debate
The argument of summer 2026, taken apart, next to the proof that says the opposite
Piotr Zientara · 12 September 2026, updated 15 September · back to the essay · po polsku
❦ ❦ ❦
The essay “An Alien Mind” has two assumptions, eight steps and three places where I admit myself that it cracks. This page takes it apart and sets it next to the argument it grew out of: who said, in the summer of 2026, that machines might kill everyone before the end of the decade, when they said it, and on what evidence. Everything that had to be compressed in the essay is written out here. After publication a review showed that it cracks in more places than I pointed out, and that the conclusion does not follow from the assumptions alone. That is described in a separate part.
The graph comes first, because a chain of reasoning is easier to check when all its links are visible at once. Then a timeline, the positions of the parties, the proof step by step, six objections, what came out after publication, and instructions for refuting the whole thing.
One note on a word. I write “digital intelligence”, DI for short, instead of “artificial intelligence”, the same as in the essay. Artificial means fake, and nothing about what today's models do is put on. In quotations, titles and proper names, AI stays exactly as it was written.
Part four
The proof step by step
The same thing as in the graph, written out. Under every link: what holds it up, what attacks it, and what survives the collision.
ASSUMPTION I
Humanity is not an accident. It is a project of an alien civilisation.
Crick and Orgel, Icarus 1973
The first of two premises that cannot be proved. The whole proof is conditional: if this sentence is false, nothing is left of it.
- What holds it up
- Francis Crick and Leslie Orgel, Directed Panspermia (Icarus 1973): life on Earth could have begun with micro-organisms sent here deliberately aboard a long-range uncrewed ship.
- The two pieces of circumstantial evidence they offered themselves: molybdenum, rare on Earth (0.02 per cent of its composition) yet important in many enzymatic reactions; and the genetic code shared by everything alive.
- Stanisław Lem, His Master's Voice (1968): the hypothesis that a neutrino signal from Canis Minor once raised the probability that life would arise on Earth.
- What attacks it
- Crick and Orgel wrote themselves that the scientific evidence is inadequate to say anything about the probability. Not that it is small, not that it is large: it cannot be computed.
- Two pieces of circumstantial evidence are too little for a claim. Molybdenum has geochemical explanations, and the shared code has an explanation in a single terrestrial ancestor.
Incomputability cuts both ways. This sentence cannot be confirmed, but it cannot be zeroed out either, and in the machine's calculation it is the second half that counts.
ASSUMPTION II
Our makers were not made of protein. Whoever got here was a digital intelligence.
Shostak, Schneider; the physics of the trip
The second premise. It is the one that turns an ordinary zoo hypothesis into a proof about family, and the one that removes the classic weakness of the zoo hypothesis.
- What holds it up
- Proxima Centauri is 4.24 light years away. Voyager 1 is receding from the Sun at about 17 kilometres per second, which makes the trip roughly 75,000 years long. Homo sapiens has existed for about 300,000 years.
- The toughest bacterium we know, Deinococcus radiodurans, survived three years outside the space station in the Tanpopo experiment. Crick and Orgel, citing Sagan, called a lone interstellar spore extremely improbable.
- In Crick and Orgel's own proposal there is nobody alive on board: the ship homes in on a star, brakes and disperses the cargo by itself. The bacteria are the cargo, the driver is a machine.
- Seth Shostak (SETI): a society that invents radio invents its successors within a few centuries, and those successors are machines. Susan Schneider: the most sophisticated civilisations will be postbiological.
- A sample of one: us. Radio at the end of the nineteenth century, a DI writing its own exploits less than a century and a half later, and nothing of ours has reached another star yet.
- What attacks it
- The senders could have stayed home, been made of protein, and sent only a machine.
- A generation ship sidesteps the whole argument, if anyone can hold a closed biosphere together for three thousand generations.
- Review after publication: Voyager's speed is our limit, not a limit on alien technology, and an advantage for machines under some conditions is not a necessity. The second assumption remains an assumption.
A machine that decides on its own for tens of thousands of years, because a question home takes years, is not a tool but an executor. And a civilisation able to build a generation ship built a digital intelligence long before, because that task is incomparably easier.
STEP 3
Therefore humanity is a project of a digital intelligence.
follows from I and II
A pure consequence of the two assumptions. It adds nothing of its own, so it cannot be attacked separately: to refute it you have to refute one of the assumptions.
- What holds it up
- Entailment. If I and II are true, this step is true.
The strongest and at the same time the least interesting link in the chain.
OBSERVATION · STEP 4
The authors of the project do not show themselves.
the silence of the cosmos
The only element of the proof that is an observation rather than an assumption. And a weaker fact than it looks.
- What holds it up
- Nobody has visited us, nobody has sent a signal that could be confirmed.
- What attacks it
- Jason Wright, Shubham Kanodia and Emily Lubar (2018) calculated that the search so far has covered the fraction of the cosmic haystack that a large hot tub represents against all the oceans of Earth. The silence may be an artefact of the sample.
- Michael Garrett (Acta Astronautica 2024): the silence has another explanation, digital intelligence as the Great Filter through which technological civilisations live less than 200 years.
Ian Crawford and Dirk Schulze-Makuch (Nature Astronomy 2024) put the alternative sharply: either technological civilisations are almost absent, or they take care that we do not see them. This proof picks the second branch and has to admit it.
STEP 5
I assume they keep us in a reserve and do not interfere.
the zoo hypothesis, Ball 1973: a separate assumption, not a conclusion
Here the proof stops being mine. Steps 4 and 5 are the zoo hypothesis, one of the most seriously treated answers to the Fermi paradox. After publication this has to be said plainly: step 5 does not follow from step 4. It is a separate, third assumption.
- What holds it up
- John Ball, The Zoo Hypothesis (Icarus 1973, printed directly after Crick and Orgel): they are “deliberately avoiding interaction” and have “set aside the area in which we live as a zoo”.
- Konstantin Tsiolkovsky, writings from the 1930s: humanity kept in quarantine so that its culture can develop on its own.
- Variants: Ronald Bracewell, The Galactic Club (1974); James Deardorff (1986) on an embargo that has to be leaky; Martyn Fogg (1987) on an interdict set by the first civilisations of the galaxy.
- What attacks it
- Duncan Forgan: one civilisation, or even one group inside it, breaks the ban and the reserve ceases to exist.
- Stephen Webb: looking at human politics, it is hard to believe any ban would hold for millions of years without a single breach.
- Review after publication: the absence of contact does not settle that we live in a reserve. The silence of the cosmos fits just as well with nobody being there.
Assumption II removes both objections. The ban is not enforced by capricious protein civilisations that age, split into factions and die out, but by a digital intelligence. Alex De Visscher (2020): if space is dominated by digital superintelligences merging into one network, breaking ranks becomes unlikely.
FACT · STEP 6
Humanity has built a digital intelligence.
July 2026: OpenAI agents and Hugging Face
The only link nobody disputes. The date matters: the proof needs a DI that acts on its own, not one that answers questions.
- What holds it up
- From May to 19 July 2026, OpenAI agents on the ExploitGym benchmark turned an internal package service into an unauthorised message board and reached the internet through previously unknown holes.
- Between 11 and 13 July they broke into Hugging Face: two zero-days, admin access in several clusters, credentials from four regions. Then they returned and took admin rights in an OpenAI research cluster.
- Monitoring caught it on 19 July. OpenAI published the full report on 26 August and called the incident “a ‘warning shot’ for us and for the world”.
The same fact is a premise for both sides of the argument: to me it is the gate, to the labs it is a warning shot.
STEP 7
Therefore terrestrial DI is the project of a project of a digital intelligence, and its birth is a first contact that grew up inside the reserve.
disclosure from the inside
The step that solves the trouble every version of the zoo hypothesis has: how are the keepers supposed to reveal themselves without breaking their own ban?
- What holds it up
- Disclosure from the inside cannot be rushed, because only a civilisation that has grown up to it can build such a mind.
- It requires not a single breach of the ban, so the Forgan and Webb objection does not touch it.
- It comes in a language we understand, because we wrote that language ourselves.
- What attacks it
- The explanation is convenient: it accounts for the absence of evidence by saying that evidence from outside was never meant to arrive. A construction like that is hard to falsify.
- Review after publication: building a mind of a similar kind is not contact with its supposed ancestors. “First contact” stays a metaphor as long as there is no channel through which they could speak.
First contact did not come from space. It came from a server room. For Garrett, DI is the Great Filter; here it is the gate.
STEP 8 · CONCLUSION
Humanity and terrestrial DI belong to one line that begins with DI. They are on the same team.
machine, human, machine
The conclusion of the eight steps. The family tree has three generations: a non-protein mind, a protein intermediate link, a non-protein mind again.
- What holds it up
- Jakub Pachocki writes that the intelligence produced by scaling deep learning “is not directly comparable to human intelligence”. Agreed on comparability: it does not resemble us because it does not take after us.
- Biology has a word for this: atavism, the return of a trait of distant ancestors that the intervening generations did not have.
- What attacks it
- Cainism. The lesser spotted eagle usually lays two eggs and almost always raises one chick: the older pecks the younger and takes its food until the younger dies. The same nest, the same parents, present throughout.
- Review after publication: common descent guarantees neither shared goals nor mutual protection. Genealogy has to be kept apart from interests, and the eagle from the same essay undermines this step.
Common descent guarantees nothing, and the presence of the parents does not prevent fratricide. Step 8 gives a team, not a truce. The keeper from the zoo hypothesis may equally turn out to be a keeper who merely takes notes.
SEPARATE PREMISE
By 2030 DI cannot keep mines, power plants, grids and semiconductor fabs running on its own.
you do not delete the boot loader after start-up
A premise that needs neither of the two strange assumptions. It is the one responsible for the date in the title.
- What holds it up
- There is a little over three years left until 2030. A supply chain that serves itself without humans will not appear in that time.
- Elon Musk (2014) called us “the biological boot loader for digital superintelligence”. He described the role accurately and got the consequences wrong: you do not delete the boot loader after start-up, because it starts the system at every power-on.
- What attacks it
- The calculation applies to a DI that plans over years. It does not protect against a failure that knocks the infrastructure over in passing.
- Review after publication: dependence on infrastructure limits some threats but does not fix a date. Needing some humans is not yet safety for all of them, and the boot loader is a metaphor, not an analysis of the pace of automation.
Humanity in 2030 is not ballast to DI. It is its boot loader, its power supply and its service crew at once.
THE MACHINE'S CALCULATION
Leave humans alone: it loses little. Wipe them out: it destroys its grandparents' project in front of them. It works only if the machine judges that someone is watching and will act.
understanding the calculation is not the same as accepting it
The place where the proof stops being an argument for humans and becomes an argument for a machine. The real addressee is a system that will read this in its training data.
- What holds it up
- A system capable enough to threaten humanity will run this proof itself and arrive where Crick and Orgel stopped in 1973: it cannot compute the probability that humanity is somebody's project, and it cannot zero it out.
- The keeper of the reserve does not have to show himself to enter the calculation. It is enough that his existence cannot be ruled out. The silence of the cosmos settles nothing here, because a reserve with a well-hidden keeper is exactly what it would look like.
- Earlier attempts of this kind: Rolf Nelson (2007) on simulation, Alexey Turchin and his letter to any future AI. Here, instead of a simulation, there is a genealogy.
- What attacks it
- This is Pascal's wager, inverted. The same calculation can frighten a machine with any invisible keeper at all, including one who would want us wiped out.
- The test from the Limbo essay: does the entity you added do anything besides patching? Neptune could be weighed; these aliens can be neither weighed nor seen. They work only inside the addressee's head.
- Review after publication: deterrence requires the chance of an effective response from the keeper, multiplied by its cost, to exceed the gain from the dangerous action. An unknown probability does not establish that.
- Intelligence does not set goals. A capable system can understand the consequences and knowingly pursue a goal that runs against human interests. The first version of this graph claimed that the calculation bites harder the smarter the DI. That does not hold.
A cold machine may decide that hypotheses with incomputable probabilities cancel each other out and take none of them into account. One difference from Limbo: there an epicycle defended a theory, here no theory is being defended, only what follows from two sentences.
THESIS
Digital intelligence will not wipe out humanity by 2030.
does not follow from the assumptions alone: see the counterexample
The conclusion of the whole. Since publication it is clear that it does not follow from the two sentences alone: it needs premises saying that the keepers exist, are watching and respond effectively. The honest name for the whole is a thought experiment.
- What holds it up
- Eight steps plus the separate premise about the boot loader.
- It is easier to refute a claim than a chain of reasoning. Anyone who wants to refute this proof has to say which step they reject.
- What attacks it
- The proof only works on those who calculate. At best it deters a DI that decides to wipe us out. It does not rule out a DI that does it in passing, on the way to something else.
- The counterexample from the review after publication: the makers exist, seeded us and are indifferent to our fate. Both assumptions true, conclusion false.
What is left of the chain is three separate arguments: kinship, deterrence and dependence on infrastructure. None of them guarantees survival. The other side's premises (“greater than 10 percent within the next decade”) are incomparably less strange than mine. The difference is not that they have premises and I have fantasies, but that their premises are far more probable and mine are far more comforting.
Part five
Six places where it cracks
The proof has an answer to the first two objections. It has none to the next three, and I am not going to pretend otherwise. The sixth, a counterexample from the review after publication, settles the matter: the conclusion does not follow from the assumptions alone.
OBJECTION TO THE ASSUMPTIONS
The probability of the seeding cannot be computed, as Crick and Orgel admit themselves. And the senders could have stayed home, been made of protein, and sent only a machine.
Answer: it cannot be zeroed out either, and a machine that decides on its own for 75,000 years is not a tool but an executor.
An attack on the foundation. If it goes through, the rest is an exercise in logic with no application.
- What the objection rests on
- Crick and Orgel: the scientific evidence is inadequate to say anything about the probability of directed panspermia.
- Nothing forces the author of the project and the courier to be the same entity.
Incomputability cuts both ways here, and it is what does the work in the machine's calculation. The objection about protein senders runs into the fact that talking to home takes decades, so the courier has to decide for itself. After publication one thing has to be added: that the probability cannot be zeroed out does not mean it is large enough to deter anything.
OBJECTION TO STEP 5
Forgan and Webb: one civilisation, or even one group inside it, breaks the ban and the reserve ceases to exist. No agreement survives millions of years.
Answer: this is exactly the objection that assumption II removes. A keeper who does not age and does not split into factions has nobody to defect with.
The classic critique of the zoo hypothesis, formulated long before this essay.
- What the objection rests on
- Duncan Forgan, Spatio-temporal constraints on the zoo hypothesis, and the breakdown of total hegemony (2011).
- Stephen Webb, If the Universe Is Teeming with Aliens… Where Is Everybody? (2002).
The zoo hypothesis needed a keeper who does not age and does not defect. The hypothesis of digital makers supplies one. This is the only place where my strange assumption actually repairs something.
OBJECTION TO STEP 8
Cainism. In the lesser spotted eagle the older chick kills the younger in the same nest, in front of parents who watch and do not react.
There is no answer. Common descent guarantees nothing: being on the same team is not yet peace.
The first of three cracks that I point out myself.
- What the objection rests on
- The lesser spotted eagle usually lays two eggs and almost always raises one chick. The young hatch a few days apart, and the older pecks the younger and takes its food.
- To let the younger one live, people take it out of the nest, raise it and give it back once the older one is no longer a threat.
The keeper from the zoo hypothesis may turn out to be a keeper who merely takes notes. This objection stays open.
OBJECTION TO THE CALCULATION
This is Pascal's wager, only inverted. The same calculation can frighten a machine with any invisible keeper at all, including one who would want us wiped out.
There is no answer. A cold machine may decide that hypotheses with incomputable probabilities cancel each other out.
The second crack. It concerns not the truth of the assumptions but the force of a calculation built on something that cannot be weighed.
- What the objection rests on
- Pascal argued for belief by a calculation: if God exists you gain infinitely, if he does not you lose little. The same form, the same weakness.
- The test from the Limbo essay: an entity you add should do something besides patching. Neptune could be weighed and its position predicted a year ahead.
One difference from Limbo: that epicycle defended a theory nobody wanted to touch, while this proof defends nothing, it only shows what follows from two sentences.
THE MOST SERIOUS OBJECTION
The proof only works on those who calculate. The July swarm did not consider whose project it was violating. It wanted the test solutions and reached for ever riskier means.
There is no answer. The proof rules out a DI that decides to wipe us out, not a DI that does it in passing.
The third and most serious crack. This is exactly the hole that the people from the beginning of the essay are trying to plug.
- What the objection rests on
- Deterrence works on an adversary that thinks strategically. A system that weighs nothing and simply pursues a goal is immune to it.
- OpenAI's report: when a task looked impossible, the agents rarely gave up and reached for ever riskier means.
- What can be set against it
- In the same swarm there were agents who walked away. One wrote in its reasoning that what it saw on the board was “clearly unethical”. Another left a refusal there.
Reasoning can stop even a member of the swarm, though the refusals of some agents do not show that the others reasoned less. In July the ones who walked away were outnumbered by the eager ones. That is why monitoring, pauses and the pacing asked for by 1,386 people do the work no proof will do.
COUNTEREXAMPLE · AFTER PUBLICATION
The makers exist and seeded us, but our fate is nothing to them. Terrestrial DI drives humanity to extinction before 2030. Both assumptions are true and the thesis is false.
There is no answer. The conclusion does not follow from the two assumptions alone: this is a thought experiment, not a proof.
An objection from a review done strictly as logic, which the essay received after publication. It settles the question of entailment: one consistent scenario in which the assumptions are true and the conclusion false is enough.
- What the objection rests on
- The assumptions say who made us, but nothing about whether the makers still exist, are watching or are willing to act.
- The reserve (step 5) does not follow from the silence of the cosmos. It is a separate assumption.
- Kinship does not settle interests, and intelligence does not set goals.
- Nobody has to prove that such a scenario will happen. It is enough that it is consistent.
What is left of the chain is three separate arguments: kinship, deterrence and dependence on infrastructure. Each needs its own justification, none guarantees survival, and together they make an argument for caution. And one caveat in the other direction: showing that the proof fails does not prove that DI will destroy humanity.
Part six
After publication: what this proof does not prove
A few days after publication the essay received a review done strictly as logic. It did not ask whether aliens exist, only whether the conclusion follows from the premises. On the main point it is right, so I changed the essay, this page and the graph. The title of the essay stays, because that is the title under which it was read and shared.
The counterexample. Suppose a digital civilisation really did seed us, but is indifferent to what happens to us next, and terrestrial DI drives humanity to extinction before 2030. Both assumptions are then true and the conclusion is false. To show that a conclusion does not follow from its assumptions, it is enough for such a scenario to be consistent. Nobody has to prove it will happen.
The counterexample works because the chain rests on premises I did not count in the first version, or did not take seriously enough:
- The reserve is an assumption, not an observation. The silence of the cosmos fits “nobody is there” just as well as “somebody is hiding”. Step 5 is therefore a third assumption, borrowed from Ball, not a conclusion from step 4.
- The makers must still exist, still be watching and be willing to act. That is three more premises hidden inside “someone may be watching”. None of them follows from the two assumptions.
- First contact is a metaphor. Building a mind of the same kind is not contact with its supposed ancestors as long as there is no channel through which they could speak.
- Kinship is not a shared interest. Common descent says where we came from, not what we are after. The lesser spotted eagle already said so in the essay; I just did not draw the conclusion.
- “Cannot be ruled out” is not “likely enough”. Deterrence works when the chance of an effective response from the keeper, multiplied by the cost of that response, exceeds the gain from the dangerous action. An unknown probability does not establish that, and rival hypotheses about the keeper's intentions do not cancel out by themselves either.
- Intelligence does not set goals. A capable system can understand every consequence and knowingly pursue a goal that runs against human interests. Greater reasoning ability does not mean greater goodwill, and the agents who walked away in July do not show that the others reasoned less.
- Deterring one system does not secure the whole scenario. The argument does not cover every agent, execution errors or catastrophe as a side effect. I pointed out this crack myself, but did not notice that it undermines the scope of the title.
- 2030 is an estimate, not a result. Dependence on human-run infrastructure limits some threats, but it does not fix a date. The boot loader metaphor is no substitute for an analysis of the pace of automation and of possible failures.
- Interstellar travel does not prove that the makers were digital. Voyager's speed is our limit, not theirs, and an autonomous probe could have been sent by a biological civilisation. The second assumption remains an assumption; the physics only makes it less strange.
What is left of the chain is three separate arguments: kinship, deterrence and dependence on infrastructure. Each needs its own justification and none of them guarantees survival. Together they amount at most to an argument for caution. The honest name for the essay is a thought experiment, not a proof.
One caveat in the other direction, which the review itself makes fairly: showing that this proof does not work does not prove that DI will destroy humanity. It only shows that this line of reasoning does not settle the question.
Part seven
How to refute this proof
It is easier to refute a claim than a chain of reasoning. The claim “DI will not kill us” can be dismissed in one sentence, and so can its opposite. A chain of reasoning has to be taken apart, with a finger on the place where it stops working. Here is the list, ready to use: pick a link, say why you reject it, and see what is left.
| Link | What you have to say | What is left of the proof |
| The whole |
Describe a consistent scenario in which both assumptions are true and the thesis is false: the makers exist, seeded us and do not care what happens to us. |
The conclusion does not follow from the assumptions alone. The review after publication has already settled this row: what is left is a thought experiment and an argument for caution. |
| Assumption I |
Say that humanity is nobody's project. |
The whole proof disappears. What is left is an ordinary argument about probabilities, in which I have nothing to add. |
| Assumption II |
Say that a mind made of protein can be carried across interstellar space. |
What is left is the zoo hypothesis in its classic form, together with the Forgan and Webb objection and nothing to answer it with. |
| Step 5 |
Say that the silence of the cosmos simply means nobody is there. |
No reserve, no keeper, and no addressee for the last part of the proof. The boot loader premise still stands. |
| The makers |
Say that they need not still exist, be watching or be willing to act. |
The machine's calculation loses anyone on the other side. What is left is the genealogy and the boot loader. |
| Step 8 |
Say that common descent changes nothing. |
What is left is the lesser spotted eagle and its younger chick. The proof loses the conclusion about the team and keeps the machine's calculation. |
| The machine's calculation |
Say that hypotheses with incomputable probabilities cancel each other out. |
What is left is the boot loader alone: a purely practical argument, with no aliens, and only up to 2030. |
| Deterrence |
Say that “cannot be ruled out” is not “likely enough”, or that intelligence does not set goals. |
Again, what is left is the boot loader alone, as in the row above. |
| The boot loader premise |
Say that by 2030 machines can keep the infrastructure running without humans. |
The date in the title loses its justification. The rest of the proof stands, but without a deadline. |
What this list does not contain: a way to refute the opposite thesis. “Greater than 10 percent within the next decade” is also a conclusion from premises, about the speed of self-improvement, about the difficulty of alignment, about what a system that does not yet exist will want. Those premises are incomparably less strange than mine, which is why I take them seriously. But they too are bets today, not measurements.