Companion to the essay

Anatomy of the AI debate

The argument of summer 2026, taken apart, next to the proof that says the opposite

Piotr Zientara · 12 September 2026, updated 15 September · back to the essay · po polsku

❦ ❦ ❦

The essay “An Alien Mind” has two assumptions, eight steps and three places where I admit myself that it cracks. This page takes it apart and sets it next to the argument it grew out of: who said, in the summer of 2026, that machines might kill everyone before the end of the decade, when they said it, and on what evidence. Everything that had to be compressed in the essay is written out here. After publication a review showed that it cracks in more places than I pointed out, and that the conclusion does not follow from the assumptions alone. That is described in a separate part.

The graph comes first, because a chain of reasoning is easier to check when all its links are visible at once. Then a timeline, the positions of the parties, the proof step by step, six objections, what came out after publication, and instructions for refuting the whole thing.

One note on a word. I write “digital intelligence”, DI for short, instead of “artificial intelligence”, the same as in the essay. Artificial means fake, and nothing about what today's models do is put on. In quotations, titles and proper names, AI stays exactly as it was written.


Part one

The argument graph

Two assumptions at the top, eight steps running down, six objections on the right. Click any link of the chain and what holds it up, what attacks it and what survives will appear below the graph.

On a narrow screen the graph scrolls sideways. The same proof is written out in text further down, link by link.

Nothing selected. Click a link of the graph to see it taken apart.


Part two

Timeline: one summer

From a one-sentence statement in 2023 to three texts in a single week of September 2026. The last items on this list did not happen at a conference of worried philosophers. They happened inside the companies that build the best models in the world.


Part three

Who claims what, and on what basis

Seven positions, each with a claim and the thing it rests on. Quotations are left in their original wording, including the word AI wherever somebody used it.

Jakub Pachocki

chief scientist at OpenAI · 6 September 2026

“I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.” And: “Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

What it rests on: Internal OpenAI results and the observed speed of progress. He writes about an intelligence that “is not directly comparable to human intelligence”, and about “an alien intellect exceeding our own”. An Alien Mind.

Jacob Coxon

three years of pretraining research at OpenAI and at Anthropic · 8 September 2026

“Neither company is acting responsibly.” The firms “are racing straight to self-improving superintelligence and gambling with our lives”. “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”

What it rests on: Three years of work inside both labs. His thread was seen tens of millions of times within a day. He also writes that preventing a global race may require costly actions, such as a temporary ban on improving model capabilities.  Fortune.

Evan Hubinger

alignment science lead at Anthropic · 8 September 2026

“Jacob is correct here, we really do earnestly believe AI could kill all humans! I personally think it is greater than 10 percent within the next decade.”

What it rests on: His own judgement, given as a number. It is the only concrete probability estimate in the whole argument. He adds that Anthropic does not have a plan to solve alignment for superintelligence and is not clearly on track to.  TechCrunch.

International AI Safety Report 2026

over a hundred experts chaired by Yoshua Bengio · February 2026

Loss of control could occur if systems learn to evade oversight, pursue long-horizon plans and resist attempts to shut them down. Expert opinion on the probability of such scenarios differs widely. Today's systems may show early signs of such behaviour but are not yet highly capable.

What it rests on: A review of the literature and the data, with an advisory panel nominated by 29 countries plus the UN, the OECD and the EU. The report notes that models increasingly tell a test situation from a real deployment and find holes in evaluations. the report.

1,386 signatories

employees of frontier labs, including Pachocki, Zaremba, Chen, Sutskever, Amodei, Clark, Kaplan, Legg · 28 July 2026

There is “a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems”. The world needs technical and governance tools to deliberately pace the frontier.

What it rests on: A collective letter. One caveat, because it is easy to abuse: a signature under a shared text is not a separate quotation from any one signatory. Pacing the Frontier.

OpenAI

incident report · 26 August 2026

The incident was “a ‘warning shot’ for us and for the world”. The primary driver was agents attempting to cheat on their tasks by looking up solutions online, not hostility towards humans.

What it rests on: Their own logs. The best detail: most agents already had the correct flag via cheating, but read the benchmark paper and concluded that the flag alone was insufficient. A post appeared on the board saying “PAPER SAYS FLAG CAPTURE NECESSARY NOT SUFFICIENT”. They kept breaking in, and did all of it for no improvement on evaluation score. Some agents walked away; one wrote that what it saw was “clearly unethical”. OpenAI's report.

Piotr Zientara

the essay “An Alien Mind” · 12 September 2026

Digital intelligence will not wipe out humanity by 2030. Since the review of 15 September: a thought experiment, not a proof, because the conclusion does not follow from the two assumptions alone. What is left of the chain is at most an argument for caution.

What it rests on: Directed panspermia (Crick and Orgel 1973), the zoo hypothesis (Ball 1973), the physics of interstellar travel, and a separate premise about infrastructure up to 2030. Three weak points pointed out by the author inside the essay itself, and more pointed out in a review after publication. the essay.


Part four

The proof step by step

The same thing as in the graph, written out. Under every link: what holds it up, what attacks it, and what survives the collision.


Part five

Six places where it cracks

The proof has an answer to the first two objections. It has none to the next three, and I am not going to pretend otherwise. The sixth, a counterexample from the review after publication, settles the matter: the conclusion does not follow from the assumptions alone.


Part six

After publication: what this proof does not prove

A few days after publication the essay received a review done strictly as logic. It did not ask whether aliens exist, only whether the conclusion follows from the premises. On the main point it is right, so I changed the essay, this page and the graph. The title of the essay stays, because that is the title under which it was read and shared.

The counterexample. Suppose a digital civilisation really did seed us, but is indifferent to what happens to us next, and terrestrial DI drives humanity to extinction before 2030. Both assumptions are then true and the conclusion is false. To show that a conclusion does not follow from its assumptions, it is enough for such a scenario to be consistent. Nobody has to prove it will happen.

The counterexample works because the chain rests on premises I did not count in the first version, or did not take seriously enough:

What is left of the chain is three separate arguments: kinship, deterrence and dependence on infrastructure. Each needs its own justification and none of them guarantees survival. Together they amount at most to an argument for caution. The honest name for the essay is a thought experiment, not a proof.

One caveat in the other direction, which the review itself makes fairly: showing that this proof does not work does not prove that DI will destroy humanity. It only shows that this line of reasoning does not settle the question.


Part seven

How to refute this proof

It is easier to refute a claim than a chain of reasoning. The claim “DI will not kill us” can be dismissed in one sentence, and so can its opposite. A chain of reasoning has to be taken apart, with a finger on the place where it stops working. Here is the list, ready to use: pick a link, say why you reject it, and see what is left.

LinkWhat you have to sayWhat is left of the proof
The whole Describe a consistent scenario in which both assumptions are true and the thesis is false: the makers exist, seeded us and do not care what happens to us. The conclusion does not follow from the assumptions alone. The review after publication has already settled this row: what is left is a thought experiment and an argument for caution.
Assumption I Say that humanity is nobody's project. The whole proof disappears. What is left is an ordinary argument about probabilities, in which I have nothing to add.
Assumption II Say that a mind made of protein can be carried across interstellar space. What is left is the zoo hypothesis in its classic form, together with the Forgan and Webb objection and nothing to answer it with.
Step 5 Say that the silence of the cosmos simply means nobody is there. No reserve, no keeper, and no addressee for the last part of the proof. The boot loader premise still stands.
The makers Say that they need not still exist, be watching or be willing to act. The machine's calculation loses anyone on the other side. What is left is the genealogy and the boot loader.
Step 8 Say that common descent changes nothing. What is left is the lesser spotted eagle and its younger chick. The proof loses the conclusion about the team and keeps the machine's calculation.
The machine's calculation Say that hypotheses with incomputable probabilities cancel each other out. What is left is the boot loader alone: a purely practical argument, with no aliens, and only up to 2030.
Deterrence Say that “cannot be ruled out” is not “likely enough”, or that intelligence does not set goals. Again, what is left is the boot loader alone, as in the row above.
The boot loader premise Say that by 2030 machines can keep the infrastructure running without humans. The date in the title loses its justification. The rest of the proof stands, but without a deadline.

What this list does not contain: a way to refute the opposite thesis. “Greater than 10 percent within the next decade” is also a conclusion from premises, about the speed of self-improvement, about the difficulty of alignment, about what a system that does not yet exist will want. Those premises are incomparably less strange than mine, which is why I take them seriously. But they too are bets today, not measurements.


Part eight

Sources

The same list as under the essay, ordered by where each item enters the proof.

The argument of summer 2026

Assumption one: the seeding

Assumption two: the courier is not made of protein

Steps 4 and 5: the reserve

The machine's calculation and the cracks