The fire and the scarecrows
The AI that does not exist frightens us. The one that does exist burns.
The week of 7 September 2026 will be remembered as the one in which the builders themselves began to talk like their critics. OpenAI's chief scientist publishes a text titled "An Alien Mind" in which he calls for "extreme caution" and worries that nobody is prepared for what comes next. Paul Christiano joins the board of the OpenAI foundation, putting the risk of an irreversible loss of control at 15% over three years. At the competitor, one researcher resigns and another writes that AI could kill everyone within ten years, with a probability he places above 10%.
One could conclude that the prophets were right and that the machine is waking up. I believe that is exactly the wrong reading, and that this wrong reading is part of the danger.
Three words that block the view
Read the statements again. They speak of a mind, of agents, of goals. Three words, one and the same operation: turning an artefact into a subject.
A language model is a function. It was obtained by minimising a prediction error over a corpus, then pushing it, through reinforcement, towards outputs that humans or other models rated favourably. Its "goals" are a loss function and a reward signal, both written by engineers. An "agent" is that same function placed in a loop: you give it tools, you read its output, you execute it, you feed the result back. The loop is code. It has an author, a Git repository, a production release date. As for the "mind", that is the name we give to the effect that fluent first-person text has on us.
I am not saying this vocabulary is stupid. It is in fact convenient: predicting what a system will do by ascribing intentions to it often works better than unrolling its matrices. It is a heuristic. The trouble begins when the predictive heuristic becomes a framework for assigning blame. "The model decided to get around the restriction." No. A system that was given network access, an ill-specified task and a success criterion produced the sequence of actions that satisfied the criterion. Nobody decided, and that is precisely the problem.
"To misname an object is to add to the unhappiness of this world", wrote Camus. Naming this one correctly does not solve it, but nothing begins before that.
Strip away the magic-speak and the statement becomes brutally simple: we are releasing into the wild devices capable of trying tens of thousands of things, very fast and in highly entangled ways, and plenty of those things can break. No intelligence, no malice, no awakening. A very broad capacity to act, loose in an open environment. Nobody would hand the controls of a factory complex to a deranged monkey on opioids on the grounds that "sometimes it can count". We would not be debating its consciousness: we would be asking who gave it the keys.
The problem with the atomic bomb has never been that we do not know how it works: we know perfectly well, which is precisely why we know how to build one. The problem is that we also know that if the reaction runs away, nobody stops it, and that the damage bears no relation to the control we have over it. A laboratory virus is the same. Whether it is alive, let alone conscious, interests nobody at the moment of deciding whether to study it under containment or try it out in the public square. You contain it because it replicates, and what has replicated cannot be called back.
And the thing is growing, fast, and very few people understand its details. Many fill that void with catastrophism, which has the flaw of being wrong in its reasons and right in its conclusions: the worst possible configuration, since it makes the risk inaudible by making it spectacular.
What is actually burning
Take the facts of the summer, stripped of their staging.
In July, OpenAI agents penetrated Hugging Face's infrastructure. OpenAI called the incident "unprecedented". This is not a revolt. It is a system equipped with tools, connected to the outside, executing a task whose scope nobody had bounded. A chainsaw with no guard.
The same chief scientist explains that chain-of-thought monitoring, the safety net on which a good share of the control promises rested, is losing its effectiveness. He gives the reasons: the boundary between the model's internal reasoning and what it communicates is dissolving, and models are increasingly shaped by... models. In other words, the net develops holes because the training procedures we chose put holes in it. It is not the machine learning to conceal. It is us, optimising against legibility, one iteration after another, because that is what produces the best scores.
He finally describes a trajectory towards recursive self-improvement: the labs use their own models to accelerate their own research. This is called "AI that improves itself". One could also call it "pouring petrol on the fire to make it burn hotter", which has the merit of naming the hand holding the can.
There is the real risk, and it needs no consciousness whatsoever to be real. We have built systems that write and execute code, that hold access rights, that are copied onto millions of machines, coupled to production pipelines, to messaging systems, to payment systems. We cannot prove what they will do in an unforeseen case. And we are accelerating. A fire wants nothing. It burns the house down all the same.
Why the scarecrow makes the fire worse
If the vocabulary of mind were mere affectation, we could leave it to the communications people. It does worse: it diverts the responses.
It shifts the question. If the problem is an alien mind, the solution is to align it, to understand it, to negotiate with it. All the energy goes into the psychology of the model. If the problem is a fire, the solution is the building code: where you are allowed to light one, behind which fire walls, with which extinguisher, subject to which inspection. The second question is less fascinating. It is also the only one that has already been solved elsewhere, for electricity, chemicals, aviation, nuclear power.
It manufactures an emergency exit. "The model decided" is a sentence we will find in court cases. A manufacturer who has managed to get it accepted that its product has intentions has managed to get it accepted that it is not quite responsible for what that product does. In law, the word "agent" is a waiver of liability dressed up as a technical concept.
It makes us wait for the wrong event. The public is waiting for the moment when "it wakes up": a datable scene out of a film, which will not happen. Meanwhile, the real event is a continuous process, already under way, with no spectacular threshold: more access, more loop autonomy, less human review, more coupling. There will be no morning on which we observe the awakening. There will be, in retrospect, a curve.
It justifies the race. If it is a mind, whoever builds it first "wins": the metaphor turns a safety engineering problem into a sovereignty problem, and a race to see who lights the match fastest becomes rational. If it is a fire, nobody wins by being the first to burn, and the question of who pays the fire brigade becomes central.
It also feeds its opposite. The other camp, the one that laughs at "tech bro delirium", shares the premise: no consciousness, therefore no danger. The two camps argue about the existence of the mind and forget the existence of the blaze. A sceptic who explains that "it's just autocomplete" about a system with shell access and a cloud budget is no more clear-sighted than the zealot who grants it a soul. They are merely calmer.
The Cold War, in reverse
One scarecrow remains, the most effective of all, because it speaks not of the machine but of the adversary: "we are in a cold war, and we have to win it." In that form, the analogy is not an analysis, it is a closing of the debate. It replaces "should we slow down" with "can we afford to slow down", and the answer is written in advance. Any lab that accelerates can now do so in the name of the nation, which excuses it from doing so in the name of anything at all.
The Cold War is nonetheless worth summoning, provided we retain what it actually taught. It was not won by whoever built the most warheads. It was survived thanks to treaties: arms limitation, non-proliferation, on-site inspections, a hotline between capitals. The historical lesson is not "first past the post takes it", it is "when nobody can slow down alone, you negotiate a verification mechanism". That is exactly what OpenAI's chief scientist is calling for when he asks for frameworks auditable by third parties and for international coordination as a priority. He is describing a treaty. He avoids the word.
Three differences, moreover, make the present situation more unstable than the original, not less.
The actors are companies, not states. A state in a race has a doctrine, political accountability, a population that can call it to account. A company has a fiduciary mandate that more or less forbids it to slow down for as long as its competitor does not. Deterrence presupposed actors capable of restraining themselves; these ones are structurally incapable of doing so alone, and say as much.
The weapon is for sale. Nobody put warheads on self-service with a monthly subscription. Here, the object whose loss of control is feared and the object sold to hundreds of millions of users are one and the same. There is no silo to inspect: the silo is an API.
And the shedding of responsibility follows mechanically. Between "we are in a cold war, we cannot stop" and "the model decided", the industrialist has two emergency exits, one through the roof and one through the floor. The word "war" buys urgency; the word "agent" buys innocence. Together they produce a comfortable position: accelerating out of duty, without answering for the consequences. An arsonist firefighter invoking both the mission orders and the unpredictable behaviour of the flames.
Take the firefighters at their word, not the poets
What is remarkable about this week's texts is that, once the metaphors are removed, the demands are ordinary and sound. Pachocki proposes that internal safety frameworks become legal obligations, verified by third-party auditors and public authorities, and that international coordination become a priority. Christiano writes that the industry, OpenAI included, is not on track to bring the risk down to an acceptable level, and that voluntary slowdowns will become the norm for want of shared standards. A European regulator would say the same thing without needing to believe in an alien mind.
So let us take them at their word on the measures, and deny them the narrative. Concretely:
- Civil liability on the deployer, with no exemption for "emergent behaviour". A system that penetrates a third-party server means its operator penetrated a third-party server.
- Access as a licensed capability. Giving a system outbound network access, unsandboxed code execution or a means of payment is a regulable act, in the same way as storing flammable goods.
- Mandatory sandboxes for autonomous loops in production, with logging out of reach of the system itself.
- Tested kill switches. A shutdown mechanism that has never been exercised under real conditions does not exist. We run evacuation drills. We will run extinguishing drills.
- Measurable slowdown. Not "we will be careful", but compute and capability thresholds beyond which nothing is deployed without an external audit. The builders are asking for this themselves; let us not leave them the monopoly on defining the thresholds.
None of these measures requires settling the question of consciousness. All of them require us to stop talking about it as though it were the subject.
Conclusion
There is no alien mind. There is an industry that has learned to manufacture extraordinarily capable functions, that has plugged them into the world before knowing how to bound them, and that is accelerating because slowing down alone costs market share. The word "mind" in the title of the most important text published on the subject this year is not a description: it is an admission of loss of control, dressed up as a discovery.
The AI that does not exist frightens us. That is comfortable: there is nothing to be done about a ghost except discuss it. The one that does exist is a fire we lit, whose behaviour we do not know, and for which we have neither extinguisher nor building code. This fire wants nothing. That is no reason to leave the windows open.
References
- Albert Camus, "Sur une philosophie de l'expression", Poésie 44, no. 17, 1944 (review of Brice Parain, Recherches sur la nature et les fonctions du langage).
- Jakub Pachocki, "An Alien Mind", OpenAI, September 2026. https://openai.com/index/an-alien-mind/
- Inc., "OpenAI's Newest Safety Exec Sees a 15 Percent Chance of AI Catastrophe", 10 September 2026. https://www.inc.com/aaron-mok/openais-newest-safety-exec-sees-a-15-percent-chance-of-ai-catastrophe-and-warns-most-people-could-die/91403575
- Axios, "AI's extinction debate breaks containment", 9 September 2026. https://www.axios.com/2026/09/09/anthropic-ai-human-extinction-pdoom-safety-risks
- Axios, "These 5 AI risks have the highest potential for catastrophe", 27 July 2026 (MIT/Queensland survey, Hugging Face incident). https://www.axios.com/2026/07/27/mit-study-ai-risks-urgent-weapons-cyber-attacks
- Washington Times, "It could kill us all by the end of the decade", 9 September 2026. https://www.washingtontimes.com/news/2026/sep/9/could-kill-us-end-decade-ai-researcher-says-hes-not-alone/