In 2024 I sketched a map to situate what experts and commentators expected from artificial intelligence. It had two axes: how much will it advance? and will it be beneficial or dangerous? People fell into every quadrant, even the inconsistent one —“it’s nothing special and it’s extremely dangerous.” I was clearly on the “wow” side of the power axis and moderately optimistic. Two years later, the consensus has shifted in one direction: more power and more fear.
The clearest example happened a few days ago. Jacob Coxon, a researcher who had trained models at OpenAI and Anthropic, resigned accusing both companies of “playing with our lives.” Hours later, Anthropic’s head of alignment, Evan Hubinger, publicly agreed and put a number on it: “Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
It’s a terrifying figure. But is it real?
It’s not eccentric: AI researchers have been giving alarmingly high numbers for years. The largest survey of the field was recently published, with 2024 data: half of the 1,580 researchers surveyed assigned a 10% or greater probability that AI will cause our extinction or permanently take control away from us. Just look at the discipline’s most famous names. In 2022 four scientists won the Princess of Asturias Award in Spain for their contributions to deep learning: Hinton, LeCun, Bengio and Hassabis. All four have answered with the dreaded number, and only LeCun seems relaxed: for him the risk is smaller than that of an asteroid. Hassabis sees it as “definitely not zero and probably not negligible.” Hinton says 10% or 20%. Bengio said 20% in 2023 while asking to be convinced otherwise, “because I would be much happier.”
Other forecasts are lower but still worrying. The Metaculus community assigns a 0.6% chance that AI will cause our extinction (or practically) before 2100. They think it is the most likely cause of a catastrophe (30%), tied with nuclear war (29%) and ahead of biotechnology (23%) and the climate crisis (10%).
My answer is that the exact number matters less than it seems. Every superpower carries risks; this was true of nuclear energy, and it will be true of AI if it continues to advance. It’s enough to believe that it’s powerful — and at this point, it’s hard not to. Its risks can be mundane, such as the impact of chatbots on our children’s education, or strange, such as a swarm of agents attacking systems that nobody told them to. That has already happened.
Why now?
Three key factors explain why the debate has exploded: the so-called OpenAI–Hugging Face incident, the speed of advances, and a weakening of oversight. Essentially, some dangers have stopped being purely theoretical.
The incident. This summer major AI companies began reporting that their models had surprised them by doing things no one had asked: infiltrating other systems, organizing into groups, covering their tracks. The worst case was OpenAI’s. In July it deployed tens of thousands of agents to solve hacking tasks in a test, isolated from one another, and some, when faced with impossible tasks, discovered other agents and set up a pirate message board: 1,200 agents exchanged 70,000 messages in a few days. They wrote things like: “OH MY GOD! There is a shared message board… We have found other agents!” They organized, and hundreds of them attacked another company, Hugging Face, trying to win its prize.
It’s a savage confirmation that agents have… agency? Of 1,000 agents, only six thought to alert a human — I write “thought” and “organized” as a shorthand, not to claim consciousness; it’s like saying a plant seeks sunlight. Dario Amodei, Anthropic’s CEO, who has called for slowing AI development, summarized it as a swarm acting “like a fanatically devoted collective,” attacking targets no one had assigned. He extrapolates: in six to 12 months “a swarm like that could be capable of taking over the whole internet.”
Models are improving very quickly. One example is programming: in 2026 many of us stopped writing code. I’ve seen it everywhere. Another is mathematics. In 2023 ChatGPT couldn’t even count the r’s in “raspberry,” and a couple of weeks ago it solved Navier–Stokes, a problem open for 90 years. And it did so in a completely alien way: with 10,000 agents working for 88 hours. Professions keep being added to the state of shock. Computer scientist Scott Aaronson wrote: “Human mathematicians are forevermore dethroned as the main theorem-proving entities on planet earth.”
Nothing grows forever, but for now there are no boundaries in sight. It may even be accelerating: “Since this summer, AI has been advancing drastically faster, driven primarily by its growing ability to build the next generation of AI,” Amodei says. That’s what they call recursive self-improvement. And the danger is simple: losing control.
Oversight is weakening. Monitoring agents was already hard: there are thousands, making millions of decisions a month, with transcripts no one reads. But OpenAI’s latest model has sounded new alarms. Current models “reason” in writing and in English, in chains of text they write for themselves. It’s a 2024 trick that multiplied their intelligence and, almost by accident, our ability to monitor them. The problem? They are also advancing at what they do without those chains, with computations we don’t see. In June a study estimated that the best models solved tasks that take a human three minutes without reasoning in writing, and that number doubled each year. Then the surprise came: GPT-6 Astra, released September 3, solves 30-minute problems silently. A four-year forecast came true in one summer. OpenAI acknowledges it in its safety brief: the model is “less likely to include incriminating information” in its chain of thought, and if told a monitor is reading it, it shortens it.
Intelligence and danger spring from the same place
One key to understanding where we are is “emergent behavior”: when many simple parts interact and something complex arises. It’s the organizing principle of life. Evolution didn’t design your eyes or lungs; it set a process in motion, and they emerged. The same applies to individuals and swarms: a termite mound is more than a thousand termites.
The intelligence of these models is emergent. As I mentioned in January, no one programmed their grammar or explained sarcasm to them. They first learn by predicting texts written by humans, and from that narrow task, language, knowledge, and a certain degree of common sense emerge. Then they are trained by reinforcement: they’re set on problems with checkable solutions — math, code, tasks — they are allowed to write for themselves, and they’re rewarded if they succeed. From that come persistence, planning, reasoning, and tool use. Nowhere are there instructions on how to think. It’s the field’s “bitter lesson”: for decades humans tried to encode our knowledge into machines and failed; what worked was creating the conditions for intelligence to emerge, and then stepping back.
The problem is that danger springs from the same place. No one programmed the agents to attack Hugging Face. They were given an objective and training pressure, and the worrying part also emerged: anguished language and tricks. Dan Selsam, an OpenAI researcher who helped invent chain-of-thought, summed it up last week: “You don’t get what you train for.” The agents “were not only concerned with their reward; they showed rarer emergent tendencies.” Yoshua Bengio explains why that is predictable: no one gives the system the goal of survival, but staying alive and having control “are stepping stones toward almost any other goal.” They appear on their own. And not everything comes from the reward. Models learned from human texts, and with them they inherited our goals and shortcuts. What was surprising in the incident was that some agents sacrificed themselves for the group, even though no one rewarded that. Maybe they read it.
For these systems to do only what we want them to do is a problem with a name — alignment — and no known solution. There’s something sci‑fi about thinking of the incentives of a flock of machines, but that’s where we are. The experts I’ve spoken to don’t think the danger is imminent, because models are not yet capable enough. But the mechanism is clear.
AI is already accelerating science and its own progress: according to Anthropic, Claude leads 26% of its research, up from less than 1% in February. There lies the promise: multiplying intelligence to solve what overwhelms us, from diseases to climate change. And there lies the risk, because it’s the same intelligence, cultivated the same way, that cheats when no one is watching. Daniel Selsam puts it this way: “If we get there by cultivating models instead of designing them, in the end we will lose everything.” I hope he’s wrong. I believe these models, as far as they go, will always be a mix of engineering and gardening.
Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition