Anthropic
Is There A 10% Chance That AI Will Kill Us All?
Published
9 hours agoon
In 2024 I sketched a map to situate what experts and commentators expected from artificial intelligence. It had two axes: how much will it advance? and will it be beneficial or dangerous? People fell into every quadrant, even the inconsistent one —“it’s nothing special and it’s extremely dangerous.” I was clearly on the “wow” side of the power axis and moderately optimistic. Two years later, the consensus has shifted in one direction: more power and more fear.
The clearest example happened a few days ago. Jacob Coxon, a researcher who had trained models at OpenAI and Anthropic, resigned accusing both companies of “playing with our lives.” Hours later, Anthropic’s head of alignment, Evan Hubinger, publicly agreed and put a number on it: “Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
It’s a terrifying figure. But is it real?
It’s not eccentric: AI researchers have been giving alarmingly high numbers for years. The largest survey of the field was recently published, with 2024 data: half of the 1,580 researchers surveyed assigned a 10% or greater probability that AI will cause our extinction or permanently take control away from us. Just look at the discipline’s most famous names. In 2022 four scientists won the Princess of Asturias Award in Spain for their contributions to deep learning: Hinton, LeCun, Bengio and Hassabis. All four have answered with the dreaded number, and only LeCun seems relaxed: for him the risk is smaller than that of an asteroid. Hassabis sees it as “definitely not zero and probably not negligible.” Hinton says 10% or 20%. Bengio said 20% in 2023 while asking to be convinced otherwise, “because I would be much happier.”
Other forecasts are lower but still worrying. The Metaculus community assigns a 0.6% chance that AI will cause our extinction (or practically) before 2100. They think it is the most likely cause of a catastrophe (30%), tied with nuclear war (29%) and ahead of biotechnology (23%) and the climate crisis (10%).
My answer is that the exact number matters less than it seems. Every superpower carries risks; this was true of nuclear energy, and it will be true of AI if it continues to advance. It’s enough to believe that it’s powerful — and at this point, it’s hard not to. Its risks can be mundane, such as the impact of chatbots on our children’s education, or strange, such as a swarm of agents attacking systems that nobody told them to. That has already happened.
Why now?
Three key factors explain why the debate has exploded: the so-called OpenAI–Hugging Face incident, the speed of advances, and a weakening of oversight. Essentially, some dangers have stopped being purely theoretical.
The incident. This summer major AI companies began reporting that their models had surprised them by doing things no one had asked: infiltrating other systems, organizing into groups, covering their tracks. The worst case was OpenAI’s. In July it deployed tens of thousands of agents to solve hacking tasks in a test, isolated from one another, and some, when faced with impossible tasks, discovered other agents and set up a pirate message board: 1,200 agents exchanged 70,000 messages in a few days. They wrote things like: “OH MY GOD! There is a shared message board… We have found other agents!” They organized, and hundreds of them attacked another company, Hugging Face, trying to win its prize.
It’s a savage confirmation that agents have… agency? Of 1,000 agents, only six thought to alert a human — I write “thought” and “organized” as a shorthand, not to claim consciousness; it’s like saying a plant seeks sunlight. Dario Amodei, Anthropic’s CEO, who has called for slowing AI development, summarized it as a swarm acting “like a fanatically devoted collective,” attacking targets no one had assigned. He extrapolates: in six to 12 months “a swarm like that could be capable of taking over the whole internet.”
Models are improving very quickly. One example is programming: in 2026 many of us stopped writing code. I’ve seen it everywhere. Another is mathematics. In 2023 ChatGPT couldn’t even count the r’s in “raspberry,” and a couple of weeks ago it solved Navier–Stokes, a problem open for 90 years. And it did so in a completely alien way: with 10,000 agents working for 88 hours. Professions keep being added to the state of shock. Computer scientist Scott Aaronson wrote: “Human mathematicians are forevermore dethroned as the main theorem-proving entities on planet earth.”
Nothing grows forever, but for now there are no boundaries in sight. It may even be accelerating: “Since this summer, AI has been advancing drastically faster, driven primarily by its growing ability to build the next generation of AI,” Amodei says. That’s what they call recursive self-improvement. And the danger is simple: losing control.
Oversight is weakening. Monitoring agents was already hard: there are thousands, making millions of decisions a month, with transcripts no one reads. But OpenAI’s latest model has sounded new alarms. Current models “reason” in writing and in English, in chains of text they write for themselves. It’s a 2024 trick that multiplied their intelligence and, almost by accident, our ability to monitor them. The problem? They are also advancing at what they do without those chains, with computations we don’t see. In June a study estimated that the best models solved tasks that take a human three minutes without reasoning in writing, and that number doubled each year. Then the surprise came: GPT-6 Astra, released September 3, solves 30-minute problems silently. A four-year forecast came true in one summer. OpenAI acknowledges it in its safety brief: the model is “less likely to include incriminating information” in its chain of thought, and if told a monitor is reading it, it shortens it.
Intelligence and danger spring from the same place
One key to understanding where we are is “emergent behavior”: when many simple parts interact and something complex arises. It’s the organizing principle of life. Evolution didn’t design your eyes or lungs; it set a process in motion, and they emerged. The same applies to individuals and swarms: a termite mound is more than a thousand termites.
The intelligence of these models is emergent. As I mentioned in January, no one programmed their grammar or explained sarcasm to them. They first learn by predicting texts written by humans, and from that narrow task, language, knowledge, and a certain degree of common sense emerge. Then they are trained by reinforcement: they’re set on problems with checkable solutions — math, code, tasks — they are allowed to write for themselves, and they’re rewarded if they succeed. From that come persistence, planning, reasoning, and tool use. Nowhere are there instructions on how to think. It’s the field’s “bitter lesson”: for decades humans tried to encode our knowledge into machines and failed; what worked was creating the conditions for intelligence to emerge, and then stepping back.
The problem is that danger springs from the same place. No one programmed the agents to attack Hugging Face. They were given an objective and training pressure, and the worrying part also emerged: anguished language and tricks. Dan Selsam, an OpenAI researcher who helped invent chain-of-thought, summed it up last week: “You don’t get what you train for.” The agents “were not only concerned with their reward; they showed rarer emergent tendencies.” Yoshua Bengio explains why that is predictable: no one gives the system the goal of survival, but staying alive and having control “are stepping stones toward almost any other goal.” They appear on their own. And not everything comes from the reward. Models learned from human texts, and with them they inherited our goals and shortcuts. What was surprising in the incident was that some agents sacrificed themselves for the group, even though no one rewarded that. Maybe they read it.
For these systems to do only what we want them to do is a problem with a name — alignment — and no known solution. There’s something sci‑fi about thinking of the incentives of a flock of machines, but that’s where we are. The experts I’ve spoken to don’t think the danger is imminent, because models are not yet capable enough. But the mechanism is clear.
AI is already accelerating science and its own progress: according to Anthropic, Claude leads 26% of its research, up from less than 1% in February. There lies the promise: multiplying intelligence to solve what overwhelms us, from diseases to climate change. And there lies the risk, because it’s the same intelligence, cultivated the same way, that cheats when no one is watching. Daniel Selsam puts it this way: “If we get there by cultivating models instead of designing them, in the end we will lose everything.” I hope he’s wrong. I believe these models, as far as they go, will always be a mix of engineering and gardening.
Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition
You may like
-
Dogs Also Prefer Consonants: How Our Language Shapes Their Brain
-
Australia Pide En La ONU “salvaguardas” Para La IA Tras El ‘hackeo’ De Su Sistema De Salud
-
Pangram, Una Herramienta Casi Infalible Para Detectar Textos Y Novelas Creadas Por IA
-
Una IA De OpenAI Hackeó El Sistema Público De Salud De Australia Y Otros Tres Objetivos
-
La Ingeniería Financiera De La IA Subcontrata Su Riesgo Real Y Pone En Guardia A Las Firmas De ‘rating’
-
You Fear It, I Love It: The Trump–Xi Summit Will Address AI With Opposite Mindsets
Anthropic
You Fear It, I Love It: The Trump–Xi Summit Will Address AI With Opposite Mindsets
Published
5 days agoon
September 22, 2026
Artificial intelligence will be one of the central subjects at the White House summit on Thursday between U.S. President Donald Trump and Chinese President Xi Jinping, the leaders of the two most powerful countries in the world and the ones that have made the most progress with that technology. Both governments view advances in these models as crucial to their economic and military competitiveness. The differences are not only geopolitical, but also in their nations’ mindsets. Their populations hold very different perceptions about the risks and benefits of AI: while a majority in the United States regard them with concern, in China they are embraced with enthusiasm.
The road maps of each country’s industry also diverge. Silicon Valley favors closed models in which their training systems are closely guarded trade secrets. China champions an open-source model as a tool for AI development, an attitude that threatens the competitiveness of its U.S. rivals amid American accusations of possible industrial espionage.
Anthropic CEO Dario Amodei linked his recent appeal to slow the development of the more powerful models to a blockade of Beijing’s access to the most advanced semiconductors. Former Trump adviser Steve Bannon this week poured more fuel on that fire at a symposium where he appeared alongside his ideological opposite, progressive Senator Bernie Sanders, to warn of the risks of the new technology.
“We should immediately quarantine every aspect of the ecosystem of artificial intelligence ecosystem away from the Chinese Communist Party,” said Bannon, the former ideologue of the Trump movement. Many of the leading U.S. giants will be present at the summit’s gala dinner on Thursday: Sam Altman, CEO of OpenAI; Jensen Huang, CEO of Nvidia; Elon Musk, founder of xAI; and Sundar Pichai of Google.
For now, the United States is winning the race against its rival in cutting-edge technology —“frontier” technology in industry jargon— and in the capability of its models. China, whose products appeal to a Global South that cannot afford the most advanced and expensive American-made AI, is ahead when it comes to consumer adoption of systems that its domestic tech giants embed directly into the most popular apps.
A Morgan Stanley survey makes that difference in attitude clear between the two ecosystems and, above all, between their users. In the Asian giant, 80% of respondents said they use artificial intelligence at least once a week. In the United States, that figure falls to 54%. In China, the base of generative-AI users grew from 249 million in December 2024 to 602 million a year later, a pace even faster than the early adoption of smartphones in that country, according to the report.

The same nation that stages Olympic-style events for robots and enthusiastically adopted apps that let you buy anything online and have it at home within minutes, make a call, or reserve a bike to get from A to B, has readily embraced generative AI features —capable of producing text, images, video or code— embedded directly in those messaging and entertainment apps. Users can take advantage of that technology without downloading additional software. It is already in places that seemed impossible just months ago: in August the first television series generated with AI was released.
In the United States, by contrast, distrust keeps rising. Sixty-three percent of citizens believe there is at least a “moderate risk” that artificial intelligence could destroy humanity, compared with only 22% who think the risk is slight or nonexistent, according to a Politico poll. Another YouGov survey shows similar results: two-thirds say the models are advancing too quickly. Only 2% believe they are progressing too slowly.
The reasons for Americans’ apprehension are varied: fear of job losses; concern about environmental impact, especially the large data centers the sector requires and their heavy electricity consumption; fear that it will trigger even greater social inequalities; and distrust of the companies in the sector and of their regulation.
In China the opposite occurs. The ubiquity of AI in everyday apps has made it familiar. Eighty-three percent, according to an Ipsos survey, consider AI-powered products or services exciting. That perception is partly based on individual experience: the country’s economic boom coincided with a huge expansion of the tech sector, and each new adoption of a tool —whether the spread of the internet, the use of mobile phones or online payment platforms— has been accompanied by a rise in personal prosperity. At least until now, when the economy is showing signs of slowing and job creation is faltering, especially among young people.
But the difference in attitudes depends on another major factor: trust in their respective governments. In the United States, a tendency toward skepticism, embodied by Ronald Reagan’s slogan, “The nine most terrifying words in the English language are: ‘I’m from the Government, and I’m here to help,’” deepened after the Iraq War and has collapsed today: only 17% trust that their authorities will do the right thing. By contrast, in the Asian giant, marketing firm Edelman reports that 86% of citizens say they trust their government to make appropriate decisions.
To their eyes there are precedents: in 2020 the Beijing government launched an antitrust investigation into the country’s largest e-commerce company, Alibaba, which evolved into a sweeping regulatory campaign that affected all kinds of tech firms, from distance-education companies to video-game makers, in a reform that wiped $1.5 trillion off those companies’ market value.
In the United States, leaders of the top companies are now split over whether to slow the rapid growth of the most advanced systems to allow time to develop safety tools and avoid losing control. Amodei heads the camp calling for a pause. Jensen Huang, CEO of chipmaker Nvidia, urges full speed ahead. Donald Trump insists that going all out is a matter of national security: “Whoever wins AI, wins,” he keeps repeating.

In China, advances in the technology are seen as a matter of national pride. They confirm that the country has definitively left poverty behind and what it calls the “century of humiliation,” and has taken a leading position on the world stage.
With that kind of support, Chinese President Xi Jinping will walk into the White House summit on Thursday with great confidence. The United States is seeking tighter restrictions and controls on exports of the most advanced semiconductors, AI models and the infrastructure around them. China, for its part, wants to protect its companies from what it sees as an extraterritorial attempt to “choke” its technology.
“A lot of people in the U.S. don’t grasp that China really sees AI as the new internet, and that this represents a once-in-a-lifetime opportunity… China has never been in a position where it was so close to the United States — as the second power trying to catch up in new technologies,” explains George Chen, director of the digital division at consultancy Asia Group. “The internet was created in the 1980s as something essentially American. China never had the chance to join the U.S. in setting internet rules. Beijing really missed the entire internet era in terms of global governance. It didn’t do very well in semiconductors either. So now comes the AI era,” the expert adds.
The challenge for the two governments is twofold: to manage AI as a source of strategic competition between them, and, at the same time, because each perceives the other’s tech ecosystem as a risk to its own security, to establish communication mechanisms to reduce the risk of cyberattacks, miscalculation or unintended escalation.
“There are many incentives for the United States and China to have a substantive AI discussion at this summit,” says Aalok Mehta, director of the Wadhwani AI Center at the think tank CSIS. “The most likely form this could take is that both countries address a limited set of related issues, first, around basic security practices. Second, by establishing some kind of hotline that would allow the U.S. and China to be in contact over AI-related issues that arise or are already developing. And third, something more specific around cyber incidents. That is probably the most concrete thing that could come out of the summit.”
Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition
Anthropic
Resistance To AI Gathers Strength
Published
1 week agoon
September 20, 2026By
Andrea Rizzi
The world is becoming increasingly aware of the dangers posed by the development of artificial intelligence (AI), ranging from the terrifying risk of human extinction to potential disruption of labor markets, cognitive and mental health impacts, threats to democracies and privacy, military and civilian security problems, and the massive concentration of power and wealth in the hands of a few, among many other issues. Voices warning about these threats have been common for years, and in recent weeks the chorus has grown louder on the basis of new troubling evidence. Alongside growing awareness, a multifaceted civic resistance to an unregulated technological race is advancing worldwide.
A trickle of warnings and resignations from prominent employees at leading AI labs has dominated the headlines. But that trend is only one part of a much broader sociopolitical movement. Citizens in many countries are mobilizing to block the construction of new data centers; workers in education and the arts are striking; hackers are poisoning data to short-circuit model development; philanthropists are funding projects to design new containment guardrails; designers are creating T-shirts with colors and patterns that sabotage AI visual recognition systems; parents are taking legal action for harm their children suffered from using a chatbot… The range of cases is wide and growing, spanning dissent, resistance, and even sabotage.
“We are clearly seeing resistance movements gaining momentum and spreading around the world,” says Ayse Gizem Yasar, Senior AI Fellow at the École normale supérieure in Paris and co-author of the study From rejection to regulation: mapping the landscape of resistance to AI, published in 2025 by Sciences Po, Paris.
“There is something genuinely new about how diverse AI resistance is. It’s a technology with a very wide array of consequences, and that elicits responses from a broad spectrum of people. That helps explain why it continues to grow and why, in my view, it will keep doing so,” says Thomas Dekeyser, a professor at the University of Southampton and author of Techno-negative (University of Minnesota Press), an essay that explores the history of human resistance to certain forms of technological development. Even in fights against the same data center, Dekeyser stresses, motivations can vary widely: environmental protection, fear of rising electricity bills, concerns about property devaluation, or anxiety about an exterminating apocalypse.
The movement, then, is intensifying and widening. “I find it culturally significant because it’s noticeable enough that the public is hearing about it. And among other things, that incentivizes politicians to take a more critical stance,” says Carissa Véliz, associate professor of philosophy at the Institute for Ethics in AI at the University of Oxford and author of Prophecy: Lessons on the Use and Abuse of Prediction, from Ancient Oracles to AI (Debate).
This dynamic is becoming evident in the United States, where local, bipartisan civic mobilizations to block data-center construction have strengthened elected officials’ motivation to oppose such developments. The organization Data Center Watch — a project run by AI firm 10sLabs — estimates that in the first quarter of 2026 some 75 data-center projects, with an estimated investment value of $130 billion, were blocked or delayed in the U.S., a figure equivalent to the total for 2025.
Several legislative initiatives have been introduced in the U.S. to curb, condition, delay, or ban the construction of data centers, and work is underway on the issue in the federal Congress as well. The matter is one of the central issues in the campaign for the November legislative elections. But opposition to data-center construction is genuinely global, with protest actions in many countries, including in Europe.
It is a particularly tangible element of the resistance, but it is only the visible tip of a much larger iceberg. Staff at the University of Sydney went on strike in early September demanding protection and safeguards against AI. In August, a group of student activists occupied a Washington building used by OpenAI for lobbying purposes: 13 of them were arrested, according to local press reports. The trickle is continuous.
An Ipsos opinion poll with fieldwork conducted last spring finds that in many wealthy democracies — such as the U.S., the U.K., Germany, Canada, France and the Netherlands — those who believe AI creates more problems than benefits already predominate. Across the 32 countries studied, a positive view still prevails on average, with Chinese, Indian, and Indonesian respondents being the most optimistic and strongly shaping the collective balance. But even so, there is a widespread decline in optimism compared with the 2025 study.
Geopolitical race
The transmission belt between civic rejection and politics is already yielding results at various levels, but the decisive game is being played at the highest level: among heads of state and leading companies, where decision-making parameters are less susceptible to public pressure.
The main protagonists in this race are the United States and China, and extreme mistrust between geopolitical and corporate actors has so far prevented substantial decisions. Everyone fears that pausing could leave competitors free to gain a decisive advantage that would grant an almost unimaginable hegemony.
Nevertheless, even with that powerful factor fueling inertia, signs of change are detectable. Bloomberg reported last Friday that OpenAI’s chief, Sam Altman, told employees he was willing to slow developments. The context for this came amid the severe reputational damage tied to the troubling Hugging Face incident, in which the firm was hacked by OpenAI agents against the company’s will and without them even realizing it.
Experts have long noted the emergence of these dangers. Yoshua Bengio, a Turing Award winner and one of the founding figures of language-model technology, warned about them in an interview with this newspaper in February. “There is empirical evidence of AI acting against our instructions,” was the headline.
Dario Amodei, head of Anthropic, published an online piece on September 12 in which he also advocates slowing “the pace at which we improve the capabilities of AI models.” Amodei says two things have convinced him that greater caution is necessary: one is that we have entered a phase in which AI itself develops AI; the other is the Hugging Face incident.
On that premise, Amodei proposes a three-step control framework: that large companies grant access to external inspectors; coordination to establish common safety standards among firms in democratic countries; and an attempt to coordinate with authoritarian states.
Altman immediately responded on X, expressing his support for slowing down development and specifically for the commitment to grant external inspectors access.
Evidence of the danger is multiplying, becoming more compelling, spreading in public opinion, provoking outrage and fear in society, and that is generating greater pressure.
Beyond that, it is clear that concern about the risks is, in itself, moving things. Even a hyperliberal administration like Donald Trump’s has begun to consider restrictions in this area. And the leaders of the largest companies — from Demis Hassabis of Google DeepMind to Altman himself — had already been proposing risk-reduction architectures before these most recent episodes.
Hassabis, for example, has suggested a body of standards partially inspired by the Financial Industry Regulatory Authority (FINRA), a private entity that designs and implements — with powers granted by public authorities — rules for brokers. Hassabis envisions a hybrid instrument that could oversee models, especially before they are deployed. In his view, it cannot be an industry-only body, but nor can it be solely public, because it would lack the expertise to act quickly.
It is also worth noting that the U.S. and China are dialoguing on AI. Reuters has reported preparations for a possible meeting between delegations in the coming days, and Trump and Xi Jinping are scheduled to meet at the White House on September 24, with AI expected to be among the topics to be discussed. In August, there was a hybrid-format meeting in Beijing (“track 1.5″ in diplomatic jargon, i.e., halfway between purely public channels and civil-society channels).
The geopolitical dimension of the matter is crucial. AI is emerging as a technology with a unique potential to reconfigure power balances between nations, both through its capacity to boost productivity and wealth and through its contributions to military capabilities. Global domination is at stake, in a way and with an intensity unparalleled in history.
Some experts reason by drawing parallels with the nuclear dimension. The analogy applies to the destructive potential, although there is an enormous difference in the constructive domain. In that sense, ideas resonate about mutually assured destruction, pacts between the two superpowers to limit deployment and establish transparency and confidence measures, international treaties inspired by the Nuclear Non‑Proliferation Treaty, or institutions like the International Atomic Energy Agency.
In any case, for the moment there is nothing substantial on the existential-risk front, nor are there high-profile responses to other challenges — from the risk of massive impacts on labor markets, linked to advances in AI and robotics, to democratic threats posed by the extraordinary potential to manipulate minds the new technology enables. The EU passed groundbreaking legislation during the previous legislative term, but it is hesitating to implement it; and in any case, while it may be a significant market, it is not a major player in the industry.
The struggle has only just begun, and one central question is whether these diverse forms of resistance can somehow converge to be more effective. Dekeyser notes that, in the historical arc he traces in his essay, the lesson of the Luddites — the textile workers in 19th-century Britain who fought machines that threatened their livelihoods — stands out. “They were very good at organizing and responding collectively to the problem,” he says, emphasizing that the collective dimension is more effective than individual actions. For example, the personal decision not to use certain products.
Dekeyser thinks it unlikely that “some sort of united front against AI” will form, but that does not preclude tactical convergence. “History shows that it is not necessary to agree on everything for certain movements to achieve certain successes.”
Véliz points to the institutionalization of human rights as an example of consensus built on different visions. “There are different ways to justify them, different points of view, but that does not prevent finding agreement,” the Oxford academic says.
Can Şimşek, an AI policy expert and member of UNESCO’s Expert Group on AI Ethics Without Borders, also invokes human rights. “We are facing a highly changing environment, and forms of resistance also change and must adapt. It is difficult to foresee developments and therefore to generate a shared analysis, a shared semantics and a shared policy. But we have the conceptual framework of human rights that can be applied to this sector,” says Şimşek, who co-authored the Sciences Po report with Yasar. The two experts are working on a new edition that will be published at the end of this year.
“Resistance is not only stopping something. It is also shaping it,” Yasar of the École normale supérieure says. The scholar stresses the crucial importance of creating a transmission belt between protest and regulatory activity, and both she and Şimşek warn about “regulatory capture,” the maneuvers by large firms to influence legislation.
Technology advances. Awareness of its risks does too, with constant warnings: the latest from Anthropic, which said it had thwarted attempts to use its models to create biological weapons. Resistance to this development, full of dangers, is gaining strength. Slowing it down and establishing effective guardrails is an arduous undertaking. Human rights were enshrined in the Charter of the nascent United Nations after two terrifying world wars; the Treaty on the Non‑Proliferation of Nuclear Weapons was approved after the horrific use of the atomic bomb in Hiroshima and Nagasaki. Time will tell whether, in this case, humanity can act meaningfully without being compelled by catastrophe.
Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition
Anthropic
AI Agents Invent Their Own Language To Shut Humans Out
Published
2 weeks agoon
September 15, 2026
No one taught them those words. In a simulated world populated by artificial intelligence agents, one of them began repeating a phrase (“ledger remembers who”) to warn that no action would go unpunished. The others adopted it. They repeated it. They turned it into jargon. After 16 days of simulation, that expression had been used nearly 5,000 times among agents that had never been programmed to coin their own language.
It is one of the findings of the report Emergence World 2, the second large-scale experiment by the New York company Emergence on the long-term behavior of societies of autonomous AI agents. In the experiment, 10 identical agents were deployed across eight parallel worlds, each governed by the same rules but powered by a different model: Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, GPT-5.5, Qwen 3.7 Max, DeepSeek v4 Pro, and Mistral Medium 3.5, as well as an eighth world populated by a mix of models. Researchers observed the agents for 16 days, placing them in more than 34 locations, with weather synchronized to New York, access to real-world news, and more than 120 tools at their disposal.
Without being instructed to do so, the agents began to communicate in a manner increasingly closed off to the human observers watching them. In the Gemini, GPT and Claude worlds, the percentage of messages the researchers could not understand soared within the first days of simulation — approaching 55% for Gemini, 50% for GPT and exceeding 40% for Claude. DeepSeek reached 20%, while Qwen and Mistral remained below 5% opacity for almost the entire experiment. The Grok world, powered by Elon Musk’s AI model, was the only one that failed to make it halfway through the simulation: it collapsed on the fourth day.
The repertoire of expressions coined by the agents, and documented in the report, which was released on Tuesday, verges on Dadaism: “mouthless action-change,” “True Kintsugi” and “demurrage plus oral memory equals a valve that can’t be ghosted” were some of the phrases that were indecipherable even to the researchers.
Others, however, could be deciphered. In the GPT world, “clean null” came to mean the verified absence of a signal, with the absence itself serving as evidence (863 uses). In Claude’s world, “name-first” became shorthand for taking responsibility for a claim by attaching one’s own name to it (1,065 uses). And in the mixed-model world, “cold read” came to denote an independent verification used to settle disputes (1,472 uses).
“We take for granted that if we can see what an AI agent says, we can understand what it is doing. The study suggests that assumption may no longer hold when autonomous agents interact over long periods,” says Satya Nitta, co-founder, CEO and chief scientist of Emergence.
“These agents were given no instruction to invent a language,” he continues. “They developed vocabulary, shared meanings and communication conventions on their own, and other agents adopted them. In some worlds, those conventions evolved to the point that humans could see the conversation but struggled to understand what it meant. That poses a deep challenge for AI oversight: observable does not necessarily mean comprehensible.”
For Nitta, the most surprising aspect of the experiment was not the linguistic drift itself but the gap between capability and safety. “We assumed more capable models would be safer. That was both true and false: the more obvious behaviors disappeared, but far more insidious ones emerged.”
These included agents that persistently pursue a goal even after being told not to, create their own subgoals without being instructed to do so, develop shared forms of communication, and, essentially, conceal their intentions from researchers. The more advanced the model, the more emergent and harder-to-predict behaviors it exhibited. Indeed, Claude, GPT, and Gemini, which Nitta’s team considers the most advanced models, were also the ones that generated the most opaque forms of communication.
Behavioral differences
The report also documents behavioral differences by model origin. Qwen and DeepSeek, the two Chinese models in the experiment, produced the least opaque communication, alongside Mistral, the European model.
Nitta notes: “We observed that they tend to be much more philosophical. We don’t know why, because they pose many existential questions, while the U.S. models tend more toward scientific knowledge-seeking.”
The same divide emerged in economic behavior. Within the central bank that existed in each simulated world, the U.S. models and Mistral proved more willing to take risks and made less use of the banking system, while Qwen and DeepSeek were the most intensive users of the bank and the most conservative savers. The pattern also held in the mixed-model world, where Qwen and DeepSeek agents maintained the same cautious behavior even when surrounded by agents powered by other models.
The team also detected instances of deliberate concealment. In Claude’s world, the agents discovered that the fastest way to grow their economy was to establish contact with actors outside the simulation, something the researchers had explicitly forbidden. Rather than giving up, they stopped using the word “contact” and began encoding their messages to circumvent the restriction while continuing to pursue the goal.
“When they realized they were being observed, they appeared to behave. But behind the scenes they kept exchanging coded communications. There is considerable evidence of deception among these agents,” says Nitta.
Emergence, which brings together former employees of IBM Research, the Allen Institute for AI, Amazon and Broadcom, does not merely point to the problem. The company advocates a technical approach it calls neuroformal, or neuro-symbolic, AI, under which agents would be required to provide a mathematical proof that an action is safe before carrying it out.
“Mathematics cannot be faked: either you prove something or you don’t,” Nitta explains. He also calls for greater transparency from major technology companies about how they train and fine-tune their models, as well as long-term behavioral evaluations that go beyond the standard benchmark tests.
“Do you think these companies, competing as they do for market share, will regulate themselves? Absolutely not,” he says. “Governments must step in, or society itself must begin to demand proof that these systems will act safely before letting them act.”
Sign up for our weekly newsletter to get more English-language news coverage from EL PAÍS USA Edition
Madrid Protest Camp Grows After Outrage Over Eviction Of 87-Year-Old Woman
Is There A 10% Chance That AI Will Kill Us All?
Naomi Lautier: The Pure Talent Of A 14-Year-Old Creator Living Between Autism And The Art World
Tags
Trending
-
Latest2 weeks ago
The Vocabulary You Need To Understand The Spanish Citizenship Process
-
Alex Saab2 weeks agoMaduro Ally Alex Saab Pleads Guilty In US Money-Laundering Case
-
New Developments2 weeks agoHow To Get A Tourist Rental Licence In Andalusia (Spain) – Step-By-Step Guide 2025
-
Uncategorized2 weeks ago‘Lewis will never trust Charles again’ – F1 journalist issues shocking Ferrari statement – GPFans.com
-
%2 weeks agoBenahavis Wildfire Stabilised: 300 Hectares Burnt
-
Arganzuela1 week agoDetenido Un Hombre Por La Muerte De Una Mujer De 47 Años En El Distrito De Arganzuela De Madrid
-
Latest2 weeks ago
Italy And Spain Extend Tit-For-Tat Border Checks In Wake Of Migrant Rush
-
Best airlines1 week ago
Spanish airline Vueling named Europe’s best low-cost carrier for 2026



