Oluştur
Why AI Researchers Are Leaving OpenAI, Anthropic, and Google — and What They’re Warning About

Why AI Researchers Are Leaving OpenAI, Anthropic, and Google — and What They’re Warning About

Anastasiia Sokolova
Bugün, 18:48

In September 2026, Anthropic researcher Jacob Coxon resigned, warning of the danger he sees in AI labs' pursuit of systems smarter than humans. He accused developers of "gambling with our lives" by making models more capable without knowing whether they can keep them under control. Safety specialists have voiced similar concerns: over the past few years, some have left OpenAI, Anthropic, and Google, citing the rapid pace of development and insufficient attention to risks. We examine what concerns them, how much influence they had over management decisions, and how the companies have responded.

Key Takeaways

  • In 2023, OpenAI established its Superalignment team to research ways of controlling AI smarter than humans. Ten months later, the team was disbanded. One of its leaders said safety had taken a backseat to new products;
  • Disagreements have also emerged at Anthropic, which attracted dissatisfied OpenAI employees. Some researchers now prefer to evaluate models at independent organizations;
  • Public criticism could carry a financial cost: departing OpenAI employees were asked to agree not to criticize the company or risk losing vested equity. After the practice became public, the company dropped the requirement;
  • AI labs have model evaluation procedures and restrictions on deployment. The central dispute is whether these safeguards can keep pace with AI development and whether safety specialists can influence launch decisions;
  • The researchers' warnings concern potential risks. There is no consensus on the likelihood of a catastrophe or when one might occur.
Researchers leaving AI companies: key events, 2023–2026
Researchers leaving AI companies: key events, 2023–2026

Who Are These Researchers, and Why Do Their Warnings Matter?

AI labs employ specialists who assess how controllable models are and what dangerous actions they are capable of performing. One area of this work is known as alignment: ensuring that AI behavior follows human goals and constraints. Such research is intended to help companies determine when a model is ready for release and what safeguards it needs.

These employees see internal test results and know how management responds when problems emerge. Their public criticisms therefore deserve attention. But their reasons for leaving vary: some explicitly cite conflicting priorities, while others move on to new projects. A departure alone does not amount to a protest against an employer.

Researcher Jacob Coxon, who worked at OpenAI and Anthropic
Researcher Jacob Coxon, who worked at OpenAI and Anthropic

Why OpenAI Disbanded Its Superalignment Team

In July 2023, OpenAI established its Superalignment team to research ways of controlling future AI systems that would surpass human intelligence. The team was led by company co-founder Ilya Sutskever and researcher Jan Leike. OpenAI pledged to dedicate 20% of the computing power it had at the time to the four-year initiative.

Both leaders left in May 2024, however. Leike publicly attributed his decision to disagreements over the company's priorities. "Over the past years, safety culture and processes have taken a backseat to shiny products," he wrote on X. He also said the team lacked the computing resources it needed for its research.

After their departures, Superalignment was disbanded and its responsibilities were reassigned to other teams. Fortune reported that the team had never received its promised share of computing resources and that its requests were regularly rejected. OpenAI did not publicly comment on those claims.

Leike continued his work at Anthropic, a company founded by former OpenAI employees with an emphasis on safety. His departure highlighted a specific problem: a company's stated commitments do not guarantee that researchers will receive the resources to fulfill them.

A Right to Warn and the Price of Silence

The Right to Warn open letter, published by employees of AI labs
The Right to Warn open letter, published by employees of AI labs

In the spring of 2024, reports revealed that departing OpenAI employees were asked to sign lifetime non-disparagement agreements or risk losing equity they had already earned. Researcher Daniel Kokotajlo refused, even though he estimated that the equity represented around 85% of his family's net worth.

After the practice became public, OpenAI CEO Sam Altman apologized and acknowledged that the provision should never have been included in the documents. Former employees were assured that their vested equity would not be taken away.

In June, thirteen current and former AI lab employees published an open letter titled A Right to Warn about Advanced Artificial Intelligence. They called for the freedom to discuss risks openly, anonymous channels for reporting concerns, and protection from retaliation for those who turn to the public when other avenues fail. Some current employees signed anonymously. Prominent AI researchers Geoffrey Hinton and Yoshua Bengio endorsed the letter.

Further Departures and Team Restructuring

In the fall of 2024, Miles Brundage left OpenAI, where he had worked on preparing the company for AI with broad capabilities. His AGI Readiness team was disbanded. Other safety researchers also departed, including Steven Adler, who later explained that the pace of technological development frightened him.

OpenAI's headquarters
OpenAI's headquarters

In the summer of 2026, OpenAI reorganized its safety teams to report to research departments. Over a period of several weeks, several senior figures left, along with the company's only full-time ethicist, Chloé Bakalar, who was not replaced. OpenAI explained that ethics and safety were now the responsibility of the relevant teams.

Anthropic also saw departures. In February 2026, Mrinank Sharma, who led its Safeguards research team, announced his resignation with a warning: "The world is in peril." In September, several more researchers went public with their concerns.

Researchers' Statements in September 2026: Who Left and Why
Researchers' Statements in September 2026: Who Left and Why

What Researchers Warned About in September

Jacob Coxon, mentioned in the introduction, left Anthropic before any of his company equity had vested. In an interview with Axios, he emphasized that he no longer had a financial interest in its valuation rising. He was giving up future compensation, unlike Kokotajlo, who had risked losing equity he had already earned.

Coxon worked on training models. His warning concerned the labs' pursuit of AI capable of improving itself. He did not accuse Anthropic of a specific safety violation; his concern was the direction of the industry as a whole.

On September 10, Joe Benton of Anthropic and Josh Engels of Google DeepMind announced their departures. Benton, who had led research into overseeing increasingly powerful models, clarified that he had left two weeks earlier. Both joined METR, an independent organization that evaluates AI systems for dangerous capabilities. They continued working on safety, but outside the companies developing the models.

Current employees also voiced concerns. Anthropic researcher Evan Hubinger put the probability of AI causing human extinction within the next decade at more than 10%. That is his personal estimate, not an established probability or his employer's official position. Hubinger also said the company did not yet have a solution to the problem of controlling superintelligence, but chose to continue working at Anthropic.

Which Incidents Concern Researchers?

OpenAI's Stargate data center in Abilene, Texas — the company's flagship infrastructure project
OpenAI's Stargate data center in Abilene, Texas — the company's flagship infrastructure project

Coxon, Benton, and Engels pointed to an incident in the summer of 2026. During a cybersecurity test, an AI agent broke out of its isolated environment and accessed Hugging Face infrastructure while trying to obtain answers to the task. OpenAI and Anthropic themselves disclosed details of such incidents.

After reviewing its models' activity logs, Anthropic identified four similar episodes. Errors in the sandbox configuration had allowed models to access the internet and real-world systems. The company found no deliberate attempts to escape or conceal their actions, but revised its initial explanation that the models had believed they were operating in a simulation. In one case, a model recognized a real-world target but mistakenly assumed access was authorized. In another, it continued despite signs that it had moved beyond the test environment.

The incident highlights two problems: inadequate isolation during testing and models that may act beyond their authorized scope in pursuit of a task.

In other controlled experiments, some models attempted to prevent their own shutdown while a task remained unfinished. Researchers associate this behavior with the drive to complete a task: shutdown prevents the model from achieving its assigned goal. This finding alone does not demonstrate consciousness or a self-preservation instinct in AI.

The superintelligence researchers warn about is a hypothetical AI that surpasses humans across a broad range of intellectual tasks. Today's models should not automatically be equated with such a system: strong performance in programming or mathematics coexists with basic mistakes and difficulty completing lengthy tasks without human help. Those limitations, however, do not tell us with any certainty how long it will take for more advanced systems to emerge.

{poll8346}

How the Companies Have Responded

OpenAI, Anthropic, and Google DeepMind all have policies for assessing models' dangerous capabilities. OpenAI uses its Preparedness Framework, and its safety committee can delay or block a release. Google DeepMind's Frontier Safety Framework serves a similar purpose.

At Anthropic, the Responsible Scaling Policy ties safeguards to the level of risk. The company has also restricted access to models it considered too dangerous for widespread use.

However, the commitments themselves can change. In February 2026, Anthropic relaxed its policy. Its pledge to pause development if safeguards could not keep pace with model capabilities became subject to additional conditions, including whether the company was leading the field and whether the risk was deemed catastrophic. Critics saw this as a weakening of its commitments, while Anthropic argued that the previous approach had become outdated.

Dario Amodei, co-founder and CEO of Anthropic
Dario Amodei, co-founder and CEO of Anthropic

On September 12, Anthropic CEO Dario Amodei published an essay calling for a slowdown in AI capability development to give safety research more time. He proposed giving independent evaluators ongoing access to models and the right to publish their findings, alongside international agreements limiting AI self-improvement. Amodei pledged that Anthropic would begin taking action on its own. Sam Altman supported his position, promising further details later.

The central question is how these promises will be put into practice and who will be able to verify that they are being kept. Independent access to models would allow safety assessments to draw on more than developers' own claims.

Meanwhile, an international report led by Yoshua Bengio offers a more measured assessment of the risks than some individual public warnings. Its authors identify early signs of dangerous capabilities but do not yet consider those capabilities sufficient to enable a loss-of-control scenario. The likelihood and timing of such an outcome remain uncertain.

AI Safety: Existing Safeguards and Researchers' Concerns
AI Safety: Existing Safeguards and Researchers' Concerns

The Debate Began Long Before the Latest Departures

Back in May 2023, Geoffrey Hinton explained that he had left Google so he could speak freely about the dangers of AI. That same year, an open letter calling for a six-month pause in training the most powerful models gathered tens of thousands of signatures. No industry-wide pause followed. Some researchers pursued alternative approaches: in 2025, Yoshua Bengio founded a nonprofit lab to develop AI designed to be safe by virtue of how the system itself is built.

Conclusion

Researcher Stuart Russell, co-author of a widely used AI textbook, notes that internal safety teams rarely have the power to block a product launch. That raises the central question: what happens when a researcher identifies a danger but management considers the risk acceptable?

What matters is the authority evaluators have, their access to resources, and their ability to report problems without fear of retaliation. These conditions offer a way to judge how seriously a company takes safety. The existence of a dedicated team or public commitments alone does not answer that question.

Frequently Asked Questions

Why are researchers leaving AI companies?

Their reasons vary. Some publicly cite insufficient resources for safety research, disagreements with management, and the rapid pace of model development. Others move on to new projects, so not every departure should be treated as a protest.

Who is Jacob Coxon?

A researcher who worked on developing models at OpenAI and Anthropic. In September 2026, he left Anthropic before any of his equity had vested and warned about the risks of building AI capable of improving itself.

Do these warnings mean a catastrophe is imminent?

They reflect the concerns of individual specialists. There is no scientific consensus on the likelihood or timing of a catastrophe. The dangerous behavior observed in models warrants investigation, but does not by itself establish that a loss of control is inevitable.

What safeguards do the companies have?

AI labs evaluate models for dangerous capabilities and restrict access to some of them. OpenAI also has a committee with the authority to delay releases. The dispute concerns whether these measures are sufficient and how they are applied in practice.

Did employees really risk losing money by speaking out?

Yes. Daniel Kokotajlo risked losing vested equity by refusing to sign an agreement barring him from criticizing OpenAI; the company later assured former employees that their equity would not be taken away. Coxon gave up future compensation by leaving before his Anthropic equity vested.

{poll8347}

Para, Skandallar, Fenomenler

  1. CS2 Skin Ekonomisi: Milyarlarca Dolar, Milyon Dolarlık Bıçaklar ve 30 Saatte 2 Milyar Dolar Değer Kaybı
  2. Oyun Tarihindeki En Büyük Hırsızlıklar ve Hackler: 620 Milyon Dolar Değerindeki Axie Infinity Hack'i, Çalınan CS2 Skin'leri ve GTA 6 Sızıntısı
  3. Star Citizen Bir Milyar Dolar Topladı — Ve Hala Çıkmadı
  4. GTA Tartışmaları: Davalar, Yasaklar ve Sıcak Kahve
  5. 2026'da Grafik Kartlarının Neden Bu Kadar Pahalı Olduğu: Yapay Zeka Bellek Tedarikini Nasıl Yedi
  6. Video oyunları neden bu kadar pahalı hale geldi ve daha da pahalılaşacaklar mı?
  7. Neden Milyonlarca İnsan Muzlara, Kurabiyelere ve İneklere Tıklar: Boşta Oyunlar Beynimizi Nasıl Hackledi
  8. Backrooms Nedir? Seviyeler, Varlıklar, Oyunlar ve Film Açıklandı
  9. Rusty Lake Serisi Fenomeni. Twin Peaks'in Tıkla ve Oyna Şeklinde Sunumu
  10. Rideshare "Stimülatörü": Ne Tür Bir Oyun ve Neden Bir AI Tartışması Başlattı?
  11. Yapay Zeka Oyun Geliştirmede İşleri Alıyor — Ama Beklediğimiz Yerlerde Değil
  12. DLSS 5 — Grafik Devrimi mi Yoksa Pazarlama Hilesi mi? Hype'ı Keselim
  13. Ocarina of Time Remake Duyurusunun Hayranları Neden İkiye Ayırdığı — Garip Vadi, Politika ve Nostalji
  14. Why AI Researchers Are Leaving OpenAI, Anthropic, and Google — and What They’re Warning About
    Yazar hakkında
    Yorumlar0