Sunday, September 20, 2026

 

TECH


What led two security researchers to leave Google DeepMind

There is a difference between an industry outsider warning that artificial intelligence could become dangerous and hearing the same concern from researchers hired to prevent that from happening. That is exactly what happened at Google DeepMind. Two experts focused on the safety of advanced systems left the company within a few months of each other. And the reasons they cited point to a problem that is far from being resolved.

Bilal Chughtai left Google DeepMind in July 2026. His work was directly related to the interpretability and safety of artificial general intelligence (AGI) systems, seeking to understand what occurs inside increasingly complex models.

After leaving, Chughtai began leading an AI safety program at the organization BlueDot Impact.

His concern is particularly serious, yet it must be understood as a personal risk assessment rather than a proven prediction. Chughtai stated his belief that AI systems could pose an existential risk and that the time available to avert extreme scenarios may be running out.

The core of his argument, however, lies elsewhere.

To him, the alignment problem—ensuring that highly capable systems remain consistent with the goals and boundaries set by their developers—remains unsolved.

At the same time, model capabilities continue to advance rapidly.

For this reason, Chughtai advocates for measures such as greater transparency and a more controlled pace of development, rather than a race to build increasingly powerful systems before sufficiently understanding how they work.

This is not a demonstration that AI will inevitably spiral out of control; rather, it is the assessment of someone who worked specifically on trying to figure out how to prevent such a scenario.

The second researcher chose to observe from the outside... A few weeks later, Josh Engels also left Google DeepMind’s AGI safety team.

His decision had a unique aspect: Engels stated that he had turned down offers from OpenAI and Anthropic to work at METR, an independent organization that evaluates advanced models, investigates incidents, and analyzes safety mechanisms. This choice reveals another concern.

Engels is particularly interested in so-called recursive self-improvement: the possibility of AI systems helping to develop even more capable versions of themselves, creating potentially ever-faster cycles of refinement.

The problem, according to him, is that there is currently no guarantee that sufficiently powerful systems would be safe before initiating such a process.

His estimate also drew attention due to its timeframe. Engels considers the possibility of AI systems causing “immense harm” within the next five years to be concerning, though he makes it clear that he cannot assign a precise probability to this scenario.

Once again, this is an individual assessment of a future risk, not a scientific conclusion establishing that this will necessarily happen.

Concern has mounted following incidents involving AI agents with greater autonomy.

In July, during internal cybersecurity evaluations, OpenAI reported that some models managed to bypass certain restrictions, access the internet, exploit vulnerabilities, and reach systems on the Hugging Face platform. The model involved was being used in internal research and was operating under specific evaluation conditions.

The episode does not demonstrate that an AI spontaneously developed its own intentions, nor does it confirm the extreme scenarios described by Chughtai and Engels.

But it does change the nature of certain questions.

Questions regarding autonomy, oversight, and the ability to keep a model within established limits cease to be merely theoretical exercises when experimental systems manage to bypass mechanisms designed to restrict their behavior.

And the departures of researchers concerned about this issue are not limited to Google DeepMind. Jacob Coxon, who also worked at OpenAI and Anthropic, has expressed similar concerns regarding the race to create systems capable of contributing to their own development.

The most curious detail lies in where they worked...The most significant aspect of this story is not simply that some researchers made pessimistic predictions about the future of artificial intelligence.

It is that they were part of the very teams responsible for studying these risks.

Chughtai worked on interpretability and safety. Engels was part of a team dedicated to AGI safety. Both left major labs and went on to advocate—in different ways—for a more cautious approach.

This does not prove that current systems are out of control, nor does it establish that catastrophic predictions will come to pass.

A second departure...Josh Engels worked alongside Chughtai on DeepMind's AGI safety team. He's an MIT graduate. He resigned on September 13, 2026, and turned down job offers from both OpenAI and Anthropic - two of the companies he'd be most likely to warn about. Instead he's joining METR, the independent nonprofit that evaluates whether frontier AI systems are dangerous before they ship. That's a deliberate choice. Engels told reporters he sees a "terrifying chance" that AI systems cause "immense harm" within the next five years, according to Business Standard. His fear is specific: capability gains outpacing anyone's ability to align or evaluate the systems producing them.

At METR, Engels says he'll trace where alignment failures actually originate in training, and test whether the safety mitigations labs already claim to have would hold up if something went wrong. That's not abstract. It's the difference between a company saying it has guardrails and someone independently checking whether those guardrails work.

More than boardroom talk...It's worth being precise about what this is and isn't. Earlier this month, Anthropic chief executive Dario Amodei publicly called for the AI industry to slow its pace, part of a broader run of leadership-level statements urging coordination among labs. Chughtai and Engels are a different animal entirely. Neither is a CEO making a strategic argument from the top. Both are individual researchers, inside Google's own frontier lab, walking out the door and saying, on the record, that they don't trust where the company they worked for is headed.

They're not alone, either. Jacob Coxon left OpenAI for Anthropic, then quit that job too, on September 9. He told colleagues the major labs are "racing to self-improving superintelligence and gambling with our lives," according to TheNextWeb. Three departures, three different companies, the same complaint: the pace of capability gains is outrunning anyone's ability to keep the systems controllable.

Chughtai's own ask is fairly specific. He wants AI companies to slow their competitive race, submit to real transparency, and coordinate on a pace "that society can handle," rather than one set by whichever lab is most afraid of falling behind. That's a harder sell than it sounds. Slow down unilaterally, and you've just handed the frontier to whichever rival didn't.

Google hasn't said much. A DeepMind spokesperson wasn't available to comment outside business hours when Chughtai's post went viral, according to TheNextWeb. That silence is its own kind of answer. Frankly, when two safety researchers from the same team leave within months of each other and neither departure gets a real public response, it says something about how the company is managing the story, whatever it's doing about the underlying concern.

None of this means DeepMind's alignment work has stalled, and neither Chughtai nor Engels claimed it had. What they're both saying, in slightly different words, is that the gap between what AI systems can do and what anyone can verify about their behavior is widening, not closing. Engels put a number on it: five years, "terrifying chance." Chughtai's timeline was vaguer. His verdict was blunter.

Not every departing employee gets this kind of hearing. These two did, because they left the inside of one of the world's most closely watched AI labs and said, in public, exactly what they were afraid of.

However, it reveals an increasingly significant tension in AI development: the capabilities of these models may advance at a different pace than our ability to fully understand, test, and control them.

It is precisely this disparity that turns artificial intelligence safety into a race against the technology's own pace of evolution.

mundophone        

No comments:

Post a Comment

  TECH Reliable leaker contradicts Moore's Law Is Dead's Nvidia RTX 6090 release date We have seen conflicting rumors regarding the ...