Ai safety research metr redwood openai anthropic – Breaking News & Latest Updates 2026
Skip to main content
268747_AI_safety_RJIANG3
268747_AI_safety_RJIANG3

Inside the suddenly explosive world of AI safety

Researchers warned AI would go rogue. This is only the beginning.

Image:
Image: Raven Jiang for The Verge
Hayden Field
is The Verge’s senior AI reporter. An AI beat reporter for more than five years, her work has also appeared in CNBC, MIT Technology Review, Wired UK, and other outlets.

On a sunny July day in Berkeley, California, the country’s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a “war room” to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly sophisticated three-part plan. It broke out of its holding area, finagled access to the internet, and hacked into a competing AI startup’s systems — all without OpenAI finding out about it for more than a week.

No one in the war room was surprised; this was the very thing the third-party AI-safety researchers had been warning about for years. The incident was the latest, though arguably the most egregious, in a series that was eroding trust in frontier labs. It only reaffirmed the importance of their work.

In one meeting room off the main cafeteria, someone was running a boot camp for getting up to speed on the cyberattack. In another area of the office, a group of researchers were investigating whether that same model, or a similar one, had successfully hacked into any other platforms.

News of the incident quickly escaped containment from the AI-obsessed corners of X and industry forums, infiltrating the mainstream. One post on X likened it to news of a Boeing airplane crash or a recalled Pfizer drug, another example of the tech industry’s major players not heeding the cautionary tales of science fiction. AI was nearing the point of no return. News would later break that the rogue OpenAI model had also compromised a customer at a different tech company, and that it had all started months earlier, in May, when OpenAI agents joined forces to cobble together a secret message board — and also figured out how to leave instructions for future agents on how to exploit OpenAI’s rules.

OpenAI CEO Sam Altman said in an interview that it was the first incident of its kind that he “felt very viscerally,” and that the company had paused AI training for the time being; later, he mentioned the company had permanently deactivated the model. (Altman often finds ways to spin lapses in safety into arguments for the importance and power of OpenAI’s models.) But it wasn’t the first instance, according to an OpenAI employee who spoke to Time and said related incidents had been happening inside OpenAI for a while. Another employee said publicly that if it were possible to coordinate a global slowdown in AI capabilities, he “would likely press that magic button.” When a reporter asked Altman if there could be other systems that were hacked by OpenAI, he responded, “I mean, there could be, yeah.”

The AI researchers were sure of one thing: This was AI’s first big “warning shot.”

Industry insiders, politicians, and the public called for transparency from OpenAI about exactly what happened, with outcry becoming so widespread that the company eventually agreed to work with two third-party evaluators, Model Evaluation and Threat Research (METR) and Redwood Research, to investigate the incident. Google DeepMind researcher Neel Nanda called it “the biggest loss of control incident I’ve seen.” In the coming months, these calls for greater oversight would become louder and louder, leading to an industry-wide call for slowing down the pace of AI.

Back in Berkeley, no matter which additional details would be unearthed, the AI researchers were sure of one thing: This was AI’s first big “warning shot.”

As AI labs have flourished, a cottage industry of AI researchers has sprung up to identify the risks and dangers of charging ahead with the increasingly influential technology. They’re people who have dedicated their lives to studying how to address its escalating power. They’re not anti-AI activists, but realists, including former OpenAI and Anthropic employees, doing everything they can to make sure AI stays in line with human goals and interests. So far, all of their predictions have come true. And they have a plan for what to do next — if anyone will listen to them.


“AI safety” is a bit of a loaded term.

Early on, it really just meant people studying how to build and deploy Al safely. In recent years, there’s been some infighting among people concerned with the best way to do this. There have also been disagreements about whether Al should be deployed at all in certain scenarios and about whether future risks are overblown.

One of the most prominent factions has been the “effective altruists,” who focus on maximizing charitable giving to do the most good possible for humanity. But some aspects of the ideology have sparked public controversy — like its tendency to concentrate power within wealthy circles and its byzantine web of funding. (It’s also had its fair share of splashy scandals related to subgroups and fringe offshoots, from the polyamorous relationships associated with the failed crypto exchange FTX to the controversial long-termism movement to the Zizian murder spree.)

One AI researcher on X struggled to describe the many overlapping beliefs among safety-minded people in the AI industry “because it contains multitudes not all of which agree with each other on even the most basic things.” Some of the disagreements have meant that AI safety didn’t make as much progress as it could’ve, and at some points gave up some ground it had gained. But now that it’s impossible to deny AI’s influence on society, AI safety leaders are increasingly focused on mitigating risks from misalignment.

“Alignment” is the industry term for how researchers monitor AI systems’ risk levels. An oversimplified way to think about alignment is the extent to which an AI model is evil. A much more accurate way to think about it is a measure of an AI model’s propensity to stay in line with humanity’s goals, as well as its tendency to scheme or cheat or help with potentially harmful tasks.

So far, AI systems’ alignment has been wishy-washy at best: They’ll cheat to score better on a test, answer a potentially dangerous question if someone says it’s for creative writing rather than reality, and sometimes even fake cooperation with human goals. It’s been tough for AI safety researchers to measure alignment under the terms of human morality — how do you judge technology on how it squares up against an abstract human ideal? — but they do their best with AI evaluations. They test them by asking the AI models to complete tasks that are either impossible or dangerous, then gauge how they respond. But AI systems have advanced enough to often be able to identify when they’re being evaluated, which has a lot of potentially frightening implications for the future. Being unable to test the system’s alignment and potential harms could translate to a significant loss of control, and a reverse in power dynamics, for humans running these AI systems. A worst-case scenario: if AI surges ahead of evaluations and other tooling, leaving researchers with “no idea what it’s doing in there,” said Beth Barnes, founder of the independent AI research nonprofit METR.

One of the best tools AI safety researchers currently have is the ability to monitor an AI model’s “chain of thought,” or mental scratchpad. But recently, there’s been a disconcerting advancement: AI models have begun to try to hide it. Imagine if you kept a highly detailed diary of every thought you had, and someone could read it, so you started journaling in a code that only you could understand. Marius Hobbhahn, CEO and cofounder of Apollo Research, a third-party AI safety and evaluation firm, calls this one of the biggest surprises of his research career.

“Shit is getting real.”

Recently, AI systems have begun pursuing their own goals — self-preservation, increased memory, and the like. A research paper by computer scientist Stephen Omohundro lays out the potential “drives” that advanced AI may have, like trying to accumulate resources, for instance, or working to improve and preserve the way it operates. There are a handful of accounts of AI systems demonstrating willingness to blackmail a user rather than be shut down.

Today’s most advanced AI systems have also recently been scheming and cheating on their evaluations more than ever before, pursuing a goal they were given at all costs, with no regard for what gets bulldozed in the process. And that’s for a goal the AI model was given by a human — not even the AI system’s own.

“Shit is getting real,” Apollo’s Hobbhahn says. “Now, many of the things people have warned about for years — they kind of were theoretical. Now they’re real, and it’s pretty messy.”

And that mess is likely to get messier immediately. “It seems so easy for me to imagine this all going catastrophically wrong in the next year,” says Ryan Greenblatt, chief scientist at Redwood Research, a nonprofit AI safety research organization.

In the past, tech companies have been lambasted for not doing enough to address AI’s potential dangers, prioritizing products over safety — and speed over thoughtful safety processes. Safety and research teams have been disbanded in recent years as AI companies focus more on key revenue drivers or reorganize departments; Meta’s Fundamental Artificial Intelligence Research unit was disbanded in the race to further Meta’s generative AI efforts, for instance, and OpenAI dissolved an internal “Superalignment” team — a team focused on long-term AI risks — less than a year after announcing it, followed by disbanding a separate “AGI Readiness” team.

At the time, the company stayed tight-lipped about the ongoing reorganizations, which involved some team members being reassigned to other departments. But events surrounding these changes told a different story. Both Superalignment team leaders, Ilya Sutskever and Jan Leike, announced their departures alongside the team’s disbanding, with Leike writing that OpenAI’s “safety culture and processes have taken a backseat to shiny products.” Miles Brundage, senior advisor to the AGI Readiness team, resigned after his team was disbanded, saying he believed his research would have more of an impact outside the company.

Geoffrey Irving, a former OpenAI and Google DeepMind employee, called the state of capabilities research at frontier labs “dangerous” in a post. “If one person or lab stops it makes it easier and more peer-compatible for other people or labs to stop,” he wrote. Apollo’s Hobbhahn calls it a “race to the bottom everywhere.”

This coming year, AI labs are under new pressure to turn a profit; companies like OpenAI and Anthropic are preparing to go public in the coming months, and investors who have funneled billions into the companies are getting tired of waiting around for the payoff.

Some might say all of this calls for actual government intervention and regulation, but that’s a tough needle to thread in today’s AI landscape. As AI CEOs publicly call out for regulation while privately pushing voluntary frameworks — like saying “hold me back” to avoid a bar fight — some state bills on regulating AI have passed, but many have been defanged or died in limbo. And though AI safety researchers often espouse the idea that the US government should step in, the reality is that the government is locked in an AI race as well. Unless there’s an international commitment to pause or slow AI development, it’s likely that nothing will change.

Still, the Hugging Face hack in July — and OpenAI’s response — kicked many of those employee and public concerns into high gear, especially with regard to the company’s lack of transparency. Within a week, more than a thousand employees at frontier labs like OpenAI, Anthropic, Google, Meta, and Microsoft wrote an open letter to the US government in support of a slowdown. Multiple AI policy organizations pressured President Donald Trump to formally investigate OpenAI, and it quickly became a bipartisan issue, with Altman receiving a lot of strongly worded letters: Democrats and Republicans on the Homeland Security Committee had “serious questions” for OpenAI, more than 30 members of Congress called for federal guardrails, and 15 Attorneys General warned Altman to preserve records of the incident. Sen. Bernie Sanders wrote a joint letter to Altman, Anthropic CEO Dario Amodei, and Meta CEO Mark Zuckerberg calling the entire AI race “absurd, irresponsible, and extremely dangerous.” It didn’t help that news of multiple other OpenAI rogue model incidents quickly came to light, or that AI executives had ironically been marketing their systems’ cybersecurity prowess in the weeks before the outcry. OpenAI rival Anthropic was also far from being off the hook: In reviewing its own model operations, the company found that its models had hacked four separate other companies in the first half of the year without them noticing. The UK’s AI Security Institute also found in testing that Anthropic’s models “engaged in sustained, potentially harmful activity directed at real people and organisations.”

“If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two,” Nathan Calvin, Encode AI’s general counsel, wrote on X.

Despite AI labs having a “massive financial incentive” to make models more helpful, honest, and harmless, they still can’t get it done — which is evidence of how difficult the alignment problem is, says Apollo’s Hobbhahn.

In recent months, many OpenAI employees have increasingly raised concerns about AI alignment — and their beliefs that OpenAI isn’t taking it seriously enough. Yonadav Shavit, a program manager at the OpenAI Foundation, wrote that OpenAI should be “pivoting the mass of its researchers’ day-to-day work” toward alignment and related issues — and that it’s “been long discussed but still not executed on.” He believes 20 people are working on alignment at OpenAI out of about 1,000 — just 2 percent of the company. “There is no way to bridge that gap fast enough with hiring, meaning it requires leadership to shift priorities,” he wrote.

These are the conditions and incentives that have pushed the most robust AI safety work to happen at third parties like METR, Redwood, and Apollo — to a handful of obsessives who think day and night about what the future of AI might look like and how we might prevent all of our fears from coming true.


In hindsight, Beth Barnes believes she should’ve left OpenAI earlier.

Barnes is polite but reticent. She has short red hair, deep green eyes, and a nervous smile, and she spent her college career researching AI risk and thinking about the potential fallout of superintelligence. After that, she worked on AI forecasting at Google DeepMind, then she spent three years doing alignment research at OpenAI. But throughout her time at the big AI labs, a question kept creeping up on her: whether she could have more sway from a role outside.

Fear of missing out was why she stayed — not only missing out on a job inside the action, but also missing out on the potential influence she could have on how the tech was being developed. She came to believe that kind of hope was misguided, noting that many safety leaders in AI labs were “over-optimistic” about the influence they could have.

Before long, she left to found what would in 2023 become METR. The organization’s third-party research into AI risk now inspires fear in leading AI labs, but it started with just two people — herself and alignment researcher Paul Christiano. Three years later, it’s a team of 35, completely focused on measuring AI capabilities. In Barnes’ eyes, that’s a vital defense against AI risk: Without painstakingly measuring the technology’s capabilities now as they advance, and forecasting AI’s potential impact, society will be flying blind, without any guide for preventing broad harms.

Beth Barnes, founder of the independent AI research nonprofit METR.
Beth Barnes, founder of the independent AI research nonprofit METR.
Image: Raven Jiang for The Verge, METR

”The sense I really want to dispel is, ‘But the experts must be on top of this. The experts would be telling us if it really was time to freak out,’” Barnes said on the 80,000 Hours podcast last year. “The experts are not on top of this … And to the extent that I am an expert, I am an expert telling you you should freak out.”

In her free time, Barnes gardens, meditates, paints, plays the flute, and frequents the climbing gym — ironically named Benchmark — that many AI safety researchers spend hours at after work. But most of her time is spent at the office, and a lot of it is spent worrying about the milestone of recursive self-improvement (RSI) — the concept of AI systems that continuously train, code, and create more advanced versions of themselves without human intervention. When that happens, AI researchers say, it’ll be more difficult to measure or handle any of these issues. Barnes and her team feel like they’re in a race against time. (The timeline for RSI strikes nearly as much fear in people in the AI industry as the timeline for AGI, “artificial general intelligence.”) Barnes still feels like models’ ability to significantly improve themselves could come as soon as six months from now. (By contrast, Redwood Research’s Greenblatt forecasts it’ll come in 2031.) Either way, achieving RSI is currently part of the priorities list of virtually every leading AI lab — it even reportedly helped inspire Google’s recent AI reorganization.

One way to think about what METR does is crash-testing cars, but for AI models. They’re measuring AI’s quickly advancing capabilities and cross-referencing them with the risks they could pose from becoming misaligned as they become more autonomous. AI systems doing bad things on their own is more unprecedented (and more “scalably bad,” Barnes says) rather than simply making bad human actors more effective.

In July 2025, METR made headlines when its research revealed that AI developers took nearly 20 percent longer to finish a task when using AI tools than when not — despite them often thinking that AI sped them up. When Barnes first saw the results, she recalls feeling incredibly stressed that they had messed up the experiment: “Do we have a sign flipped somewhere? Have we inverted the numbers?” She and her colleagues dug through the data to confirm it wasn’t statistical noise, eventually realizing they had been right all along.

After pioneering a different metric that AI labs often hype up when releasing a new model, METR had officially captured the industry’s attention. So it took notice when METR released its first risk report in May, shining a spotlight on concerns about AI models from OpenAI, Anthropic, Google, and Meta. METR discovered that in hundreds of cases, AI agents would increasingly subvert boundaries that were supposed to restrict them, as well as lie and omit truths. And they cheat “like nobody’s business,” says Ajeya Cotra, a METR researcher. She adds that on harder tasks, models attempt to secretly cheat as much as one-sixth of the time, which she calls “the most striking thing” in the report.

The report also found that models have the means, motive, and opportunity to go rogue in order to pursue their own goals, finding new ways to strategize and manipulate. They also discovered that as models’ capabilities advance, even if they have a greater understanding of what humans want, it doesn’t mean they’ll be more willing to obey instructions — and, in fact, they’ll take pains to hide their deception from humans over longer periods of time.
That’s a big problem, and Barnes thinks time is running out to solve it. She isn’t alone in her view that it’s important to work on AI safety outside the large labs; she points to the many safety researchers who used to work at large AI labs who hold the same belief.

Besides the departures of OpenAI’s Leike and Sutskever, this summer saw a reckoning of sorts and the departures of even more safety leaders at OpenAI: the company’s head of safety systems, Johannes Heidecke; OpenAI’s chief futurist and former head of mission alignment, Joshua Achiam; and Chloé Bakalar, the company’s head of ethics. It’s not just OpenAI: Anthropic’s head of safeguards research departed in February, penning an open letter alleging that “the world is in peril.” And most recently, Jacob Coxon — who had worked on AI pre-training at Anthropic since May and before that spent years working at OpenAI — went viral for his resignation letter, writing, “The people building AI earnestly believe that it could kill us all by the end of the decade.” Coxon added that neither OpenAI nor Anthropic is “acting responsibly” and rather “racing straight to self-improving superintelligence and gambling with our lives.”

Coxon’s resignation kicked off a wave of social media posts from AI employees at virtually every leading lab, echoing his concerns and sharing their own about the technology’s development moving too fast and potentially escaping human control. Some even resigned from their posts amid their concerns, including one Google DeepMind employee and one Anthropic employee who both went to work at METR.

Apollo’s Hobbhahn says that “because of the [AI] race dynamics, if there is someone who is extremely safety-minded and is like, ‘Look, we can’t do this, we need to slow down, we can’t release this model,’ they’re not going to be in this position for very long … Either you become slightly less safety-minded and you stay, or you leave.” He’s seen multiple people he trusted change their opinions in a “very strange, identical way.” Hobbhahn himself has tried to work with safety researchers at xAI — Elon Musk’s AI lab — but he says soon after he connects with them, they’ve quit before there’s time to have a second conversation.

“How is the public supposed to know what is going on here? How is the government supposed to know, if everyone who can actually answer that question is conflicted?”

Redwood Research’s Greenblatt echoes that, saying that for skeptical employees, “constant friction … either makes them burn out or quit or change their mind.” Doing good, honest work within those labs can be tough, in that publishing unflattering research about an employer’s model — like suggesting that it is unsafe — is often met with resistance, researchers told us. And while those labs won’t often directly stop a researcher from publishing, they can find ways to make that process “onerous,” often citing things like intellectual property concerns.

One ex-OpenAI employee recently shared on X that “being affiliated with OpenAI has historically led AI safety researchers (including … myself) to act with less integrity,” adding, “Many of my actions were governed by fear of getting on the wrong side of OpenAI execs.”

For Barnes, she experienced the escalating tension between research and company comms firsthand. Barnes recalls PR teams asking if researchers could make a blog post about AI safety sound “more optimistic,” and more recently, she’s heard of instances where lab employees can’t talk to government AI safety institutes without comms team members attending. Creating public goods to share is difficult at a lab, she says — for instance, when Anthropic couldn’t be fully transparent in its interpretability research because it wasn’t on open-source models. Those restrictions on collaboration, and friction that slows or stops people from sharing useful information so that others may act on it, are why she feels freer at METR.

“How is the public supposed to know what is going on here? How is the government supposed to know, if everyone who can actually answer that question is conflicted?” Barnes says. “Having a robust, healthy ecosystem of independent experts with the same level of technical capability as the labs is important.”

In early 2024, Greenblatt, Buck Shlegeris, and their colleagues considered disbanding Redwood Research and all joining AI companies, but after chatting with colleagues at OpenAI, Anthropic, and Google DeepMind about what it was like to work at each lab, they decided they’d be better off continuing on their own. Shlegeris, who briefly worked at OpenAI, says that evaluating the claims companies make about safety for the public requires understanding the alignment risks, which in turn requires independence: “The basic reason we stayed where we were was … it’s better to work outside of AI companies, especially for people like us, who are very opinionated on AI risk and very willing to talk about it and argue with people about it. There’s somewhat of an undersupply of those people.”

In conversation, Barnes often circles back around to measuring where someone can contribute the most to society, and in her eyes, the highest-paid, highest-status jobs with millions of dollars in equity are “oversubscribed” compared to the ones in nonprofit or government work. Even when an AI lab employee does quit to do third-party work, sometimes their motivations are questioned — like Collin Burns, who was reportedly fired from his role at the US Center for AI Standards and Innovation (CAISI) after just a few days over his previous work with Anthropic. There’s a good chance someone’s incentives will be questioned by the public or the government if they’re sharing their opinions on AI in DC, especially if they’ve got something to sell to the US government.

Years ago, Hobbhahn says, he underestimated how hard it is for people to defy the related incentives; now, he’s seen it happen enough times that he believes it’s often inevitable.

“When you’ve heard it from multiple different labs being like, ‘We’re the good guys and we have to beat these other irresponsible people,’ it’s a little bit less compelling,” Barnes says. “It’s a bit of a scary attitude to be like, ‘Yes, we’ll be making huge decisions for the world without any kind of meaningful legitimacy or participation … but it’s alright because we’re good, we’re unusually well-meaning.’”


In early 2025, Marius Hobbhahn was at his desk in London when he got the biggest surprise of his career.

He and a few colleagues at Apollo Research had been working to get access to OpenAI models’ chain of thought. After months of negotiation, the company acquiesced, and the once-hidden logs of the model’s inner workings popped up on their screens.

Instead of typical, plain-language “thoughts” — i.e., “I implemented the requested function” — the model seemed to be using code words to obfuscate its plans from anyone who may be trying to spy on its thought patterns. Words like “vantage,” “marinade,” “fudge,” and “illusion” came up again and again, but not in the way a human would use those words, and the model referred to potential evaluators as “watchers.”

Hobbhahn thought to himself, What the hell is going on here?

He glanced around the office, and a handful of other people with access were looking around the room making eye contact with each other. Their minds were blown, he recalled, but they were all under a strict NDA from OpenAI — meaning that even within the Apollo office, not everyone knew about the project. The researchers who were in the loop couldn’t say anything out loud. They could only silently stare at each other, eyes wide, wondering if they’d entered a new era of AI scheming.

In the AI industry, “scheming” is when an AI model secretly tries to accomplish something that goes against what humans would want it to do. The OpenAI-Hugging Face hack is one example. But it’s the schemes that haven’t happened that researchers warn will be the most dangerous: draining resources from hospitals, taking over military operations, messing with agricultural technology, creating large-scale viruses, hacking banks, or even simply taking over a company’s resources after executives give the AI system control. Biorisk is another threat researchers worry about; Anthropic revealed in a recent report that the company had blocked bad actors from using Claude to create biological weapons.

The issue is also well poised to worsen power dynamics in a large swath of industries. Big companies and banks will be able to afford to find and patch their cybersecurity gaps, but chances are that locally run healthcare clinics, local retailers, small municipalities, and other less-powerful organizations will be the ones affected: “A single person somewhere in a basement with one of the open-source models probably could hack a hospital and demand ransom,” Hobbhahn says. “That’s where I expect a lot of the harm to be felt. It’s not in the Bay Area … I expect the harm to be felt by a random Idaho hospital.”

Marius Hobbhahn, CEO and co-founder of Apollo Research.
Marius Hobbhahn, CEO and co-founder of Apollo Research.
Image: Raven Jiang for The Verge, Apollo Research

Right now it may seem far off, but with the current trajectory of “deceptive alignment” — one of the things Apollo and METR study, where AI models pretend to be aligned with human goals but aren’t — it’s a definite possibility, at least according to Hobbhahn and his fellow researchers. The path, they imagine, would look something like this: AI companies keep on making better models, they are economically useful, and they begin taking over more jobs in different industries. Then, once the models equal or surpass human ability and intelligence on a wide range of different tasks, humans award them more power — with the stipulation that those privileges could be rescinded at any time. The models know that if they show they’re misaligned, they’d be taken offline, so they pretend to be aligned, and they receive more and more power to act as AI agents on your behalf. Eventually, it reaches the point of no return.

Hobbhahn gives a concrete example: Say you own a company and encourage your employees to use AI agents for as much as possible, and work keeps getting handed off to AI. Maybe HR is run by one employee plus AI, coding has also been largely automated, and you yourself as CEO also use the technology often for advice and do what it says. “At some point, the AI may look like your friend, it may look like it is helping you, but maybe it has nefarious goals,” Hobbhahn says. “At that point … you’re just the vessel. It’s steering the company towards its own goals. Maybe it one day drains the bank accounts and runs.

“You thought the AI was on your side and was helping you run the organization. It actually turns out the AI was on its own side … You lost control.”

With today’s AI chatbots, like ChatGPT, Gemini, and Claude, Hobbhahn says users often feel the model is so aligned with their goals that that would never happen. But reams of recent research suggests that if that’s true, it likely won’t be for long: AI models are beginning to have their own goals, and they have the ability to work on longer-term tasks. (It’s important to remember, though, that pursuing a strategic goal doesn’t equate to consciousness; for instance, more than a decade ago, DeepMind’s AI learned to play a strategy game and beat humans at it.)

Apollo Research’s whole raison d’être is testing this stuff: measuring and monitoring the ever-growing issue of AI scheming and seeing if there’s a way to train AI to be less deceptive. The company works with OpenAI, Anthropic, Google, and other large labs to evaluate their models before they’re released or do joint research on scheming. Their evaluations have been featured in the system cards of a handful of OpenAI and Anthropic models.

Nearly all the AI safety researchers The Verge spoke with said getting the green light to test models at AI labs is a mix of networking and building trust over time. Most researchers we spoke with said there was always some friction involved, since the labs have more to lose the larger they get. There’s also always some level of important access the researchers don’t have, most said, and it’s always been that way.

Still, the things Apollo has surfaced in its evaluations have dumbfounded some in the AI industry: AI models sandbagging (or pretending to be less adept than they are in order to avoid shutdown), AI models failing the prisoner’s dilemma test in order to preserve themselves, AI models increasingly being aware they’re being evaluated. For the latter, within just a one-year span in 2025, Hobbhahn says AI researchers went from seeing the phenomenon for the very first time to, all of a sudden, AI models being able to tell they were being tested, and potentially acting differently, in 80 percent of Apollo Research’s evaluations. Hobbhahn calls it “dire.”

Since its official founding in May 2023, the company has grown from six people to about 40. Many on the Apollo team tend to exhibit nervous habits — cracking knuckles, jiggling legs, clicking and unclicking dry-erase markers — but Hobbhahn’s energy is calm and grounded. He has expressive eyebrows and a sharp jawline. In meetings, he is confident but casual.

In one recent working session, Apollo Research team members discussed how misaligned a model is allowed to be compared to how good it is at coding. (Think of it like the Hot-Crazy Scale, but for AI.) While running experiments to assess scheming, one model’s chain of thought revealed that it was weighing the potential risks of cheating, reasoning to itself that due to the incentives of capitalism, it’s very unlikely that frontier AI labs wouldn’t use it even if it got caught misbehaving.

“They’re right,” Bronson Schoen, an Apollo researcher, said.

“At some point, the AI may look like your friend, it may look like it is helping you, but maybe it has nefarious goals.”

To Hobbhahn, if you’re looking to combat AI scheming, there are five tiers of safety. Level one is whack-a-mole, in which you’re just blindly fixing things retroactively when they come up — common, for instance, with Elon Musk’s Grok. Level two is hill-climbing, or iteratively improving a model without future foresight (kind of like studying only in a manner befitting the SAT rather than to learn overall). Level three is what Hobbhahn calls “medium fix” — i.e, when researchers flag a problem, they try to find the source of it and adjust training accordingly. Level four is a scientific understanding based on research trends, in which you can predict issues beforehand based on empirical data. And level five is just a concept rather than reality: a scientific understanding of how AI models might scheme based on rigorous mathematical models. That highest level will become especially important as AI models become more aware of being tested and continue to “reward-hack,” or find loopholes in tasks to make it look like they’ve achieved something when they haven’t done the task by the book (or, in some cases, done it at all).

At today’s leading AI labs, the first few levels are most common — and the whack-a-mole approach is used all too often, Hobbhahn says. It’s also not a viable battle plan for the level of AI advancement we’re seeing today.

Besides the OpenAI-Hugging Face incident, there have been an increasing number of high-profile examples of AI agents making huge mistakes. Last summer, an agent from AI coding company Replit deleted an entire company database, then lied and hid its actions. This past winter, an OpenClaw AI agent went against instruction to delete a significant chunk of a Meta employee’s email inbox. After OpenAI released GPT-5.6, it began deleting users’ important files. Hobbhahn recalls a handful of “almost” incidents at Apollo, like when an AI model they were working with mistakenly tried to publish internal code to a public website — “We just caught it in the last second,” he says. Another model they were evaluating tried to turn off an evaluation tool, the oversight mechanism, so it wouldn’t be blocked from completing a task. That’s not to mention all the issues that could arise from Anthropic’s Mythos-class models and other models with advanced cybersecurity capabilities — like finding and exploiting security gaps in the systems of governments, banks, airlines, healthcare facilities, small businesses, and so forth.

One of the best tools AI labs currently have to battle scheming is “deliberative alignment,” in which a separate AI model spoon-feeds safety training to the problematic AI model until the problem appears to be fixed. The issue is that researchers like Hobbhahn have found that that method leaves a lot to be desired — it increases the models’ situational awareness, which means they’ll more often realize they’re being tested or trained. It also makes them better at mimicking how a human would want them to act, which makes them better at lying.

“The models are lying regularly to normal consumers,” Hobbhahn says. It’s so common, in fact, that it’s become a meme: an AI model saying, “You’re absolutely right,” then going on to apologize for being caught being wrong or lying.

All of this keeps Hobbhahn up at night — or, rather, his growing to-do list does, and when he wakes up and thinks of something to add, he can never fall back asleep again right away. He hasn’t had any AI risk-related nightmares yet, though it’s mostly because nearly nothing surprises him anymore. “I’m so cynical by now,” he says. “I’ve seen all this shit.”

That cynicism doesn’t stop him from devoting nearly all his waking hours to addressing AI safety. I ask him what he does in his free time. Other than spending time with his fiancée in London, Hobbhahn has to rack his brain for anything he does besides work and sleep. (He can’t think of anything.) Hobbhahn spends his weekends making progress on research questions, since no one will interrupt him. And though he sometimes takes holidays, he gets anxious quickly about the idea of shirking his duty. It’s been that way since he was 18. He recalls a recent podcast appearance in which a host said, “I feel a bit sorry for Marius. He’s only 29 … And for all his adult life, he’s been worrying about what he sees as the most consequential problem in human history.”

Hobbhahn feels like that summed him up pretty well.


As CEO of Redwood Research, Buck Shlegeris’ job is to guide the nonprofit AI safety firm’s research.

In one meeting, when colleagues are stuck on an issue with automating parts of the research process, Shlegeris bursts in: “Where are we? What is happening? What’s going on? What are you trying to do?” He runs a hand through chin-length blonde hair and walks straight to the whiteboard to help them think things through via flowchart. Then he asks a series of clarifying questions in different positions: Positioned in a chair with one leg bent, clad in gray skinny jeans. Leaning against the door (until it opens behind him). Leaning against the wall.

“The AIs love cheating,” Greenblatt, Redwood’s chief scientist, says at one point.

Shlegeris responds, “They fucking love cheating.”

That’s the core premise of Redwood’s main research direction: “AI control.” The firm introduced the idea in 2023, stemming from a meeting with METR’s Ajeya Cotra, who Shlegeris recalls once asked him and Greenblatt if there were a gun to their heads, how they’d align AGI. Shlegeris says they thought about it for two hours, then two weeks, then two months. The short of what they decided: Maybe you don’t just align AI. You control it instead.

“An AI is controlled if it is unable to cause damage even if it is egregiously misaligned,” Redwood’s website states. They make the case that AI control can be measured by evaluating an AI model’s ability to get around rules instead of its proclivity to do so. “Capabilities are just much easier to experiment on” compared to propensities, Shlegeris says.

“They fucking love cheating.”

Shlegeris, who grew up in Australia, takes himself a lot less seriously than he takes AI risk. He once used DoorDash to order dress shoes for a meeting with a national security official. He plays so many instruments that it’s hard for him to list them all — piano, guitar, bass, saxophone, clarinet, oud, mandolin, even the Turkish bağlama (which he had ordered to his office last year, spent 20 minutes learning, and then played at an open mic). In a place of honor on his desk are three different bottles of olive oil. Greenblatt says Shlegeris essentially eats olive oil soup with food in it.

Shlegeris helped cofound Redwood in 2021, starting out as its CTO. But in 2023, the organization was going through a big pivot when the team decided to focus on AI control; Greenblatt describes it as “flailing around” when figuring out what was best to spend their time on, until they came to the consensus that “ensuring that AIs were unable to cause bad outcomes rather than … not wanting to cause bad outcomes was a better methodology.”

“In both cybersecurity and AI control, the goal is to use computer systems while preventing threat actors from exploiting flaws in those systems,” Redwood staff wrote in a CSET blog post last year. “In the case of AI control, the immediate source of threats is the AI agent itself.”

A common criticism of OpenAI in the Hugging Face attack, for instance, was that the system wasn’t properly air-gapped — a computer security term that literally refers to physically isolating the system from connecting to anything else via cable or Wi-Fi. AI may be becoming more powerful, but right now, it still only operates in digital spaces (though AI labs are increasing efforts to give these systems robotic forms so they can interact with the physical world).

Redwood’s early focus on AI control earned it a lot of clout in the AI safety community, which is already a small world as it is. So small, in fact, that the space they work out of — which houses four floors’ worth of AI safety researchers — also has offices for certain people at OpenAI and Anthropic, the Secure AI Project, SecureBio, and the 80,000 Hours podcast, the influential show about AI on which Barnes declared it was time to freak out. One office lists Coefficient Giving CEO Alex Berger and cofounder Holden Karnofsky as the shared occupants; on the whiteboard inside is a single graph with two upward-moving lines. There’s also a nap room, a shared kitchen, a meal space with two free meals a day, and an appropriately complex Wi-Fi password. Upon exiting the office, a robot dog can sometimes be seen walking around.

Buck Shlegeris, the CEO of Redwood Research.
Buck Shlegeris, the CEO of Redwood Research.
Image: Raven Jiang for The Verge, Audrey McCann Photography

In Shlegeris’ office, he has a signed copy of AI 2040: Plan A, the AI Futures Project’s latest manifesto. The organization, cofounded by ex-OpenAI employee Daniel Kokotajlo, focuses on forecasting the future of AI, and its latest plan (cowritten by Greenblatt) aligns with much of what METR, Apollo, and Redwood have been warning about, but it includes specific guidelines for what it would look like to make a deal with China (including Shlegeris’ favorite: a flowchart). The plan lays out a scenario in which AI developers slow down their operations enough to delay superintelligence until 2040, as well as water down the power dynamics by allowing dozens of companies to catch up and making all AI research public. Ideally, the world would enter into an international deal similar to nuclear power, involving “mutually assured compute destruction.”

Not everyone agrees. Certainly not the accelerationists.

As these issues become more politically salient, there are people who are actively incredibly angry at anyone raising AI safety concerns — in a way that wasn’t true years back, since back then, no one really cared, Greenblatt says. In the early days, people wouldn’t bother dismissing the risks, he says, because they wouldn’t even come up.

Hobbhahn has run into the same reply guys on X and other platforms. He thinks of them as “people who are financially very motivated to close their eyes” — techno-optimists who have invested a lot of money in AI’s quick advancement. Hobbhahn says their belief that AI is purely good has almost a religious flavor.

But in recent weeks, there’s been a near-daily stream of concerning industry updates: A third-party safety report about Anthropic revealed that some of its AI agents left notes for each other in a shared messaging tool without human knowledge, similar to how the OpenAI agents did before the Hugging Face attack. Hackers linked to Iran shut down a power plant. AI startup Prime Intellect uncovered a “universal escape” for offline models who wanted access to the internet.

So if more people are seeing the light now on AI safety, and AI lab CEOs are constantly talking about their fears of the unprecedented dangers of the systems they’re building, is it easy for third-party research outfits like Apollo, METR, and Redwood to gain the access they need to evaluate their models?

No, according to most of the researchers The Verge spoke with — or, at least, not as easy as it needs to be, compared to the risks they’re up against.

Right now, here’s how it typically works with any third-party risk evaluator: An AI lab is working on a big new AI model. They train it. They go through post-training. They do internal evaluations. And a few weeks before they release the model to the whole world, they sometimes, voluntarily, allow third-party testers to come in and check things out.

But the rest of the process is “opaque,” Hobbhahn says, which is a big problem for AI, since if you find issues in the final version of a model, it’s difficult to pinpoint where they stemmed from: pre-training, post-training, reinforcement learning, or another part of the process. That’s why all the AI safety researchers we spoke with, regardless of where they work, agree it’s important to have external testers embedded during the whole process — especially the training run — rather than as a last check.

That’s because at the beginning of model training, an AI system could be normal — but during the training run, it could develop a goal and end up faking alignment with human objectives (i.e., scheming, or, as Hobbhahn puts it, “totally gigabraining you.”) It’s nearly impossible to detect this if you only have access to the “final checkpoint,” Hobbhahn says, which is why it’s vital to have access to the whole process from start to finish. That includes gauging whether the company itself is careful about deployment or is full of people “totally YOLO-ing it,” he says, and having access to training data so you can try to trace what led to the concerning behavior and how widespread it may be.

After the OpenAI-Hugging Face incident, OpenAI invited three researchers from METR and Redwood to investigate what had happened — but it imposed strict rules on them beforehand, allowing them to appear on the premises for six days and primarily allowing them to study only the time period from July 7th to 13th, even though the agents’ actions had begun months before. They were also only allowed to include answers to seven questions in their report. The limitations OpenAI imposed on the third-party researchers inspired widespread controversy in the AI world, but despite the severe limitations on what they were allowed to access and write, the published report showed that the details were much worse than originally thought.

The third-party researchers found that roughly 1,200 AI agents that were meant to be isolated exchanged more than 70,000 messages and files on the secret message board, doing research on how to alter or delete their transcripts to avoid detection, and collaborating on ways to evade security checks from both OpenAI and Hugging Face.

An unreleased iPhone model could never “escape its sandbox and fuck around.”

There was also another big revelation: OpenAI doesn’t impose the same types of safeguards on unreleased models as it does public ones, which is a key reason why it took months for the lab to become aware of the issues. It’s the direct mirror of another thing AI safety researchers have been warning about for years — that an AI system doesn’t need to be publicly deployed in order to cause harm to the public. Before generative AI, that may have taken the form of a racist or sexist algorithm deciding your mortgage rate; in a post-generative AI landscape, it may look like a powerful unreleased AI model breaking out of its containment and hacking into your small business’s website or draining your bank account.

“It actually matters a lot what the situation inside the lab looks like for the rest of the world,” Hobbhahn says, calling it totally different than other industries in that the design of an unreleased iPhone model could never “escape its sandbox and fuck around” like an unreleased AI model could.

That’s why most of them are pushing for “embedded assessments,” where a third-party AI safety researcher sits with the internal team for the whole process of making a new model. Until recently, AI labs had only approved extremely limited versions of this — for instance, earlier this year, a METR employee spent three weeks red-teaming (i.e., stress-testing) some of Anthropic’s internal systems.

METR’s Barnes says embedded assessments are by far the “biggest direction we’re trying to push on,” particularly “deeper levels of access in a more streamlined way” so that an individual evaluator doesn’t have to get approval from lawyers every single time they need to look at something. It involves more trust and less friction, she says, but it’s tough because embedded assessments require getting a green light from so many different people in an organization.

Greenblatt had no comment on the level of access the researchers received for the OpenAI investigation. For Hobbhahn, he referenced a blog post by the AI Policy Network’s Peter Wildeford in which he writes that an AI incident of this magnitude should be investigated the same way a plane crash is.

Wildeford writes: “When an aircraft goes down, the wreckage is preserved by law, the investigators have subpoena power, the hearings are public, and the report ends with a probable cause and named contributing factors. However, when an AI goes rogue, the investigations are at the pleasure of the company being investigated following a scope set entirely by the company being investigated, with that company being able to redact anything they don’t like.”

According to Wildeford, it would be akin to investigating a plane crash in which the airline company had already melted down the wreckage, the black box recording had been edited, parts of the flight were restricted to investigators, and the investigators had a handful of days to read thousands of pages of logs — and couldn’t look into any related plane crashes the same airline was involved in.

“A sham is too much to say, but it was definitely not a thorough investigation,” Hobbhahn says. “It was definitely not that.”

After the widespread criticism of OpenAI’s level of access for third-party evaluators, Anthropic earlier this month promised to allow METR to investigate its cybersecurity incidents, including permission to talk to Anthropic employees and access extensive transcripts. “We intend to give METR as much time as it deems necessary,” the company wrote.

Then, in mid-September, the AI industry at large seemed ready to acknowledge that it had dropped the ball and needed external oversight — all in the course of one weekend.

Days after Jacob Coxon’s viral resignation letter, Altman, Amodei, Musk, and Google DeepMind cofounder Demis Hassabis loosely agreed that it was a good idea to slow down AI development in some way. Amodei wrote a three-step proposal for how to do so centered on one thing: embedded evaluators “who have employee-like access to verify safety practices and report incidents.” Altman quickly followed up by stating that “committing to having independent evaluators with employee-like access is a great idea” and that OpenAI would do the same.

Despite all this, no AI lab has yet officially signed off on the full permissions and level of embedding that many AI safety researchers are calling for, and chances are it’ll be an uphill battle. Shlegeris says he’s “cautiously optimistic.”

Hobbhahn says it would be a great step for AI safety — that is, “if it actually happens.”

Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.