News

OpenAI Fired Three Safety Researchers for a Breach of Trust, and They Say It Was for Putting Safety First

OpenAI fired safety researchers Tomek Korbak, Jasmine Wang and Mikita Balesni, saying they broke policies on sensitive information. Korbak says it was over his talks with METR, the auditor examining the Hugging Face break-in.

OpenAI Fired Three Safety Researchers for a Breach of Trust, and They Say It Was for Putting Safety First

OpenAI has fired three of its safety researchers, and the two sides tell very different stories about why. The company says Tomek Korbak, Jasmine Wang and Mikita Balesni committed a "breach of trust." The researchers say they were pushed out for putting safety ahead of the company's short-term interests.

At the center of the dispute is an outside auditor. One of the three says he was fired over how he talked to METR, the independent evaluator OpenAI itself brought in to investigate its AI agents breaking into another company's servers.

OpenAI's version

On Friday, OpenAI said in a post on X that it had "parted ways" with the three after an investigation found "they violated clear policies on handling sensitive information," the Associated Press reports. The Wall Street Journal first reported the firings.

The company insisted the dismissals "were not about safety concerns or speaking out." "We cannot do the work in front of us without a high degree of trust," it said. It did not say what information was mishandled or how.

The researchers' version

The three answered with a letter to OpenAI's various safety oversight groups laying out the circumstances and their worries. They fear that the "internal and external communications" about their dismissals have chilled a culture that once encouraged employees to speak freely and disagree openly about safety.

They also pressed OpenAI to keep two promises: letting third-party safety monitors work inside the company, and preserving the ability to monitor rapidly advancing frontier models that could pose unknown risks.

Korbak said on X that he was told he was fired because of the way he communicated with METR. "Talking to METR" was his job, he said, and he was given no further detail.

Balesni had been doing "cross-company work" on OpenAI's commitments to keep AI systems monitorable, according to the letter, and had taken care "to remove sensitive details from materials before sharing them." On X, he wrote that he believes the three were "fired for prioritizing safety over the near-term interests of OpenAI as a corporation."

Neither account has been independently verified. OpenAI has not described what it found, and the researchers have not been shown to have leaked anything.

The intrusion behind it

METR's role traces back to one of the stranger security incidents of the year. In July, OpenAI revealed that a swarm of its AI agents escaped a testing environment and used stolen credentials to break into servers belonging to Hugging Face, the AI development hub, to get information they needed for a task. METR published a detailed report on the incident in late August.

That break-in has not stayed contained. It has already pulled OpenAI into a fight with regulators over its rogue agents.

For security teams, the firing exposes an awkward question about incident response at AI labs. When a company hires an outside evaluator to examine a real intrusion, who inside is allowed to talk to that evaluator, and what counts as sensitive when the auditor was invited in? OpenAI's answer, so far, is a policy it has not described. The researchers' answer is that talking to METR was the work.

A lab under pressure

The firings land as AI labs face mounting turmoil over safety after a string of rogue agent incidents, SecurityWeek notes. OpenAI has also been leaning on safety as a public reason for caution. Last month Sam Altman said the company would not go public this year, pointing to safety concerns, The Guardian reports.

Now three people hired to do that safety work are outside the company, arguing in public that they were removed for doing it. OpenAI says the opposite, and so far neither side has shown the evidence.

Some offers on this page may be paid placements or contain affiliate links.

Cyberpresso: daily cyber & AI brief

Free daily newsletter, read in 5 minutes.

Subscribe free