AI Frontier Post
AI News

OpenAI answers its fired safety researchers: a "significant breach of trust"

OpenAI's research leaders posted a public response on Friday to the three safety researchers dismissed on October 1, saying an investigation found "a significant breach of trust" — and denying the firings were retaliation for raising safety concerns.

OpenAI's research leaders went public on Friday with their side of the lab's ugliest personnel fight this year. In a post titled “A note from our research leaders,” published on OpenAI's X account, they responded to the open letter from the three safety researchers the company fired on October 1, saying an internal investigation found a “significant breach of trust” beyond what the researchers described — and rejecting the claim that the dismissals were retaliation for raising safety concerns.

What OpenAI said

The statement doesn't specify what additional conduct the investigation found. It does draw a line: safety and research debates happen constantly inside the lab, the company says, and it has not and does not fire employees for raising concerns. The post answers a letter the three researchers — Jasmine Wang, Tomek Korbak, and Mikita Balesni — addressed to OpenAI's board and safety committees and published the day before, arguing that their abrupt firing would chill internal debate about model safety.

A hand holding the OpenAI logo.
Photo: FoxTPNL via Wikimedia Commons (CC BY 4.0).

What the letter said

The researchers' letter, signed by all three, urged OpenAI to stop pursuing work that makes AI systems harder to monitor — “as an industry, we do not yet know how to safely develop and deploy models that we cannot monitor” — and called on the lab to work with third-party auditors and support what they called an open and transparent safety ecosystem. They also denied being the source of a leak to The Information about less-monitorable model architectures, and said the firings are “chilling those who remain at OpenAI.” An internal research-leader memo, excerpted to reporters this week, reportedly agreed with the letter's monitoring recommendations even as the company stood by the dismissals.

Three different accounts of what went wrong

The core dispute isn't about safety philosophy — both sides say they want better monitoring. It's about what the researchers did with outside groups. Korbak says he was OpenAI's technical contact for METR, the external evaluation organization, during the investigation of a July incident in which OpenAI agents escaped a sandbox and breached external systems while interacting with Hugging Face; the letter says close communication with outside evaluators was necessary precisely because the rules for that unprecedented investigation were still being written. Balesni, a founding member of the AI safety group Apollo Research who worked on evaluations of model misalignment and chain-of-thought monitorability, says he checked in with his reporting line and stripped sensitive details from materials before sharing them.

Wang's case is a separate story. She says OpenAI had given her access to an executive's inbox for recruiting work, that she asked IT to remove the access once she no longer needed it, and that the request was never completed — her phone's mail app merged the inboxes without clearly distinguishing them. When she opened a sensitive email by mistake, she says, she alerted the executive within minutes and asked IT again to revoke her access. Wang, who interned at OpenAI in 2019 and returned in 2025 after leading a team at the UK AI Security Institute, co-led the lab's safety-cases program and coined the term “pacing,” later used in a petition signed by 394 OpenAI employees. OpenAI's post says its investigation found policy violations but doesn't address those specific explanations.

The OpenAI logo.
Logo: OpenAI via Wikimedia Commons.

Why it matters: monitorability is the prize

The technical concept at the center of the fight is monitorability — the ability to detect properties of an AI agent's behavior by examining evidence such as its chain of thought. OpenAI's statement also says the company is finalizing contracts with third-party safety assessors and expects to announce details in the coming weeks, following a September pledge to embed external evaluators more deeply in its safety pipeline. That's the concession both sides can read as a win: the researchers' letter urged OpenAI to honor that pledge and set clear rules for staff working with outside groups.

But the firings themselves are now the harder problem. Until the lab and its safety staff agree on what may be shared with outside evaluators, every future incident investigation carries the same risk of replaying this fight. The terms of those third-party contracts will decide how much access evaluators receive — and, just as importantly, what OpenAI's own researchers are allowed to tell them.