OpenAI is facing criticism after dismissing three AI safety researchers who allege they were pushed out after raising concerns about the monitoring of AI systems and collaboration with external safety organisations, according to media reports.
Mikita Balesni, Tomak Korbak and Jasmine Wang, who worked on AI safety and alignment, have questioned the reasons given for their terminations.
OpenAI, however, said the dismissals followed violations of company policies on accessing and handling sensitive information, denying that the decisions were linked to raising safety concerns.
Researchers raise safety concerns
Korbak said on X that he was told his dismissal was related to his communications with AI safety organisation METR, which partnered with OpenAI to investigate an incident involving AI agents that autonomously hacked into AI company Hugging Face during testing.
Korbak, who described himself as OpenAI’s “main technical point of contact” with METR, said he believed his dismissal was linked to concerns he had raised about monitoring AI systems.
“For months, I’d been raising safety concerns that we’re losing the ability to monitor what AI agents think, one of our best tools for catching when they misbehave.”
Balesni said he was told he was “speaking too much to third party safety organizations”. He said he understood the company was implying he had leaked intellectual property, an allegation he denies.
Wang said she was dismissed for having “accessed an executive’s email,” which she said had been granted for recruiting purposes. She said she had repeatedly asked for her access to be revoked and notified the executive and OpenAI’s IT team after accidentally accessing a sensitive email.
“The reasons that we were provided for our terminations are simply not adding up,” Wang wrote on X.
OpenAI denies punishing safety concerns
OpenAI told CNN that the researchers were not dismissed for raising safety concerns but for conduct that breached its policies on handling sensitive information.
The company said its investigation found multiple violations, including issues beyond the handling of information shared with an external evaluation group.
In a statement, OpenAI said it parted ways with the three “for violating our policies on accessing and handling sensitive company information.”
“Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work,” the OpenAI spokesperson added.
An OpenAI research leader also told staff in a memo that the decisions were unrelated to employees speaking out about safety.
“I want to be very clear that these decisions were not about raising safety concerns or speaking out. We have always encouraged that and always will. We do not terminate employees for raising concerns,” the research leader wrote, according to the memo.
Fears over workplace safety culture
The three researchers wrote to OpenAI leadership outlining their concerns about the dismissals and the importance of maintaining close collaboration among safety researchers and external organisations. The Wall Street Journal first reported on the firings and the letter.
“We do not believe the path to superintelligence can be navigated safely if the people closest to the risks can no longer work in high-trust, high-bandwidth ways with each other and with third parties,” the researchers wrote. “It is that culture we’re trying to defend.”
Balesni said former colleagues had expressed confusion and fear about speaking openly or communicating with external parties.
“My former colleagues are telling me they are confused about what to believe. They also are afraid to speak, and worry their personal phones will be searched for messages to us and third parties,” Balesni said.
“I worry the pervading fear to speak up and engage with third parties will mean OpenAI will cut corners on safety behind closed doors.”
The dispute comes amid wider debate across the AI industry over safety oversight, transparency and the role employees play in raising concerns about the risks posed by increasingly capable AI systems.
