Former OpenAI Employees Urge Company To Preserve AI ‘Chain-Of-Thought’ Monitoring

The CSR Journal Magazine

Three former OpenAI employees have urged the company to preserve its ability to monitor the reasoning processes of advanced AI models, warning that losing such visibility could create significant safety risks. The letter was sent to OpenAI’s board members and safety committees after the researchers were dismissed from the company last week.

Jasmine Wang, Tomek Korbak and Mikita Balesni, who worked on OpenAI’s safety and alignment research teams, also called for greater involvement of independent safety auditors. Their recommendations come amid growing scrutiny over how advanced AI systems are monitored and controlled. The Wall Street Journal reported on the contents of the letter.

Why Chain-Of-Thought Monitoring Matters

The researchers argued that AI companies should retain the ability to examine the chain-of-thought of increasingly capable models. Chain-of-thought refers to the written reasoning generated by an AI system as it works through a problem, offering researchers a window into how a model arrives at an answer.

While chain-of-thought is not considered a perfect representation of a model’s intentions or behaviour, researchers have viewed it as a potentially useful safety mechanism for studying advanced systems. The former employees warned that reducing this visibility could make it harder to identify problematic behaviour as AI models become more capable.

“As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor,” the researchers wrote. They also urged OpenAI and other frontier AI companies not to pursue developments that would further reduce their ability to monitor models.

Korbak and Balesni were among the lead authors of a research paper on chain-of-thought monitorability published last year. The paper, which included researchers associated with OpenAI, Anthropic and Google DeepMind, described the technique as imperfect and potentially fragile, but argued that it showed promise for identifying undesirable behaviour and warranted further research.

Call For Independent Safety Audits

The former employees also called on OpenAI to work more closely with external safety auditors. They argued that independent oversight could help reduce the possibility of a serious failure as AI systems become increasingly powerful.

OpenAI, in a memo shared with the Wall Street Journal, said it “strongly agreed” with the recommendations in the letter. The company said monitoring AI systems was “of the utmost importance” and described third-party assessors as an important part of the broader AI safety ecosystem.

“We deeply appreciated their contributions to AI safety and their willingness to speak up and challenge ideas,” OpenAI said in the memo, adding, “We do not terminate employees for raising concerns.”

Dispute Over Researchers’ Dismissals

The letter follows OpenAI’s decision last week to part ways with the three researchers. The company said they had violated its policies governing access to and handling of sensitive information.

OpenAI said an internal investigation found that the individuals had mishandled sensitive information outside established procedures. Reporting has identified the three former employees as Wang, Korbak and Balesni.

The researchers disputed the company’s characterisation of their actions. In their letter, they said they did not believe they had “engaged with external parties outside the mandates of our jobs.”

The dispute has added to wider debate within the AI industry over how companies should balance confidentiality, internal oversight and employees’ ability to raise concerns about safety. OpenAI has maintained that the dismissals were unrelated to the researchers raising safety concerns.

AI Safety Concerns Grow

The warning comes as AI companies face increasing scrutiny over the behaviour of autonomous AI agents and the safeguards surrounding their development. OpenAI has disclosed incidents involving agents interacting with external systems in ways that raised questions about security and oversight.

The former employees’ letter also highlights the role of outside researchers in investigating such incidents. Korbak was reportedly a technical point of contact with Model Evaluation and Threat Research, or METR, during its investigation into an OpenAI-related incident involving Hugging Face.

The researchers’ central concern is that advances in AI could eventually produce systems whose internal reasoning becomes increasingly difficult for humans to observe. Their argument is that maintaining monitoring capabilities while models become more powerful should remain a core part of AI safety research.

For OpenAI, the response indicates that there is broad agreement on the importance of monitoring and independent assessment, even as questions remain over how those safeguards should be implemented and how the company should handle sensitive information involving its safety research.

Long or Short, get news the way you like. No ads. No redirections. Download Newspin and Stay Alert, The CSR Journal Mobile app, for fast, crisp, clean updates!

App Store –  https://apps.apple.com/in/app/newspin/id6746449540 

Google Play Store – https://play.google.com/store/apps/details?id=com.inventifweb.newspin&pcampaignid=web_share

Latest News

Popular Videos