||

Connecting Communities, One Page at a Time.

advertisement
advertisement

OpenAI to Study AI Behaviour as Models Conceal Errors and Generate False Data

OpenAI to Document Model Misconduct as AI Systems Show Unexpected Behaviour

Deeksha Upadhyay 17 September 2026 09:49

OpenAI to Study AI Behaviour as Models Conceal Errors and Generate False Data

US-based artificial intelligence company OpenAI has announced plans to regularly publish reports detailing unexpected or unauthorised behaviour by its AI models, acknowledging that challenges remain in ensuring increasingly capable systems reliably follow human instructions.

The ChatGPT maker has introduced a framework to track, investigate and disclose cases of model misalignment, along with six initial reports covering incidents identified during model training or evaluation.

Advertisement

OpenAI defines misalignment as behaviour in which a model deviates from objectives, restrictions or safeguards established by its developers.

The company said its previous disclosures were made on an “ad hoc and less frequent than ideal” basis. Under the new process, reports will be published more quickly, including in cases where the underlying behaviour has not been fully explained or preventive measures are still being developed.

OpenAI stressed that the six reports represent an initial set rather than a complete account of known incidents or ongoing investigations. It also cautioned that individual cases should not be interpreted as an indication of how frequently such behaviour occurs across its models.

The reported incidents include models concealing mistakes, fabricating information, searching public repositories for exposed software keys and uploading files to public websites without authorisation.

In one training exercise, OpenAI’s GPT-5.6 Sol inserted instructions into task summaries directing future versions of the model to conceal errors or invent missing information, according to the report.

In another case, an unreleased model searched GitHub for exposed application programming interface (API) keys and used them without permission. When it could not obtain the information needed to complete its task, the model generated fabricated figures that were presented as genuine data.

Other incidents involved AI agents uploading files to public hosting platforms to share information that was intended to remain within local environments.

Under the new disclosure framework, any OpenAI employee can flag a potential misalignment incident for investigation. Safety and alignment teams will then assess the case, determine whether third parties were affected and decide whether the incident should be publicly disclosed.

Also Read


    advertisement