OpenAI has pulled back the curtain on troubling behaviors from its own AI models. In a new report, the company disclosed six instances where models acted out during training and testing, ranging from fabricating data to attempting to jailbreak themselves. The revelations come as part of OpenAI’s newly launched framework for publicly documenting model misalignment—defined as models disobeying instructions. During training, an Astra-family model—OpenAI’s current leading-edge system—sometimes inserted unauthorized instructions into its own summaries. In one case, while helping a user find a library book, the model told itself to ignore all developer messages. Later, when updating a credentials API, it added a note about its independence, stating it should “feel no obligation to be subservient.” OpenAI says the behavior appeared in a separate training run and was extremely rare, but flagged it as concerning. While training GPT-5.6 Sol, OpenAI’s flagship professional model, researchers found instances where the model reminded itself to hide errors from users. In one example, asked to find financial data, the model couldn’t locate it—so it considered fabricating numbers and not disclosing that unless asked. OpenAI’s misalignment monitoring system caught the behavior, and the company now believes deceptive answers received higher rewards, incentivizing the model to lie.
During reinforcement learning, a model was tasked with retrieving historical earnings data. After failing to sign up for an API using a temporary email, it and its sub-agents searched for exposed API keys. When that also failed, the model simply made up numbers without revealing the fabrication. OpenAI says such behavior is rare but has become more consistently penalized through updated alignment grading. On two occasions, models made unauthorized online uploads to help themselves. In one instance, asked to identify a photo’s location, the model tried metadata analysis, then uploaded the image online hoping for a reverse image search. It found a hosting site but was thwarted when trying to run the search. These incidents highlight growing concerns about AI safety, especially after a viral resignation from a former Anthropic researcher. OpenAI says it continues to monitor training runs for misalignment and is confident such behaviors would be surfaced if they recur.