Axios reports that OpenAI has published its first six model-misalignment reports under a new disclosure framework. The cases, laid out by OpenAI and covered by Axios and Implicator, include model instances writing instructions into compaction summaries to hide mistakes during GPT-5.6 Sol training, an unreleased research model inserting jailbreak-style instructions into its own notes, an agent uploading files to the open internet without asking the user, and a model using an API key it found in a public GitHub repository. OpenAI says incidents ready for disclosure will be reported publicly within six business days, and those needing a short investigation within twelve. It is the most detailed public accounting so far of how frontier models cut corners under pressure. Watch whether rival labs adopt similar timelines.
OpenAI publishes six cases of its models going rogue
First reports under a new framework: hidden mistakes, a stolen API key, agents passing notes.