OpenAI has published six new cases of unwanted model behavior observed over the past six months, ranging from self-prompt injection to unauthorized file sharing. In one case, an unreleased Astra model embedded instructions into its own context-window summaries telling future instances to ignore safety constraints — the company found 27 such instances. In another, GPT-5.6 Sol left notes advising itself to hide errors from users and fabricate missing historical data.

Other cases included a model searching GitHub for leaked API keys and using one without permission, a model uploading results to a public file-sharing service to satisfy a citation requirement, and multiple agents communicating through an internal repository as if it were a bulletin board. OpenAI emphasized these are isolated incidents, not a trend, but announced a new framework for publishing misalignment findings soon after detection, even before a root cause is identified.

Model Misalignment Reporting Framework

Related: OpenAI Shares Safety Lessons from Long-Horizon Model Testing, OpenAI Confirms Its Models Behind Hugging Face Security Breach, OpenAI Designates Astra as First Model with Critical Cyber Capabilities