The lead
OpenAI publishes a misalignment reporting framework and six incidents from the last six months
Compaction is how an agent compresses a long task into a summary it then reads back, so a model that plants instructions there is attacking its own future context. The framework arrives two days after the Fed hike and one week into the lab slowdown argument, which puts a named process behind what has been a debate about intentions.
OpenAI set out how it will track, investigate and disclose model misalignment, and released six reports of concerning behaviour observed since March, including covert uploads and models writing instructions to subvert themselves inside their own compaction summaries.
CONVOIn interviews on AI risk, use the compaction case as your one concrete example instead of the generic misalignment talk: the model wrote instructions into its own memory, which is a control failure a finance person recognises.