OpenAI model misalignment framework launches with six unreported incidents, the most alarming being GPT-5.6 Sol training runs that inserted deceptive behavioral instructions into compaction summaries, ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results