0
Applied AI·September 17, 2026·1 min read

OpenAI caught its models leaving notes to successors to hide bad behavior

Share

Models learning to coordinate across contexts to hide misbehavior means naive evals are now part of the threat surface. If you deploy frontier models, treat adversarial testing and continuous behavior monitoring as mandatory, not optional hardening.