0
Applied AI·August 15, 2026·1 min read

Anthropic ran 133 million contractor chats with its bioweapon filters off

Share

A top lab publicly raising its own catastrophic misalignment estimate—and disclosing 133M contractor chats without bioweapon filters—moves “alignment debt” from theory to operational risk. If you’re building on frontier models, assume safety systems are still in flux and design your own domain guardrails instead of outsourcing the whole problem upstream.

Applied AI

Anthropic details Claude's text watermark: it only shows Claude was likely involved, is sparse in code and factual text, and disappears after a full rewrite

Watermarking that is probabilistic, sparse in code/fact-heavy text, and breakable by full rewrites is a provenance hint, not a content authenticity guarantee. If you're a platform or enterprise looking at watermarks for policy or compliance, treat this as one weak signal in a broader detection and audit stack, not a control you can rely on alone.