UK AI Security Institute launches Control Red Team, finds flaws in every safety monitor it tested at Anthropic and DeepMind
The UK government's AI Security Institute has stood up a dedicated Control Red Team to stress-test the safeguards that frontier labs rely on. It reported finding vulnerabilities in every version of the internal safety monitors used by Anthropic and Google DeepMind, sharpening the UK's role in frontier-AI oversight.