Skip to content
uk-ai.news

Red teaming

Deliberately attacking or probing an AI system to find its weaknesses before someone else does. Red teamers try to make a model misbehave — produce harmful content, leak data, or bypass its safeguards — so those flaws can be fixed.

The term comes from security, where a “red team” plays the role of the attacker to test a system’s defences. In AI it means much the same thing: rather than waiting to discover a model’s failings in the wild, a team sets out to provoke them on purpose. They craft tricky prompts, look for ways around the guardrails, and try to coax the system into doing things it is meant to refuse.

The point is to find problems while they can still be fixed cheaply and quietly. If a model can be talked into giving harmful instructions, revealing private training data, or producing biased output, it is far better that a friendly tester surfaces this before release than that a malicious user finds it afterwards. What red teaming uncovers feeds back into stronger safeguards and further training.

It has become a standard part of releasing powerful models, and a formal one: the UK’s AI Safety Institute and its counterparts red-team frontier models as part of assessing whether they are safe to deploy.

Articles using this term