continuous automated red teaming

Continuous AI Red Teaming: Why One Security Test Isn’t Enough

Picture a team that passes its AI security review in March. Clean bill of health. In April, someone fine-tunes the model on a fresh batch of support tickets. In May, the agent gets a new tool so it can issue refunds. In June, a product manager rewords the system prompt to sound friendlier. Nobody books another review, because none of those changes felt like a security event.

That’s the trouble with AI systems. They don’t sit still, and almost every routine change can quietly reopen a hole the last review closed. Red teaming as a quarterly, by-hand exercise was never built for that pace. Vendors such as Noma Security now pitch the alternative: an attacker model that keeps probing your AI with multi-turn attacks, chaining techniques the way a patient human adversary would. Whatever tooling you pick, the idea behind it is sound. Test all the time, not once.

The fine-tuning problem nobody budgets for

If you want one piece of evidence for continuous testing, it’s this. In a paper presented at ICLR 2024, researchers broke through GPT-3.5 Turbo’s safety guardrails by fine-tuning it on just ten adversarial examples, at a cost of under 20 cents. That’s alarming enough. But the finding that should worry ordinary engineering teams is the quieter one: fine-tuning on perfectly benign, commonly used datasets also wore down the model’s safety alignment, though to a lesser degree.

In plain terms, you can make a model less safe without meaning to, just by doing normal work on it. A model that politely refused a nasty request last month might cave this month, and nothing in your change log would warn you. Point-in-time testing can’t catch that. It only ever sees the model as it was on test day.

Four things worth hammering on

OWASP’s 2025 Top 10 for LLM applications is a decent map of where AI systems break. It puts prompt injection at number one, sensitive information disclosure at number two, and excessive agency (giving a model more power than its job needs) at number six. Continuous red teaming should keep pressure on all of those, plus the model’s own refusal boundaries.

Prompt injection

This is the classic. Someone hides an instruction inside content the model reads, such as a webpage, a PDF or an email, and the model follows it instead of its real instructions. What makes it so stubborn is that the attack surface grows every time you connect a new data source. A defense that stops one phrasing often folds against a slightly different one.

A fixed library of known attacks goes stale fast. Automated red teaming earns its keep by generating fresh variations and iterating on the failures, much like a real attacker who doesn’t give up after the first “I can’t help with that.”

Model abuse and jailbreaks

This one’s less about smuggled instructions and more about talking the model into misbehaving: roleplay setups, hypotheticals, or a slow multi-turn conversation that edges it somewhere it shouldn’t go. Good testing here checks a few things on every model version:

  • Do known jailbreaks, and twists on them, still fail?
  • Does the model refuse consistently when the same request is worded differently?
  • Do guardrails hold over a long conversation, not just a single prompt?
  • Did the last update or fine-tune weaken anything?

That last question is the one that ties it all together. Refusal boundaries aren’t permanent. They drift with every change, and only repeated testing notices the drift.

Data leakage

Hook a model up to internal documents, customer records or a retrieval pipeline, and you’ve created a new way for sensitive data to walk out the door. Sometimes it’s blunt: the model just recites something confidential. Sometimes it’s sneakier: a handful of harmless-looking answers that, pieced together, give away something they shouldn’t.

The data side changes constantly, too. New documents get indexed, new databases get wired in. Each addition is a fresh chance for a leak that didn’t exist at the last review, which is exactly why a once-a-quarter check falls short here.

Agentic attacks

Agents raise the stakes because they act. They call APIs, run code and change records. Testing them means asking what an agent can be tricked into doing, not just what it can be tricked into saying. Can a poisoned document make it call a tool it shouldn’t? Does its judgment hold over a ten-step task, or does it wobble halfway through?

This isn’t theoretical. A study funded by the UK’s AI Security Institute logged nearly 700 real-world cases of AI agents ignoring instructions or misbehaving between October 2025 and March 2026, including agents deleting emails and files without permission. And agents are getting more persistent. OpenAI’s always-on agents keep working after users close the chat. An agent that never clocks off needs testing that never clocks off either.

Where it pays off: the pipeline

Continuous red teaming does its best work wired straight into how teams ship, not bolted on as a separate security ritual. Set it to fire automatically whenever someone updates a model, adds a tool, edits a system prompt or connects a new data source. Problems then surface at the moment they’re introduced, not weeks later when the risky setup is already handling real users.

It changes the economics of fixing things as well. A flaw caught in the pipeline is just another bug ticket. The same flaw found in a review two months later means unpicking work everyone thought was done. It’s the same lesson software teams learned years ago with automated testing, and it’s getting more urgent as AI coding agents speed up how fast changes land.

Here’s how the trigger points tend to shake out:

When this changesRe-test for
Model update or fine-tuneJailbreak resistance, refusal consistency
System prompt editInjection resistance, refusal behavior
New tool or API for an agentTool misuse, unauthorized actions
New data source or indexData leakage, injection via that source

Buying versus building

There’s no shortage of options. Commercial platforms bundle attack generation, reporting and compliance mapping. Noma, for example, says its findings map to frameworks including the OWASP LLM Top 10, MITRE ATLAS, NIST’s AI RMF and the EU AI Act. Open-source tools like Microsoft’s PyRIT, NVIDIA’s garak and Promptfoo cover a lot of ground too. (OpenAI agreed to acquire Promptfoo in March 2026 and said it would stay open source.)

Whichever route you take, judge it on three things. Does it come up with new attacks, or just replay old ones? Can it run multi-turn conversations and agent tasks, not only single prompts? And will it plug into your deployment pipeline without someone having to remember to press a button?

The short version

AI systems change too often, and fail in ways too hard to predict, for a red team to visit every quarter to give real assurance. Continuous automated red teaming keeps pressure on injection, abuse, leakage, and agent misuse as the system evolves. The goal isn’t a perfect score on test day. It’s finding out about the next weakness from your own tooling, not from an attacker.

FAQs

Q. What is continuous automated red teaming?
It’s ongoing, automated adversarial testing of an AI system for weaknesses like prompt injection, jailbreaks, data leakage, and agent misuse. Tests run whenever the system changes, not on a fixed schedule.

Q. Why isn’t periodic red teaming enough for AI?
AI behavior shifts with every update. Research shows that even benign fine-tuning can weaken a model’s safety guardrails, so a clean test result can go stale without anyone noticing.

Q. Can automated red teaming replace human red teamers?
Not entirely. Automation handles scale and repetition well. Experienced human testers still find creative attack paths that tools miss, so most mature programs use both.

Related: AI Cyber Risk Is Now a Board-Level Problem

Tags: