It’s Friday afternoon. A vendor advisory drops about a new loader making the rounds, and somebody on the team knocks out a rule before the weekend. It parses, it looks sensible, and it goes live. Nobody tests it against an actual sample, because there isn’t one handy and there’s a backlog of other work.
Multiply that by a few hundred rules and a few years, and you have most detection libraries. They’re full of rules written by different people, under different deadlines, tested to wildly different standards. Some catch things. Some just make noise. And some quietly catch nothing at all, which is the worst kind, because a rule that never fires looks exactly like a rule guarding against a threat that never showed up.
Now add AI. Language models can turn a threat report into a draft rule in seconds, and attackers have started using the same tech to reshape their malware on the fly. Malware-analysis vendor VMRay lays out detection engineering as a full lifecycle, running from threat research and data checks through design, rule writing, testing, and upkeep. That kind of structure used to be nice to have. It’s quickly turning into the only thing standing between a team and a library nobody trusts.
The trouble with writing rules on the fly
Talented people working without a shared process still end up with a mess. It isn’t a skills problem. It’s a consistency problem. One analyst runs every rule against live samples. Another ships off a quick skim of a blog post. Both are doing their best. A year on, though, nobody can vouch for which rules actually work.
Then there’s the sheer volume. Threat feeds, sharing groups and vendor reports pile up faster than any team can turn them into tested detections by hand. Something gives. Either the backlog swells, or people start cutting corners on testing just to keep their heads above water. Usually both.
AI makes writing rules easy. That’s the catch
Researchers have been putting LLMs through their paces on detection rules, and the headline numbers look great. In one study presented at the ACM Web Conference in 2025, a framework called LLMCloudHunter turned real cloud threat reports into Sigma rules, and 99.18% of them compiled cleanly into Splunk queries.
Look closer, though, and that figure answers a pretty narrow question: does the thing parse? Whether it catches the attack is a separate matter entirely. A 2026 study that built Snort rules from live exploit traces found 93.4% stayed quiet on a pile of harmless web traffic, which is encouraging. Even so, the authors were upfront that proving the rules catch real attacks needs its own replay-based testing.
So AI hasn’t gotten rid of the bottleneck. It’s just moved it. Writing rules used to be the slow bit. Now proving they work is. VMRay’s guide says much the same about LLM-assisted rule writing: someone still has to check the draft uses the right fields, matches the behavior it’s meant to catch, and actually works with the logs the organization collects. It’s the same story across engineering more broadly. As AI coding agents take over first drafts, reviewing their work becomes the real job.
Meanwhile, the other side has AI too
In late 2025, Google’s Threat Intelligence Group reported something new: malware that calls a language model while it’s running.
PROMPTFLUX, an experimental dropper, pings Google’s Gemini API to rewrite its own code and dodge antivirus tools. PROMPTSTEAL went further. Russia’s APT28 used it in Ukraine, and it queried a model hosted on Hugging Face to cook up commands as it went. Google called it the first malware it had seen querying an LLM during live operations.
Before anyone panics: PROMPTFLUX wasn’t capable of real damage yet, and Google shut off its Gemini access. But the writing’s on the wall. Malware that rewrites itself makes short work of any signature built from a single sample. Detections have to go after behavior, and they have to be tested against variations rather than whichever file happened to land in a report.
What a repeatable pipeline actually looks like
Think of it as five stages, each with a job the next one depends on.
It starts with intake. Intelligence turns up from vendor feeds, ISACs, your own incidents and open research, and the quality is all over the place. Before anyone touches a query, check each item against your actual exposure. Does it target tech you really run? Would your current defenses miss it? Does the source have a decent track record? Skip this and whatever shouts loudest gets worked on first, which is rarely what matters most.
Next comes the hypothesis, the step everyone wants to skip (and AI makes skipping it tempting). A report describes attacker behavior in broad strokes. A rule needs specifics: which data source would show the activity, what exactly to look for, and what harmless activity looks similar enough to cause false alarms. Without that, you’ve got a guess in valid syntax, no matter how polished the draft looks.
Then validation, which is where ad hoc programs usually fall apart. This means running real or carefully simulated malicious activity in a safe environment and watching what telemetry it throws off. Run the rule against the threat itself, against variations of the technique (non-negotiable now that malware can rewrite itself), and against benign activity that looks like it. Then write down exactly what you tested, so whoever inherits the rule in two years knows where its edges are.
Deployment comes after that, and it shouldn’t mean flipping straight to full alerting. Plenty of teams run new rules in monitor-only mode for a few weeks first. The rule logs what it would have flagged without waking anyone up, and it only graduates to real alerts once its live false-positive rate looks sane.
Last is maintenance, the unglamorous bit. Rules rot. Attackers change their playbooks, and your own environment shifts under your feet: new software, new logging, new ways the business works. Put reviews on the calendar. Retire rules that don’t earn their keep. Keep everything in version control so you can see what changed and roll it back when an update goes sideways.
Stop counting rules
Rule count was always a vanity metric. With AI it’s close to meaningless, because generating a thousand rules now costs next to nothing. Better measures look like this:
Metric | What it really tells you |
| Coverage against MITRE ATT&CK | Which attacker techniques you can genuinely spot |
| Time from intake to validated rule | How quickly intel turns into protection |
| Production false-positive rate | Whether analysts trust the alerts |
| Share of rules with documented testing | How much of the library is proven versus assumed |
Keep a close eye on that last one as AI-drafted rules start flowing in. If it drops while your rule count shoots up, you’re going backwards, however busy everyone looks. And when budget season rolls around, numbers like these make a far stronger case than “things feel more organized these days.”
The upshot
None of this is glamorous. Pick the threats that matter, spell out what you’re hunting for, prove the rule works against the real thing, and keep it fresh. AI speeds up the drafting and hands attackers some new tricks, but the core discipline doesn’t change. Teams that put their effort into validation will get real mileage out of AI-drafted rules. Teams that don’t will end up with a big, impressive-looking library they can’t actually rely on.
FAQs
Q. Can AI write detection rules?
It can draft them fast, and most drafts compile without trouble. Compiling isn’t the same as catching anything, though, so every AI-written rule still needs testing against real malicious and harmless activity.
Q. What is AI-enabled malware?
It’s malware that talks to a language model while it runs, for instance, to rewrite its own code or generate commands. Google first documented families like PROMPTFLUX and PROMPTSTEAL in 2025.
Q. How often should detections be reviewed?
On a regular schedule, plus whenever false positives spike, attacker techniques shift, new logging comes online, or an incident exposes a gap.
Related: Where Should Manufacturers Start With AI Automation?
