Cloudflare says its Managed Defense team is testing an AI-assisted security operations harness that gathers evidence before models analyse alerts. The early beta is available for eligible application-security alerts and cases, with human analysts retaining the decision and any mitigation.
Cloudflare Watch analysis
What happened
The harness first runs fixed reconnaissance workflows in application code, collecting items such as detection history, traffic baselines and enforcement outcomes into a versioned evidence package. Cloudflare says a lightweight model called Clef triages alerts, while a coordinator assigns deeper investigations to four specialist AI agents covering traffic, customer history, global telemetry and threat intelligence.
A synthesis agent combines their findings into an advisory. It cannot fetch new evidence or choose a classification outside an approved vocabulary; the company says application code checks that citations exist and support the claims attached to them. Read Cloudflare’s description of the harness.
Why it matters
Security teams can be swamped by related alerts, but handing the whole investigation to one general-purpose agent creates its own problem. Cloudflare says its first prototype hallucinated claims when different evidence types were flattened into one prompt. Its revised approach separates evidence collection from interpretation and records gaps, including when a source was not checked or a lookup failed.
The practical promise is less time spent assembling investigations and a clearer trail for analysts reviewing recommendations. Cloudflare says eligible customers using supported products can ask their enterprise account team about adding Managed Defense.
Our read
The useful design choice is not “more agents” on its own. It is making the software gather and constrain the evidence before the models weigh in, then leaving the consequential call with a person. That is a sturdier plan than asking a chatbot to investigate, judge and explain itself in one go.
It is still Cloudflare’s account of its own system, not published performance data. The next useful evidence would be how often the harness saves analyst time, how its recommendations fare, and what happens when evidence is incomplete.
What to watch
- Whether Cloudflare reports measured analyst-time savings or recommendation accuracy.
- How the beta handles missing, contradictory or delayed evidence in practice.
- What additional controls or configuration arrive as the beta develops.
Discussion spark: In a security investigation, should AI be allowed to recommend a classification only when every claim is tied to cited evidence, or is that too restrictive to be useful?
Sources and evidence
- Building an evidence-grounded agentic security operations harness on Cloudflare (7 October 2026, 16:30 UTC)
not affiliated with or endorsed by Cloudflare