
Cloudflare is rolling out an evidence-grounded, multi-agent security operations harness designed to help Managed Defense analysts keep pace with alert floods without losing sight of what evidence actually supports. The company says the system combines deterministic recon, specialist AI agents, and human review to turn noisy security alerts into consolidated cases with cited findings, visible gaps, and recommended next steps.
Why Cloudflare built an agentic security operations harness
Security alerts rarely arrive neatly one at a time. A single event can trigger a burst across an environment, forcing analysts to sort out which alerts are related, which are meaningful, and which are just noise. Cloudflare describes this as the “alert paradox”: the more signals arrive, the harder it becomes for even experienced defenders to keep up.
Cloudflare’s answer is a built-in security operations harness for Managed Defense that speeds up evidence gathering, aggregates related detections, and keeps track of missing sources while new alerts continue to arrive. The company says the goal is not to replace analysts, but to reduce the repetitive work that slows them down and make it easier to focus on decisions and mitigations.
Why Cloudflare moved away from a single AI agent
Cloudflare says its first prototype exposed the limits of using one general-purpose AI agent to run an entire investigation. The system produced useful analysis, but it also hallucinated claims that the evidence did not support. More importantly, the company found that flattening telemetry, detector descriptions, policies, and threat intelligence into one prompt caused their roles to blur together.
The company cites three recurring problems with the single-agent approach:
- Context became authority. A detection is a hypothesis, not proof that an exploit succeeded or an attack occurred.
- Scope drifted. An agent could query the wrong account, time range, or source.
- Failure disappeared. A timeout might be impossible to distinguish from “checked and not found.”
To fix that, Cloudflare shifted evidence collection and scope enforcement into application code before model analysis begins.
Recon first, inference second on Cloudflare
The front half of the harness is deliberately non-agentic. Before any model inference runs, deterministic code executes a fixed recon workflow using versioned API calls. That process collects the customer identity, detection history, traffic baseline, enforcement outcome, and network observations.
Each item is stored with its source, version, and timestamp. Cloudflare says that matters because it can see both the request that triggered an alert and the action applied to it. The result is a reproducible snapshot: if the same investigation is replayed later, any differences should come from interpretation, not from changing inputs.
Filtering obvious noise early
Not every alert deserves a deep investigation. Cloudflare says many alerts are repeat fires from known traffic patterns, and paging analysts every time would make it easier to miss a true incident.
For lightweight triage, the company uses Clef, its open-source decision model running on Workers AI. Clef compares each alert with its reconnaissance data, including whether the event has been seen for the customer before, what analysts decided previously, and whether the traffic resembles normal human behavior. High-confidence false positives are deterministically classified as passive and stay available as context without entering the active queue.
Specialist AI agents handle the deeper investigation
When an alert needs more review, a coordinator AI agent dispatches four specialist AI agents in parallel. Each one has a narrow task, which Cloudflare says makes unsupported claims easier to catch and recommendations easier to audit.
- Traffic analysis reviews request behavior, historical changes, and enforcement.
- Customer context reviews earlier alerts, dispositions, and analyst decisions.
- Global telemetry compares the activity with privacy-preserving Internet-wide signals.
- Threat intelligence checks indicators already admitted to the alert or case.
A synthesis AI agent then combines those typed findings into one advisory. Cloudflare says this final model cannot fetch new evidence or pick a classification outside the approved vocabulary. The system is designed so application code, not the model, enforces boundaries and tenant isolation.
Global context without exposing customer data
Cloudflare says one of the advantages of its platform is the ability to compare an alert against patterns seen across its global network. An IP address might be probing a single site, scanning thousands, or appearing for the first time, and those scenarios carry different risk signals.
To preserve privacy, the global telemetry specialist works only with aggregates and never receives another customer’s individual records or identity. Cloudflare says the view draws on data and features from its CDN, WAF, DDoS, Turnstile, Rate Limiting, and Cloudforce One threat intelligence.
That global perspective is then weighed against customer history. The company stresses that broad activity can inform an investigation without automatically implying a campaign against every customer.
History is evidence, not just background
Cloudflare’s system keeps each evaluation aware of what came before it: the alert, the pattern, and the customer. The recon dossier includes the alert’s history, such as how many times a service alert has fired, how often it was marked as false positive, and what a Managed Defense analyst concluded.
Approved context from previous alerts and cases is pulled forward so earlier findings do not have to be reconstructed from scratch. Cloudflare also says related alerts are aggregated into a consolidated case that stores evidence, findings, and recommendations, while a human analyst still confirms the actual scope.
From evidence to decision
Before analysis starts, the system assembles a versioned evidence package that includes the subject, scope, time anchor, admitted evidence, policy versions, sources, and any coverage gaps. Specialist agents must cite items from that package, and application code checks that every citation exists, belongs to the investigation, and supports the claim.
If a finding is invalid, Cloudflare says it is corrected or recorded as a limitation. Clef is then used again to judge whether the evidence is sufficient for a decision and whether anything contradicts the collected record. Based on that evidence, it chooses from a deterministically reduced set of attack classifications and dispositions.
The workflow runs on Cloudflare’s developer platform. Workers admits and validates evidence, Workflows coordinates each stage and saves work before the next begins, D1 stores investigation and advisory state, R2 holds bounded context and evidence artifacts, and case-chat state persists in Durable Objects and is enriched with AI Search.
What happens when evidence is incomplete
Cloudflare is careful to separate missing data from negative evidence. If a lookup times out, a metadata field is absent, or a threat intelligence search returns no match, the system keeps the evidence already collected and records the gap.
The advisory distinguishes between three states:
- Not checked
- Checked, with no matching result
- Checked, with evidence supporting absence
If global telemetry is unavailable, the system can describe what looks unusual for the customer, but it cannot claim the pattern is widespread. When evidence is insufficient, it makes no classification or disposition recommendation.
Human analysts still make the call
Cloudflare says the final recommendation is intended to lead to remediation, not just another ticket. Suggested actions may include a rate limiting rule for an abusive path, a WAF custom rule for a signature, or a DDoS protection change. For fully managed customers, analysts can apply the suggested rules; other customers receive recommendations in the dashboard and through their chosen alert path.
Even so, the company says Managed Defense analysts remain responsible for the decision and any mitigation. The system gives them the evidence behind the AI-generated recommendation so they can accept it, revise it, or group alerts into a case.
What’s next for the Managed Defense system
Cloudflare says the early beta is available in Managed Defense for eligible application-security alerts and cases. Over the next few quarters, the company plans to add a Custom Managed level with more flexibility for organizations, and it also wants to explore continuous AI agents that can watch traffic for patterns fixed rules and thresholds may miss.
For now, the pitch is straightforward: keep the model narrow, keep the evidence visible, and keep the human in charge. In a security operation overwhelmed by alert volume, Cloudflare is betting that structured AI assistance can help analysts move faster without giving up rigor.
Source: Original report
Was this helpful?
Explore more: Application Audit & Review More Cybersecurity Tech News
Last Modified: October 7, 2026 at 10:33 pm
3 views

