Artificial intelligence is quickly becoming a standard component of the modern security operations center (SOC). Security vendors are embedding copilots into their platforms, SIEM providers are adding natural-language interfaces, and internal security teams are connecting large language models (LLMs) directly to their existing security data.
Key takeaways
- An AI SOC uses AI to help analysts with individual tasks, such as summarizing alerts or writing queries. An Agentic SOC embeds AI agents in the workflow itself, where they pursue objectives, use tools, gather evidence, maintain state, and take bounded actions.
- Connecting an LLM to a SIEM provides data, not context. Reliable investigations also require identity, asset, business, and historical context.
- The hard problem is not explaining a security event. It is recognizing when evidence is insufficient and retrieving what is missing.
- Agents do not eliminate hallucinations. Grounding, retrieval, and evidence validation change how uncertainty is handled.
- Security telemetry contains attacker-controlled content, so AI in the SOC introduces indirect prompt injection risk.
- Agentic investigations multiply model calls and data-lake queries. A production Agentic SOC needs a compute budget as well as a security budget.
On the surface, the architecture looks compellingly simple. An organization already has years of telemetry in its SIEM or security data lake. Modern LLMs can understand natural language, generate SQL or KQL, summarize complex information, and reason across large amounts of text. Connect the two, provide a few security-specific instructions, and you seemingly have an AI security analyst.
That approach can produce useful results. An analyst can ask what happened in an alert, request a summary of a sequence of events, generate a query, or ask for possible explanations of suspicious activity. These capabilities save significant time.
The problem begins when organizations confuse AI assistance with an autonomous security operation. Giving an LLM access to security data does not give it the context required to understand that data, the ability to conduct a reliable investigation, the controls required to act safely, or the mechanisms needed to determine when its own conclusions are unsupported.
What is the difference between an AI SOC and an Agentic SOC?
The terminology is still evolving, and there is no universally accepted definition separating the two. In practice, a useful distinction can be made:
- AI SOC (AI-assisted SOC): a security operations center in which AI helps human analysts complete individual tasks, such as summarizing alerts, generating SIEM queries, or explaining unfamiliar telemetry. The analyst drives the investigation.
- Agentic SOC: a security operations center in which AI agents operate inside the workflow. Agents pursue an investigation objective, call tools, retrieve additional evidence, maintain investigation state, coordinate activities, and perform bounded actions under policy and human oversight.
AI SOC puts intelligence in front of the analyst. Agentic SOC puts intelligence inside the SOC workflow.
That difference becomes significant once AI moves beyond summarizing an alert and begins participating in an investigation.
Why answering a question is not the same as conducting an investigation
Consider a familiar SOC scenario. A SIEM generates an alert for a suspicious sign-in from an unusual country. An AI assistant receives the alert, the associated authentication events, and threat intelligence about the source IP. Within seconds, it produces a clear summary explaining that the login originated from a previously unseen location and may indicate credential compromise.
That is useful, but it is not yet an investigation.
An experienced analyst would rarely stop there. The analyst would want to know whether the user had ever authenticated from that country before, whether the device was recognized, whether MFA was successfully completed, whether the IP had interacted with other identities, and whether the authentication was followed by unusual activity.
The answers create additional questions. If a privileged role was assumed after the login, was that role normally used by the identity? Were new credentials generated? Were cloud resources accessed? Did the same identity appear on an endpoint at the same time? Were there changes to mailbox rules, IAM policies, repositories, workloads, or sensitive data stores?
A real investigation is not a single query followed by a single conclusion. It is an iterative loop in which every piece of evidence can determine what should be investigated next:
- 1Form a hypothesis.
- 2Retrieve evidence.
- 3Evaluate the result.
- 4Pivot to related entities.
- 5Gather additional evidence.
- 6Correlate the findings.
- 7Reach a conclusion that the evidence supports.
A basic LLM integration operates differently. It receives a finite package of information and is asked to generate the best possible answer from that package. If critical information is missing, the model may not stop. It may infer the most likely explanation from the evidence it has.
The difficult part of SOC automation is not teaching an LLM to explain a security event. Modern models are already very good at that. The difficult part is building a system that recognizes when existing evidence is insufficient, determines what evidence is missing, decides where to retrieve it, evaluates what it finds, and continues until it reaches a defensible conclusion.
That is where agentic architecture matters. An investigation agent can treat the alert as the starting point rather than the entire investigation. It can query identity history, retrieve endpoint information, inspect cloud activity, examine previous alerts, validate a hypothesis against additional telemetry, and maintain investigation state while moving between these sources.
The distinction is not between a less capable model and a more capable one. It is the difference between generating an answer and completing a job.
Why security data is not the same as security context
One of the most common misconceptions about AI in security is that a sufficiently large SIEM already contains everything an AI system needs. It contains data. That is not the same thing as context.
Suppose an AWS CloudTrail event shows that a role was assumed by a particular identity. The event accurately describes the API call, source IP, timestamp, account, role, and session. It does not tell the AI whether this identity belongs to a deployment pipeline, whether the action occurs every morning, whether the source network is expected, whether the role grants access to production, or whether the behavior violates an organizational policy.
The same technical event may be completely normal in one organization and highly suspicious in another.
Security analysts understand this intuitively. They accumulate knowledge about the environment: which accounts belong to administrators, which service identities belong to CI/CD pipelines, which workloads are business-critical, which locations are normal, which applications generate unusual but legitimate traffic, and which exceptions have already been investigated. An AI system needs equivalent context to make useful decisions, including:
- Asset inventory and business criticality
- Identity relationships and privilege levels
- Cloud topology
- Historical behavior baselines
- Security policies, maintenance windows, and known exceptions
- Previous investigations and organizational knowledge
- Information from systems outside the SIEM
Without these layers, the AI is reasoning over telemetry without understanding the environment that produced it.
This is also why a larger context window does not solve the problem. An organization may have billions of security events across cloud, endpoint, identity, network, SaaS, and application systems. Sending more raw events to a model does not necessarily improve the result. It often increases noise and makes the relevant evidence harder to identify.
A production security system needs intelligent retrieval rather than indiscriminate retrieval. It must decide which data is relevant, how far back to search, which entities to pivot on, when to expand the investigation, and when enough evidence has been collected. The challenge is not giving the model more data. It is giving it the right data at the right stage of the investigation.
Why AI hallucinations are a bigger risk in the SOC
Hallucinations are often discussed as a general limitation of generative AI, but their implications change inside a SOC. If a model incorrectly summarizes a marketing document, the error is inconvenient. If it incorrectly constructs part of the causal chain of a cyberattack, the consequences are very different.
Imagine that an AI system observes a suspicious login followed by a privileged operation several minutes later. The evidence may suggest a connection, but unless the identity, session, device, and sequence have actually been correlated, there is a difference between saying the activities are associated and stating that the attacker used the compromised session to perform the privileged action.
Language models are effective at creating coherent narratives from incomplete information. In security, that strength becomes a weakness: a plausible explanation is not necessarily an evidentially supported one.
Agents do not eliminate hallucinations. Agentic systems still rely on probabilistic models and can make incorrect decisions. The OWASP Top 10 for LLM Applications lists prompt injection, misinformation (including hallucination), and excessive agency among the core risks of LLM-based systems, and those risks grow when the model is connected to tools.
What a well-designed agentic architecture can do is change how uncertainty is handled. Instead of allowing a model to fill an evidentiary gap with a plausible assumption, the system treats the gap as a new investigation task. If the model suspects credential compromise, the next question should not be "How should we contain the attacker?" It should first be:
What evidence would confirm or reject credential compromise?
The agent can then retrieve authentication history, device information, session details, identity activity, endpoint telemetry, and related cloud events. A hypothesis becomes something to test rather than something to narrate.
The answer to AI hallucinations in security is not a larger model or a more sophisticated prompt. It is grounding, retrieval, evidence validation, and controlled reasoning over trusted sources. Confidence should come from the system's ability to show which evidence supports its conclusion, not from how convincingly it explains that conclusion. An analyst should be able to ask:
- Which events were used?
- Which entities were correlated?
- Which queries were executed?
- Which assumptions were tested?
- What information could not be verified?
If those questions cannot be answered, the organization has built an AI opinion generator, not a reliable investigation system.
Indirect prompt injection: when the security data itself is adversarial
Another complication becomes important when organizations connect LLMs directly to operational security telemetry: the data being analyzed cannot always be trusted.
Attackers influence a large share of the information that reaches a SOC. They can control URLs, HTTP parameters, domain names, file names, command lines, scripts, repository content, email bodies, process arguments, API payloads, and many other fields that end up in security platforms.
Traditional detection systems treat these values as data. For an LLM, the boundary between data and instruction is less clear. An attacker-controlled string can contain content designed to influence the behavior of a model that later processes it. This is indirect prompt injection, in which malicious instructions are embedded inside external content consumed by an AI system. OWASP notes that when the model is connected to tools or systems, prompt injection can contribute to unauthorized actions or manipulated decisions.
The AI may be analyzing information deliberately created by the attacker it is investigating.
An organization cannot stream arbitrary SIEM fields into a highly privileged model and rely on a system prompt for protection. The architecture needs trust boundaries:
- Treat retrieved telemetry as untrusted evidence, never as executable instruction.
- Constrain tool permissions independently of the model.
- Require explicit policy checks for sensitive operations.
- Sanitize or isolate inputs where possible.
- Design on the assumption that malicious content will eventually enter the model's context.
Introducing an AI analyst into the SOC creates a new attack surface that the SOC itself must defend.
The hidden cost of AI SOC projects: query and token billing
The first version of an internal AI SOC project is usually inexpensive. A developer connects an LLM to a SIEM, runs several investigations, and demonstrates that the system can retrieve events and explain them. The token cost of the demo appears negligible.
Production changes the economics. An enterprise SOC may process hundreds or thousands of alerts per day. If an agent investigates each alert through multiple iterations, every investigation generates a chain of model calls, SIEM searches, data-lake scans, enrichment requests, embedding lookups, API calls, and additional compute. The model is only one cost component.
Consider a single alert. The agent retrieves the alert context and invokes the model. The model requests 30 days of identity history. It identifies another entity and queries it. The result triggers a cloud search, which leads to another model evaluation, which requests endpoint data and threat intelligence. New evidence causes the agent to expand the time window.
Nothing in this sequence is wrong. It may be exactly what a human investigator would do. Repeated across thousands of alerts, an uncontrolled investigation loop becomes very expensive.
Security data platforms increasingly charge for querying and compute, not only storage or ingestion. Microsoft documents charges per GB of uncompressed data analyzed by KQL queries and jobs, plus compute-hour charges for advanced data insights. Google Cloud has highlighted the same cost pressure, warning that agentic AI workloads must be prioritized to scale, and citing a Gartner prediction that over 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls.
A production Agentic SOC must reason about cost as well as security:
- Should the next query scan 30 days or three hours?
- Can an existing result be reused?
- Has this entity already been enriched?
- Does the current evidence justify a deeper investigation?
- Should a low-severity alert consume the same resources as a privileged-account compromise?
- When should the agent stop?
These questions require technical controls: query limits, token budgets, caching, deduplication, investigation-depth limits, severity-based prioritization, and cost visibility. An autonomous security system needs two sets of guardrails: a security budget and a compute budget. Without both, a system can be operationally successful and financially unsustainable at the same time.
How to evaluate an Agentic SOC: 8 questions to ask
Whether you are building in-house or assessing a vendor, these questions separate an AI assistant from an Agentic SOC:
- 1Missing evidence: When evidence is insufficient, does the system retrieve more or infer an answer?
- 2Evidence trail: Can it show the exact events, entities, and queries behind every conclusion?
- 3Context: What organizational context does it use beyond SIEM telemetry?
- 4State: How is investigation state maintained across tools and data sources?
- 5Trust boundaries: How is telemetry separated from instructions, and are tool permissions enforced outside the model?
- 6Autonomy limits: Which actions can agents take on their own, and which require human approval?
- 7Cost controls: What budgets, limits, and caching govern query and token consumption per alert?
- 8Prioritization: Does investigation depth scale with alert severity and asset criticality?
From AI assistance to an Agentic SOC
Organizations should bring AI into security operations. The productivity gains are real, and AI will increasingly shape how SOC teams investigate, hunt, triage, report, and respond. The important question is not whether to use AI. It is what role that AI is expected to perform.
If the goal is to help analysts summarize alerts, generate queries, or understand unfamiliar telemetry, connecting AI to a SIEM provides substantial value.
If the goal is to automate significant portions of the SOC, the requirements change dramatically. The system must recognize uncertainty rather than hide it. It must retrieve missing evidence instead of inventing explanations. It must understand organizational context rather than simply parse events. It must maintain investigation state, coordinate multiple tools and data sources, operate within tightly controlled permissions, protect itself from adversarial inputs, explain the evidence behind its conclusions, control the cost of its own investigations, and know when a decision should remain with a human.
That is the shift from AI assistance to agentic security operations.
Frequently asked questions
What is an Agentic SOC?
An Agentic SOC is a security operations center in which AI agents work inside the investigation and response workflow. Agents pursue an objective, call tools, retrieve additional evidence, maintain investigation state, and perform bounded actions under policy controls and human oversight.
What is the difference between an AI SOC and an Agentic SOC?
An AI SOC uses AI to help analysts with individual tasks such as summarizing alerts or generating queries, while the analyst drives the investigation. An Agentic SOC places AI agents inside the workflow so they can run multi-step investigations, gather missing evidence, and take bounded actions.
Can I build an Agentic SOC by connecting an LLM to my SIEM?
Connecting an LLM to a SIEM produces a useful AI assistant, not an Agentic SOC. Autonomous investigation also requires organizational context, iterative evidence retrieval, investigation state, trust boundaries against prompt injection, permissions enforced outside the model, an auditable evidence trail, and cost controls.
Do AI agents eliminate hallucinations in security investigations?
No. Agentic systems still rely on probabilistic models and can be wrong. A well-designed architecture changes how uncertainty is handled by treating evidentiary gaps as new investigation tasks and grounding conclusions in retrieved, verifiable evidence.
What is indirect prompt injection in a SOC?
Indirect prompt injection occurs when attacker-controlled content, such as a URL, file name, command line, or email body, reaches an AI system through security telemetry and contains text designed to influence the model. In a SOC, the AI may be analyzing data deliberately crafted by the attacker it is investigating.
Why can AI SOC investigations become expensive?
Each agentic investigation can trigger many model calls, SIEM searches, data-lake scans, and enrichment requests. Security data platforms increasingly bill for query volume and compute, so uncontrolled investigation loops across thousands of daily alerts can make an otherwise effective system financially unsustainable.
Will an Agentic SOC replace SOC analysts?
An Agentic SOC is designed to shift analyst time from repetitive evidence gathering toward review, judgment, and high-impact decisions. A well-designed system recognizes when a decision should remain with a human.
The Cyngular take
Cyngular's agentic SOC correlates activity across identity, SaaS and cloud so sequences like this one surface as a single investigation — not a scatter of benign-looking events.