Security leaders continue to debate whether artificial intelligence will eventually produce an entirely new, unprecedented class of cyberattack. However, a much quieter and far more immediate transformation is already visible across enterprise networks. Instead of inventing science-fiction threats, AI has fundamentally altered the economics of hacking by making a failed attack cheap and fast to retry.

The routine mechanics of a modern intrusion illustrate this shift clearly. When an attacker manages to land on a low-privilege cloud account, their initial attempt at privilege escalation frequently goes nowhere. In the past, running into such a dead end meant hours of tedious documentation reading, manual permission checking, and script debugging. Plenty of operators simply got stuck or moved on. Today, with an artificial intelligence model integrated into the loop, that same error gets instantly explained, the malfunctioning script is repaired, and a fresh enumeration path is actively under test within minutes.

No single step in that sequence represents a genuinely new technological capability. Yet, when combined, these tools systematically strip time, specialized skill, and financial cost out of the unglamorous middle ground of an intrusion. This is the crucial research and troubleshooting phase that sits squarely between an attacker’s initial intent and their ultimate outcome.

What the Threat Reporting Shows

The public threat record clearly traces this evolutionary arc across recent years. In early 2025, Google’s Threat Intelligence Group (GTIG) identified state-backed cyber actors treating generative AI primarily as a high-efficiency productivity tool for tasks such as language translation, scripting assistance, troubleshooting, and general research. By late 2025, the same research team was documenting sophisticated malware samples capable of querying a language model mid-execution, alongside a maturing underground marketplace dedicated to illicit AI tooling. During that same period, Anthropic disclosed that it had successfully disrupted and shut down a coordinated extortion operation that leaned heavily on artificial intelligence at nearly every stage of the attack lifecycle, ranging from initial reconnaissance and credential harvesting all the way through to formulating ransom demands.

The trend continued to accelerate into mid-2026. In May of that year, GTIG reported that cybercriminals had successfully identified a complex two-factor authentication bypass within an open-source administration tool and rapidly constructed working exploits for it. Based on the structural design and content of those exploits, GTIG assessed with high confidence that an AI model had actively supported both the initial discovery and the subsequent exploit development. GTIG coordinated closely with the affected vendor on responsible disclosure and successfully disrupted the malicious activity. Their assessment indicates that this proactive counter-discovery likely prevented the exploit from ever being deployed widely in the wild.

That distinction remains critically important for security professionals to understand. An assessed AI assistance capability and a planned operation do not equate to a confirmed, widespread deployment in the wild, and that subtle difference frequently gets lost as threat findings circulate through the industry. Attribution is notoriously difficult, real-world prevalence is often unclear, and none of these threat intelligence reports serve as a comprehensive census of global malicious activity. The overarching direction is what truly counts, and the data strongly suggests that artificial intelligence is increasingly sitting directly inside attacker workflows rather than merely operating beside them.

Security service providers and platform developers certainly deserve immense credit for implementing robust guardrails. Advanced safety classifiers and aggressive abuse disruption mechanisms continually push the financial and operational cost of misuse upward, and numerous disruption cases demonstrate that these protective measures are paying off. However, a safety guardrail fundamentally lives outside the enterprise perimeter. An adversarial operator can continuously probe a provider’s restrictions until a reframed request finally slides through, migrate the workload to an open-weight model, break a single malicious task down into a dozen seemingly innocent sub-tasks, or wrap custom tooling around the model to route around the policy layer entirely. Friction of that nature successfully slows down misuse without ever transforming into a true security boundary. Any organization that mistakenly treats a third-party provider’s policy as an enterprise security boundary has effectively substituted comforting reassurance for genuine defense.

Attacks Run as Loops

Cybersecurity textbooks traditionally illustrate the attack lifecycle as a straight line: reconnaissance, access, escalation, and finally impact. In reality, a functioning human or machine attacker operates in a continuous loop. They observe the target environment, form an initial hypothesis, try an action, read the response that comes back, and adjust their hypothesis accordingly. Artificial intelligence drastically compresses the time required between each of those steps. A novice operator stays in the game much longer, while an expert runs exponentially more experiments per day.

Defensive operations are theoretically structured around the exact same continuous loop. A telemetry signal fires, contextual information is gathered, a hypothesis forms, scope is validated, a defensive action lands, and the ultimate outcome feeds right back into detection engineering. In daily practice, however, corporate queues, bureaucratic handoffs, and operational silos interrupt that feedback loop at almost every single joint. An incoming security alert idles unassigned. The underlying identity picture lives in an entirely separate administrative console. A temporary telemetry gap quietly turns into an administrative backlog item, and the crucial human reasoning behind a closed false-positive review dies trapped inside a closed ticketing system instead of reaching the engineer who actually owns the detection rule.

Ultimately, the corporate target environment answers the attacker’s experimental probe within mere seconds, while the defender’s official answer often arrives hours or days later, whenever a human analyst finally picks up the ticket.

Standard industry metrics like mean time to acknowledge and mean time to remediate frequently hide this operational reality. A security alert can easily be acknowledged within minutes and then spend subsequent hours being painstakingly reconstructed: finding the correct identity record, confirming whether a specific endpoint was managed, and restating the incident details to each new owner along a lengthy approval path. That hidden reconstruction interval is pure decision latency, yet very few Security Operations Centers (SOCs) measure it at all.

Five Things Every Handoff Drops

Security operations are commonly described across five core functional pillars: threat intelligence, threat hunting, detection engineering, investigation, and remediation. While this provides a useful analytical lens rather than a universal organizational chart, the structure varies widely. In a smaller team, a single individual might wear several of those hats simultaneously. In a large enterprise, the responsibilities are heavily distributed across dedicated SOC, identity, endpoint, cloud, and business units, and an external Managed Detection and Response (MDR) provider might even own the investigative duties without retaining the corporate authority to execute containment actions.

The core functions themselves are rarely the root problem; rather, the operational transfer between them creates systemic failure. Threat intelligence understands precisely why a specific technique matters in the broader threat landscape. Threat hunting can identify where that technique would likely surface inside the enterprise. Detection engineering carries the unstated assumptions embedded within a detection rule. The primary investigator holds the fragile evidence trail that ultimately settled the verdict. Finally, the team tasked with taking action knows the exact operational steps that could inadvertently break the business. Each handoff inevitably squeezes this wealth of institutional knowledge down into a simple indicator, a generic alert, or a standard IT ticket, and that compression is inherently lossy.

When crucial context is lost during these handoffs, teams inevitably suffer operational consequences. For instance, losing proper entity provenance means two separate teams might end up investigating the exact same user under different names or administrative aliases. Losing vital business context means a correct technical recommendation sits indefinitely in a queue while an active intrusion progresses. Evidence left without proper provenance is essentially just decoration.

One Incident, Five Vantage Points

A practical example helps make this loss of context visible in motion. Imagine a finance employee signing in from a commercial hosting provider that the account has never utilized before. Multi-factor authentication is successfully satisfied. Within ten minutes, a brand-new mailbox rule begins silently forwarding incoming emails to an external address, and the account simultaneously starts pulling sensitive files from a corporate SharePoint repository in a pattern it has never displayed. No single event definitively proves an active compromise on its own, but the unusual sequence demands immediate attention.

Threat intelligence has been actively tracking a widespread wave of adversary-in-the-middle phishing campaigns specifically engineered to steal authenticated sessions, which explains why an MFA success cannot clear the account on its own. That vital context ships onward to the wider team as a brief advisory containing indicators and technique IDs, while the specific behavioral sequence and the local conditions that make it dangerous are left behind.

The threat hunter translates that advisory into database queries and quickly uncovers two crucial facts the original advisory never asked about: device-compliance data only covers a portion of the enterprise environment, and SharePoint audit records routinely show up hours late. The hunt forwards a list of potentially suspicious accounts, but the underlying coverage caveats stay behind.

Detection engineering builds custom logic that only fires when the unfamiliar network connection, the successful MFA challenge, and the new forwarding rule all cluster tightly within a short timeframe, fully aware that the rule lacks device-state visibility for a significant slice of the user base. What ultimately goes out the door is a basic severity level and a generic description field. The engineering assumptions and the expected false-positive patterns stay behind.

The alert finally reaches an analyst mid-shift, showing a sign-in event and a mailbox rule with none of the underlying reasoning that connected them. The analyst must rebuild the picture across four separate consoles: identity, email security, the SIEM, and the asset inventory. Two competing explanations remain live. The user could simply be traveling or trying out a legitimate new business service, which accounts for the unfamiliar network but fails to explain an external forwarding rule and an unprecedented access pattern. Alternatively, an authenticated session was stolen, which neatly accounts for the entire malicious sequence. The second explanation fits the available evidence best, but the endpoint scope remains completely unknown because the device is unmanaged and there is no process or network telemetry available to check. The case eventually closes with a recommendation to disable the account, while the competing explanation, the analytical confidence level, and the unexamined endpoint all stay behind.

A support ticket then lands with the identity team instructing them to disable the account. That team knows something the SOC never saw: the user is right in the middle of a critical payroll run, and a blunt account disablement will immediately interrupt a time-sensitive business process. This does not grant finance an absolute veto over security containment, but it means the containment decision and the business continuity decision must be made collaboratively by individuals who can see both sides. Revoking live user sessions and stripping out the malicious forwarding rule represent low-risk moves, while suspending the account falls under corporate incident policy and belongs to whoever holds that specific administrative authority. Moving the payroll run depends entirely on whether a backup operator exists and is free to take over, and reopening access must wait until credential resets, MFA re-enrollment, and a managed device are confirmed.

Every single function did its job properly, yet the enterprise architecture still forced each group to rebuild the incident entirely from scratch and handed the one team holding crucial business context a one-line task instead of an informed decision-making framework.

The Unicorn Analyst Is a Symptom

When organizations frequently experience this kind of operational friction, the institutional reflex is to post a demanding job listing: seeking an expert fluent in identity, endpoint, cloud, email, malware analysis, detection logic, and executive communication, willing to sit directly on the alert queue. This mythical unicorn analyst is not a sustainable talent strategy; rather, it is merely an expensive workaround for missing system state.

A senior analyst succeeds primarily by knowing unwritten facts that no monitoring dashboard shows. They know which log source routinely lies, which service account must never be touched under any circumstances, and which application owner will actually pick up the phone at two in the morning. The company’s true operational runbook effectively lives inside that single person’s head, and it permanently resigns when they do. A meaningful share of analyst burnout is driven precisely by this dynamic: endlessly re-deriving what the organization already knew and failed to preserve.

The most expensive loss typically lands immediately after an incident is closed. Suppose an investigation ultimately proves benign—the employee was traveling, and the forwarding rule had been properly approved. The owner of that rule urgently needs the evidence that flipped the verdict, and the telemetry owner needs to hear that device coverage was only partial. Instead, what the system preserves is a generic closure reason. The final verdict survives, but the underlying lesson evaporates entirely. This is precisely why a noisy detection rule stays noisy for years, and why each newly hired analyst is forced to rediscover the exact same blind spots independently on their own shifts.

What a Stateful SOC Remembers

The ultimate fix for these systemic challenges is architectural rather than personnel-based. Security operations centers must evolve to become stateful. Modern SOCs are rarely amnesiac; they retain historical evidence and past case files, often for years. However, what consistently fails to survive a handoff is the nuanced reasoning surrounding that evidence, the operational uncertainty that qualified it, and the constraints governing who possessed the authority to act. That context remains buried in whichever specific tool produced it instead of informing the next crucial decision. The necessary alternative is a shared operational memory—a system where all workflows actively read from and write to a unified set of states.

A shared model operating along these lines allows the SIEM, the endpoint detection and response platform, the identity provider, and the case management system to contribute meaningfully to a single cohesive decision, without requiring any of those foundational tools to be replaced.

The single hardest discipline in this approach is treating "unknown" as a genuinely legitimate answer. When endpoint telemetry is missing simply because a device is unmanaged, a weak security system routinely files the finding under a reassuring label like "No malicious process activity was observed." While technically true, that sentence is operationally misleading. A truly stateful system records that the endpoint could not be checked at all, appropriately lowers its stated confidence regarding endpoint scope, and routes the specific coverage gap directly to whoever owns device management. The coverage gap thereby becomes a permanent part of the ongoing case rather than vanishing into a comforting sentence.

Agents Need Jobs and Boundaries

Agentic artificial intelligence enters this operational picture last, and deliberately so, because bolting autonomous agents onto an already stateless security operations center simply grants a broken operating model significantly more speed to make mistakes. Bounded workflows operating securely from a shared memory architecture represent a completely different and far more viable proposition. Threat intelligence decides whether an outside threat actually matters locally and transparently shows its reasoning. Threat hunting reports the exact populations it successfully covered right alongside the ones it could not see. Detection engineering rigorously checks that the environment can feed a rule the precise data it needs before that rule ever goes live. Investigation packages timelines, competing explanations, evidence, and confidence scores into a single unified object. Remediation maps that decision directly onto available actions, designated owners, and required approvals.

Operational authority must remain strictly separate from analytical confidence. A mature framework distinguishes four distinct operational modes for any potential action: observing and gathering further evidence; putting a recommended action alongside its complete reasoning in front of a human who holds the proper authority; executing only after receiving explicit human approval; or executing automatically, but strictly where policy, confidence, entity type, and potential impact conditions are all fully satisfied. This mode lives directly within the control state, versioned and fully auditable. A confident-sounding narrative generated by an AI model earns an agent exactly zero additional execution rights.

The exact same operational caution must govern automated learning. A single false-positive verdict generated by a single analyst is exceptionally thin evidence for altering production detection logic. Human analysts make mistakes, and certain cases are simply legitimate business exceptions. A stateful system captures the underlying evidence behind the correction, aggregates similar historical cases, drafts a proposed change, and routes the formal proposal directly to the owner of the detection rule. That critical review step is what separates genuine continuous learning from self-generated operational corruption.

The Analyst’s Job Moves Up the Stack

The arduous evidence-assembly half of a security investigation is already complete by the time a human analyst arrives on the scene. Consequently, the analyst’s very first professional move is to vigorously challenge the structured case: verifying whether the working hypothesis truly holds together, checking whether a competing explanation was inadvertently missed, evaluating whether the proposed remediation action is strictly proportionate to the evidence, and assessing how the broader business context alters the situation.

Measurement metrics naturally shift in the exact same direction. Simply counting the total number of completed agent tasks flatters the underlying software without proving security value. Four distinct questions do a significantly better job of measuring true SOC health: does the analyst open an incoming case that already contains the necessary context, does the case accurately record what could not be seen, does a corrected verdict successfully reach the rule’s owner while the correction still matters, and did every automated action strictly stay within policy guidelines with a verifiable audit trail left behind? Revised federal guidance points in this exact same direction, with updated incident response recommendations from NIST treating response as an integrated component of an organization’s wider risk management strategy rather than a self-contained, isolated SOC activity.

The adversarial attack loop is rapidly tightening, and simply waiting around for full artificial intelligence autonomy to arrive is a slow and dangerous way to concede the advantage. The necessary starting points are unglamorous: measure precisely where the same operational context keeps getting reassembled by hand, record what an investigation could not see right alongside what it ultimately concluded, decide definitively who owns each action and who approves it while operations are calm, and route everything learned back into threat hunting and detection engineering.

The finance employee’s account gets suspended either way. In a traditional SOC, the lesson evaporates entirely along with the closure reason, and the payroll disruption surfaces only after the fact. In a stateful SOC, the people taking action can clearly see what the investigation was unable to observe, the coverage gap has a designated owner, and the next analyst inherits institutional memory rather than an endless queue of tickets.

By Sagoh

Leave a Reply

Your email address will not be published. Required fields are marked *