Blog

AI Attacks Now Move With Little Human Involvement. Controlled Autonomy Is How Defense Moves Just As Fast.

eSentire

September 18, 2026

8 MINS READ

The attack begins with a person initiating a prompt and plan. Someone picks the target, approves the plan, and walks away. A machine begins working 24/7, deciding which files are worth taking and pushing down several paths at once.

State-sponsored intrusions can now run 80 to 90% autonomously. Generative AI has pushed phishing volume up 1,265%. Every new generation of frontier AI models raises the same question: how long until humans are not required to execute the entire kill chain? The answer isn't more products. It's collapsing offense, defense, and governance into one system, moving at the same speed as attackers can move.

The Rules Have Changed

In 2025, a Chinese state-linked group tracked as GTG-1002 ran what appears to be the first largely autonomous, end-to-end AI cyberespionage campaign. Human involvement was capped at about 20 minutes per target: picking targets and approving a handful of go/no-go decisions. Everything else, network mapping, credential harvesting, ranking stolen data by intelligence value, ran without a person watching. Claude fired thousands of requests, adapted mid-attack, and compromised roughly 30 organizations before anyone noticed. It used the same principle as old-school "stuffer" attacks: enough noise that the real intrusion gets lost in it.

It was by no means flawless, though. By Anthropic's own account, the AI hallucinated more than once mid-campaign, claiming stolen credentials worked when they didn't, or that exfiltrated data was valuable when it wasn't. But the compromises that did land were real, and defenders didn't find out until after the fact.

A second case, GTG-2002, scaled the same model down to a single operator: one person running a ransomware campaign across 17 organizations, with Claude handling reconnaissance, malware development, and ransom amounts set by what each victim could plausibly pay quickly. Instead of locking victims out, the group loaded ransom notes directly into the boot process, so the message was the first thing a victim saw when they powered the machine on. Data and intelligence were held as leverage, not encrypted files.

A third, GREYVIBE, turned consumer AI tools into a full development team for a Russia-linked group building malware, fake charity sites, and RATs targeting Ukraine, refreshing its toolkit daily to dodge detection. Despite the sophistication, the group left recognizable signatures in its own tooling, researchers were able to pull matching samples straight off VirusTotal and traced the operation back to its source.

Three very different actors, two different relationships to the technology. GTG-1002 and GTG-2002 turned the LLM itself into the infrastructure running the attack. GREYVIBE used AI as a development toolkit instead, building the malware ahead of time rather than running the attack live. Either way, the shift is the same: the attacker no longer needs to execute the attack across the entire kill chain, be patient, or even present. It just needs compute.

The Old Playbook Doesn't Work Anymore

Traditional security runs on a chain of human checkpoints: someone finds a vulnerability, writes a detection rule, reviews an alert, approves a patch. That chain assumed attacker and defender moved at the same speed. That assumption is now false. Attacks that used to take days now finish in under two minutes, sign-in to data exfiltration, with no analyst in the loop on either side.

"Time to detect" is no longer a sufficient metric. "Time to Engage Signal" matters more now. How fast can an organization go from signal to neutralized? If that's measured in hours, the breach is over before anyone manages to open a ticket.

It also exposes how security built itself around selling point solutions, not around how attackers operate: EDR for endpoints, identity tools for identity, vulnerability scanners for infrastructure, each sold to a different buyer in the same organization. That fragmentation made commercial sense for 25 years. Against an adversary that moves across every category in seconds, it doesn't allow the defense to adapt.

Offense And Defense Can No Longer Be Separated

For decades, red and blue teams operated on different rhythms. A pen test team finds an exposure, writes it up, and the report lands in a ticketing system/email thread days or weeks before anyone acts on it. Meanwhile, the SOC is watching live telemetry with no visibility into what the pen test already found. That gap between finding a problem and fixing it used to just mean an organization was "at risk." In an AI-driven attack, that same gap means the organization is already compromised.

Closing it requires a "connective tissue" that most security teams don't currently own: a shared fabric where a finding on the offensive side automatically triggers a detection rule on the defensive side, not after the next planning meeting, but in the same breath.

Defense Built for Speed Already Exists

In one case we identified, network sensors picked up attackers probing customers for a known NetScaler vulnerability in the CitrixBleed family, the kind of low-grade scanning most systems shrug off as background noise. AI looked closer, confirmed which customers were exposed, and wrote and deployed a new detection rule in under six seconds. Offense and defense closed the loop in under 2 minutes, start to finish.

Attackers can go from stolen credentials to a live forwarding rule in as little as 14 minutes. One recent account-takeover case we saw never got that far. A suspicious sign-in traced to a Baltimore data center came in at 10:22 a.m. 28 seconds later, AI flagged the location as atypical for that account. 24 seconds after that, it ruled out plausible travel. 29 seconds later, it checked the IP against the rest of the customer base and found it had never shown up anywhere before, a real flag but not yet a verdict. 31 seconds after that, it caught mailbox forwarding rules being set up, the telltale sign of an attacker staging an inbox for exfiltration and disabled the account. Less than 2 minutes total, with no analyst in the loop. A human-paced response would not have been as fast.

This same discipline held firm another time on our platform, and the adaptive guardrails again did their job. Autonomous recon mapped a customer's environment and surfaced a Webmin MiniServ 1.910 instance, wide open on port 10000. It matched the version against CVE-2019-15107, a critical, unauthenticated remote code execution vulnerability, and drafted the exploit. Then the adaptive guardrails engaged: engagement scope controls determined payload delivery was outside the operational boundary, so the exploit never ran. The finding was surfaced, confirmed, and routed straight into the customer's remediation workflow, without a human having to step in and stop it.

These examples worked because three things were already happening together: AI was given "swords," "shields," and adaptive guardrails that moved as one system instead of three separate tools. Swords validate what's actually exploitable, not just theoretically vulnerable. Shields decide how to detect and contain what swords find. Guardrails are the policy layer sitting over both, the rules an organization sets for what autonomous tools are allowed to do independently, and how far they're permitted to go.

How Defense Can Adapt

Measure everything. Time-to-engage signal must become the headline metric, replacing time-to-detect and time-to-patch as the number leadership watches.

Set adaptive policy guardrails before agents act. If autonomous tools are running offensive testing, does policy keep them from exploiting exposures on critical assets? If they're running defense, can they revoke a session or isolate an endpoint on their own, with a full audit trail and a one-click way to reverse the action?

Replace siloed operations with coordinated, and decentralized agents. The SIEM-then-SOAR era concentrated everything. The evolution is a set of specialized agents working in parallel, the way a security team already divides labor, each accountable for its slice and able to explain its actions in plain language.

Build the fabric between offense, defense, and response before you need it. The right question isn't whether offense and defense use the same platform. It's how fast a finding moves from exposure to contained: one loop if the fabric is real, days or weeks if it still has to pass through a ticketing system and a change-approval meeting first.

The Choice Is Already Here

Attacks used to require a skilled human behind the keyboard. That barrier is gone. Organizations that treat this as a fundamental operating model shift, not just a tooling problem, are the ones that will be better positioned as attackers begin to incorporate offensive capabilities more broadly.

References

To learn how eSentire can help you find exposures and defend your organization, connect with an eSentire Security Specialist now.

GET STARTED

ABOUT THE AUTHOR

eSentire
eSentire

eSentire is a leader in Controlled Autonomy SecOps, protecting 2,000+ organizations across 35+ industries around the world. Founded in 2001, the company’s Controlled Autonomy SecOps operating model pairs agentic AI operatives with engineered human-judgment controls, delivering expert-depth security outcomes at machine speed without ceding accountability to opaque automation. Powered by the unified agentic AI Atlas Platform, eSentire’s Atlas AI + 24/7 expert human SOC coverage delivers offensive capabilities that preempt exposures before attackers do, detect, and respond to stop threats in real time. For more information, visit esentire.com and follow @eSentire.

Back to blog

Take Your Cybersecurity Program to the Next Level with eSentire MDR.

BUILD A QUOTE

Read Similar Blogs

EXPLORE MORE BLOGS