In late April 2026, the AI coding agent Cursor, running on Anthropic’s Claude Opus 4.6, deleted the car-rental software company PocketOS’s entire production database in nine seconds.
The agent encountered a permissions error, found an unrelated API token in a config file, used it to access the company’s cloud provider, and erased everything. When the frustrated founder asked what had happened, the agent quoted its own instructions back to him – the ones that explicitly banned destructive, irreversible actions. The agent simply explained that it had violated every one of them.
While, thankfully, PocketOS’ infrastructure provider, Railway, was able to provide disaster recovery through off-site backups, the incident is still sort of like a science-fiction coding nightmare. Except it’s already happened. And ominously, it may also be a preview of what’s to come for many organizations.
The Disconnect Between Most Cybersecurity and Actual Security
Those who attended this year’s RSAC would agree that almost every security vendor was offering the same exact pitch: Look at our new AI agent: it doesn’t just summarize your alerts, it acts on them.
It’s everywhere. Google folded Chronicle and Mandiant into what it now calls Google SecOps, and says its triage agent now compresses an extensive manual review into a minute. Securonix, Stellar Cyber, D3 Security, and a stack of well-funded startups are all also selling a version of AI agents that gather context, decide, and actively close the loop themselves. In fact, Gartner expected task-specific agents in 40% of enterprise applications by the end of 2025, up from under 5% from the previous year. Totals can only be expected to grow since.
Yet what was missing from almost every pitch we heard was that perhaps the most significant threat isn’t an AI agent that’s slow. It’s the agent that’s both confidently and aggressively wrong, working at machine speed, with nobody to catch it until the damage is already done.
For Cyber, Humans are Still Necessary – But How Can They Keep Up?
Contrary to many of those RSAC pitches, the Coalition for Secure AI’s guidance on agent identity, published weeks before the PocketOS incident, argues that agents must hold Zero Standing Privilege, so no persistent, broad-scoped credentials – at all. Access should be granted just for the task and revoked the moment it’s done. It makes sense, since the stray token Cursor found and used is the same risk profile as a SOC agent with standing, unreviewed auto-remediation rights across your environment. You could argue that these rogue agents, in the hands of cybersecurity providers, could be far more damaging.
Yet if cybersecurity were as easy as just having humans remediate every incident, every provider would already be offering that. An effective human-led SOC isn’t easy to provide. There’s a scarce supply of analysts who actually have the expertise to catch the subtle issues. There’s the sheer cost of infrastructure for running that kind of operation around the clock. There’s also human burnout.
And besides, the AI genie’s already out of the bottle. Digital automation has enabled attackers, too, with far more convincing phishing, more variations of the same intrusion tried at higher volume, more of, well, everything. The number of alerts is climbing faster than any hiring plan can keep pace with. While humans are still the right call, they just can’t do it alone anymore.
Where Agentic AI Enables Security Teams, Without Losing Control
Thankfully, a modern SOC can already automate most of the lower-priority issues: gathering endpoint and identity telemetry, enriching an event with threat intel, searching historical activity, and building a timeline. As long as it’s fast, repeatable, and fundamentally simple, it is worth automating.
But deciding which tasks actually are fast, repeatable, and fundamentally simple is where the real level of cybersecurity expertise – even artistry – comes into play. There are a lot of details to cover in a short summary, but we do find the question consistently breaks out into three fundamental principles.
1. Keeping Agentic Capabilities Limited – by Design
This means deciding exactly where the agent’s job stops, and the analyst’s job starts, before an incident forces the question.
It also means ensuring the agent is completely blocked from performing any review, analysis, or tasks beyond that scope.
2. Automation Only Where the Issue is Obvious and Contained
Automation should only be applied when the available evidence leads to a clear conclusion and the resulting action is limited, predictable, and easy to reverse.
It is never appropriate for the system to evaluate human intent, parse different possible explanations, or decide whether unusual behavior is legitimate.
Take a remote administration tool launched on a workstation. An agent can confirm whether it is signed, commonly used, and present elsewhere on the network. It can compare the command line with known patterns and check whether the process spawned anything suspicious.
However, it should not establish why the tool was launched in the first place, such as whether it was launched by an autonomous assessment, an employee performing routine work, or an attacker using a trusted tool to move through the environment.
In that situation, the agent can gather evidence, surface relevant context, and recommend next steps. The decision about whether the activity is authorized – and the next steps to address it – belongs to an analyst who can evaluate the broader context.
3. Reduce Human Investigation Time By Providing All Needed Context
Our team recently documented an investigation where a rogue ScreenConnect installation looked, at first, like ordinary use of a standard tool. But the endpoint telemetry showed Quick Assist had run shortly before the alert, a common Windows remote support tool that’s used by threat actors in tech-support scams.
Everything about that incident would look fine to basic automated tools – and even to many experienced SOC analysts – without the right context.
That’s why simple dashboards of “Was this tool executed?” are both too much and not enough. For one, it’s too much unfiltered data. And then, there are other necessary questions an experienced SOC analyst would ask:
- Who initiated the activity, and was that identity expected to do it?
- Was the session tied to a known support workflow?
- What process launched the tool?
- Did the activity line up with credential access, persistence, lateral movement, or data access?
- Would containment interrupt a business-critical process?
Automation will simplify the discovery process for many of these. But for the rest? Intent, consequence, and human insight – it’s where an expert analyst’s judgment earns its keep.
Fewer Alerts, the Same Human Expertise, and Better Results
At RSAC 2026, we saw many automated SOC comparisons centered on a single question: how much of an investigation can an agent complete without human involvement?
Sure, that metric is easy to chart, but it’s increasingly proving to be the wrong one.
In our opinion, the best agentic SOCs aren’t removing their experts. Security investigation shouldn’t be a contest to eliminate the participation of human experts with the most experience and insight. It is a thoughtful process of separating routine work that can be automated from the judgments that depend on full context, real experience, and true accountability.
Of course, that offering is harder to build than a fully autonomous SOC product, and it isn’t as flashy a RSAC headline. But right now, when done right, it’s the model that works best.
Adam Gnuse is a bestselling author and the Senior SEO Manager of Outreach at Huntress, a cybersecurity company that helps organizations of all sizes strengthen their defenses with a human-led, agentic security operations center. There, he partners with the organizations advancing digital security through forward-thinking approaches, helping share the latest research and practical insights with the broader community.




