Decoding the world of cybersecurity

Autonomous pentesting is solving the wrong half of the problem

Jay Kaplan, CEO and Co-founder of Synack, argues autonomous pentesting can expand coverage, but without expert human validation it risks shifting security teams from untested assets to untriaged findings.

Autonomous pentesting is solving the wrong half of the problem
Summary
  • Agentic AI can expand pentesting coverage across environments that annual testing often leaves untouched.
  • Jay Kaplan argues autonomous findings still require expert human validation to maintain accuracy, credibility, and accountability.
  • DORA's threat-led penetration testing regime reinforces the role of accredited human testers and professional responsibility.

Contributed article

Jay Kaplan

CEO and Co-founder, Synack

Ask an enterprise what share of its attack surface gets tested and you’ll get a number. Ask how they arrived at it, and the room goes quiet.

Our research with Omdia, The 2026 State of Agentic AI in Pentesting, puts the average at 32%. Ninety-five percent of organisations rank pentesting a top priority, and fewer than a third of their environment sees a tester. But 32% is the flattering read. It measures coverage against the assets an organisation knows it owns. It does not count the ephemeral cloud workloads, the acquired subsidiary’s forgotten estate, or the internal LLM someone stood up last quarter with a key sitting in the front end. The denominator is bigger than the spreadsheet says. The gap is worse than the number.

Most CISOs already suspect this. The harder question is why the gap has widened while budgets have grown.

The answer is unflattering to my own industry. For twenty years we sold the annual penetration test as a security control when it was really a receipt, a point-in-time artefact produced to satisfy an auditor, delivered six weeks after the code it described had already shipped. Organisations bought the receipt because the receipt was what the framework asked for. Nobody was rewarded for measuring what remained untested.

Agentic AI arrived promising to fix exactly this, and on the coverage question it genuinely can. Agents don’t get bored running reconnaissance across 40,000 hosts. They don’t ration attention across a quarter. Pointed at the sprawl that has historically gone untouched between annual tests, they are the first credible answer to the breadth problem we’ve had.

But breadth was never the hard half.

The failure mode of autonomous testing isn’t that agents find nothing. It’s that they find everything and stand behind none of it. Our data shows 69% of security leaders would need agentic AI to reach at least 85% accuracy against manual testing before they’d trust it in production. I’d argue the accuracy threshold is the wrong frame entirely. At enterprise scale, 85% accuracy means roughly one in seven findings is wrong. The first time an engineering team burns a sprint on a phantom, the credibility of the entire feed is gone. Security programmes rarely die of false negatives. They die when engineering stops reading what we send them.

So an unvalidated firehose doesn’t close a coverage gap. It relocates it, from untested to untriaged, and every security team I know is already drowning on the triage side of that ledger. Moving a risk from unknown to unread is not risk reduction. It’s a change of filing cabinet.

There’s a commercial question underneath this that the market has been slow to ask out loud: who is accountable for the finding? Autonomy is a capability claim, but it is also, conveniently, a liability structure. When no human validated a result, no human is answerable for it, and the customer inherits the judgement call they thought they were buying. It’s worth noticing that half the products currently marketed as agentic pentesting are a scanner with a language model bolted onto the report generator, sold at a multiple of what the scanner used to cost.

European regulators, to their credit, worked this out before the vendors did. DORA’s threat-led penetration testing regime doesn’t merely require testing; it specifies who may perform it, demanding accredited testers with demonstrable experience and professional indemnity behind their conclusions. That is a regulator writing human accountability into law. You cannot satisfy TLPT with a fleet of agents and a dashboard, and the reason isn’t technophobia. It’s that someone has to be answerable when the assessment is wrong.

This reframes the talent debate too. We don’t have a shortage of security talent. We have a shortage of security talent doing work worth their salary. Every hour a senior researcher spends re-running reconnaissance or grooming scanner output is an hour not spent on the business-logic abuse that actually breaks the system: the chained flaw across three services that no model has been trained to imagine, because it has never appeared in a corpus.

Automate the first hour. Protect the second.

That’s what the practitioners themselves are converging on. Some 64% name agent-led, human oversight as their preferred operating model over full autonomy, and organisations that have actually deployed agentic pentesting are 1.6 times more likely to say human oversight is a permanent requirement. Note the direction of travel: proximity to the technology increases the appetite for humans in the loop rather than reducing it. The people selling autonomy are, almost without exception, further from the work than the people rejecting it.

Closing the coverage gap requires two things. Test more of the environment, far more often. Then confirm that what you found is genuinely exploitable, with someone’s name against the verdict. Agents give you the first. Only expert humans give you the second.

Skip the second half, and “we deployed AI pentesting” becomes a line in next year’s post-mortem, a few paragraphs above the part where someone asks why nobody read the findings.

×