Back to Blog

Is AI pentesting as good as a human pentester?

Daniel Thatcher
Daniel Thatcher
Security Research Engineer

Key Points

With the rise of AI pentesting, a lot of people may be wondering if it’s as good as a human. Having started as a pentester twelve years ago, and having moved through similar roles to working on our own AI pentester, I've come to understand the advantages and disadvantages of both. Here’s how they compare.

The strict methodology was never the good bit

You may expect AI pentesters to follow strict methodologies in the way that many human pentests do. We experimented with this approach early on - structured systems where each agent had a specific vulnerability type to look for. What we found was that agents would miss things outside their remit entirely, or miscategorize what they did find as the one thing they'd been told to look for, causing confusing reports.

A complete lack of structure also failed to produce good results, as completely free agents without any guidance couldn’t break down and manage large applications properly. They could find interesting things, but at the cost of missing a load of other vulnerabilities in our targets. We’ve put a lot of testing into finding the balance between these approaches, and settled on giving agents freedom with broadly defined roles to help them approach and manage large applications, and stricter guardrails to stop them causing damage. 

This does mean that our AI pentester doesn’t tick off the parts of a methodology document that you’d typically see in a human pentest. That may make some people uncomfortable - a strict methodology gives people assurance that someone has gone through your application and tested everything on the list. An unauthenticated SQL injection needs to be found, and a checklist is there to make sure it does.

The methodology gives a lot of people the idea that a human pentest is a perfectly repeatable product, as any tester will be able to follow the methodology, and everything important will get found. The reality is a lot more nuanced than that though, as pentesting is a skilled job, and many of the most impactful findings on a pentest will come from the tester being experienced in noticing weird behaviours and following up on them.

For our AI pentester, we ended up with much more interesting and impactful results by not trying to guide our agents through specific vulnerability types. A good example is a complex authentication bypass the agents found that had been sitting in a codebase through five years of pentests without any human testers picking it up. The mechanics of that issue mean that you’d never think to write a testing guide that could be strictly followed to identify it, and if we’d tried to limit our AI pentester too much it never would have found it either.

You could say that only finding the common things would still be pretty good, but there's such an array of vulnerabilities that aren't common that you find during a pentest. And just because they’re not common, it doesn’t mean that an attacker can’t find them too, especially now AI is being used to enable attacks more.

A strict methodology being carried out by a competent tester only ever provided a low baseline of reliability, and it didn’t even do that perfectly. What makes a pentest good is the part that a skilled tester does outside the checklist, and that's where AI thrives - following its nose through your application and finding the things that were never going to make a list.

What they know, and what they can’t ask

Our AI pentests take a white box approach, which means the agents get the source code. They work through the entire codebase and build a picture of not just what every page is, but how the application is meant to function and what requests can be made. Doing this to the level of detail that the AI is able to isn’t feasible for a human pentester during a typical time window for a pentest.

The AI pentester also builds up a sense of what the application is for, and a good example of this is access control - it doesn’t need to be told that low privileged users shouldn’t be able to reach admin functionality, or that a particular bit of data is sensitive.

You can give them extra information as well, just like a human tester. For example, people often specify that they’re worried about their customers being able to access each others’ data, and the pentester has done a great job of focusing on these issues and identifying a lot of them.

What the agents can't do is ask you for anything you didn't already give or tell them. If I don't understand what a part of an application does while I’m doing a pentest, I can ask my contact, and those conversations can be very useful for testing the application properly, and assessing risk in the final report. I can also clarify the scope with my contact if needed, or ask for permission to verify something when I’m unsure if there’s a risk of disruption. We’ve had to deliberately make our AI pentesting agents cautious on this front to avoid going out of scope.

Neither finds everything

The freedom that lets AI pentesting agents find something like a complex auth bypass is the same thing that makes them less predictable. They decide where to dig, and that won't be identical every time.

But pentest results have always varied with the person doing them. It's why many consultancies will rotate your tester each year - different people find different things on the same application. This comes down to people’s experience, specialisms, and incidental knowledge. I’ve found issues because of a blog post or specification I happened to read years earlier and was left sitting in the back of my head. And even the things on a checklist get missed sometimes. People are fallible, and covering an entire application methodically over multiple days can be hard.

Consistent coverage is something we're actively working to improve, but it’s not something that was ever offered perfectly from anyone. It's why every human pentest report comes with a disclaimer saying as much, even though most people read it as a formality.

The remediation I couldn’t have written

Every pentester has strong opinions about what makes a good report. I’ve built up many of my own over the years, as have the rest of the current and former pentesters at Intruder. These opinions are what we built our report around.

The executive summary is the hardest page of a report to get right. For a lot of people it's their entire experience of your work, and at a good consultancy a lot of effort goes into it. We’ve worked to get to the point where the agents are now writing genuinely good ones, and I often read them first on our own reports to get a good overview before diving into detail.

Where they're still off is scoring risk. They rate things higher than I would, partly because they can't have the conversations with you that would bring a rating down, and partly because LLMs err towards caution when you start talking about cybersecurity. That's the right direction to be wrong in, but it does mean you'll see things flagged as more serious than they are, and it's something we're working on.

Remediation guidance is where the agents excel. After a long week on a pentest it's tempting for a human to write generic advice, and you catch yourself and try to do better. Unfortunately, you're usually limited by time and usually by not having the source code. The agents don't have either problem, so they'll tell you the line of code to change, and what to change it to on every issue. They also write about mitigating controls that can help reduce the impact of the issue, or provide a temporary fix.

We're seeing that speed up how quickly people actually fix things. We had a customer with 78 issues who fixed all of them within a few days, which as a human pentester I have never seen happen. I couldn't have written the level of detail that made those fixes possible in the time I'd have had.

So is AI pentesting as good as a human?

AI pentesting and human pentesting are different, and once you stop trying to rank them the comparison gets a lot more useful.

The agents are finding serious issues in applications that human testers have already been through multiple times. They'll read a whole codebase, work through every path in it, and write up every instance of what they find with the fix down to the line of code.

And you can run one whenever you want. For what a single human engagement costs you can run four AI pentests, which means you can do deep testing on a regular basis rather than once a year. That’s a key thing that is missed by human pentesting, and often applications have undergone a lot of untested changes in between pentests.

See what our AI pentesting finds in your application. Get started. 

Get our free

Ultimate Guide to Vulnerability Scanning

Learn everything you need to get started with vulnerability scanning and how to get the most out of your chosen product with our free PDF guide.