← Research
Field Notes

Why Traditional Penetration Testing Is Broken

A penetration test is a photograph: sharp, useful, and obsolete within a week.

Every year, organizations spend anywhere from a few thousand to several hundred thousand dollars on penetration tests. The report arrives, executives feel reassured, a compliance box gets checked, and everyone moves on.

The attackers don't.

The gap is what's broken: a security posture measured once, set against an attack surface that changes every day. Penetration testing itself isn't the problem, it remains one of the best ways to find complex, chained vulnerabilities. The problem is asking a point-in-time test to describe a moving target.

The snapshot problem

The snapshot problem is that a penetration test only describes the environment at the moment it was run, while the environment keeps changing after the report is delivered. A penetration test is a photograph. It captures one environment, over one narrow window of time, as assessed by one team. The picture can be sharp, genuinely useful, and obsolete within a week.

This isn't a new problem. Environments have always drifted between assessments. What has changed is the rate of drift. Engineering teams now ship code continuously and it's only accelerating. That acceleration is what turns an old, tolerable gap into an unacceptable one. By the time a report is delivered, developers have shipped new code, cloud resources have spun up and down, containers have been rebuilt, dependencies have been bumped, and a new API or two has quietly gone live. None of that is in the photograph.

Attackers don't work on your schedule

There's also an assumption baked into annual testing that the threat operates on the same cadence you do. It doesn't. Attackers aren't waiting for next year's engagement or aligning to your compliance calendar. They scan continuously, and every new internet-facing app, stale dev environment, exposed credential, or vulnerable dependency is a fresh opportunity the moment it appears. That asymmetry is measurable: CrowdStrike's 2026 Global Threat Report found that 42% of vulnerabilities were exploited before their public disclosure. That means a meaningful share of exposure never shows up as a 'known' risk in a CVE scanner and an annual test is too slow to catch it. The defender measures annually. The adversary measures constantly.

The surface got too big to photograph

Part of why the snapshot model held up for so long is that environments used to be small, a rack of servers you could enumerate by hand. A typical environment now spans public cloud, Kubernetes, internal and external APIs, SaaS, third-party integrations, CI/CD pipelines, a deep tree of open-source dependencies, and increasingly AI-enabled components. Human-led assessment is still essential for depth and judgment, but no team can manually maintain continuous visibility across a surface that large and fluid.

The largest piece of that surface is often the least visible: the software you didn't write. Modern applications are assembled from thousands of third-party libraries, and a single vulnerable one can expose the whole system. The last few years have turned supply-chain compromise into a primary attack path rather than an edge case. The uncomfortable reality is that many organizations don't know which vulnerable components are running in production until a critical advisory lands. Then the scramble begins.

From "are we secure?" to "how is our exposure changing?"

The fix isn't to abolish penetration testing. It's to stop asking it to do a job it was never built for. Periodic deep testing should sit alongside continuous validation, so the operative question shifts from a yearly "are we secure today?" to an ongoing "what changed, and what did it expose?" Continuous validation is what lets you answer, on any given day: what assets appeared this week, which of them carry known vulnerabilities, which new attack paths opened up, and how fast you can see all of it. Security becomes a process, not an event.

This is what AI is really good at. Reconnaissance, asset discovery, vulnerability correlation, attack-path analysis, software composition: these are high-volume, repetitive tasks where machine scale actually helps. AI can keep pace with a surface that humans can't manually track. What it doesn't do is replace the practitioner. Validation, prioritization, and the judgment call about what matters still belong to people. The useful framing isn't "AI vs. analysts." It's AI handling the volume so analysts can spend their attention where it counts.

How do we approach this differently?

At ADCL, we assert that security testing should be continuous rather than episodic, and we built ThreatWell around that idea. ThreatWell runs continuous offensive testing and attack-surface validation, keeping the picture current instead of updating it once a year, alongside software composition analysis that surfaces the risk buried in supply chains and open-source dependencies before an advisory forces the issue. Both capabilities serve the same goal to find the exposure before someone else does.

The question is no longer whether to run penetration tests. You must. It's whether a periodic test, on its own, can describe a surface that never stops moving. As environments keep expanding and organizations ship faster, continuous validation stops being a luxury and becomes the baseline.

The organizations that adapt get visibility, lower risk, and faster response. The ones that don't, will learn about their vulnerabilities after the attackers do.

Get the research
before the attackers do.

We publish what we find so defenders can fix it first. Want this applied to your environment? Talk to us about ThreatWell.