Why CVSS Is Breaking Down in the Mythos Era of Validation

Defenders are facing a new problem: vulnerability discovery has exploded thanks to AI, but only a sliver of those findings matter in practice. In the first six months of 2026, over 35,800 CVEs were recorded—almost 50% more than the same period in the previous year. Yet just 495 of those were confirmed as exploited in the wild, and 116 were already under attack as soon as they were made public. That gap between risk potential and real threat is growing as AI models begin spitting out thousands of potential vulnerabilities for every one that’s truly dangerous.

That surge is being driven in part by “Mythos-class” models. These systems have flagged 26,153 potential vulnerabilities in open-source software, but only 421 have been patched upstream so far. In short: the volume of vulnerability candidates is skyrocketing, while follow-through isn’t keeping pace.

The Falling Shortcomings of CVSS and Untested Assumptions

The Common Vulnerability Scoring System (CVSS) provides a helpful baseline for categorizing risk. But it wasn’t built to factor in the specifics of an organization’s environment. An exposure on a disconnected, heavily firewalled server may score high, but pose zero risk in practice. Meanwhile, a low-scored flaw in a critical internet-facing asset might present far greater danger. With AI rapidly generating more CVEs, relying on generic severity scores becomes not just inefficient but misleading.

Automated pen testing (pentesting) helps—these tools can launch real exploits, chain together credentials, and map out how far an attacker might reach. But in practice, they only cover about one-third of an organization’s attack surface annually. Dead zones remain: systems that can’t be reached safely, or CVEs without working exploits yet. These are exactly where Mythos-like models and AI threat tools introduce the most alarm—and where traditional CVSS or severity labeling gives no clear action path.

Validating What Actually Matters: A New Trinity of Approaches

To get ahead, defenders need three interconnected capabilities to validate exposures in meaningful ways:

  • Exploitability validation: confirming whether a vulnerability can actually be used against your specific stack—including cases with no known exploit yet.
  • Security control validation: testing whether your prevention and detection tools are working properly in practice.
  • Agentic pentesting: running real exploits in a controlled, end-to-end fashion to show how far an attacker could penetrate.

These three form a framework that shifts security teams from responding to every High or Critical CVE toward focusing on the ones that pose real, timely threats. They must fit together in a platform that lets evidence flow between them so that new risks reprioritize automatically and fixes get revalidated rather than just closed out.

This model aligns with recent industry guidance emphasizing validated attack paths, decision-enabled response, and exposure reduction baked into day-to-day operations. It turns vulnerability management from a reactive checklist into a proactive risk validation practice.

Picus Security is doubling down on this trend. Its upcoming Validation Summit ’26 (October 14–15) will showcase what mature exposure validation looks like in real enterprise environments. Leaders from Chanel, Atlassian, and the NFL will share how they’ve restructured programs, tools, and workflows around all three validation pillars, including practical sessions showing what to do when an exploit exists, when it doesn’t, and how to handle patches and re-verification.

Analysis: The disparity between what AI tools discover and what actually threatens your organization makes clear that security teams must be more discerning than ever. CVSS alone is no longer enough. Putting energy into building exploitability evidence, control validation, and controlled, agentic penetration testing will separate noise from risk. If teams can master this three-part framework, they won’t just patch more—but fix smarter. Watch for platforms that can unite these capabilities without exponential effort, and for vendors increasingly judged by what they validate—not just what they detect.