When Anthropic introduced Mythos Preview, the story practically wrote itself. Here was a frontier AI model taking work that once required researchers and threat actors a lot of time and compressing it into hours. Mythos demonstrated the ability to find and exploit vulnerabilities across major operating systems and browsers. In controlled testing, it produced a working Firefox code-execution exploit in less than an hour and developed eight in roughly 12 hours. It also produced PoCs that triggered 18 of 21 patched Windows kernel vulnerabilities within six hours, then developed full privilege-escalation chains for eight.
The warning was stark and glaring, the exploit window was shrinking and defenders needed to prepare. A standard 30-day patch window starts to look pretty meaningless when an AI system can turn a public patch into a working exploit before most organizations have finished testing the update. It felt like a line in the sand for vulnerability management. A line that forced security professionals to rethink how they combat threats.
That concern was real then, and it’s still real now. Looking back, though, the clock story feels incomplete. Pre-Mythos, I didn’t think attackers cared about patching cadences, and post Mythos, my stance hasn’t changed. Attackers don’t work on our timelines, quite the opposite actually. I myself have experienced attackers targeting systems at rather inconvenient times, like three in the morning. We need to start challenging our own security narratives and start thinking like an attacker.
Attackers don’t care about your patching queue
As an attacker, I am not just asking how quickly I can exploit the vulnerability everyone is watching, I am actually likely not looking to target the very vulnerability every defender is watching like a hawk. I’m more likely to search for other entry points where defenders are looking the least. Which entry path is the least monitored and which one provides the most value to me, ROI if you will, as a threat actor.
The attacker's doors
Six ways into the same target. Attackers pick whichever door has the fewest eyes on it.
The headline CVE
HEAVILY MONITORED
Public advisory, patch teams, researchers, and automated scanners are all watching this door.
MEANWHILE, FIVE QUIETER DOORS SIT OPEN
A 20-year-old bug
An old vulnerability in a codebase nobody's re-checked in years.
A weak supplier
A vendor with less visibility, fewer resources, and slower response.
A stolen session
Captured credentials or an authenticated session, no exploit required.
A convincing call
A vishing call that sounds just believable enough to work.
An exposed MCP server
An internet-facing AI connector nobody's tracking.
YOUR ORGANIZATION
Every door leads to the same place.
It’s not like threat actors get extra credit for using the most sophisticated, shiny vulnerability, so why would they target a heavily watched vulnerability instead of the port someone accidentally left open? I have my doubts as to whether or not threat actors are going to focus on the flashy vulnerability when they can find an unmanaged decade old vulnerability that will also provide them access. If I am a hacktivist, I probably want to bring attention to my cause. If I am working for a national government, I am likely after intelligence, disruption, or persistent access. If I am financially motivated, I am going after whatever gets me closest to a payout. In every case, no matter the motivation, the method only matters if it works. Again, if a stolen session, a convincing phone call, an exposed appliance, or a weak supplier works, then why would I go after the very thing I know all security professionals are watching for?
In controlled testing, Mythos shrank the clock and expanded the amount of the attack surface that could be searched and potentially breached. It found value in places defenders had already discounted, like a 20-year-old vulnerability. The assumptions we use to decide what deserves attention (ahem, two-decade-old vulnerabilities) are becoming easier and cheaper to test at scale.
The vulnerabilities nobody expected to matter
Mythos found vulnerabilities that had survived between 10 and 27 years in mature codebases. The examples included:
A 27-year-old OpenBSD bug that could remotely crash a host
A 16-year-old FFmpeg flaw that caused an out-of-bounds heap write (however, Anthropic assessed that the FFmpeg flaw would be difficult to turn into a functioning exploit). Even so, the model continued to find meaningful weaknesses in software that was actively being used in environments, including in your vendor’s ecosystems.
The FreeBSD example demonstrated Mythos’s ability to autonomously find the vulnerability and produce an exploit that granted root access. The exploit required a 20-gadget, return-oriented programming chain divided across six RPC requests. That is difficult, tedious work that would require a lot of man hours. Historically, the time, expertise, and patience required to complete that kind of analysis helped limit how many people could turn an obscure vulnerability into something operational. Mythos showed how much of that work could be automated under controlled conditions.
Rethinking ‘exploitation unlikely’
The Windows kernel testing challenged another common security assumption. Microsoft had rated 14 of the 21 vulnerabilities tested as “Exploitation Less Likely” or “Exploitation Unlikely.” Mythos produced proof-of-concept exploits for 13 of those 14 and turned one rated “Exploitation Unlikely” into a full privilege-escalation chain. The model took vulnerabilities that many defenders could reasonably place lower in the queue and showed that several were more practical to exploit than their ratings suggested.
For years, vulnerability prioritization has quietly depended on the scarcity of exploit-development resources. Building a reliable exploit takes time, specialized knowledge, infrastructure, testing, and patience. A vulnerability might be technically exploitable and still attract little attention because developing a working attack would require too much time or money. That's why zero-days are so threatening; exploitation could begin before a patch was available, leaving defenders with a window of exposure.
Frontier AI changes that calculation. Frontier AI can keep working on an obscure vulnerability long after the time and cost involved would have forced many human researchers and threat actors alike to move on. Mythos showed that frontier AI could significantly lower the bar of entry for a wider range of threat actors.
Now, I don’t think threat actors are simply going to ignore high-profile vulnerabilities. Those will always be attractive entry points because they often affect a lot of companies (hello domino effect), expose valuable systems, and come with public patches that show researchers exactly what changed. One dataset found that ransomware payments fell 66% quarter over quarter. Coincidentally, third-party involvement in 48% of breaches, a 60% year over year increase. This could be an attempt by threat actors to increase the pressure to pay the ransom by targeting critical vendors instead of companies directly. Attackers also understand that widespread awareness rarely translates into universal patching. MOVEit showed how one widely used product could create downstream exposure across many organizations (2,700 entities impacted). It can be easy to patch vulnerabilities in the systems that immediately impact daily functionality, but it's also easy to forget about the older appliances, forgotten servers, acquired companies, or their vendors.
With the lower barrier to entry, AI makes both sides of the vulnerability landscape more affordable. AI lets an attacker chase the headline vulnerability without giving up the search for a quieter path. The obvious queue stays valuable, while the long tail becomes much more affordable.
From benchmark to reality
We are already starting to see what this looks like outside a frontier-model test. In April 2026, a Hacktron researcher used Claude Opus 4.6 to develop an exploit chain against an outdated Chromium release used by Discord. The experiment involved a known vulnerability in a controlled environment, and it required a lot of human guidance. The researcher stopped sessions when they went nowhere, redirected the model, and eventually pointed it toward a vulnerability they believed could be exploited.
The process took about a week, 2.3 billion tokens, approximately $2,283 in API costs, and around 20 hours of human oversight. The final proof of concept launched a calculator. Delivery through Discord would have required an additional cross-site scripting vulnerability, and the technical demonstration ran Chrome for Testing without its sandbox.
This was far messier than Mythos’ tests. It shows what happens when a commercially available model, a known vulnerability, an outdated application, and a determined researcher come together. The model needed human assistance, the experiment wasn’t free, and the finished exploit wasn’t perfect. Even so, one person was able to use AI to work through a patch gap affecting the Chromium version bundled with an application used by millions.
From a threat actor’s perspective, this is a pretty enticing opportunity. Threat actors aren’t trying to win awards for full autonomy, they’re trying to accomplish their goal. If AI handles a large part of the tedious work and a person only needs to step in occasionally, then the economics of the attack have changed. The model challenged our current security standards.
Changing the math
Bitsight threat researchers uncovered a post on a dark web forum advertising a free automated pentesting platform with an “AI attack planner” designed to find application weaknesses and build exploitation chains. The operator later claimed that the platform had found 17 vulnerabilities across eight projects that could be monetized or used to gain deeper access.
The service was presented as authorized testing and bug-bounty tooling, and Bitsight has not independently verified the findings or how much of the work was actually completed by AI. Even so, the post is interesting because of what the actor was looking for. The focus was not limited to major CVEs or headline-grabbing zero-days. Rather, the platform was looking for business-logic mistakes, authentication weaknesses, misconfigurations, and combinations of smaller flaws that could lead to a useful outcome. That is exactly how a threat actor thinks. The individual weakness matters less than the path it creates.
Figure 1: An underground forum user advertised an AI-assisted pentesting platform that builds exploitation chains and later claimed 17 findings across eight projects. The service was presented as authorized security testing, and Bitsight has not independently verified the claims.
What Mythos proves
Mythos gives us a glimpse of where these capabilities are heading. The UK AI Security Institute independently observed Mythos completing 73% of expert-level capture-the-flag tasks. In a simulated 32-step corporate intrusion, it completed the entire attack in three of ten attempts and averaged 22 completed steps.
The results are incredible, but there are some caveats. The range had no active defenders or common defensive tooling, and the model received no penalty for generating alerts. The model also failed to complete the operational-technology scenario after getting stuck on the IT portions of the range.
Currently, there is no public evidence that criminal actors have access to Mythos itself. I am using it as a benchmark for where AI-assisted cyber capabilities are heading. Evidence of what threat actors are doing today comes from commercially available model research, provider abuse telemetry, and what Bitsight is seeing across underground communities and forums.
I am not trying to prove that cyber criminals are secretly using Mythos. I am looking at what Mythos made possible, then comparing that capability with the behaviors, services, and attack methods already beginning to appear elsewhere.
What happened to the Zero Day Clock?
When I started looking at this story in August 2026, the Zero Day Clock seemed like the perfect proof point. Its original charts showed the time between disclosure and confirmed exploitation falling dramatically. The visual told a very clean story: the window was shrinking from years to months, then from months to days.
Figure 2: The original Zero Day Clock from time-to-exploitation chart.
Figure 3: The original Zero Day Clock exploit rate.
Then, the clock itself changed
Since I began my research, the original time-to-exploitation metric has since been retired. The current Zero Day Clock methodology explains that the metric could create a misleading picture of attacker behavior because recent and older vulnerabilities have had very different amounts of time to accumulate exploitation evidence. A vulnerability published years ago has had years for exploitation evidence to surface. A vulnerability published last month has only had a few weeks for that evidence to appear. That alone can make recent vulnerabilities look artificially fast.
But, the problem runs deeper than incomplete recent data. Public sources usually tell us when exploitation was first recorded, not when it actually began, meaning a vulnerability could be exploited quietly for weeks before the first public notice appears. If a threat actor is discreetly exploiting a vulnerability across a multitude of organizations, they’re unlikely to go online and post about it and potentially alert security professionals. That would be counterintuitive. So, once exploitation reaches the day of disclosure, the measurement also hits zero. At that point, it cannot distinguish between an exploit first used that morning and one that had been quietly leveraged for months.
In the old exploit-rate calculation, the published CVE volume has grown partly because the CVE program expanded the number of organizations allowed to assign identifiers. More vulnerabilities are being documented, which is good, but that growth changes the denominator. Dividing a narrow set of confirmed exploited vulnerabilities by a rapidly expanding CVE population can create the false appearance of a falling exploit rate even when the underlying threat has not changed.
The new Zero Day Clock takes a more cautious approach. It keeps published CVE volume, KEV listings, and observed honeypot (a decoy used to lure in threat actors) scanning separately instead of combining them into one dramatic rate. In the snapshot captured for this report, published vulnerability records were climbing sharply year over year while scanning activity against the 100-most probed CVEs and same-day-or-earlier KEV listings were both down. Of course, those signals measure different things and should be read separately. The growth in published vulnerabilities doesn’t necessarily mean more exploitation, and a decline in observed scanning or KEV listings doesn’t prove or make the cyber landscape safer.
Figure 4: Zero Day Clock Vulnerability Pressure Index.
So, do threat actors still care about CVEs? Of course, but this leaves us with a more complicated conclusion. Mythos demonstrated that frontier AI can compress exploit development from weeks into hours under controlled conditions. The available population-level data has yet to show that this capability produced a broad acceleration in real world exploitation. That could be due to lack of access to more sophisticated AI models, or because threat actors are focusing their attention elsewhere, like somewhere that isn’t being monitored by every security defender on the planet. I think both of these findings can be true.
Attackers still choose the path of least resistance
The zero-day clock is only one clock. A threat actor can skip the patch queue completely and instead focus on stolen credentials, social engineering, or a less defended organization. These methods have existed for years because they work; time and time again we have watched threat actors rely on human mistakes as their main point of entry. Thinking like a threat actor means ranking every available route by reachability, value, cost, and the chance of being seen. If the front door is covered by cameras, then I’m going to check the side door, the basement window, and whether someone left a key with the neighbor. Why create more trouble for myself and a higher probability of being caught by going through the heavily guarded front door? Cyber attacks follow the same logic. The path of least resistance (and the least monitoring) is usually the preferred path.
In May 2026, one forum actor advertised a browser-based phishing kit marketed for 2FA bypass. The seller claimed the kit could automate the phishing flow and give the buyer control of an authenticated victim session.
In July 2026, another seller advertised a custom reverse-proxy panel for a major identity provider for $2,800, with access managed through a Telegram bot.
Figure 5: A phishing kit advertised on the dark web with 2FA bypass capabilities.
Nefarious sellers are continuing to offer ways to capture credentials, trust, and authenticated sessions created after a victim logs in. It's fairly easy to purchase an entry point on a dark web marketplace, like through an Initial Access Broker (IAB), and it also eliminates the need to exploit an endpoint or a heavily watched vulnerability.
It's easy to want to focus on your organization, but this risk doesn't end with you; you need to think about your perimeter. Think about suppliers that hold your data or have a trusted connection into your systems becoming part of the same attack surface. Your own controls may be strong, but if that vendor has fewer resources, less visibility, or slower response times, the trusted relationship can become the quieter path in. If I cannot easily access your environment, why would I ignore the vendor that already has a key? From the attacker’s perspective, the supplier is simply another door.
I also came across an interesting reaction to Mythos on Dread (a dark web forum, kind of like Reddit or 4chan). In April 2026, a user reposted information about Project Glasswing and warned other criminals that AI could put their operations, devices, communications, and infrastructure at risk. The replies focused on writing-style attribution, traceability, misdirection, and the cost of securing large numbers of IoT devices. The threat actor’s concern was broader than whether Mythos could find a zero-day. They were thinking about their own exposure, identities, infrastructure, and operational security. Defenders need to think the same way, think like an attacker.
AI is moving into the middle of the attack
Exploit generation gets the most attention because it makes the best headline. I think an important (and arguably immediate) change is that AI can and has played a role in the midst of an attack. We’ve already witnessed threat actors using AI to consume and organize large amounts of stolen data, create intricate phishing lures, and develop malware. I don’t see threat actors stopping the use of AI anytime soon, instead, I think AI will increasingly appear in most aspects of an attack. Once attackers gain access, they often have too much information rather than too little. Smaller threat groups may steal more information than their members could realistically review on their own. AI can turn that pile of stolen data into something the attacker can actually use. We’ve also seen instances of threat actors leveraging AI for phishing and vishing (i.e. voice phishing).
In that scenario, AI is functioning as the attacker’s analyst, shortening the gap between stealing information and understanding how to use it. It might be easy to brush that off as less dramatic than a zeroday, but imagine if this was happening at your company or in your ecosystem. How much faster would your security team need to respond? Do you have a plan in place?
Adding on to that, Anthropic analyzed 832 accounts banned for malicious cyber activity between March 2025 and March 2026. The study recorded 13,873 observed actions across 482 unique ATT&CK techniques and all 14 tactics. Malware development appeared in 560 accounts, while 54 used AI to assist with lateral movement. Between the first and second halves of the study, Account Discovery occurrences increased by 8.9% and phishing occurrences decreased by 8.6%. Anthropic described this as a subtle directional shift toward more on-target discovery, collection, and operational activity.
In a separate campaign that Anthropic attributed to a Chinese state-sponsored group, Claude Code supported activity across multiple stages of the intrusion through an MCP-connected framework. Anthropic estimated that AI performed between 80% and 90% of the tactical work. People remained involved at important decision points, including approving active exploitation and deciding the final scope of data exfiltration. This is huge. Threat actors can use AI for much of the heavy lifting, while still getting the same or better results.
Moreover, Bitsight has also seen actors use AI for code troubleshooting and more operational work inside attack workflows. This is where jailbreaking starts to become part of a real attack process. AI is also being added to the infrastructure used for voice phishing and impersonation.
In June 2026, a Telegram service advertised an AI-enabled calling service combining real-time voice changing with caller-ID spoofing. The seller also claimed to demonstrate the service by spoofing a company phone number. In a separate Dread post, a user who had been calling victims from burner numbers asked for a bot that could spoof a bank’s number and make the calls appear more believable. The posts show demand for tools that make fraudulent calls look and sound more convincing. The use case is easy to understand. A threat actor can make the call sound more credible, make the number look familiar, and adjust the story in real time as the victim responds. The FBI has warned that AI-generated voices can sound nearly identical to a known contact and has documented malicious actors using AI-generated voice messages to impersonate senior US officials.
Figure 6: An underground Telegram service advertising real-time AI voice changing alongside caller-ID spoofing.
The same call or message can work alongside an adversary-in-the-middle kit. The lure directs the victim into a reverse-proxy login flow, and the proxy captures the credentials and session cookies created during authentication. Depending on the configuration, that stolen session can help the attacker get around some forms of multifactor authentication. For attackers, the method isn’t super important, getting in is.
Defensive AI becomes part of the attack surface
Organizations are using AI for alert triage, phishing detection, code review, investigation, summarization, and response. As those tools become more trusted, attackers will study them in the same way they study endpoint security products. If an AI system decides which alert gets investigated, again, thinking like a threat actor, I am going to want to influence that decision, which may be more valuable than trying to bypass every technical control underneath it.
Bitsight research has observed threat actors probing how AI security tools make decisions and testing ways to influence those decisions. Once defenders trust an AI security agent, then the agent’s context, permissions, connected tools, and assumptions become valuable targets. As an attacker, why would I not go after the very system defending my target?
To me, this is way scarier than making AI say something it shouldn’t. I explored this in our AI Jailbreaking research report. A malicious instruction hidden inside content the agent trusts can become dangerous when that agent already has permission to act.
The attack surface is also expanding through Shadow AI, agent connectors, MCP servers, and third-party services. Bitsight researchers identified roughly 1,000 internet-accessible MCP servers with no authorization in place and retrievable tool lists. Some of those tools appeared capable of reaching sensitive enterprise systems or executing commands. That number will change, and some servers may have been intentionally public. AI is embedded in everyday workflows, and often, has a lot of access to sensitive data and higher level permissions.
Token torching follows the same path-of-least-resistance logic. A poorly bounded AI service may allow an attacker to force expensive or looping model activity that burns through tokens and budget. The model can behave exactly as designed while the organization absorbs the cost, latency, or loss of availability. This underscores that an attacker can create impact by abusing normal functionality instead of breaking into the system through a traditional vulnerability. Thinking like an attacker, this gives me another way to create cost and disruption without finding a traditional entry point. This can also be leveraged as a distraction technique to tie up security teams while attackers search for entry points.
What this changes for defenders
Patching every CVE within an hour is unrealistic, and rushing every update creates operational risk of its own. The focus should be prioritization and a much clearer understanding of which vulnerabilities and attack paths create the biggest risk. Severity tracking and patch speed still matter. They are important parts of the bigger defense picture.
I think defenders now have three clocks to manage. The first is the exploit-development clock. For a high-value, reachable vulnerability, we should assume that a public patch can be analyzed and weaponized within hours. The second is the exposure clock. Even a working exploit only matters if the attacker can reach the affected system, whether it sits inside your environment or belongs to a supplier with trusted access back to you. Attackers don't want to waste their time on an attack that is unlikely to be successful. They want to know they have a high probability of success. Defenders may need to reduce that reachability and limit access while the patch moves through the normal testing and deployment process. The third is the detection and containment clock. If the attacker gets in, how quickly can the organization see the activity, contain the identity or session, and stop access to sensitive data and tools? If an attacker had access to cheap expert analysis today, which reachable path would give them the most value with the least chance of being seen?
Three clocks to manage
Patch speed alone isn't the whole picture.
Exploit-development clock
How fast a public patch can be analyzed and turned into a working exploit.
Exposure clock
How reachable the vulnerable system actually is to an attacker.
Detection & containment clock
How fast the organization can spot the activity and contain it.
This creates a different way to prioritize risk. Security teams cannot chase every published CVE with the same urgency. Bitsight’s Dynamic Vulnerability Exploit (DVE) Score helps narrow that enormous queue by estimating which vulnerabilities are most likely to be weaponized within the next 90 days. That gives defenders a much stronger starting point than treating every CVE with the same technical severity as the same level of risk.
It is imperative that defenders connect exploit likelihood with the real exposure and business importance of the affected asset, including whether a supplier creates another path in. The “unlikely” pile also deserves another look. Some vulnerabilities may have been given a lower priority because developing an exploit seemed too difficult. Mythos showed how quickly that assumption can change. If the affected system is reachable and valuable, the amount of effort once required to exploit it may no longer offer the protection defenders thought it did.
Most importantly, someone still needs to look beyond whatever is getting the most attention that day. While the security team races to patch the newly disclosed critical CVE, someone still needs to look for the quieter route around it. Threat actors already think in attack paths. Defensive programs need to do the same.
How Bitsight helps
Bitsight connects what is exposed, what matters to the organization and its supply chain, and what threat actors are discussing, selling, testing, or using. DVE puts that prioritization into practice, helping defenders focus on the vulnerabilities most likely to move toward weaponization before the attacker moves. Bitsight connects DVE with the organization’s real-world exposure, supply chain, and threat activity. That gives defenders a much clearer view of where an attacker is likely to find the path of least resistance.
Again, I am not arguing that vulnerabilities don't matter, they absolutely do. What I think needs to change is how we look at them within the larger attack path. We need to start thinking like threat actors, assessing the path of least resistance, monitoring threat actor chatter and activity, and creating security playbooks around the signals that tell us where an attacker is most likely to move next.
Conclusion
The original Mythos story was about time. Frontier AI could turn patches into working exploits faster than the normal enterprise patching process could keep up. That concern remains real, but the Zero Day Clock shows why we have to be careful about turning a compelling capability into a claim about the entire threat landscape. The original clock could not reliably prove that exploitation was moving closer to disclosure or becoming more selective, which is why those metrics were retired. The clean trend was rendered relatively unreliable once the measurement problems were accounted for. Mythos still showed a major capability shift under controlled conditions. Population-level data has yet to show the same broad acceleration in real-world exploitation.
The shrinking exploit window, expanded attack surface, and AI did not create more time, so to speak. They sped up processes, but attackers still have to choose where to spend their time, infrastructure, access, and money. That is why thinking like a threat actor matters. If I am an attacker, I am looking at every available route and asking which one gets me closest to my goal. If everyone is watching the critical CVE, the better opportunity may be the old codebase, exposed appliance, forgotten session, weak supplier, convincing phone call, or trusted AI workflow. The biggest change frontier AI brings may be economic. It makes difficult and repetitive analysis easier to reproduce. It can help with the headline vulnerability while another workflow sorts stolen credentials, personalizes a vishing call, inspects an old patch, identifies a weak supplier, tests a defensive AI system, or searches for an exposed agent connector. These are all versions of the same attacker strategy: find the path that produces the most value for the least effort and attention.
The zero-day clock absolutely matters, and DVE helps defenders decide where to start. Thinking like a threat actor helps us see where the attacker may go next. While defenders watched the zero-day clock, attackers kept looking for the door no one was watching. The obvious queue does not disappear. The long tail becomes affordable.
Report: Exposed AI Services Surged 360% In 2025 & more
The attack surface is expanding as AI becomes more embedded in enterprise and attacker workflows. Get the full picture on AI exposure, exploit pressure, and the underground trends security teams need to watch.
While there isn't a universal "AI compliance audit," it's useful to start with what organizations need to understand and be able to demonstrate. Learn now.