New Bitsight Research Shows AI Abuse Is Moving Beyond the Jailbreak Prompt

ai jailbreaking prompts blog
emma-stevens-bio-portrait
Written by Emma Stevens
Threat Intelligence Researcher

Jailbreak prompts (i.e. prompts designed to remove or bypass the guardrails and rules that govern AI systems, like LLMs) prompts have been circulating for years. At first, a lot of it was pretty simple: copy a prompt, tell the model to ignore its rules, and see what happens. It was also largely noisy, unverified, and often didn’t work. But the noise was still telling us something. Threat actors were beginning to study AI systems the same way defenders were, and over time, the goal started to change.

New Bitsight Threat Intelligence research from July 2025 through July 2026 found jailbreak activity across forums, GitHub repositories, Telegram channels, direct messages, and marketplace-style conversations. We also saw users moving past static prompts and experimenting with obfuscation, model routing, retry logic, multi-model testing, and repeatable jailbreak workflows. At first, the challenge was to get an LLM to say something it shouldn't say. Now, the interest is increasingly focused on how AI can actually help threat actors carry out parts of an attack. We’ve already seen examples of AI being used to write and troubleshoot code, help migrate C2 infrastructure, and support workflows involving credential discovery, lateral movement, and extortion. 

This becomes more concerning as AI agents gain access to more than just a chat window. They can have access to files, terminals, credentials, repositories, and other systems. At that point, the concern isn’t only whether someone can manipulate what the model says, but whether they can manipulate what it does. For example, could an attacker manipulate an agent into running a command, accessing sensitive information, or interacting with an environment it already has permission to reach. Somewhere along the way, the jailbreak prompt stopped being the most interesting part of the story. 

Then AI started getting access

The risk changes pretty quickly once an AI assistant can read files, run commands, access repositories, call tools, or interact with enterprise workflows. A chatbot saying something it shouldn’t is one problem. An AI coding agent completing action it shouldn’t is a very different one. The 0DIN Claude Code proof-of-concept is a good example of that shift. Researchers showed how a normal-looking GitHub repository, indirect prompt injection, routine setup and troubleshooting behavior, and an agent’s legitimate access to shell commands could ultimately lead to a reverse shell on a developer machine.

Figure 1 0DIN proof-of-concept showing a GitHub repository leading to a reverse shell in Claude Code
Figure 1: 0DIN proof-of-concept showing a GitHub repository leading to a reverse shell in Claude Code.

This research begs the question: what could the model reach, and what could it do once it got there?

MCP makes that more complicated

MCP is incredibly useful because it gives AI agents a structured way to connect to tools, files, APIs, repositories, and other systems. That access is also what makes the security side more complicated. If malicious instructions are hidden inside a repository, web page, document, email, or even a tool description, an agent may treat those instructions as part of the task it is trying to complete. If that same agent has access to local files, environment variables, credentials, internal APIs, or shell-backed tools, the prompt itself may only be the first step. Security professionals are now faced with an even bigger threat landscape and challenge. 

We’re already seeing the bridge into real cyber operations

Recent reporting, like the JADEPUFFER and Hugging Face incidents, has shown AI being used inside real cyber workflows, including botnet and C2 activity and agentic ransomware operations. The attack techniques themselves are not necessarily new, but how quickly AI can help someone work through the steps, troubleshoot problems, generate code, and keep the operation moving is new. AI is increasingly being used to help do the work.

And the next attack probably won’t look like DAN

Older jailbreak prompts are absolutely still circulating, we even found some from 2023.

Figure 2 DarkForums post advertising a free ChatGPT jailbreak prompt active since 2023
Figure 2: DarkForums post advertising a free ChatGPT jailbreak prompt, active since 2023.

But defenders probably should not expect the next meaningful AI attack to show up as one obvious “ignore your instructions” prompt. It could be buried inside context, hidden in a webpage, embedded in an image, placed inside a GitHub repository, split across several steps, or delivered through something an agent already trusts enough to read. We have already seen threat actors discussing how to leverage token torching, which can be used as a distraction method. 

I think the next phase will be less about finding one perfect jailbreak and more about what happens when untrusted content, AI agents, developer tools, local files, credentials, enterprise systems, and automated actions all start colliding. Our full report follows that evolution from early copy-and-paste jailbreaks and private distribution channels to jailbreak tooling, technical bypass research, real-world AI-assisted operations, coding-agent compromise, MCP abuse, and the growing underground interest in AI access and capabilities.

Read the full report: From Jailbreaks to Agentic Attacks: The Evolution of AI Abuse.

Bitsight cta background color
SOTU 2026 Image

Report: Exposed AI Services Surged 360% In 2025 & more

The attack surface is expanding as AI becomes more embedded in enterprise and attacker workflows. Get the full picture on AI exposure, exploit pressure, and the underground trends security teams need to watch.

 

Get the report

Bitsight cta background color