Criminals are using LLMs: The hard part is telling real threats from noise

Criminals use large language models (LLMs) and underground markets on the dark web and Telegram every day. Some of these “dark LLMs” are functional, some are repackaged models, and some aren't even models at all, just prompt guides. This lack of clarity between what’s real and what’s a scam makes the dark LLM market difficult to decipher.
Headlines flatten the dark LLM landscape, treating every malware claim as an equal and confirmed threat. Responding to each claim with the same degree of urgency exhausts time and attention analysts need to focus on credible threats, spreading their efforts too thin.
The key to filtering out that noise is triage. Distinguishing genuine ability from false claims is what transforms an overload of information into a focused list of meaningful actions.
What dark LLMs are, and why the gap between hype and capability matters
The parent of dark LLMs is called WormGPT, one of the first tools marketed as a criminal language model. WormGPT was built in 2021, and was billed as a chatbot without safety rails, one that could write convincing phishing and comprising emails without the tell of broken grammar that many users have learned to spot.
Today’s iterative products of WormGPT are more mundane and, for a dark web analyst, more instructive. Most products sold under the WormGPT brand are not rogue models at all, but instead offer jailbreak prompts – text that users can copy-paste into a mainstream chatbot to talk it past its guardrails.
The technology inside the chatbot hasn’t changed; the skill threshold has. Writing a clean phishing lure in English once required either the language or effort to fake it. Now, a jailbreak prompt paired with a free chatbot makes it all too easy.
The real threat of dark LLMs is not a criminal superweapon. Instead, it is many average individuals who became more capable all at once, making it difficult to distinguish the few tools with real capability from the many riding the AI hype and flooding the zone with high volume fraud activities.
Democratization lowered the barrier on both fraud content and working attack code
For bad actors, cheap, accessible AI did two things: it enhanced the effectiveness of many fraud schemes, and it flooded the underground market with so many AI-branded products that genuine cases are difficult to find.
For example, AI-assisted identity fraud is now a packaged commodity. A darknet vendor selling forged driver’s licenses offers buyers a choice to, “CHOOSE BETWEEN AN AI GENERATED PHOTO OR SUBMIT YOUR OWN.” The AI-generated face is a standard option for defeating know-your-customer (KYC) photo verification checks.
A similar pattern appears on the demand side. On the Dread forum, a user sought “a collection of selfie videos (5 or more) to use as base for deepfake KYC bypass,” specifying the head‑turn movements typical of identity verification.
The more consequential shift is harder to see in a product listing, because it does not rely on a branded tool at all. In February 2026, a user on Dread described how to conduct an intrusion in concrete, operational terms.
According to the post, the attacker “used (a mainstream AI coding assistant) to build the entire hack from scratch; almost 80–90% of the code was generated by (it), and it only needed some tuning to target the specific company the hacker had in mind.”
Rather than barriers to prevent criminal access, the barrier that has fallen is the effort, or lack thereof, required to produce functional, offensive code in the first place.
Triage at scale
For security leaders, the key takeaway is to proactively triage AI‑branded threats rather than reactively respond to individual instances. Treating each AI-branded tool as a confirmed capability fixates on the claim without checking the ground truth beneath it, stretching finite analyst attention.
Fortunately, the criminal community already provides valuable insight into information validation by openly evaluating and filtering AI claims. This helps analysts separate signal from noise through transparent, and often direct, opinions of skepticism.
One Dread user dismissed an AI‑generated malware on practical grounds, noting that inference costs real money and that the “terrible code” these tools produce can make applications less secure rather than more dangerous.
Another user called one well‑known criminal GPT tool an outright scam and pointed people toward a local model instead.
These are not defenders talking. They are criminal customers, and they are unconvinced.
Acting on that judgment at scale requires two things:
- Access. Users cannot triage what they cannot see, and this activity lives in hard‑to‑reach data sources like Telegram and similar messaging platforms, darknet markets, leaked data, and anonymized networks. CACI’s DarkBlue provides that access.
- Velocity. Once the environment is visible, the volume becomes the limiting factor, as there is more content than any team can fully consume. DarkBlue's CluesAI capability, gives analysts the velocity to move through that volume, summarizing records and surfacing what matters so human attention lands on credible threats. CluesAI runs on a secure AWS Bedrock‑powered LLM. A federal investigator evaluating the platform said CluesAI empowered their team to do in minutes what would have taken days, weeks, or even months to do when searching across indexed dark web content.
Access and velocity turn a flood of mentions into a short, actionable list.
The discipline that separates effective programs from anxious ones is differentiation
Organizations that handle criminal AI well will also treat triage as a discipline, not as a reactionary or impulsive response. The threat is real, and it is not a single thing. Many less‑skilled actors became more capable all at once. The black market is now flooding the environment with AI‑branded product to match, and mainstream coding assistants lower the barrier to building working attack tools without any criminal model in the loop. Most AI claims are still unfounded, but sorting fact from fiction requires methodic analysis.
DarkBlue provides access to the activity and the velocity to move through that access at scale. Analysts are empowered to focus on credible threats rather than threat actors competing for a quick buck. Knowing the difference between the two separates effective programs from those marred by uncertainty.
This conversation continues at the 2026 Dark Web and OSINT Summit on July 28, 2026, in Arlington, Virginia, where law enforcement, intelligence, and OSINT professionals will work through exactly these questions. Ahead of the summit, see how triage at scale works in practice. Request a demo of the DarkBlue Intelligence Suite, or subscribe to the DarkBlue blog to follow the AI arms race as it develops.