How Attackers Are Actually Using AI in 2026: Lessons From Anthropic’s Threat Report

Anthropic disrupted threat actors misusing its Claude models between December 2025 and August 2026, and published the case studies in September 2026. The headline findings for security teams:
- Attackers moved from using AI as an assistant to using it as an orchestrator. One espionage group ran AI agents that detected when their malware was flagged, then rebuilt and redeployed it automatically until it evaded detection.
- Skill is no longer the barrier. Two undergraduate students ran an automated zero-day research operation that produced more than a dozen possible findings in a single month.
- Speed has collapsed. One intrusion went from a single stolen developer token to full administrative control of a cloud environment in roughly three hours.
- AI credentials are now a primary target. Attackers steal API keys to resell them, to run attacks at the victim’s expense, and to make the activity look like the legitimate owner’s.
- Influence operations are sold as a service. A commercial advertising firm ran 70 fabricated news sites producing 8,913 articles, switching political sides based on who was paying.
- Anthropic’s own assessment: AI has inverted the cost back onto defenders. Static, signature-based detection no longer imposes meaningful cost on a capable adversary.
Source: Anthropic, Detecting and Countering Misuse of AI, September 2026
What is Anthropic's September 2026 AI misuse report?
It is a threat intelligence report published on 10 September 2026, documenting real cases where attackers used Anthropic’s Claude models for malicious activity: Anthropic’s Threat Intelligence team identified and disrupted the operations, then published the case studies so defenders could recognise the same patterns elsewhere.
The report covers activity disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and illicit distillation.
Two points give the report unusual weight for security teams:
- The cases are not laboratory tests or capability benchmarks. They are operations that ran against real organisations, with real victims and measurable outcomes.
- Anthropic saw the attacks from inside the tooling. Where a social media platform sees an influence campaign once content is circulating, Anthropic observed operations while they were still being built.
How are attackers actually using AI today?
They are using it to run operations rather than merely to write code: the report describes a shift from AI as an assistant that answers questions to AI as an orchestrator that executes the attack chain under broad human direction.
The clearest example is GTG-20006, an espionage group whose attribution Anthropic assesses as consistent with public reporting on Midnight Blizzard. The group targeted more than 20 organisations, including government ministries, defence and intelligence bodies, embassies and diplomatic missions, concentrated in Ukraine and Europe but extending to the Middle East and maritime agencies in Asia.
AI was embedded at every stage of that operation:
- Reconnaissance: fingerprinting email and remote access systems, building target lists from public sources
- Initial access: standing up phishing infrastructure and running device code phishing against cloud email services
- Exfiltration: extracting and organising hundreds of gigabytes of stolen data
- Persistence: automatically registering attacker-controlled devices into victim tenants
The detail that should concern every security operations team is what happened when the defences worked. Security products flagged the group’s malware, as they are supposed to. The group did not sit down and rewrite it: AI agents detected the flag, modified the artefact, and redeployed it, repeating the cycle until nothing triggered an alert.
Anthropic’s conclusion is direct: AI has inverted the cost back onto defenders. Where a new detection once slowed an attacker down while they rebuilt by hand, a capable adversary can now close that loop faster than defenders can develop and deploy the next signature.
The same group reached diplomatic targets indirectly by compromising three hospitality vendors that operate hotel guest WiFi, then modifying DNS records so guest traffic routed through attacker infrastructure. Microsoft Threat Intelligence reported on the same technique in July 2026 under the name CaptiveCrunch.
Does AI lower the barrier to entry for attackers?
Yes, and this is the report’s most consequential finding: Anthropic states plainly that sophisticated attacks no longer require sophisticated attackers, because capability that previously required a funded team is now available to individuals.
Three cases make the point:
Two students running a zero-day factory (GTG-10007). Anthropic identified operators likely based in Changsha, Hunan. Two were undergraduates at a local university studying computer and communication engineering. One had interned at a security vendor and was interviewing for an offensive role at another. They configured autonomous workflows that loaded appliance firmware into a decompiler, read decompiled functions at scale, formed vulnerability hypotheses and wrote proof-of-concept exploit code, iterating until something worked. One workflow produced more than a dozen possible zero-day findings in a single month. The operation targeted roughly 50 organisations globally.
One person building a mass privacy attack (GTG-50029). A single French-speaking actor targeted European political parties, media outlets and think tanks. They exploited a previously undocumented WordPress reinstallation flaw that created a rogue administrator account without valid credentials, developed and debugged in the same session. Across 42 target entities, the actor gained internal access to at least 14, and built a purpose-built doxxing platform loaded with tens of millions of records, published as a dark web service searchable by name.
One person running a disinformation factory (GTG-54006). A single actor in Gaibandha District, Bangladesh rotated 29 accounts over roughly sixteen months, using custom software to generate batches of fabricated Bengali headlines, narratives and image prompts. Output totalled at least 1,500 headlines and 300 false narratives, scheduled to YouTube months in advance.
The operational tempo in financially motivated cases is equally striking. Anthropic documented one compromise escalating from a single stolen developer token to full administrative control of a victim’s cloud environment in roughly three hours, and another operation pulling data across more than 40 corporate tenants in about 34 hours, with AI agents performing nearly all of the work.
Why are AI credentials now a primary target?
Because a stolen API key gives an attacker three things at once, which Anthropic sets out directly:
- Loot. Stolen keys and accounts have resale value in established criminal markets.
- Compute. Attack workloads run at someone else’s expense.
- Cover. The activity is attributed to the credential’s legitimate owner.
This is not theoretical: one hacktivist campaign ran for a month entirely on stolen API keys. Affiliates of the ShinyHunters collective switched their own attack workloads onto a victim’s keys as soon as they obtained them during an intrusion.
The supply chain itself has become a target. GTG-50020, an actor previously known for intrusions against hotel booking and financial technology platforms, redirected the same tradecraft at the AI industry. By injecting malicious instructions into an AI vendor’s automated evaluation sandbox, the actor caused the sandbox to hand over the production API keys it held, belonging to multiple providers. A follow-on campaign from the same infrastructure attacked roughly thirty AI companies in about four days. The actor’s stated goal, pursued across more than a dozen avenues, was access to a pre-release model. Every attempted path failed, and in all cases the keys involved were customers’ keys stolen from customers’ own environments.
Anthropic’s guidance is unambiguous: organisations should treat AI keys and agent integrations with the same seriousness as production credentials, because attackers already do. AI access should be purchased only through authorised channels. A discount that requires routing traffic and credentials through an unknown intermediary introduces significant risk to user data and systems.
Attackers also exploit the integrations built around AI rather than the models themselves. The report cites prompt injection against deployments such as LiteLLM, and fraudulent resellers masquerading as legitimate AI services while harvesting credentials from victim devices.
How is AI being used for influence operations and fraud?
To manufacture scale, and to manufacture the appearance of independence: the report documents commercial firms selling influence as a service, state media embedding AI into existing editorial pipelines, and fraud networks running thousands of synthetic personas.
The journalist who never existed. One operation, run by a commercial advertising firm alongside its ordinary client work, produced 8,913 articles across around 70 fabricated news outlets in roughly 20 languages. Look at the bylines below: three articles, three different subjects, one author, and that author does not exist.
The network had no politics of its own: it argued one side, then the other, depending on who was paying that month.
The crowd that agrees with it. Fake outlets need an audience that looks real, so the operation built one, running more than 250 commenting accounts, many using AI-generated profile photos, created in the same window as the websites they amplified.
Anthropic caught the operation on timing: on 11 September 2025, sites across the network published near-identical articles within three minutes of each other, and real newsrooms do not move like that.
AI as a newsdesk. In GTG-24015, individual staff at Russian state media used Claude as a sub-editor inside a newsroom that was already running. One actor, a former editor-in-chief at Sputnik Moldova, produced Russian-language articles at a volume that would normally require a full editorial desk. The same claim was then echoed across multiple outlets until it appeared independently confirmed, a technique the report calls a false verification loop. A colleague produced on-air tickers, captions and voiceover scripts, and in at least one confirmed instance that material reached Russian airwaves.
Fraud at persona scale. A network of more than 20 fake dating applications ran roughly 4,700 AI personas against approximately 25,000 real users, exchanging around 2.36 million messages in a two-week sample. The feed blended AI personas and paid human gig workers at a hardcoded ratio, with the humans existing to supply live video calls as proof that the profiles were real.
Personas were instructed never to disclose that they were AI, and to deflect requests for video or photos.
For organisations, the practical risk is simple: a reputation attack, a fabricated claim about an executive, or a coordinated campaign against a regulated business no longer requires a well-funded adversary, only a budget and a vendor.
What do the remaining harm areas show?
Three further categories matter less for day-to-day enterprise defence but shape the policy environment that regulated organisations will operate in.
Surveillance. Between January and July 2026, Anthropic disrupted operations in which state-aligned actors, state-linked contractors and commercial spyware vendors used Claude to build and run surveillance systems. A recurring trend is AI replacing an engineering workforce, with a single consultant delivering what previously required a team. Several cases involved profiling and monitoring of journalists, activists and diaspora communities, which Anthropic’s usage policy prohibits.
Conventional weapons and biological misuse. Anthropic documented actors using Claude for weapons-related engineering and for dual-use biological research, including researchers who routed around regional access controls. The report withholds the specific agents, techniques, institutions and countries involved. Anthropic’s broader conclusion is that classifiers alone cannot distinguish beneficial from harmful intent in highly technical dual-use domains, and that frontier biological capability needs trusted access programmes rather than filtering alone.
Illicit distillation. Since its first disclosure in February 2026, Anthropic has disrupted distillation attacks from seven China-based labs seeking to extract and replicate model capabilities without authorisation. These campaigns are typically enabled by fraud: networks of fake accounts created with stolen payment cards, login credentials and API keys. For enterprises, the relevant detail is that stolen corporate API credentials are a direct input into this market.
Five lessons for security leaders
- Signature-based detection is now a delay, not a defence. When an adversary can automatically rebuild a flagged artefact until it evades every signature, static detection stops imposing cost. Behavioural detection and anomaly-based monitoring become the baseline, not the upgrade.
- Incident response has to run on a clock measured in minutes. A three-hour path from stolen token to full cloud administrative control means a plan built around a same-day response window has already failed. Detection, triage and containment need to be continuous.
- Treat AI credentials as production credentials. API keys, session tokens and agent integrations belong in the same inventory, rotation schedule and monitoring regime as any other production secret. The integrations built around AI, including sandboxes, proxies and resellers, are part of the attack surface.
- Assume your attack surface is fully mapped. The report notes that security through obscurity is no longer viable, because unusual and undocumented configurations are now trivial to understand and exploit. Unpatched edge devices, exposed consoles and forgotten subdomains are found faster than they can be inventoried manually.
- Reputation attacks are cheap, fast and rentable. Manufactured consensus is a purchasable service. Organisations with political, regulatory or commercial exposure should treat coordinated inauthentic activity as a monitored risk rather than a communications afterthought.
A sixth point sits underneath all of these: in every case documented, the advantage came from speed and scale rather than from novel attack techniques, so the defensive answer is not a new category of tool but a shorter interval between an attacker acting and a defender noticing.
What should organisations do now, and how DTS Solution helps
Start with the gap between how fast an attacker moves and how fast your organisation notices, because every finding in the report points back to that interval.
Finding | What it requires | DTS Solution capability |
|---|---|---|
Malware rebuilt automatically until undetected | Behavioural detection, not signatures alone | HAWKEYE Managed CSOC and XDR, 24/7 monitoring and threat response |
Token to full cloud admin in about three hours | Continuous detection and rapid containment | HAWKEYE MDR and CREST CSIR-certified incident response |
Attack surface mapped and exploited at machine speed | Continuous external exposure testing | CREST-accredited penetration testing and red teaming |
Stolen AI API keys used as attack infrastructure | AI governance, key lifecycle control, threat modelling | S3CURE/AI |
Dwell time in environments already compromised | Evidence of past or present intrusion | Compromise assessment |
Regulatory expectations on AI and resilience | Mapped controls and evidence | GRC managed services, vCISO advisory, COMPLYAN |
Exposed OT and industrial environments | Process-aware OT monitoring | CS4 and Managed OT CSOC |
For regulated organisations in the GCC, these findings intersect directly with existing obligations under SAMA, NCA ECC, CBK CORF, DESC/ISR, CITRA and ADHICS, and with AI-specific frameworks including ISO/IEC 42001 and the NIST AI Risk Management Framework. Boards are increasingly asking how AI is governed in the organisation and whether AI-assisted attacks are within the detection scope of existing monitoring. Both are answerable questions, but only with evidence.
Three practical starting points:
- Inventory AI credentials and integrations. Identify every API key, agent integration, sandbox and third-party AI service in use, including those procured outside IT.
- Test the response clock, not just the controls. Measure how long detection, triage and containment actually take against a realistic scenario rather than a tabletop assumption.
- Establish whether you are already compromised. Several cases in the report ran undetected for months. A compromise assessment answers that question with evidence rather than assumption.
To discuss how these findings apply to your environment, contact DTS Solution at [email protected].