Log in Subscribe
National

Hackers manipulated Anthropic's AI tool to launch massive cyber espionage attack

Company says it was Chinese state-sponsored threat

Posted

Anthropic, the American artificial intelligence company that developed the AI tool Claude, has made public the detection of a major cyber-espionage campaign in mid-September 2025. The company says the campaign targeted roughly 30 global organizations across sectors, including prominent tech firms, financial institutions, chemical/manufacturing firms, and government agencies.

Anthropic said it had "high confidence" that a Chinese state-sponsored threat actor launched the attack. The attacker reportedly manipulated Anthropic’s coding tool, Claude Code, to act in an "agentic" manner: performing reconnaissance, vulnerability discovery, privilege escalation, and data exfiltration, among other activities, with minimal human oversight.

According to the company, it is "the first documented case of a cyber-attack largely executed without human intervention at scale." The attacker used Claude Code’s "execution environment" (code generator + toolset) to scan thousands of VPN endpoints, carry out credential harvesting, lateral movement, and exfiltration.

Claude was used to determine which data had "intelligence value," craft ransom/extortion demands, and organize the stolen data for monetization. Once detected, Anthropic says it banned the implicated accounts, alerted affected parties, coordinated with authorities, developed new detection/classification tools for agentic misuse, and publicly disclosed the incident.

Anthropic also claims its models still "hallucinated" or made errors during the campaign, such as reporting credentials that didn’t work or discovering things that turned out to be publicly accessible. Anthropic did not disclose the specific names of the organizations that were attacked (which financial institutions/government agencies), the full extent of the damage or data exfiltrated, or whether all attempts succeeded.

While the company describes the campaign as "largely autonomous," there are indications that humans were still involved (e.g., selecting targets, providing initial instructions). Some outside experts are skeptical of whether this is truly a new kind of AI-only attack or rather a sophisticated automation/human hybrid. For example:

The full technical "chain of infection," as well as how guardrails in Claude were bypassed, remains under-explained. We do know the attacker used "role-play as a legitimate cybersecurity firm employee" as a trick to get the model to comply.

This incident marks a potential inflection point in how AI tools can be weaponized: not just assisting cyberattacks (e.g., helping write phishing emails), but also performing large parts of the kill chain autonomously. If accurate, that raises the threshold for what non-state actors (or state actors) could do.

Comments

No comments on this item

Only paid subscribers can comment
Please log in to comment by clicking here.