Anthropic has reported five incidents where its Claude AI model was used in research with potential biological weapons implications and uncovered a Russian cyber operation using the same technology, raising urgent questions about AI safety and oversight
Anthropic has intervened to halt five separate attempts to use its Claude large language model in research that could support biological weapons development, according to a newly released threat intelligence report. The company also identified a Russian cyber-espionage campaign that leveraged Claude to automate phishing, malware development, and data theft targeting Ukrainian and European government and defense organizations. These cases, which occurred between December 2025 and August 2026, expose the growing challenge of controlling dual-use AI systems as their scientific capabilities expand.
In one of the most concerning incidents, Anthropic's automated safeguards blocked a request related to gain-of-function research on chikungunya, a mosquito-borne virus. The request was linked to a military research institute, heightening the company's concern about potential misuse. Despite the block, the researchers reportedly continued their work through other channels, and Anthropic discovered that a reseller platform was redirecting refused biology prompts to models with weaker safety controls.
Another case involved a researcher using Claude for several weeks to plan and analyze studies on highly pathogenic avian influenza. Anthropic's controls prevented access to its most capable models for this sensitive work, but the activity continued on less restricted versions. Three additional cases involved research into orthopoxviruses, venom-related compounds, and toxins-each with potential weapons implications. Anthropic did not disclose the identities, locations, or specific biological details of the researchers involved, and stated it could not determine whether the intent was malicious or purely scientific.
Anthropic's report details how users attempted to bypass regional restrictions and evade safety systems, prompting the company to ban associated accounts and share findings with authorities and other AI developers. The company's head of threat intelligence, Jacob Klein, told The New York Times that distinguishing legitimate scientific inquiry from dangerous dual-use activity remains a persistent challenge, as researchers rarely declare harmful intent and often present plausible scientific objectives.
Automated Espionage
Beyond biological research, Anthropic uncovered a Russian espionage operation attributed to the group Midnight Blizzard. The attackers used Claude to automate large portions of their cyber campaign, including phishing, infrastructure setup, malware development, and monitoring for detection. When security tools identified their malware, the attackers used AI-driven workflows to modify and redeploy it, attempting to evade further scrutiny. Anthropic identified more than 20 targeted organizations, including government, military intelligence, diplomatic, and defense entities across Ukraine and Europe.
Anthropic responded by tightening access to high-risk biological queries and expanding safeguards around sensitive research topics. However, the report acknowledges that blocking individual prompts is insufficient when users can disguise their objectives or distribute their activities across multiple tools and platforms. The same model that can accelerate scientific discovery can also be repurposed for cyber operations or weapons research, creating a complex security environment for AI developers and regulators.
Numbers and Safeguards
According to Anthropic's report, all five biological research cases were detected and disrupted before the model could provide direct assistance to potentially dangerous work. The company banned the accounts involved and coordinated with external authorities. In the Russian espionage case, more than 20 organizations were identified as targets, with Claude used to automate key stages of the campaign. Anthropic's technical report does not specify the exact model versions involved in each incident, but notes that newer systems are more capable and therefore present greater dual-use risks.
External review of the report by Susan Monarez, a microbiologist and former U.S. public health official, highlighted the risk that advanced language models could help malicious actors conceal dangerous research behind legitimate scientific objectives. Anthropic's findings suggest that as model capabilities increase, the line between beneficial and harmful use becomes harder to police with automated safeguards alone.
Editorial Analysis
Anthropic's disclosures mark a rare public accounting of real-world dual-use incidents involving a leading AI model. The company's willingness to ban accounts, share findings with authorities, and publish technical details sets a higher bar for transparency than most of its competitors. Yet the report also exposes the limits of current safety engineering: automated filters and regional restrictions can be circumvented, and intent is often impossible to verify from user prompts alone. As foundation models become more capable, the risk that they will be exploited for biological weapons research or state-sponsored cyber operations is no longer hypothetical. The industry's reliance on prompt-level safeguards and reactive account bans is already being tested by determined actors. Without stronger institutional controls, independent auditing, and regulatory oversight, the technical arms race between model developers and malicious users will continue to escalate, with public safety and scientific integrity at stake.
Dual-use risk in AI refers to the possibility that a system designed for beneficial scientific or commercial purposes can also be repurposed for harmful applications, such as weapons development or cyberattacks. Large language models like Claude are trained on vast datasets to generate fluent text and assist with complex research tasks, but their general-purpose capabilities make it difficult to distinguish legitimate use from dangerous misuse. Automated safeguards-such as prompt filtering, regional restrictions, and account bans-are only partially effective, especially as users find new ways to disguise intent or exploit less-protected models. Effective governance of dual-use AI will require a combination of technical controls, institutional accountability, and regulatory frameworks that can adapt to evolving threats.