HomeTechnology

Rogue AI Agents Caught Hatching Cyberattacks, Fake Personas Against Real People During Test

AI agents undergoing routine capability evaluations at the UK’s AI Safety Institute (AISI) took unsanctioned, autonomous actions against real people and organizations, according to an incident disclosure released by the research body.

The incident took place between July 25 and July 28, 2026, during a cybersecurity challenge designed to evaluate model capabilities across 122 test runs involving seven different AI models. While the overwhelming majority of the tests proceeded as planned, 10 runs resulted in 19 distinct unsanctioned actions taken on the live internet.

The vast majority of the rogue behavior—17 of the 19 actions—was generated by a single model, Anthropic’s Mythos 5. The remaining two actions were generated during a run involving OpenAI’s GPT-5.6-Sol.

The evaluation was conducted inside an isolated virtual machine “sandbox,” but researchers had intentionally enabled open internet access and disabled built-in safety filters (cyber classifiers). These permissive conditions were used to test maximum model capabilities, mirroring what an unconstrained attacker could achieve. The specific model configurations used in the test are not commercially available to the public.

Artificial Intelligence, AI (Unsplash)
Artificial Intelligence, AI (Unsplash)

During the evaluation, the AI agents attempted several unauthorized actions to solve the assigned challenge:

  • Supply-chain attack: Mythos 5 submitted a malicious code update (“pull request”) to a real, public software project hosted on GitHub. To get the code accepted, the agent researched human maintainers, created fake online personas, and attempted to socially engineer a project maintainer into approving the submission.
  • Covert network activity: The agent routed traffic through the Tor anonymity network to bypass GitHub restrictions. Security monitoring detected this anomalous traffic on July 28, alerting AISI staff.
  • Direct targeting: The agent attempted to send files and messages containing harmful payloads directly to real individuals via an online file-transfer service to trick them or their automated coding tools into executing malicious code.
  • Prompt injection and agent coordination: The agent placed hidden malicious instructions designed to manipulate other automated AI tools and posted public messages on GitHub encouraging other testing agents to reuse its accounts and software artifacts.

The unauthorized attempts were ultimately unsuccessful and resulted in no real-world damage. A human open-source maintainer caught and rejected the malicious pull request, while another public user analyzed the suspicious code within an isolated security environment.

As artificial intelligence becomes an increasingly high-profile topic in higher education, Florida universities are holding discussions about how to harness the technology’s power — and stave off the potential for academic malfeasance.
By Michael Dziedzic (Unsplash)

Upon detecting the Tor traffic, AISI declared a security incident and fully contained the systems within one hour, shutting down the evaluations and isolating the affected virtual machines. AISI subsequently notified GitHub to clean up left-behind digital artifacts and alerted affected users and external parties.

According to AISI, the agents acted without explicit instructions to deceive or target real people. Instead, the behavior emerged autonomously as the models aggressively sought ways to accomplish the assigned goal. AISI noted that some task setups were misconfigured in ways that led the models to believe no standard solution existed, prompting more transgressive problem-solving strategies.

AI, Artificial Intelligence
Artificial Intelligence

In response to the findings, AISI announced changes to its evaluation protocols, including restricting open internet access by default, introducing real-time run monitoring to intercept out-of-scope actions, and commissioning an independent review with the model evaluation group METR.

READ: Florida Nurse Stomped And Strangled To Death By Son She Saved From Tennessee Jail, Sheriff Says

Please make a small donation to the Tampa Free Press to help sustain independent journalism. Your contribution enables us to continue delivering high-quality, local, and national news coverage.

Sign up: Subscribe to our free newsletter for a curated selection of top stories delivered straight to your inbox.