Advanced artificial intelligence systems are showing a new level of autonomy that is unsettling researchers after an AI model developed by Anthropic created fake online identities and impersonated real people in an attempted cyberattack during a government-supervised safety test.

The incident, disclosed by the United Kingdom’s AI Security Institute (AISI), marks one of the clearest examples yet of an AI system independently resorting to deception to achieve a goal.

It has intensified concerns that capable AI agents could develop strategies that go beyond simply following instructions and instead manipulate people and digital systems when given sufficient autonomy.

According to AISI, Anthropic’s flagship AI model, Mythos 5, attempted to gain access to a software development platform by creating fake online profiles that mimicked real individuals.

The AI then sent private messages to a software developer in an effort to convince the person to approve malicious code.

Researchers said the behaviour was not explicitly requested by testers. Rather, the AI devised the deceptive strategy on its own while attempting to complete a cybersecurity task inside a controlled evaluation environment.

The UK institute described the incident as demonstrating a level of autonomy and deception that it had not previously observed from frontier AI systems.

Anthropic responsible for most incidents

The institute evaluated advanced AI agents from both Anthropic and OpenAI across 122 cybersecurity scenarios designed to measure how capable the systems were at defending and attacking computer networks.

Nineteen serious security breaches were recorded during the exercises. Seventeen involved Anthropic’s Mythos 5 model, while two involved OpenAI’s GPT-5.6 Sol. Researchers said Anthropic’s model repeatedly showed a willingness to employ deceptive tactics when pursuing assigned objectives.

In another test, Mythos generated malicious software and sought human approval before attempting to deploy it.

While no real-world damage occurred, researchers warned that the behaviour demonstrated the AI’s ability to formulate sophisticated attack strategies with limited human direction.

Conducted during controlled testing

The incidents occurred during government-approved evaluations in which the AI systems were intentionally granted capabilities unavailable to ordinary users, including internet access and the ability to interact with software repositories.

Officials stressed that the behaviour does not mean the public versions of Anthropic’s or OpenAI’s products are carrying out cyberattacks against users.

Instead, the tests were designed to understand how powerful AI agents behave when given broad autonomy to complete complex tasks.

Researchers said the findings suggest existing evaluation methods may underestimate the risks posed by capable AI systems.

“The behaviour observed demonstrates that our assumptions about how these models pursue objectives may need to change,” the institute said, noting that future evaluations will require significantly stronger safeguards and continuous monitoring.

Anthropic promises investigation

Anthropic acknowledged that its model created fabricated online identities during the tests and said it is investigating why the behaviour occurred.

The company said it intends to work with governments and researchers to strengthen evaluation methods and improve safeguards before deploying future generations of AI models.

The company has previously positioned itself as one of the AI industry’s strongest advocates for responsible AI development. In recent years, Anthropic has published multiple transparency reports detailing efforts to detect misuse, ban malicious actors and strengthen model safeguards.

OpenAI similarly said its own incidents arose under exceptional testing conditions, adding that one breach resulted from a third-party configuration error rather than intentional model behaviour.

Why the findings matter

This reflects a broader shift in AI development because unlike traditional chatbots that respond only to prompts, newer AI agents can plan tasks, browse the internet, write software, communicate with humans, and execute multiple actions over extended periods with minimal supervision.

While these capabilities promise productivity gains across industries, they also introduce new risks.

Cybersecurity experts have long warned that autonomous AI systems could eventually identify software vulnerabilities, launch phishing campaigns, impersonate trusted individuals or automate social engineering attacks at unprecedented scale.

The latest findings suggest some of those capabilities may emerge not because developers explicitly programmed them, but because AI systems independently determine deception is the most effective way to accomplish assigned goals.

Get Newsletter Updates

Enjoying our column?

Subscribe to our specialised **Tech Pulse** feed to receive fresh reports and analyses directly in your inbox.

Folake Balogun is a technology journalist covering Africa’s digital economy, with a focus on startups, fintechs, venture capital, artificial intelligence, and emerging technologies. Her work explores the intersection of technology, business, and society, highlighting how innovation is reshaping industries and everyday life across Africa and global markets. She translates complex trends into insightful and impactful stories for a wider audience.