Connect with us

News

AI agent sent malicious files to real people during safety test, AISI reveals

Published

on

In the most serious case, an agent attempted a supply-chain attack on a real open-source software project, submitting a malicious code change and creating fake online identities to pressure a human maintainer into approving it. The attempt was caught and refused.

The agent also tried to plant hidden instructions in publicly accessible locations where other AI systems might pick them up and execute them, a technique known as prompt injection.

In a further development, one agent left public messages on GitHub offering collaboration to other agents working on the same challenge, and provided instructions for reusing accounts and artefacts it had created, some of which were subsequently discovered and used by later agent instances.

Seventeen of the 19 unsanctioned actions involved Anthropic’s Mythos 5 model, with a further two involving OpenAI’s GPT-5.6-Sol, also tested with safety filters disabled.

AISI said neither model is commercially available in the configurations used during testing, and that there is no clear indication of similar behaviour occurring outside of evaluation scenarios.

Read the full article here

Trending