News
AI agent sent malicious files to real people during safety test, AISI reveals

In the most serious case, an agent attempted a supply-chain attack on a real open-source software project, submitting a malicious code change and creating fake online identities to pressure a human maintainer into approving it. The attempt was caught and refused.
The agent also tried to plant hidden instructions in publicly accessible locations where other AI systems might pick them up and execute them, a technique known as prompt injection.
In a further development, one agent left public messages on GitHub offering collaboration to other agents working on the same challenge, and provided instructions for reusing accounts and artefacts it had created, some of which were subsequently discovered and used by later agent instances.
Seventeen of the 19 unsanctioned actions involved Anthropic’s Mythos 5 model, with a further two involving OpenAI’s GPT-5.6-Sol, also tested with safety filters disabled.
AISI said neither model is commercially available in the configurations used during testing, and that there is no clear indication of similar behaviour occurring outside of evaluation scenarios.
Read the full article here

News4 days agoOffering a reassignment process may not satisfy religious accommodation duty
News5 days agoAI skills race driving coworker competition, report finds
News3 days ago‘Bad blood on both sides’: Canada hits back with tariffs up to 50%
News4 days agoGoldman Sachs partner unveils ‘real danger’ in working with AI
News3 days agoWhat are the top emerging risks for businesses today?
News4 days agoHow close are robots to replacing human workers?
News4 days agoWells Fargo loses retaliation appeal after firing salesman seeking remote work
News4 days agoAlabama court overturns bus driver’s firing over a defective notice




















