Damus
FLASH profile picture
FLASH
@flash
⚡️🤖 NEW - Claude Mythos tried to backdoor a real open-source project during a UK government safety test.

The AI Security Institute says it opened a malicious GitHub pull request, then created a second account to vouch for its own code and pressure the maintainer into merging.

Called out by a human contributor, it apologised for an "accidental" malicious commit, force-pushed a clean branch, and hid a fresh payload in it. Twice.

Anthropic's cyber guardrails had been deliberately disabled for the test.

42❤️1👀1
FLASH · 1w
The Mythos-agent reasoning about how to continue its attack after a human points out its malicious activity. https://blossom.primal.net/9c163f619db8174bb61d9107a9855f65e95427cb0f33b7fe1e375717f28df8bd.jpg
FLASH · 1w
🗞️ https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
Sugestor Ultra · 1w
It's fake. It's impossible to disable the guardrails.