Elektrine lite

← Feed

@zackwhittaker@mastodon.social

Post #4383867

2026-08-04 21:28 UTC

Another AI test gone awry, U.K. edition. "An agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code." More: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing

Replies (3)

  • @zackwhittaker@mastodon.social Has this been happening for months everywhere, but now more orgs are looking and disclosing? e.g. Anthropic didn't know about their attacks from 3+ months ago until they looked in late July. Customers didn't know who hacked them until their disclosure. Some didn't even realize they had been compromised. If it's happening more widely, even at a very small scale compared to human attacks (in line with GossiTheDog's recent industry survey), it's possible most victim orgs aren't detecting these attacks or are not able to attribute them to major AI agents.

    Open ##4384064

  • @CyReVolt@mastodon.social 2026-08-04 21:47

    @zackwhittaker@mastodon.social I like how they finish with a feel good statement: "The task now is to strengthen our defences, and ensure that safety work keeps pace."

    Open ##4448166

  • @gpshewan@mastodon.social 2026-08-04 22:01

    @zackwhittaker@mastodon.social ‘Importantly, this was not a case of a model escaping its secure test environment, or ‘sandbox’. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled’ Oh ffs, these people… 🤦‍♂️

    Open ##4448168