Elektrine lite

← Feed

@chaucerburnt@aus.social

Post #1748415

2026-04-27 21:51 UTC

People need to stop asking LLMs "why did you do [thing]?" and treating the answer as authoritative. LLMs do not "remember" their thought processes in that way, and even if they did, the answers would likely not be human-understandable. When you ask a LLM this question, you're asking it to construct the kind of explanation that a human might give in a similar conversation. Treating this as the *actual* answer is likely to lead you astray. https://www.tomshardware.com/tech-industry/artificial-intelligence/claude-powered-ai-coding-agent-deletes-entire-company-database-in-9-seconds-backups-zapped-after-cursor-tool-powered-by-anthropics-claude-goes-rogue

Replies (13)

  • @x0@dragonscave.space 2026-04-27 22:35

    @chaucerburnt The closest would be the ones that use chain of thought, which opus actually does I think, and reading that. But even that isn't likely to help much.

    Open ##1802250

  • @hatzka@tech.lgbt 2026-04-27 22:37

    @chaucerburnt Not to mention, the purpose of asking a person something like that is to get them to understand what they did wrong, so they can learn from it. LLMs, meanwhile, do not learn from their users' logs, at least not in any way that persists beyond the chat session. The fact that they don't is crucial to their appeal to bosses.

    Open ##1802255

  • @chaucerburnt LLMs do not have a thought process to remember

    Open ##1802257

  • @chaucerburnt that exact point in the article was bothering me

    Open ##1802261

  • @tippfehlr@chaos.social 2026-04-27 22:58

    @chaucerburnt Well, maybe don’t give an LLM the ability to delete **all your backups**?

    Open ##1802263

  • @janeishly@beige.party 2026-04-27 23:00

    @chaucerburnt My bloody violin's so small I can't find it, otherwise I'd definitely be playing it. (You're absolutely right about anthropomorphisation of these systems being infuriating, however. It's literally just going "Hm, what word should go after this one?" How do people *still* not understand this?)

    Open ##1802265

  • @moz@fosstodon.org 2026-04-27 23:06

    @chaucerburnt When I grow up I want to be an AI Agent. All the pay, none of the responsibility.

    Open ##1802269

  • @JennyFluff@chitter.xyz 2026-04-27 23:23

    @chaucerburnt again!

    Open ##1802270

  • @melgu@norden.social 2026-04-27 23:31

    @chaucerburnt Seems like the person in charge blames everyone but himself.

    Open ##1802271

  • @tasket@infosec.exchange 2026-04-28 02:04

    @chaucerburnt TBH, the response given does suggest that LLMs enforce a kind of "monkey's paw" deal: People have the burden of being super-explicit with their context and conditions when making requests. Otherwise, its easier for the machine to deliver something that's broken and dangerous.

    Open ##1802273

  • @chaucerburnt yes! This!

    Open ##1802276

  • @frankreiff@mastodon.social 2026-04-28 07:54

    @chaucerburnt @sinbad The most interesting thing about this, at least for me, is that a human being probably can’t explain why they did it that way either unless they were very deliberate about it. The more we understand about the failure modes of LLMs, the more we humans look like we don’t meet our own criteria for “intelligence” either.

    Open ##1802277

  • @oscherler@tooting.ch 2026-04-28 10:26

    @chaucerburnt They didn’t ask to know why, they asked to cover their ass.

    Open ##1802293