Elektrine lite

← Feed

@daisy@cloudisland.nz

Post #4459410

2026-08-09 05:04 UTC

@glyph@mastodon.social @mcc@mastodon.social Yeah. My personal experience is that in simpler, more interactive workflows, the LLM is confidently and sanctimoniously wrong often enough that I don’t feel I can trust it and stake my professional reputation on its output without independently verifying *everything*. And that makes me distrust more “advanced” workflows that involve less direct supervision. And while I understand that “agentic” workflows include “harnesses” and “guardrails” that attempt to automate the process of verification, I don’t necessarily trust my own ability to write those, and I think that the assumption that it is possible to exhaustively enumerate upfront every possible failure case feels like an exercise in hubris.

Replies (2)

  • @daisy@cloudisland.nz 2026-08-09 05:06

    @glyph@mastodon.social @mcc@mastodon.social Also I really dislike how it turns the dynamic into an adversarial one, where you have a thing that isn’t human but which is designed to pretend that it is in a very uncanny valley way, and it is trying to sneak mistakes past you and you are trying to stop it from doing that. But I recall you’ve already written about this dynamic at length.

    Open ##4490590

  • @glyph@mastodon.social 2026-08-09 05:06

    @daisy@cloudisland.nz @mcc@mastodon.social I could _imagine_ a future where something like an LLM agentic loop might work, but we are like 7 years out on development of the precursor tools, like pervasive and intuitive sandboxes for every developer environment. i.e. (or do I mean c.f.?) https://pyvideo.org/pybay-2024/when-arbitrary-code-execution-is-working-as-intended-what-code-is-python-supposed-to-execute.html sorry for all the self-promo here, it's just easier to reference previous works than to tediously paraphrase them :)

    Open ##4490591