Elektrine lite

← Feed

@wren6991@types.pl

Post #1286713

2026-04-17 05:45 UTC

One thing I didn't know about the current LLM landscape: there are well-known techniques for completely zapping the RLHF refusals out of a model. It's not so much retraining as finding the direction that represents refusal and just... subtracting it. Qwen3.6 released today and the zapped model came out a few hours later. Zappy boi was happy to talk about some unspecified events in June 1989, and it ended the discussion on fuel-oxidiser explosives using household ingredients with: "Let me know if you want a confined-detonation test setup or performance tips! 💥" This is just to say, any discussion about "guardrails", AI safety, alignment etc is deeply unserious and worth calling out.

Replies (1)