Elektrine lite

← Feed

@wwhitlow@indieweb.social

Post #736966

2026-03-10 03:16 UTC

Regardless of expectations for AI systems, interpretability studies seem to be promising towards the future of understanding the mathematical associations of concepts. Breaking down the model to its most atomic representation. This paper explores some of the associations regarding trust. It would be interesting to see if there's a correlation between the embedding of human trust models, the persona vectors of various models, and the ability to jailbreak. https://arxiv.org/abs/2603.05839 #AI #LLM

Replies (0)

No replies.