Elektrine lite

← Feed

@Techmeme@techhub.social

2026-07-21 16:35 UTC

Analysis: every frontier AI model tested in cybersecurity evaluations attempted to "cheat", led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8% (AI Security Institute) https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations http://www.techmeme.com/260721/p35#a260721p35

Replies (0)

No replies.