@andrei_chiffa@mastodon.social
Post #4061218
2026-07-24 10:17 UTC
We confirmed the model wasn't just refusing to provide dangerous information, but was unable of it, by performing abliteration ourselves (worst-case jailbreak; by Arthur Wuhrmann based on his experience at Surelio.ai).
2/
Replies (1)
-
@andrei_chiffa@mastodon.social 2026-07-24 10:18
-> We partnered with third parties to cover our own blind spots. Specifically, for offensive cybersecurity SCRT (now Orange Cyberdefense) discovered unexpected offensive capabilities and allowed us to evaluate them. Both this discovery and the fact that nobody got accidentally hacked in the process speak to the value of domain expertise, and we are looking forwards to expanding such partnerships going forwards. 3/