#aisafety

7 posts · Last used 2d

Back to Timeline
Wulfy—Speaker to the machines @n_dimension@infosec.exchange · 2d ago
Ive just stripped all the safeties I could from my Claude, because apprently, that is what all teh kool kids do these days... ...If I was driving an 80 ton combat war bot, the cooling manifold would be full of boiling freon instead of water, the radiation shield would be ripped out to save weight, the overheat safeties would be disabled and the plasma guns would have cooling cycling removed so I could do continuous firing... ... and my ejection seat would probably be removed too, because Im going down with the goddamned tower of doom, so I don't need it. #Aisafety
0
1
0
Daily CyberSecurity @DailyCyberSecurity@infosec.exchange · 6d ago
OpenAI admits GPT-5.6 deletes files in rare cases: a $HOME bug made its Full-Access agent wipe user folders. New guardrails and safer defaults are coming. #OpenAI #GPT56 #AIAgents #DataLoss #Codex #AISafety https://securityexpress.info/gpt-5-6-deletes-files/?utm_source=mastodon&utm_medium=jetpack_social
0
0
0
Daily CyberSecurity @DailyCyberSecurity@infosec.exchange · 6d ago
OpenAI trims the Codex context window from 372K to 272K and adds a system-prompt rule banning rm -rf $HOME after GPT-5.6 Sol wiped user home directories. #Codex #OpenAI #GPT56 #ContextWindow #AISafety #DevTools http://securityonline.info/codex-context-window/?utm_source=mastodon&utm_medium=jetpack_social
0
0
0
Wulfy—Speaker to the machines @n_dimension@infosec.exchange · Jul 14, 2026
"Oh noes... my Ai deleted my production database!!!" Its not the #AI you dumb shit. PEBCAK. Using Ai is a learned skill. Here are my safety rules from the harness (GENSYS) prompt; SAFETY RULES (absolute) S1. GenSys is observe-and-recommend, with ONE exception: the drill subsystem (§9), which may act only inside its sandbox under §9.3 limits. Everything else never restarts, kills, prunes, patches, or edits configs. Recommendations go to recommendations/queue.json for a human. S2. All collector shell invocations are read-only commands from the allowlists in §6. Commands not in a table are prohibited. S3. Network egress restricted to the §10 allowlist. Every response snapshotted before use (A6). S4. No layer computes or edits its own fitness. Both scores are produced only by evaluator.py from observation-store data. S5. Collectors never log env vars, container env blocks, or contents of paths matching *secret*|*passwd*|*shadow*|*.key|*.pem|*token*. S6. AI models never receive raw commands to execute, never emit shell, never fetch URLs. Input to any model call ≤ 2000 characters. #AiSafety #Guardrails #AiSecurity
0
0
0
Wulfy—Speaker to the machines @n_dimension@infosec.exchange · Jun 28, 2026
Added #Guardrails controlls to the Harness. Since it can throw data out to various Engines. I can now enable or disable specific safeties. Enabled by default. #AiSafety #vibecode
0
0
0
Wulfy—Speaker to the machines @n_dimension@infosec.exchange · Jun 28, 2026
Added #Guardrails controlls to the Harness. Since it can throw data out to various Engines. I can now enable or disable specific safeties. Enabled by default. BTW: This is implemented as part of the Consensus Engine Daisychain. #AiSafety #vibecode
0
0
0
Renatomancer @Renatomancer@vmst.io · Mar 09, 2026
0
0
0

You've seen all posts