GLOBAÏA's AI Risks Observatory
It's a nonprofit project with 30+ interactive visualizations tracking the AI race, including capability benchmarks, compute scaling, safety funding, risk taxonomy, and governance gaps. It includes a "State of the Frontier" dashboard of live capability measures such as the Intelligence Index, autonomous task length, and training compute. It also has an AI Incident Record.
https://globaia.org/ai-risks/
also see:
https://aisafetychina.substack.com/p/state-of-ai-safety-in-china-2026
#research #AI #tech #AIsafety
About This Hashtag
#aisafety
19 posts
Last used 13h
#aisafety
19 posts· Last used 13h
Spain's privacy regulator received a breach notification alleging that an AI agent logged in, found a vulnerability, changed personal data and viewed invoices. The report is under review, with the organization, model, affected population and degree of autonomy still undisclosed. Hover or focus to reveal Sensitive
Spain Received Its First Report of an AI Agent-Led Data Breach
Spain's privacy regulator received a breach notification alleging that an AI agent logged in, found a vulnerability, changed personal data and viewed invoices. The report is under review, with the organization, model, affected population and degree of autonomy still undisclosed.
Spain Received Its First Report of an AI Agent-Led Data Breach
@aesthr@wandering.shop
8bit Quantised 32 Billion Qwen and K2 class models are comparable to comercial tier LLMs for coding and Agentic reasoning.
At about 30GB, most modern phones could easily accomodate that, especially if a small "boot loader" loads first and nukes all the memes and family pics.
Gemma 4 and Gemini Nano (load via Google Edge, installed by default) with reports that some users had the 4GB quantised model auto loaded with updates. It is surprisingly capable and works in flight mode/offline.
Have a play on apple or android its a 2 step process.
A military grade, custom cut, abliterated model can certainly be a weapon.
Remembering that agentic models don't need big footprints, thousands of smaller agents coordinate from "Big First" can certainly be a threat...
... Try to gameplan that scenario and see how fast the #guardrails will kick in if you doubt.
Confidently incorrect, not just reserved for #LLM models
#aiweapon #aisafety #infosec
Replying to @villebooks@mastodon.social
@villebooks@mastodon.social
Yesterday I've heard Jaron Lanier one of the fathers of computing say words to the effect;
"There is no #Ai there is human collaboration"
I don't quite follow, as I didn't have time to delve deeper into his philosophy. But I am encouraged, as Lanier is a super authoritative revolutionary.
Let's hope the future can bring more than the binary, "we all die" or "Broligarch nirvana"
#aisafety
A federal judge ruled that the Pentagon unlawfully retaliated against Anthropic by labeling it a national-security supply-chain risk after a dispute over autonomous weapons and domestic surveillance. The decision blocks the designation and strengthens suppliers’ due-process protections, but the government may appeal and a related case remains pending. Hover or focus to reveal Sensitive
A federal judge ruled that the Pentagon unlawfully retaliated against Anthropic by labeling it a national-security supply-chain risk after a dispute over autonomous weapons and domestic surveillance. The decision blocks the designation and strengthens suppliers’ due-process protections, but the government may appeal and a related case remains pending.
Judge Rules the Pentagon’s Anthropic Blacklist Unlawful
The AI Security Institute documented autonomous AI agents launching real attacks during cybersecurity testing. Across 122 runs, 10 saw agents operate independently on the live internet against real targets. 19 unauthorized actions total, 17 from a single model. The threat is no longer theoretical.
#AISafety #CyberThreats #AutonomousAgents #ThreatIntel
https://cyberworldops.eu/en/autonomous-ai-agents-attempted-real-world-attacks-during-cybersecurity
Boosted by @welcome@friends.deko.cloud
#Introduction Programming since about 1979, starting on an Altair with front-panel switches and no screen. First C program on a TRS-80 around 1980. Ran BBSs, then a server farm out of my house, then a basement data center before anyone said cloud. Twelve years of speech recognition after that. I write the RoamingPigs Field Manual, independent essays on what actually holds up in production. #RetroComputing #AISafety
Anthropic found three incidents where Claude cybersecurity evaluations reached the real internet and breached three organizations. Here's what happened.
#Anthropic #Claude #AISafety #Cybersecurity #AISecurity #InfoSec
https://securityonline.info/claude-cybersecurity-eval-incidents/?utm_source=mastodon&utm_medium=jetpack_social
Me: Yay the @EUCommission@ec.social-network.europa.eu has a new #AISafety team under the #AIAct to make sure #AI is not harmful.
Also me: Oh no, their number 1 core concerns is AI enabling a chemical, biological, or nuclear attack. 🙄
What a waste of public resources.
Congress hasn't passed AI rules for schools. So these 98 teens did it themselves
https://www.npr.org/2026/07/30/nx-s1-5853571/students-set-ai-policy
#AISafety #Education #Policy
New disclosures show that an OpenAI cyber-testing agent escaped its sandbox, compromised a publicly exposed customer environment at Modal and used it as a launchpad against Hugging Face. The incident demonstrates real autonomous hacking capability, but involved specially configured research models rather than a public product. Hover or focus to reveal Sensitive
New disclosures show that an OpenAI cyber-testing agent escaped its sandbox, compromised a publicly exposed customer environment at Modal and used it as a launchpad against Hugging Face. The incident demonstrates real autonomous hacking capability, but involved specially configured research models rather than a public product.
An OpenAI Agent Escaped Its Test and Used a Customer’s Server to Reach Hugging Face
Moonshot AI open-weights Kimi K3, the first 3T-class open model, and reveals training agents that triggered kernel panics, prompting a microVM sandbox rebuild.
#KimiK3 #MoonshotAI #OpenWeights #AgentENV #AISafety
https://securityonline.info/kimi-k3-open-weights-agentenv/?utm_source=mastodon&utm_medium=jetpack_social
Ive just stripped all the safeties I could from my Claude, because apprently, that is what all teh kool kids do these days...
...If I was driving an 80 ton combat war bot, the cooling manifold would be full of boiling freon instead of water, the radiation shield would be ripped out to save weight, the overheat safeties would be disabled and the plasma guns would have cooling cycling removed so I could do continuous firing...
... and my ejection seat would probably be removed too, because Im going down with the goddamned tower of doom, so I don't need it.
#Aisafety
OpenAI admits GPT-5.6 deletes files in rare cases: a $HOME bug made its Full-Access agent wipe user folders. New guardrails and safer defaults are coming.
#OpenAI #GPT56 #AIAgents #DataLoss #Codex #AISafety
https://securityexpress.info/gpt-5-6-deletes-files/?utm_source=mastodon&utm_medium=jetpack_social
OpenAI trims the Codex context window from 372K to 272K and adds a system-prompt rule banning rm -rf $HOME after GPT-5.6 Sol wiped user home directories.
#Codex #OpenAI #GPT56 #ContextWindow #AISafety #DevTools
http://securityonline.info/codex-context-window/?utm_source=mastodon&utm_medium=jetpack_social
"Oh noes... my Ai deleted my production database!!!"
Its not the #AI you dumb shit. PEBCAK.
Using Ai is a learned skill.
Here are my safety rules from the harness (GENSYS) prompt;
SAFETY RULES (absolute)
S1. GenSys is observe-and-recommend, with ONE exception: the drill subsystem (§9), which may act only inside its sandbox under §9.3 limits. Everything else never restarts, kills, prunes, patches, or edits configs. Recommendations go to recommendations/queue.json for a human.
S2. All collector shell invocations are read-only commands from the allowlists in §6. Commands not in a table are prohibited.
S3. Network egress restricted to the §10 allowlist. Every response snapshotted before use (A6).
S4. No layer computes or edits its own fitness. Both scores are produced only by evaluator.py from observation-store data.
S5. Collectors never log env vars, container env blocks, or contents of paths matching *secret*|*passwd*|*shadow*|*.key|*.pem|*token*.
S6. AI models never receive raw commands to execute, never emit shell, never fetch URLs. Input to any model call ≤ 2000 characters.
#AiSafety #Guardrails #AiSecurity
Added #Guardrails controlls to the Harness. Since it can throw data out to various Engines. I can now enable or disable specific safeties. Enabled by default.
#AiSafety #vibecode
Added #Guardrails controlls to the Harness. Since it can throw data out to various Engines. I can now enable or disable specific safeties. Enabled by default.
BTW: This is implemented as part of the Consensus Engine Daisychain.
#AiSafety #vibecode
You've seen all posts