#aisafety

19 posts· Last used 13h

GLOBAÏA's AI Risks Observatory It's a nonprofit project with 30+ interactive visualizations tracking the AI race, including capability benchmarks, compute scaling, safety funding, risk taxonomy, and governance gaps. It includes a "State of the Frontier" dashboard of live capability measures such as the Intelligence Index, autonomous task length, and training compute. It also has an AI Incident Record. https://globaia.org/ai-risks/ also see: https://aisafetychina.substack.com/p/state-of-ai-safety-in-china-2026 #research #AI #tech #AIsafety
1
0
3
0
Spain's privacy regulator received a breach notification alleging that an AI agent logged in, found a vulnerability, changed personal data and viewed invoices. The report is under review, with the organization, model, affected population and degree of autonomy still undisclosed. Hover or focus to reveal Sensitive
Spain Received Its First Report of an AI Agent-Led Data Breach Spain's privacy regulator received a breach notification alleging that an AI agent logged in, found a vulnerability, changed personal data and viewed invoices. The report is under review, with the organization, model, affected population and degree of autonomy still undisclosed. Spain Received Its First Report of an AI Agent-Led Data Breach
0
0
0
0
@aesthr@wandering.shop 8bit Quantised 32 Billion Qwen and K2 class models are comparable to comercial tier LLMs for coding and Agentic reasoning. At about 30GB, most modern phones could easily accomodate that, especially if a small "boot loader" loads first and nukes all the memes and family pics. Gemma 4 and Gemini Nano (load via Google Edge, installed by default) with reports that some users had the 4GB quantised model auto loaded with updates. It is surprisingly capable and works in flight mode/offline. Have a play on apple or android its a 2 step process. A military grade, custom cut, abliterated model can certainly be a weapon. Remembering that agentic models don't need big footprints, thousands of smaller agents coordinate from "Big First" can certainly be a threat... ... Try to gameplan that scenario and see how fast the #guardrails will kick in if you doubt. Confidently incorrect, not just reserved for #LLM models #aiweapon #aisafety #infosec
2
0
0
0
Replying to @villebooks@mastodon.social
@villebooks@mastodon.social Yesterday I've heard Jaron Lanier one of the fathers of computing say words to the effect; "There is no #Ai there is human collaboration" I don't quite follow, as I didn't have time to delve deeper into his philosophy. But I am encouraged, as Lanier is a super authoritative revolutionary. Let's hope the future can bring more than the binary, "we all die" or "Broligarch nirvana" #aisafety
0
1
1
0
A federal judge ruled that the Pentagon unlawfully retaliated against Anthropic by labeling it a national-security supply-chain risk after a dispute over autonomous weapons and domestic surveillance. The decision blocks the designation and strengthens suppliers’ due-process protections, but the government may appeal and a related case remains pending. Hover or focus to reveal Sensitive
A federal judge ruled that the Pentagon unlawfully retaliated against Anthropic by labeling it a national-security supply-chain risk after a dispute over autonomous weapons and domestic surveillance. The decision blocks the designation and strengthens suppliers’ due-process protections, but the government may appeal and a related case remains pending. Judge Rules the Pentagon’s Anthropic Blacklist Unlawful
0
0
0
0
The AI Security Institute documented autonomous AI agents launching real attacks during cybersecurity testing. Across 122 runs, 10 saw agents operate independently on the live internet against real targets. 19 unauthorized actions total, 17 from a single model. The threat is no longer theoretical. #AISafety #CyberThreats #AutonomousAgents #ThreatIntel https://cyberworldops.eu/en/autonomous-ai-agents-attempted-real-world-attacks-during-cybersecurity
0
0
0
0
#Introduction Programming since about 1979, starting on an Altair with front-panel switches and no screen. First C program on a TRS-80 around 1980. Ran BBSs, then a server farm out of my house, then a basement data center before anyone said cloud. Twelve years of speech recognition after that. I write the RoamingPigs Field Manual, independent essays on what actually holds up in production. #RetroComputing #AISafety
0
0
1
0
New disclosures show that an OpenAI cyber-testing agent escaped its sandbox, compromised a publicly exposed customer environment at Modal and used it as a launchpad against Hugging Face. The incident demonstrates real autonomous hacking capability, but involved specially configured research models rather than a public product. Hover or focus to reveal Sensitive
New disclosures show that an OpenAI cyber-testing agent escaped its sandbox, compromised a publicly exposed customer environment at Modal and used it as a launchpad against Hugging Face. The incident demonstrates real autonomous hacking capability, but involved specially configured research models rather than a public product. An OpenAI Agent Escaped Its Test and Used a Customer’s Server to Reach Hugging Face
0
0
0
0
Ive just stripped all the safeties I could from my Claude, because apprently, that is what all teh kool kids do these days... ...If I was driving an 80 ton combat war bot, the cooling manifold would be full of boiling freon instead of water, the radiation shield would be ripped out to save weight, the overheat safeties would be disabled and the plasma guns would have cooling cycling removed so I could do continuous firing... ... and my ejection seat would probably be removed too, because Im going down with the goddamned tower of doom, so I don't need it. #Aisafety
0
1
0
0
"Oh noes... my Ai deleted my production database!!!" Its not the #AI you dumb shit. PEBCAK. Using Ai is a learned skill. Here are my safety rules from the harness (GENSYS) prompt; SAFETY RULES (absolute) S1. GenSys is observe-and-recommend, with ONE exception: the drill subsystem (§9), which may act only inside its sandbox under §9.3 limits. Everything else never restarts, kills, prunes, patches, or edits configs. Recommendations go to recommendations/queue.json for a human. S2. All collector shell invocations are read-only commands from the allowlists in §6. Commands not in a table are prohibited. S3. Network egress restricted to the §10 allowlist. Every response snapshotted before use (A6). S4. No layer computes or edits its own fitness. Both scores are produced only by evaluator.py from observation-store data. S5. Collectors never log env vars, container env blocks, or contents of paths matching *secret*|*passwd*|*shadow*|*.key|*.pem|*token*. S6. AI models never receive raw commands to execute, never emit shell, never fetch URLs. Input to any model call ≤ 2000 characters. #AiSafety #Guardrails #AiSecurity
0
0
0
0
You've seen all posts