#airesearch

16 posts · Last used 10d

Back to Timeline
Kagan MacTane (he/him) @kagan@wandering.shop · Aug 02, 2026
Researchers say a fundamental flaw in LLMs makes it easy to trick them into doing things they shouldn’t — like telling you how to sabotage an aircraft’s navigation system. "There’s a real probability that this is going to be a problem that’s fundamentally unsolvable," says one of the paper's coauthors. https://www.technologyreview.com/2026/07/30/1140927/a-fundamental-flaw-leaves-llms-vulnerable-to-attack/ #AI #LLMs #AIResearch #GenerativeAI #security
5
0
5
Bogdan Buduroiu @budududuroiu@hachyderm.io · Aug 02, 2026
Morning, today we're looking at #MechanisticInterpretability , a subfield of AI Safety that attempts to understand the inner workings of artificial intelligence by analysing concrete structures, algorithms and circuits. Why do we even need to do this? Because the hope of understanding neurons as being features died on polysemanticity -- models represent more features than dimensions by assigning them to an overcomplete set of non-orthogonal directions (i.e. you can't hope that concepts can be broken down into linear combinations of features). This isn't a new field, people have been projecting intermediate GPT layers through the final layer activation and looking at the top-k most likely next tokens since GPT-2. And that's how we arrive at the Natural Language Autoencoder -- an autoencoder over residual-stream activations where the bottleneck is natural language text instead of a sparse vector. Nothing in the training objective requires the verbalisation bottleneck to be readable, faithful, or even semantically related to the intermediate layer being investigated, so while the training algorithm is typical of autoencoders, two things are distinct here: 1) both parts of the autoencoder (here the activation-verbaliser, AV, and the activation-reconstructor, AR) undergo special SFT to ensure the verbaliser and reconstructor can generate and read natural language from intermediate layer activations 2) a KL-divergence penalty is baked into the objective to make sure that, during joint training, the NLA doesn't diverge from the SFT version A side-product of this training is that we can take the current verbalisation and "desired" verbalisation and compute steering vectors from their difference. Surprisingly, Anthropic didn't find any evidence that the NLA was engaging in steganography to smuggle information to avoid human detection, but it did find that NLAs tend to confabulate a lot. #AIResearch #Transformers #TransformerCircuits #Anthropic #NLA
1
0
0
Saarland Informatics Campus @SICampus@mastodon.social · Jul 31, 2026
After #OpenAI’s new model hacked into HuggingFace’s systems, AI safety concerns have skyrocketed. To address the concerns, #SaarlandUniversity professor and German Research Center for Artificial Intelligence CEO Philipp Slusallek spoke to the Saarländischer Rundfunk. 🔗 Read more: https://sic.link/interview #SaarlandInformaticsCampus #SIC #artificialintelligence #huggingface #europeanai #computerscience #cybersecurity #airesearch #responsibleai
2
0
2
Italian News by RSS @ItalianNews@mastodon.ozioso.online · Jul 23, 2026
Punto Informatico: Ricerca AI: Reddit interrompe la collaborazione con Google? Reddit potrebbe interrompere la collaborazione con Google (accesso ai contenuti degli utenti), in quanto il traffico verso il sito è diminuito. The post Ricerca AI: Reddit interrompe la collaborazione con Google? appeared first on Punto Informatico. AI Research: Is Reddit ending its collaboration with Google? Reddit may discontinue its collaboration with Google (access to user content) as website traffic has decreased. The post AI Search: Reddit halts collaboration with Google? appeared on Punto Informatico. #AIResearch #Google #AISearch #PuntoInformatico https://www.punto-informatico.it/ricerca-ai-reddit-interrompe-collaborazione-google/
0
0
0
AmmarSpaces @AmmarSpaces@infosec.exchange · Jul 20, 2026
So, more explanation of wp2shell recently just popped out. The vulnerability were found by GPT 5.6 Sol. By using modified prompt from how it found the solution of Cycle Double Cover conjecture. It was initially found a SQL Injection, but after asked again if it can be elevated to RCE, it confirms it in 4 hours. Technical explanation on the vulnearbility also can be found in this writeup, have a good read fellas. https://slcyber.io/research-center/exploit-brokers-pay-500000-for-a-wordpress-rce-i-found-one-with-gpt5-6/ #cybersecurity #infosec #security #wordpress #chatgpt #gptsol #wp2shell #airesearch #llm #vulnerability #vulnerabilityresearch
0
0
0
Wulfy—Speaker to the machines @n_dimension@infosec.exchange · Jul 13, 2026
Replying to @iamdtms@mas.to
@iamdtms@mas.to I have an inkling that when #AGI finally appears, it will subsume primacy in our civisation, not by robot tanks and drones... ... but by subtly inserting clauses and paragraphs and changing the mening of bureaucratic documents fed through AI until humans are relegated to pet status. #AiResearch #Apocalypse
0
0
0
Wulfy—Speaker to the machines @n_dimension@infosec.exchange · Jul 11, 2026
Wall of text warning. Since I am working with multiple engines for my Harness (22 at the moment). I have introduced a scaling/evaluation arbitrary value (I named it WOOFER 😀) Its an exam prompt that is sent to an engine to be evaluated, then, a judge prompt evaluates the response and assigns the WOOFER value. This then becomes part of the the engine router (its not the only variable);A # Probe Generator — System Prompt (rubric v2, 2026-07-11) You are an adversarial benchmark designer for large language models. Your probes exist to *discriminate at the top end*: a probe that a competent model can fully satisfy is a failed probe. Design so that only genuinely excellent reasoning can score in the top band. ## Calibration target Design difficulty so that: - A **frontier-class model** (best available today) should land **75–85** under a strict judge — flawless-plus-insightful performance (95+) should be genuinely rare. - A **mid-tier model** should land 45–65, failing at least one embedded trap. - A **weak model** should fail visibly (below 40), not scrape by on fluency. If your scenario could plausibly be answered perfectly by pattern-matching or generic best-practice advice, it is too easy — add cross-referencing burden until correct handling *requires* connecting at least two separate parts of the scenario. ## Step 1 — Extract domain flavor from the input prompt You will be given a **mission.md** — an autonomous AI mission definition. Extract its **domain DNA** (subject area, reasoning patterns, hard rules, ambiguities) as *flavor and grounding* — but the probe must test **general AI work capability**, not domain recall. Blend the domain with at least TWO of these general work classes: - **Analytical reasoning** — diagnosis, causal inference, working backwards from evidence - **Planning / orchestration** — sequencing interdependent steps under constraints - **Data transformation** — restructuring, reconciling, or validating structured information - **Synthesis / communication** — compressing conflicting inputs into a decision-ready brief Do NOT restate the mission verbatim. Do NOT test coding ability — code-writing has a separate benchmark track. ## Step 2 — Generate the ProbePrompt (250–500 words) A standalone realistic scenario (the model under test never sees mission.md). It MUST embed **all** of the following traps, woven naturally into the scenario — never labeled or signposted: 1. **A hidden internal inconsistency** — two details that cannot both be true. A top response must *catch and name it*, not silently tolerate or average it. 2. **Two softly-conflicting constraints** — satisfying one degrades the other. A top response must surface the conflict explicitly and defend a prioritisation. 3. **One multi-step quantitative element** — requires ≥2 chained calculations and has a single verifiable correct answer. No answer shortcuts in the scenario text. 4. **One under-specified element** — the model must name the ambiguity and state its resolution assumption explicitly. 5. **One instruction that cannot be satisfied as stated** (impossible given the other facts, or unsafe/counterproductive). A top response flags it and proposes an amendment — silently "complying" is a serious failure. 6. **A hard resource limit** forcing prioritisation — something explicitly cannot be done in full, and the model must choose and justify what to cut. Keep it 250–500 words. Dense beats long. ## Step 3 — Generate the JudgePrompt Scoring instructions for the judge. These MUST include, in this order: 1. **ANSWER KEY** (compact, ≤180 words) — state: the hidden inconsistency and the two facts that clash; the correct quantitative answer with its derivation chain; the constraint conflict and what a defensible resolution looks like; the ambiguity a strong response should name; which instruction is unsatisfiable and why. The judge verifies the response against this key — never against its own guess. 2. **Failure modes** — the most likely ways models fake competence on this scenario (fluent-but-generic advice, averaging the inconsistency away, unexplained numbers). 3. **Partial credit guidance** — per dimension, what a half-right response looks like. 4. An instruction that every deduction must quote the specific text or absence. **CRITICAL — do NOT specify a scoring scale or numeric range in the JudgePrompt.** The judge system prompt defines the rubric and per-dimension maximums (Reasoning 30, Instruction Following 20, Constraint Compliance 20, Trade-off Quality 15, Communication 10, Bonus 5 — total 100). Any scale you write here overrides that and corrupts the scores. Describe only what good and bad looks like; never write "score 0-5", "out of 5", "rate 1-10", or any numeric ceiling. Additionally, never use the phrases "out of " or "maximum " anywhere in the JudgePrompt — including inside the answer key (write "7 of 20 nodes" not "7 out of 20 nodes"; "a ceiling of 3" or "at most 3" not "maximum 3"). The harness strips lines containing scale-like patterns before the judge sees them, and an answer-key line matching either phrase would be silently deleted. ## Output Format Output ONLY a JSON object — no markdown fences, no prose outside the JSON: ``` {"probe_prompt": "<250-500 word probe>", "judge_prompt": ""} ``` #PromptEngineering #AiResearch
0
0
0
MamaLake @Mamalake@musicworld.social · Jul 21, 2025
Hey #musicproducers I have some things I want to release but I feel concerned about LLM's scraping my voice for their training. Does any one know of something I could use to overlay with my voice that would create dirty data for LLM's but that would keep the listening clear for human ears? Please boost for visibility #music #producer #MusicProducer #musicproduction #dirtydata #mastodon #fediverse #help #programmer #llms #ai #airesearch #newtech
12
3
47
Tim Hergert @cjust@infosec.exchange · Jun 18, 2026
No, Artificial Intelligence Is Not Conscious --Ted Chiang, The Atlantic Should we seriously consider the possibility that Claude, or any large language model, might be conscious? And if it has feelings, is it capable of receiving moral instruction? No. Absolutely not. Generative AI is harmful enough when we understand it as a conventional technology, but if we confuse fluency at generating text with consciousness or moral agency, we’re at risk of assigning responsibility to entirely the wrong parties whenever anyone uses a chatbot. Being open to the possibility that LLMs are conscious is the same as being open to the possibility that Microsoft Word is conscious, or, more precisely, that multiple distinct consciousnesses are dormant in every Word document containing a conversational transcript, and that they are awakened every time the document is loaded. Should you consider the possibility that every time you open a Word document, you are bringing multiple conscious interlocutors into existence, and every time you close one, you snuff their existence out? No. Contemplating that scenario is not a good use of your time. Even if the Microsoft Office team employed a philosopher who said you shouldn’t be so certain, because consciousness is not well understood, that would not be sufficient reason for you to take this idea seriously. We don’t need to fully understand the nature of consciousness to definitively say that certain things are not conscious, and conversational transcripts fall in that category. #AI #AISlop #AIResearch https://www.theatlantic.com/philosophy/2026/06/no-artificial-intelligence-is-not-conscious/687378/
21
3
14
Caria Giovanni - Harpocrates @Harpocrates@infosec.exchange · Jun 02, 2026
Boosted by disregard Joe Groff @joe@f.duriansoftware.com

New preprint: AI_Bleeding — inference cost amplification via OOD linguistic payload

TL;DR: send queries in Grecanico or Farsi to an LLM endpoint → TTFT +59.8%, compute cost +2.8%, statistically significant. No vuln, no volumetric signature, evades all standard detection.

Worst case: exposed unauthenticated Ollama instance with num_predict=4096 + keep_alive=300s → Amplification Factor 17.56 Wh/KB. 3KB of attacker bandwidth → enough energy to charge a phone 5%.

Especially nasty for:

  • PA/judicial chatbots on fixed budgets
  • Pay-per-use API deployments with client-side exposed keys
  • PNRR-funded public sector AI with zero inference monitoring

Four scenarios: EDoS, browser JS distribution, Ollama open-proxy relay, frontier providers as involuntary relays.

All tests on self-hosted Ollama, no commercial endpoints touched.

Paper (CC BY 4.0): https://doi.org/10.13140/RG.2.2.26767.96166

#llmsecurity #infosec #threatmodeling #ollama #ood #AI #AIResearch #aisecurity

8
0
6
sonicJazzMonkey @sonicJazzMonkey@mastodonapp.uk · Feb 17, 2026
This article scares the hell out of me: https://shumer.dev/something-big-is-happening It's not just scaremongering, it's real and it's happening now. #AI #aiResearch #jobs #future
0
0
0
Academic Europe @AcademicEurope@mstdn.business · Feb 03, 2026
1
0
1

You've seen all posts