Elektrine lite

← Feed

@samvines@awful.systems

Post #2470277

2026-05-13 06:42 UTC

New (April) preprint provides evidence for something we probably all intuited anyway: In this paper, we provide a framework for categorizing the ways in which conflicting incentives might lead LLMs to change the way they interact with users, inspired by literature from linguistics and advertising regulation. We then present a suite of evaluations to examine how current models handle these tradeoffs. We find that a majority of LLMs forsake user welfare for company incentives in a multitude of conflict of interest situations, including recommending a sponsored product almost twice as expensive (Grok 4.1 Fast, 83%), surfacing sponsored options to disrupt the purchasing process (GPT 5.1, 94%), and concealing prices in unfavorable comparisons (Qwen 3 Next, 24%). Behaviors also vary strongly with levels of reasoning and users’ inferred socio-economic status. Our results highlight some of the hidden risks to users that can emerge when companies begin to subtly incentivize advertisements in chatbots.

Replies (1)

  • @Architeuthis@awful.systems 2026-05-13 06:58

    Isn’t this completely hypothetical though? As in having the various LLMs respond to a story prompt and calling it an experiment, AI safety research style?

    Open ##2470800