Post #2839080
2026-05-21 07:01 UTC
A multi-agent LLM where each agent learns when to defer to a human, trained with GRPO on a cost-aware reward. Each defer event becomes SFT data, so the model gradually absorbs the human's expertise. Tunable cost knob trades accuracy against human-call budget at deployment, no retraining.
https://benjaminhan.net/posts/20260520-adaptive-collaboration-mapo/?utm_source=mastodon&utm_medium=social
#ICLR #HumanInTheLoop #AgenticSystems #Metacognition #RL #AI
Replies (0)
No replies.