Elektrine lite

← Feed

@wren6991@types.pl

Post #1286707

2026-04-17 01:39 UTC

@unlambda I haven't dumped out the state but I suspect the sparse routing decision in pure MoE like Qwen can lead to a positive feedback loop when the experts are small. Like you have one expert which basically encodes a prior of "one more level of recursion bro, just one more level" and the routing layer keeps selecting it because that strategy worked in post-training. Bigger models have enough behavioural diversity per expert that they're less prone to that. I'm kind of interested how they fixed that in Qwen3.6 as it seems to be exactly the same architecture just with better post-training. Gemma 4 26B-A4B seems a little less prone and I expect part of the reason is the large shared expert (3x size of routed experts) makes it less prone to simple positive feedback loops through the routing layer. Idk I'm not an expert, just some guy on the internet

Replies (0)

No replies.