Megadose AI progress, ranked and analyzed.

Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors

· ArXiv · AI/CL/LG ·
LLMs explored less when action labels carried familiar semantic cues, even when those cues pointed the wrong way.

The paper defines “semantic bandits” to test how natural-language action labels bias exploration and exploitation in LLM agents. Informative labels helped when their implied rewards matched the task, but badly hurt performance when they were misleading. The authors also report that negative rewards led models to explore substantially more than equivalent positive rewards. ArXiv · AI/CL/LG's note

score 5

Categories: Research