Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors
Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance exploration and exploitation. Unlike classical agents, LLM agents engage with tasks through natural language, exposing them to semantic information with no formal counterpart in the task structure. We introduce the semantic bandit, an extension of the multi-armed bandit setting that explicitly considers the textual labels assigned to actions, and use it to study how semantic
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%David Eric Austin →
“Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors”
- LinkedLinked via arxiv author · 85%Kaheer Suleman →
“Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors”
- LinkedLinked via arxiv author · 85%Jackie Chi Kit Cheung →
“Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors”
- FuzzySimilar title/name (fuzzy) · 59%microsoft/semantic-kernel →
“Fuzzy title match (0.73): “Semantic Bandits: In-Context Exploration-Exploitation is Bia” ≈ “microsoft/semantic-kernel””
