Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC) and present a method for learning and certifying counterfactual report mediators that hold a model's reports to a causal contract: invariant to forbidden influences (pressure, prestige, restyling) and responsive to licensed ones (genuine evidence). These two demands, resist and update, pull in opposite directions. We study them on a Bayesian-wi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 45%Researchers Demonstrate Chain-of-Thought Spoofing Against LLM Reasoners - Let's Data Science →
- LinkedLinked via arxiv author · 85%Sen Yang →
“Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs”
- LinkedLinked via arxiv author · 85%Yuen-Hei Yeung →
“Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs”
