newsGoogle News — LLMTrust 62 · AggregatorPublished 4d agoLive · 4d ago
Causal evidence that language models use confidence to drive behaviour - nature.com
Causal evidence that language models use confidence to drive behaviour nature.com
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research →
- PossiblePossibly related (embedding) · 51%Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs →
- PossiblePossibly related (embedding) · 50%Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery →
- PossiblePossibly related (embedding) · 50%Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision →
- PossiblePossibly related (embedding) · 50%When Linguistic and Internal Confidence Diverge in Large Language Models →
Covers
paperFrom Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception ResearchpaperResist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMspaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperIntrospective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed SupervisionpaperWhen Linguistic and Internal Confidence Diverge in Large Language Models
Related across the graph
paperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperFrom Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception ResearchpaperResist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMspaperWhen Linguistic and Internal Confidence Diverge in Large Language ModelspaperIntrospective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision
