newsGoogle News — LLMTrust 88 · AggregatorPublished 4d agoLive · 3d ago
Causal evidence that language models use confidence to drive behaviour - Nature
Causal evidence that language models use confidence to drive behaviour Nature
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research →
- PossiblePossibly related (embedding) · 50%Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision →
- PossiblePossibly related (embedding) · 50%Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs →
- PossiblePossibly related (embedding) · 50%Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery →
- PossiblePossibly related (embedding) · 49%Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs →
Covers
paperFrom Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception ResearchpaperIntrospective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed SupervisionpaperResist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMspaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperInference-Time Steering for Cross-Lingual Factual Consistency in LLMs
Related across the graph
paperInference-Time Steering for Cross-Lingual Factual Consistency in LLMspaperBefore the Action: Benchmarking LLMs on Prospective Hypothesis DiscoverypaperFrom Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception ResearchpaperResist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMspaperIntrospective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision
