newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago
LLMs know when they are wrong. I made a fix relating to Anthropic's new "global workspace" paper [R]
I have posted before about finding out a model's actual confidence in its answer through probes and hidden states (AUROC ~0.83–0.88 across every model I tested, 7B to 72B). This is the know-say gap. From my work and the work done by others in this space it is likely a routing problem. By making a tiny bridge from a linear probe on mid-layer sate plus ten trained weights that write the probe's estimate onto the confidence-digit logits can make the model ve
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Evaluate a model properly →
- PossiblePossibly related (embedding) · 47%algorithmicsuperintelligence/optillm →
- PossiblePossibly related (embedding) · 46%Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment →
- PossiblePossibly related (embedding) · 49%Amirhosein-gh98/Gnosis →
- PossiblePossibly related (embedding) · 50%Two Axes of LLM Abstention: Answer Correctness and Question Answerability →
- PossiblePossibly related (embedding) · 47%Confident at the moment of action: belief miscalibration in LLM play under hidden information →
Covers
Covers (incoming)
Related across the graph
paperEvil Spectra: How Optimisers can Amplify or Suppress Emergent MisalignmentpaperConfident at the moment of action: belief miscalibration in LLM play under hidden informationrepoAmirhosein-gh98/GnosispaperTwo Axes of LLM Abstention: Answer Correctness and Question Answerabilityrepoalgorithmicsuperintelligence/optillmtutorialEvaluate a model properly
