Read original ↗
newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago

Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R]

Anthropic’s Jacobian Lens work introduced a way to inspect verbalizable representations inside language models. Follow-up experiments suggested that entropy in this internal “workspace” might help identify confidently incorrect answers. I tested that hypothesis on Qwen3-4B across ~11,400 examples from seven distinct datasets, including TriviaQA, PopQA, NQ-Open, TruthfulQA, HotpotQA, GSM8K, and CommonSenseQA. Three main findings: It can

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

Related across the graph

paperAuditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation CaveatpaperHow Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLIpaperEmergent Misalignment Recruits a Pre-existing Persona SubspacepaperThinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot SeepaperWhen are likely answers right? On Sequence Probability and Correctness in LLMspaperEstimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMspaperOn the Threat Model of Weird Generalization and Emergent MisalignmentpaperWhen Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk SpacepaperThe Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMspaperHeaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language ModelspaperGraded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal SteeringpaperThe Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant ContextpaperLocal and Global Regimes of Geometric Complexity in Language Model RepresentationspaperThe One-Word Census: Answer-Choice Conformity Across 44 Language ModelspaperHow Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple MitigationpaperGotta Catch them all: the modes of SycophancypaperThe Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation ProtocolpaperWhen RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public CompaniespaperHow Temperature Shapes Ideological Discourse in Retrieval-Augmented Generation?paperLinear representations of grammaticality in neural language modelspaperLosing My Composure: Predicting Compositionality Over Time