ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models
Large vision-language models (LVLMs) often hallucinate, generating content that the input image does not support. Preventing such content during decoding calls for a candidate-specific measure of how strongly the image supports the token under consideration. The model's visual-token states offer a natural source of this evidence because projecting each state through the output head reveals which vocabulary items that position favors. These position-wise readouts cannot be pooled directly because their probability magnitudes are not comparable across visual positions. Vocabulary ranks provide a
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual” ≈ “VioletVision-3B””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual” ≈ “pytorch/vision””
- FuzzyOverlapping authors or contributors · 62%ultralytics/yolov5 →
“Shared author/contributor keys: choi”
- LinkedLinked via arxiv author · 85%Jihae Jeong →
“ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Languag”
- LinkedLinked via arxiv author · 85%Junha Choi →
“ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Languag”
- LinkedLinked via arxiv author · 85%Hwanjo Yu →
“ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Languag”
