Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. As our first key contribution, we present Symbal, which util
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%pytorch/pytorch →
“Shared author/contributor keys: varma”
- LinkedLinked via arxiv author · 85%Maya Varma →
“Symbal: Detecting Systematic Misalignments in Model-Generated Captions”
- LinkedLinked via arxiv author · 85%Jean-Benoit Delbrouck →
“Symbal: Detecting Systematic Misalignments in Model-Generated Captions”
- LinkedLinked via arxiv author · 85%Sophie Ostmeier →
“Symbal: Detecting Systematic Misalignments in Model-Generated Captions”
- LinkedLinked via arxiv author · 85%Akshay Chaudhari →
“Symbal: Detecting Systematic Misalignments in Model-Generated Captions”
- LinkedLinked via arxiv author · 85%Curtis Langlotz →
“Symbal: Detecting Systematic Misalignments in Model-Generated Captions”
