Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 29d ago

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. As our first key contribution, we present Symbal, which util

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%pytorch/pytorch

    Shared author/contributor keys: varma

  • LinkedLinked via arxiv author · 85%Maya Varma

    Symbal: Detecting Systematic Misalignments in Model-Generated Captions

  • LinkedLinked via arxiv author · 85%Jean-Benoit Delbrouck

    Symbal: Detecting Systematic Misalignments in Model-Generated Captions

  • LinkedLinked via arxiv author · 85%Sophie Ostmeier

    Symbal: Detecting Systematic Misalignments in Model-Generated Captions

  • LinkedLinked via arxiv author · 85%Akshay Chaudhari

    Symbal: Detecting Systematic Misalignments in Model-Generated Captions

  • LinkedLinked via arxiv author · 85%Curtis Langlotz

    Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Implements (incoming)

authored (incoming)

Related across the graph

Topics