Read original ↗
paperarXivTrust 82 · PrimaryPublished 2d agoLive · 2h ago

On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized. In practice, degradations such as defocus, haze, motion blur, noise, cautery smoke, and packet loss introduce structured distribution shifts which may compromise video-text alignment. We study the robustness of temporal VLMs under such shifts caused by corruptions in clip frames. We introduce Endo-C6, a compact corruption benchmark of six endoscopy-realistic pe

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “On the Robustness of Temporal Vision-Language Models for Sur” ≈ “VioletVision-3B”

  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “On the Robustness of Temporal Vision-Language Models for Sur” ≈ “pytorch/vision”

  • FuzzyOverlapping authors or contributors · 62%openai/openai-agents-python

    Shared author/contributor keys: bilal

  • LinkedLinked via arxiv author · 85%Darakshan Rashid

    On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

  • LinkedLinked via arxiv author · 85%Raza Imam

    On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

  • LinkedLinked via arxiv author · 85%Ufaq Khan

    On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

  • LinkedLinked via arxiv author · 85%Muhammad Bilal

    On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

  • LinkedLinked via arxiv author · 85%Shazad Ashraf

    On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics