Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world driving inci

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 53%vlm-starter
  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “AUTOPILOT VQA: Benchmarking Vision-Language Models for Incid” ≈ “VioletVision-3B”

  • LinkedLinked via arxiv author · 85%Siddharth Damodharan

    AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

  • LinkedLinked via arxiv author · 85%Radhika Gupta

    AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

  • LinkedLinked via arxiv author · 85%Ali Alshami

    AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

  • LinkedLinked via arxiv author · 85%Ryan Rabinowitz

    AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

  • LinkedLinked via arxiv author · 85%Jugal Kalita

    AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “AUTOPILOT VQA: Benchmarking Vision-Language Models for Incid” ≈ “pytorch/vision”

Implements

Has model

authored (incoming)

Implements (incoming)

Related across the graph

Topics