When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation
A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has to keep up with the stream and stop a harmful reply before a user reads it. Recent vision-language guardrails instead generate a chain of thought before they issue a verdict. They believe that step-by-step reasoning yields a safer guard. This design makes the guard heavy and slow, since the model must decode many tokens for harmfulness detection. We pose the question of whether a vision-language guard really needs to reason in order to screen a response. We
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “When Are Reasoning-Based Guardrails Not Efficient? ResponseG” ≈ “VioletVision-3B””
- PossiblePossibly related (embedding) · 56%New Research: AI models can give dangerous responses despite output guardrails - The AI Journal →
- PossiblePossibly related (embedding) · 46%Understanding Annotator Safety Policy with Interpretability - Apple Machine Learning Research →
- FuzzySimilar title/name (fuzzy) · 87%guardrails-ai/guardrails →
“Fuzzy title match (0.94): “When Are Reasoning-Based Guardrails Not Efficient? ResponseG” ≈ “guardrails-ai/guardrails””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “When Are Reasoning-Based Guardrails Not Efficient? ResponseG” ≈ “pytorch/vision””
- LinkedLinked via arxiv author · 85%Dongbin Na →
“When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation”
