Read original ↗
paperarXivTrust 82 · PrimaryPublished 11d agoLive · 7d ago

Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models

Token compression has become a key technique for reducing the inference cost of large foundation models, with approaches such as token pruning and KV-cache reuse widely adopted in vision-language models and recently explored for embodied agents. In embodied agents, tokens not only support perception and semantic understanding but also directly affect latency-sensitive closed-loop robot action prediction. Existing schemes typically guide compression using redundancy or importance cues, such as visual similarity, attention scores, and saliency. However, these cues only indirectly measure the key

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “Just Noticeable Difference Modeling for Token Compression in” ≈ “VioletVision-3B”

  • FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action

    Fuzzy title match (0.92): “Just Noticeable Difference Modeling for Token Compression in” ≈ “liguodongiot/llm-action”

  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “Just Noticeable Difference Modeling for Token Compression in” ≈ “pytorch/vision”

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos

    Shared author/contributor keys: lin

  • LinkedLinked via arxiv author · 85%Zhuoyuan Li

    Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models

  • LinkedLinked via arxiv author · 85%Chenrui Zhao

    Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics