Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts
Fine-grained human action recognition (FHAR) must distinguish visually similar actions that differ mainly in body configuration, timing, or local appearance. RGB representations retain visual context but often suppress joint-level geometry, whereas skeleton representations encode kinematics but discard dense spatial detail. We introduce FineX, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology. Pairwise cross-attention enables symmetric, stream-preserving information exchange, followed by a streamwise latent sparse Mixture-of-Experts that
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action →
“Fuzzy title match (0.92): “Fine-Grained Action Recognition with Cross-Attentive Latent ” ≈ “liguodongiot/llm-action””
- LinkedLinked via arxiv author · 85%Imtiaz Ul Hassan →
“Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts”
- LinkedLinked via arxiv author · 85%Tasweer Ahmad →
“Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts”
- LinkedLinked via arxiv author · 85%Nik Bessis →
“Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts”
- LinkedLinked via arxiv author · 85%Ardhendu Behera →
“Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts”
