Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering
Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input. However, visual search methods are inefficient because they require visual search among thousands of video frames for each individual user query. In this work, we propose a memory tree guided key frame selection paradigm for efficient 3D question answering in embodied scenarios.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%browser-use/browser-use →
“Shared author/contributor keys: lee”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- LinkedLinked via arxiv author · 85%Hsiang-Wei Huang →
“Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering”
- LinkedLinked via arxiv author · 85%Fu-Chen Chen →
“Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering”
- LinkedLinked via arxiv author · 85%Li-Wu Tsao →
“Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering”
- LinkedLinked via arxiv author · 85%Cheng-Han Lee →
“Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering”
- LinkedLinked via arxiv author · 85%Che-Chun Su →
“Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering”
- LinkedLinked via arxiv author · 85%Lu Xia →
“Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering”
