Read original ↗
paperarXivTrust 82 · PrimaryPublished 21d agoLive · 18d ago

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, re

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Qian Kou

    Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

  • LinkedLinked via arxiv author · 85%Xiaofeng Shi

    Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

  • LinkedLinked via arxiv author · 85%Xiaosong Qiu

    Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

  • LinkedLinked via arxiv author · 85%Jinhua Zhou

    Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

authored (incoming)

Related across the graph

Topics