Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
IAR tries to make a model carry a fixed document set in its weights, then regain broader instruction performance.
The paper frames the problem as retrieval-free document knowledge internalization: answering from a bounded corpus without fetching documents at inference time. Its three stages inject document knowledge through reconstruction-style objectives, align the model with answer-only QA supervision, and recover general ability by merging with the base instruction model. Across CC and CCI, and Llama, Phi, Qwen, and SmolLM families, it reports better domain QA and general-task results than Vanilla SFT in 7 of 8 dataset-model settings. HF Daily Papers' note
The paper frames the problem as retrieval-free document knowledge internalization: answering from a bounded corpus without fetching documents at inference time. Its three stages inject document knowledge through reconstruction-style objectives, align the model with answer-only QA supervision, and recover general ability by merging with the base instruction model. Across CC and CCI, and Llama, Phi, Qwen, and SmolLM families, it reports better domain QA and general-task results than Vanilla SFT in 7 of 8 dataset-model settings. HF Daily Papers' note
score 5