UNREAL: Unifying Retrieval and Long-Context with a Single Model
UNREAL uses the frozen LLM’s own representations to choose evidence for both retrieval and long-context inference.
The paper says the method adds under 500K trainable parameters while leaving the backbone unchanged. On a 3B-token Wikipedia index, its tested dense and hybrid backbones beat retriever-reranker baselines, with the largest HotpotQA recall rising from 49.1% to 73.2%. The same mechanism is used to strip distractors before generation in long-context tasks, improving NoLiMa and LV-Eval results at 128K and 256K contexts. The authors also report lower FLOPs and faster time-to-first-token than full-context inference from about 32K tokens onward. Source: HF Daily Papers' note
The paper says the method adds under 500K trainable parameters while leaving the backbone unchanged. On a 3B-token Wikipedia index, its tested dense and hybrid backbones beat retriever-reranker baselines, with the largest HotpotQA recall rising from 49.1% to 73.2%. The same mechanism is used to strip distractors before generation in long-context tasks, improving NoLiMa and LV-Eval results at 128K and 256K contexts. The authors also report lower FLOPs and faster time-to-first-token than full-context inference from about 32K tokens onward. Source: HF Daily Papers' note
score 6