Megadose Built for builders and researchers.

UNREAL: Unifying Retrieval and Long-Context with a Single Model

· HF Daily Papers ·
UNREAL uses the frozen LLM’s own representations to choose evidence for both retrieval and long-context inference.

The paper says the method adds under 500K trainable parameters while leaving the backbone unchanged. On a 3B-token Wikipedia index, its tested dense and hybrid backbones beat retriever-reranker baselines, with the largest HotpotQA recall rising from 49.1% to 73.2%. The same mechanism is used to strip distractors before generation in long-context tasks, improving NoLiMa and LV-Eval results at 128K and 256K contexts. The authors also report lower FLOPs and faster time-to-first-token than full-context inference from about 32K tokens onward. Source: HF Daily Papers' note

score 6

Categories: Research