Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads
For fixed token budgets, testing more questions beat spending on extra reads or search trajectories.
The paper evaluates agentic RAG budget choices on HotpotQA and MuSiQue. At roughly 34.14-34.39M model tokens, broader question coverage reduced standard error by 33% versus five reads and 12.6% versus three trajectories. Forecasting methods came within 4.0% and 3.5% of the measured allocations. Temperature zero sharply reduced answer disagreement, from 14.3% to 3.4%, without materially changing comparison precision. HF Daily Papers' note
The paper evaluates agentic RAG budget choices on HotpotQA and MuSiQue. At roughly 34.14-34.39M model tokens, broader question coverage reduced standard error by 33% versus five reads and 12.6% versus three trajectories. Forecasting methods came within 4.0% and 3.5% of the measured allocations. Temperature zero sharply reduced answer disagreement, from 14.3% to 3.4%, without materially changing comparison precision. HF Daily Papers' note
score 4