Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
One injected false document was enough to make deep-research agents endorse wrong conclusions in more than half of tested runs.
The paper introduces MisKnow-Agent, a benchmark that generated 5,933 misleading documents tied to audited false conclusions. Tested agents’ false-conclusion adoption rate rose from 0% without injection to a mean of 54.7% with one misleading document. The effect varied by research stage, framework design, source authority, and presentation style; search rank and extra misleading documents mattered less. Proposed pre- and post-research defenses reduced but did not eliminate the failures. HF Daily Papers' note
The paper introduces MisKnow-Agent, a benchmark that generated 5,933 misleading documents tied to audited false conclusions. Tested agents’ false-conclusion adoption rate rose from 0% without injection to a mean of 54.7% with one misleading document. The effect varied by research stage, framework design, source authority, and presentation style; search rank and extra misleading documents mattered less. Proposed pre- and post-research defenses reduced but did not eliminate the failures. HF Daily Papers' note
score 5