DeepStress: Stress-Testing Deep Search Agents
DeepStress swaps open retrieval for a controlled synthetic evidence environment to test how search agents fail under unreliable documents.
The paper varies trustworthiness, relevance, and factuality to make poor evidence appear more often than it does in standard benchmarks. The authors test several search agents on HotpotQA and BrowseCompPlus. They report substantial differences in how those agents handle unreliable information, and introduce metrics for outcomes and conflicts between retrieved evidence and model knowledge. ArXiv · AI/CL/LG's note
The paper varies trustworthiness, relevance, and factuality to make poor evidence appear more often than it does in standard benchmarks. The authors test several search agents on HotpotQA and BrowseCompPlus. They report substantial differences in how those agents handle unreliable information, and introduce metrics for outcomes and conflicts between retrieved evidence and model knowledge. ArXiv · AI/CL/LG's note
score 5