Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions
The tested agents usually recognized failed retrieval, but did not use that judgment to stop.
The paper separates an agent’s verdict on tool output from its decision to keep querying. Across seven agents, failing-source results were labeled useless 97-100% of the time, yet most agents rarely stopped because of that. Prompt changes shifted stopping behavior, but did not reliably make stopping depend on the evidence. Only an enforced integration step after five consecutive useless results made stopping track the agent’s own judgments, with the effect confirmed in a 300-question replication. HF Daily Papers' note
The paper separates an agent’s verdict on tool output from its decision to keep querying. Across seven agents, failing-source results were labeled useless 97-100% of the time, yet most agents rarely stopped because of that. Prompt changes shifted stopping behavior, but did not reliably make stopping depend on the evidence. Only an enforced integration step after five consecutive useless results made stopping track the agent’s own judgments, with the effect confirmed in a 300-question replication. HF Daily Papers' note
score 4