Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction
The paper argues that extraction benchmarks can pass fabricated answers unless they log the agent’s tool calls.
The authors found a model that passed a fidelity check while never opening the datasheet, after a structured-output constraint disabled tool use. They benchmarked 37 hand-curated datasheet claims and built dispatch-level instruments to attribute failures and detect silent tool-use failures. Their detector did not flag 207 clean fidelity-passing runs and recovered all 50 planted faults matching its tool-withholding rules, while leaving other failure modes unmeasured. A physical-measurement oracle covered only 2 of the 37 claims, underscoring how partial external verification can be. ArXiv · AI/CL/LG's note
The authors found a model that passed a fidelity check while never opening the datasheet, after a structured-output constraint disabled tool use. They benchmarked 37 hand-curated datasheet claims and built dispatch-level instruments to attribute failures and detect silent tool-use failures. Their detector did not flag 207 clean fidelity-passing runs and recovered all 50 planted faults matching its tool-withholding rules, while leaving other failure modes unmeasured. A physical-measurement oracle covered only 2 of the 37 claims, underscoring how partial external verification can be. ArXiv · AI/CL/LG's note
score 4