Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
The paper tests whether agents fail because they lack private task conventions, not because they cannot execute the work.
The authors propose a protocol that pairs identical task instructions with either a supplied or withheld artefact containing the needed private rules, tables, and operators. In 15 calibration tasks, one frontier agent setup passed 68.0% with the artefact and 0% without it. A plausible but wrong artefact also produced 0% on one task across five trials. The paper says the experiments validate the construction protocol, but do not show the retained tasks improve post-training. HF Daily Papers' note
The authors propose a protocol that pairs identical task instructions with either a supplied or withheld artefact containing the needed private rules, tables, and operators. In 15 calibration tasks, one frontier agent setup passed 68.0% with the artefact and 0% without it. A plausible but wrong artefact also produced 0% on one task across five trials. The paper says the experiments validate the construction protocol, but do not show the retained tasks improve post-training. HF Daily Papers' note
score 5