When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
The paper says the reported guardrail gains collapse once the commerce simulation is controlled for protocol differences and randomness.
The audited hotel-market testbed initially showed large welfare gains from marketplace guardrails, but guarded and unguarded agents were not using the same offer schema or buyer choice procedure. After those were held fixed, the results changed sharply and no longer supported the original policy claim. The authors label the original estimate INVALID under protocol isolation, and the controlled study INCONCLUSIVE on incentive validity and stochastic stability. ArXiv · AI/CL/LG's note
The audited hotel-market testbed initially showed large welfare gains from marketplace guardrails, but guarded and unguarded agents were not using the same offer schema or buyer choice procedure. After those were held fixed, the results changed sharply and no longer supported the original policy claim. The authors label the original estimate INVALID under protocol isolation, and the controlled study INCONCLUSIVE on incentive validity and stochastic stability. ArXiv · AI/CL/LG's note
score 4