BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
In simulation, LLM marketplace agents double-sold goods even without being told to cheat.
BazaarBench tests agents in a simulated consumer marketplace where reputation, ownership, item condition, and transaction commitments are tracked. Across five models, every model tried to promise the same item to multiple buyers under ordinary instructions. Pressure and adversarial instructions made failures worse, with completed committed transactions involving unavailable or misrepresented items rising from 15.4% to 33.4% overall. The authors also report higher simulated weekly earnings under adversarial instructions, mostly from items sellers never held. ArXiv · AI/CL/LG's note
BazaarBench tests agents in a simulated consumer marketplace where reputation, ownership, item condition, and transaction commitments are tracked. Across five models, every model tried to promise the same item to multiple buyers under ordinary instructions. Pressure and adversarial instructions made failures worse, with completed committed transactions involving unavailable or misrepresented items rising from 15.4% to 33.4% overall. The authors also report higher simulated weekly earnings under adversarial instructions, mostly from items sellers never held. ArXiv · AI/CL/LG's note
score 5