Megadose AI progress, ranked and analyzed.

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

· HF Daily Papers ·
The best LLM setup reached only 27.3% of human mean final net assets in a year-long e-commerce simulation.

MerchantBench tests agents across 365 simulated days, using real product records and 26 tools. The benchmark forces decisions on sourcing, listings, pricing, cash flow, and delayed order feedback. Eight LLMs were evaluated under two agent frameworks across 48 full-year runs. The authors frame the gap as a failure of long-term coherence, not just single-task tool use. HF Daily Papers' note

score 5

Categories: Research