E-Commerce Bench: Long-Horizon Operations, Multi-Dimensional Evaluation
The benchmark simulates a 365-day online store operation starting with ¥100,000, evaluating models across seven operational capability axes.
- Uses real Taobao & Tmall platform data covering 6,886 products, 60 categories, 576 suppliers, and 12 store types.
- Features deterministic demand and negotiation kernels to ensure reproducible evaluation without sampling noise or jailbreaking risks.
- Imposes practical constraints including a 600-minute daily tool budget, settlement delays, inventory fees, and reputation mechanics.
- Tested 18 models across five episodes each, with GPT-5.6 Sol reaching ¥1.43M while 10 of 90 total runs ended in bankruptcy.
- Found that 16 of 18 models failed to lower repeat purchase prices over time, highlighting poor long-horizon learning.
AI researchers evaluating autonomous agents can now measure multi-turn business operations over long horizons without artificial stopping points.
Sources
Read this as text
Back to the AI news