# E-Commerce Bench: Long-Horizon Operations, Multi-Dimensional Evaluation

The benchmark simulates a 365-day online store operation starting with ¥100,000, evaluating models across seven operational capability axes.

- Uses real Taobao & Tmall platform data covering 6,886 products, 60 categories, 576 suppliers, and 12 store types.
- Features deterministic demand and negotiation kernels to ensure reproducible evaluation without sampling noise or jailbreaking risks.
- Imposes practical constraints including a 600-minute daily tool budget, settlement delays, inventory fees, and reputation mechanics.
- Tested 18 models across five episodes each, with GPT-5.6 Sol reaching ¥1.43M while 10 of 90 total runs ended in bankruptcy.
- Found that 16 of 18 models failed to lower repeat purchase prices over time, highlighting poor long-horizon learning.

## Why it matters

AI researchers evaluating autonomous agents can now measure multi-turn business operations over long horizons without artificial stopping points.

## Sources

- [Qwen: E-Commerce Bench: Long-Horizon Operations, Multi-Dimensional Evaluation](https://qwen.ai/blog?id=e-commerce-bench)

---

Summarized by dstilled on 2026-09-03. https://dstilled.ai/story/7f71e9ca-cc11-49b3-bd9a-3cb1f1679429
