# OpenAI tests cost-effective startup agents with GPT-5.6

@hypha_ai retained 98% of GPT-5.5’s document-extraction accuracy at 1/18 the cost with GPT-5.6 Luna.

- @RogoAI used programmatic tool calling for financial research, matching evaluation quality with 21% fewer input tokens.
- Retained reasoning and compaction raised GPT-5.6 Sol’s ARC-AGI-3 score from 13.3% to 38.3%.
- GPT-5.6 Sol used roughly 6x fewer output tokens in the ARC-AGI-3 evaluation.
- The work examined model selection, reasoning, and tool calling across teams in multiple industries.

## Why it matters

Startups building agents now have measured examples of reducing token use and costs while maintaining or improving task performance.

## Sources

- [OpenAI: OpenAI tests cost-effective startup agents with GPT-5.6](https://x.com/OpenAIDevs/status/2089374207818059793)

---

Summarized by dstilled on 2026-08-17. https://dstilled.ai/story/99055c71-de0f-4e29-bd84-611874fc06a0
