OpenAI tests cost-effective startup agents with GPT-5.6
@hypha_ai retained 98% of GPT-5.5’s document-extraction accuracy at 1/18 the cost with GPT-5.6 Luna.
- @RogoAI used programmatic tool calling for financial research, matching evaluation quality with 21% fewer input tokens.
- Retained reasoning and compaction raised GPT-5.6 Sol’s ARC-AGI-3 score from 13.3% to 38.3%.
- GPT-5.6 Sol used roughly 6x fewer output tokens in the ARC-AGI-3 evaluation.
- The work examined model selection, reasoning, and tool calling across teams in multiple industries.
Startups building agents now have measured examples of reducing token use and costs while maintaining or improving task performance.
Sources
Read this as text
Back to the AI news