How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast
The new mode delivers up to 8x faster token generation than Astra Standard mode for API, ChatGPT Work, and Codex users.
- OpenAI used internal models to generate optimized inference kernels directly for Blackwell and Rubin GPU architectures.
- The acceleration targets low-latency multi-step loops, including tool calls and coding agent edit-test-debug cycles.
- Access is available immediately through the OpenAI API, ChatGPT Work, and Codex.
Developers building agentic workflows and interactive coding assistants get significantly reduced latency between iterative tool calls.

Sources
Read this as text
Back to the AI news