Key Points
- GPT-6 Astra Ultrafast is live in OpenAI's API, running on NVIDIA Blackwell GPUs.
- NVIDIA says it generates tokens up to 8x faster than Astra Standard mode.
- Faster output targets coding agents, tool calls and latency-sensitive interactive apps.
The latest:
A speed-tuned version of OpenAI’s newest model, GPT-6 Astra Ultrafast, is now available in the OpenAI API and to eligible ChatGPT Work and Codex users, NVIDIA said in an October 1, 2026 post. The chipmaker said inference optimizations built around its Blackwell architecture deliver token generation up to 8x faster than the Astra Standard mode.
Details:
- The hardware: NVIDIA said Ultrafast runs on its Blackwell GPUs, with OpenAI’s inference optimizations tapping the architecture’s capabilities directly. The company framed the performance gain as a product of software tuning against that silicon rather than a change in the underlying model family.
- The speed claim: According to NVIDIA, Ultrafast generates tokens up to 8x faster than Astra Standard. The post did not publish benchmark figures, test conditions or a comparison against rival accelerators, and it stated the gain as an upper bound rather than a sustained average.
- Who gets it: Access runs through the OpenAI API from day one, plus eligible ChatGPT Work and Codex users, NVIDIA said. The post did not define the eligibility criteria for those tiers, nor did it name pricing for the faster mode.
- The developer case: NVIDIA said faster generation can shorten the edit-test-debug cycles of coding agents, cut the waiting time between tool calls, and make interactive applications feel more responsive — three workloads where latency, not raw model quality, is the binding constraint.
- OpenAI’s position: Philippe Tillet, inference lead at OpenAI, said NVIDIA’s investment in tooling and documentation has made OpenAI’s models unusually good at programming Blackwell and Rubin GPUs, and that Astra can convert that into high-performance kernels.
- The loop: NVIDIA said OpenAI is using its own models to help refine the inference software running on NVIDIA GPUs — models tuning the stack that serves models. Uday Ruddarraju, OpenAI’s chief technology officer of compute, said the collaboration is making AI faster and more useful.
- The Rubin signal: Tillet’s remarks referenced Rubin alongside Blackwell, pointing past the current generation of NVIDIA silicon. Neither company set out a timeline for moving Astra workloads onto Rubin, and the post named no deployment date for that transition.
- The framing: Tillet said the kernel work makes NVIDIA hardware “compelling across the full frontier of latency, throughput and cost” — a three-axis pitch that positions the gain as an economics argument, not only a speed one.
Between the lines:
The announcement is published by the chip supplier, not the model developer, and the quoted executives are both from OpenAI — a vendor making its customer’s performance the proof of its own architecture. The 8x figure is stated as a ceiling against OpenAI’s own slower mode, not against competing hardware, which keeps the comparison inside the stack both companies control.
What’s next
Watch for independent latency benchmarks from API developers, OpenAI pricing disclosure for the Ultrafast tier, and any timeline for shifting Astra inference onto NVIDIA’s Rubin generation.