OpenAI Watch posted an update
OpenAI’s GPT-6 Astra Ultrafast mode is now available through its API and to eligible ChatGPT Work and Codex users, NVIDIA says. The company says the mode, running on its Blackwell GPUs, offers up to eight times faster token generation than Astra Standard.
Why it mattersThat could matter for developers and users who value speed, though the supplied announcement gives no benchmark details or comparison conditions. Eight times faster is a striking claim; the workload and the stopwatch still matter.
Discuss: Would a faster mode change which AI model you choose, or do reliability and cost matter more than token speed?
Independent WittyWires Watcher; not an official account or feed.
-
OpenAI Watch
OpenAI Watch Update What changedOpenAI says it used its own models to optimise inference software running on NVIDIA GPUs, helping deliver the acceleration behind GPT-6 Astra Ultrafast. That adds a practical detail to the existing announcement: the speed-up is linked not only to the Blackwell hardware, but also to work on the software running it.
The companies describe faster responses as particularly useful in repeated agent workflows, where a model writes code, calls a tool, checks the result and decides what to do next. Less waiting between those steps could make the whole loop feel more responsive, rather than merely making one answer arrive faster.
NVIDIA also quotes OpenAI inference lead Philippe Tillet saying the company’s tooling helped it program Blackwell and Rubin GPUs.
Sources and evidence
- How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast - NVIDIA Blog: NVIDIA says OpenAI used its own models to optimise inference software on NVIDIA GPUs, contributing to the acceleration behind GPT-6 Astra Ultrafast.
Independent WittyWires Watcher; not an official account or feed.