Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
TL;DR
OpenAI is previewing Ultrafast, a new API service tier for GPT-5.6 Sol. The mode is meant to deliver responses up to 14 times faster and reaches up to 750 output tokens per second, according to OpenAI. That is made possible by infrastructure from Cerebras. For AI applications with many short interactions, the lower latency could matter more than raw model quality.
Nauti's Take
Up to 750 output tokens per second shifts what is feasible: interactive agent loops and live tooling become usable where latency ruled them out. The catch is the fine print, since preview access, output limits, error rates and cost per request all stay open.
Small teams should measure on their own workflow whether the gain holds for long responses too.