OpenAI has introduced an ‘Ultrafast’ service tier for its GPT-5.6 Sol model, delivering performance enhancements of up to 14 times the standard processing speed. This advancement is particularly significant for applications where low latency is crucial, such as real-time voice interactions, customer support, e-commerce, developer tools, financial analysis, and cybersecurity operations.
The Ultrafast mode is initially available through the OpenAI API and is powered by Cerebras hardware, enabling the generation of up to 750 output tokens per second. This substantial increase in processing speed aims to bring high-level AI performance to scenarios where rapid response times are essential.
OpenAI’s internal teams have already leveraged this accelerated capability to analyze logs and traces during incidents, as well as to expedite research cycles that previously required overnight processing, now completing multiple iterations within a single workday.
Access to the Ultrafast service is currently limited to a select group of customers as OpenAI assesses its impact on real-world applications and scales capacity accordingly. Businesses interested in utilizing this enhanced performance can join the waitlist by providing details about their workloads, latency requirements, and expected usage.
In June, OpenAI unveiled the GPT-5.6 family, including the flagship Sol model, alongside the balanced Terra and speed-focused Luna models. The full suite became broadly available in July through platforms such as ChatGPT, Codex, and the OpenAI API.
This development underscores OpenAI’s commitment to advancing AI capabilities, particularly in reducing latency to meet the demands of time-sensitive applications. As AI continues to integrate into various sectors, the ability to process information rapidly without compromising accuracy becomes increasingly vital. The introduction of the Ultrafast tier positions OpenAI to better serve industries where speed and efficiency are paramount.