Computing power surges 14-fold! OpenAI officially launches GPT-5.6 Sol’s “Ultra-Fast” mode.
OpenAI officially launches the GPT-5.6 Sol “Ultra-Fast” mode, with peak output speeds of up to 750 tokens per second and overall processing speeds up to 14 times faster than the standard mode. Built on Cerebras wafer-scale engine architecture, the mode is designed for high-intensity scenarios such as incident response, financial research, real-time customer service, programming, and complex research.

OpenAI has officially introduced GPT-5.6 Sol in a new “ultra-fast” mode, moving the model beyond its earlier limited preview phase. The version had first been made available internally to a small group of selected customers in July, but its wider debut is now drawing attention because of its unusually high throughput.
According to public details, GPT-5.6 Sol’s ultra-fast mode can generate output at speeds of up to 750 tokens per second. OpenAI says this boost does not require users to fall back to a smaller or less capable model. Instead, the faster setting can deliver as much as a 14x increase in overall processing speed compared with the standard mode.
What is driving the speed increase
The main reason behind this jump in performance lies in infrastructure rather than a simple software tweak. OpenAI attributes the improvement to support from Cerebras and its wafer-scale engine architecture. This design includes 44GB of on-chip SRAM, which helps reduce the data-transfer bottlenecks that often limit more conventional GPU-based systems.
By keeping more data closer to the compute layer, the system can move through inference workloads far more efficiently. That architectural advantage appears to be the foundation for the model’s much higher response rate.
Benchmark performance
Test results referenced in the announcement suggest that the new mode does more than just improve raw generation speed. It is also said to outperform comparable competing offerings by a wide margin in runtime efficiency.
One of the more notable figures shared so far involves the so-called “Humanity’s Last Exam” benchmark. In that test, GPT-5.6 Sol’s ultra-fast mode reportedly processed all 2,500 questions in about 11 hours, highlighting the model’s ability to handle large-scale workloads quickly.
Workflows it is designed for
OpenAI positions this mode for tasks where low latency and sustained heavy processing matter most. The company specifically points to several use cases where the faster configuration could be especially valuable:
incident response
financial research
real-time customer support
programming
complex research workflows
In practical terms, that means the model is aimed at environments where waiting for output can slow down decision-making or interrupt a larger chain of work.
Access will expand gradually
Despite the formal launch, OpenAI is not opening the feature to everyone at once. The company says computing resources remain limited, so access will be granted progressively rather than broadly from day one.
OpenAI plans to evaluate customers based on how well their workloads match the strengths of the new mode, while also factoring in available compute capacity. In other words, adoption is expected to expand in stages as resources allow.
With GPT-5.6 Sol’s ultra-fast mode, OpenAI is clearly signaling a push toward higher-performance AI systems that can handle demanding tasks without compromising model capability. The release also underscores how much model speed is now tied to advances in underlying hardware architecture, not just improvements in the models themselves.