Cloud & Compute Infrastructure
OpenAI partnered with Cerebras for GPT-5.6 Sol UltraFast inference, reaching up to 750 tokens per second for initial customers.
· ComputeLabs Research · from the August 13, 2026 edition
OpenAI partnered with Cerebras to provide low-latency inference for GPT-5.6 Sol through an UltraFast mode. The service was released as a limited preview to an initial group of customers rather than as a general-availability product.
OpenAI said the UltraFast tier can run GPT-5.6 Sol at speeds of up to 750 tokens per second. Another announcement described the preview as delivering speeds of up to 14 times the comparison baseline, although the source does not define that baseline.
The stated performance is an “up to” inference-throughput figure, not a guaranteed speed for every prompt, model context, or customer workload. The sources do not disclose pricing, Cerebras system counts, data-center locations, service-level commitments, or the size of the initial customer group.
Additional reporting
- financialjuice(t.me)
- financialjuice(t.me)
- financialjuice(t.me)
- financialjuice(t.me)
- Financial_Express(t.me)
- OpenAI
- Cerebras
- GPT-5.6 Sol UltraFast

