Sign InOpen Brain
OpenAIOfficial ReleaseOfficial Source

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

OpenAI’s preview tier runs GPT-5.6 Sol at up to 750 output tokens per second, targeting agent workloads where generation latency is the bottleneck.

OpenAI · Aug 13, 2026
Open Source Open MarkdownOpen JSON
Source Summary

OpenAI is previewing **Ultrafast**, an API service tier for **GPT-5.6 Sol** powered by Cerebras. It claims up to **14× faster** operation and up to **750 output tokens per second**.

Practical Implication

Builders with latency-bound coding agents should test whether faster generation materially shortens end-to-end runs, especially when tool execution is already quick.

Agent-Ready Context
OpenAI is previewing **Ultrafast**, an API service tier for **GPT-5.6 Sol** powered by Cerebras. It claims up to **14× faster** operation and up to **750 output tokens per second**.

Builders with latency-bound coding agents should test whether faster generation materially shortens end-to-end runs, especially when tool execution is already quick.

Both figures are qualified maxima. The material gives no pricing, availability, workload methodology, or comparison baseline, and tool latency may still dominate many runs.
Connected Context · Feed7 Judgment

Ultrafast adds a model-specific serving option whose claimed ceiling is substantially more concrete than the gateway’s generic fast-mode control, but it does not establish faster completed agent work. It narrows evaluation to latency-bound runs where generation is material, with tool time, price, access, and workload-level outcomes still unresolved.

Context Map
infracoding#coding-agents
Uncertainty
Both figures are qualified maxima. The material gives no pricing, availability, workload methodology, or comparison baseline, and tool latency may still dominate many runs.