AI Tools
What is GPT-5.6 Sol Ultrafast and how much does it cost?
GPT-5.6 Sol Ultrafast is a new OpenAI API tier powered by Cerebras hardware that generates up to 750 output tokens per second, but it has no public price and is currently restricted to a select group of customers. Standard gpt-5.6-sol API access costs $5 per 1M input and $30 per 1M output tokens.
What matters
- GPT-5.6 Sol Ultrafast is a new OpenAI API service tier powered by Cerebras hardware, announced on August 13, 2026.
- It outputs up to 750 tokens per second, but pricing is not public and access is limited to a select group of customers.
- Standard gpt-5.6-sol API pricing is $5.00 per 1M input and $30.00 per 1M output tokens (short context, verified August 15, 2026).
- Fast mode, the fastest tier you can request today, costs $10.00 per 1M input and $60.00 per 1M output.
- Cerebras reports Ultrafast runs 11x faster than Fable 5 and 5x faster than Opus 4.8 on Fast mode, based on Artificial Analysis output speeds.
What is GPT-5.6 Sol Ultrafast?
GPT-5.6 Sol Ultrafast is a new service tier in the OpenAI API, powered by Cerebras hardware, that serves OpenAI's gpt-5.6-sol model at up to 750 output tokens per second. Cerebras and OpenAI announced it on August 13, 2026, and the Cerebras announcement calls GPT-5.6-Sol-Ultrafast "the world's fastest frontier model."
Per the announcement, Ultrafast Mode is "launching first in the OpenAI API and powered by Cerebras." It is available initially to "a select group of customers, with access expanding over time." OpenAI's preview page and a Cerebras early-access request form are both live, but the tier is not available to all API users yet.
gpt-5.6-sol is OpenAI's current flagship model: the pricing page notes the alias daybreak-blue-latest currently points to gpt-5.6-sol, and daybreak-red-latest points to gpt-5.6-cyber.
How fast is GPT-5.6 Sol Ultrafast?
The headline number is 750 output tokens per second, which Cerebras says comes "without any quality compromise." For context, the company compared Ultrafast with output speeds reported by Artificial Analysis: GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5 and 5x faster than Opus 4.8 on Fast mode.
On Humanity's Last Exam, a 2,500-question benchmark aimed at PhD-level knowledge in fields like chemistry and economics, Cerebras reports GPT-5.6 Sol Ultrafast worked through the exam "achieving comparable accuracy nearly 7x faster." The evaluation used GPT-5.6 Sol Ultrafast with Codex on xhigh reasoning (July 10) against Claude Fable 5 with Claude Code on xhigh reasoning (July 13-15). These are vendor-run numbers, not independent measurements, and they compare against models measured by a third-party tracker, so treat them as directional.
The practical pitch: real-time agents, live coding assistants, and latency-sensitive products get frontier-model quality without a multi-second wait.
GPT-5.6 Sol API pricing: all four tiers, verified
Ultrafast has no public price, but the rest of the gpt-5.6-sol rate card is published on OpenAI's pricing page and was verified on August 15, 2026. All prices below are per 1M tokens, short context.
OpenAI splits the card into short-context and long-context columns, and the two columns diverge sharply. Standard output is $30.00 per 1M for short context and $45.00 for long; Fast mode output is $60.00 short and $90.00 long. Batch and Flex both sit at half of standard pricing, $15.00 output per 1M. Regional processing endpoints add a 10% uplift for eligible models released on or after March 5, 2026.
Fast mode was renamed from priority processing on July 30, 2026. You can request it today with service_tier: fast (the legacy service_tier: priority still works) in your API calls.
Ultrafast pricing and access: the honest catch
There is no public price for GPT-5.6 Sol Ultrafast. OpenAI has not published a rate card, and the Cerebras early-access form only asks what you are building and your expected workload. "Select group" means enterprise and approved partners get in first, and a sales conversation is more likely than a checkout page.
Two things to plan around. First, the speed claims (750 tokens per second, 11x, 7x) come from Cerebras' own evaluation plus Artificial Analysis comparisons, not from a benchmark you can reproduce today. Second, the tier lives inside the OpenAI API, so you pay OpenAI per token; Cerebras does not sell gpt-5.6-sol access separately through its own inference API.
If your product needs this speed before general availability, the request form is the only door, and there is no stated timeline for wider rollout.
Which tier should you use today?
Until Ultrafast opens up, the decision is between the four published tiers:
- Standard ($5.00 in / $30.00 out): the default. Fine for chat apps and offline analysis where a few seconds of thinking time is acceptable.
- Fast mode ($10.00 in / $60.00 out): double the price for priority compute. The only published tier you can switch to today for lower latency.
- Batch ($2.50 in / $15.00 out): half price for asynchronous jobs that do not need a live response.
- Flex ($2.50 in / $15.00 out): the same discount as Batch on the current rate card, for workloads with flexible deadlines.
For a real-time agent product, Fast mode is the honest baseline for what speed costs right now, and $60.00 per 1M output is the number to compare against Ultrafast when its pricing appears.
What this means for frontier API buyers
The Ultrafast launch is a supply-side story: a chipmaker's inference stack now powers an OpenAI service tier, and OpenAI is reselling that capacity inside its own API. For buyers, the pattern is that frontier inference is splitting into speed tiers with widening price gaps, and the fastest tier is the least transparently priced one.
Alternatives still matter. Claude (Fable 5 and Opus 4.8 are the models Cerebras benchmarked against) and DeepSeek compete on price per token; DeepSeek's API moved to peak and off-peak pricing on August 16, 2026, with off-peak rates 50% below peak. None of them currently publishes a Cerebras-style ultrafast tier with a public rate card, which is exactly why the Ultrafast price gap is worth watching.
At a glance
| Service tier | Input | Cached input | Cache writes | Output | Notes |
|---|---|---|---|---|---|
| Standard | $5.00 | $0.50 | $6.25 | $30.00 | Default tier for API requests |
| Fast mode | $10.00 | $1.00 | $12.50 | $60.00 | 2x standard; request with service_tier: fast |
| Batch | $2.50 | $0.25 | $3.125 | $15.00 | 50% off standard |
| Flex | $2.50 | $0.25 | $3.125 | $15.00 | 50% off standard |
| Ultrafast | Not public | Not public | Not public | Not public | Cerebras-powered; up to 750 output tokens/sec; select group |
FAQ
Is GPT-5.6 Sol Ultrafast free?
No free tier has been announced. Pricing is not public, and access is by request to a select group of customers. Cerebras' early-access form asks for your use case and expected workload, which points to negotiated enterprise pricing rather than a free tier.
When will GPT-5.6 Sol Ultrafast be generally available?
OpenAI and Cerebras have not announced a date. The August 13, 2026 announcement says access expands as capacity grows. Requesting early access through Cerebras is the only path today.
What is the difference between Ultrafast and Fast mode?
Fast mode is a published OpenAI service tier you can request today with service_tier: fast, priced at $10.00 input and $60.00 output per 1M tokens (short context). Ultrafast is a separate Cerebras-powered tier with no public price, up to 750 output tokens per second, limited to select customers.
Related reading
GPT-5.6 Sol vs Luna, DeepSeek API pricing, Claude tool profile, ToolBistro radar index