Affiliate disclosure: ToolBistro may earn a commission from some links, at no extra cost to you. Facts come from official sources; we do not publish fabricated testing or ratings.

AI Tools

GPT-5.6 Sol pricing: how much does it cost on OpenRouter vs OpenAI?

GPT-5.6 Sol pricing changed this week: OpenRouter cut the model to $2.50 per million input tokens and $15 per million output tokens, a 50% cut that lands at exactly half of OpenAI's direct API rate of $5 and $30. The discount, live as of August 19, 2026, also halves cache rates: $0.25/M reads and $3.125/M writes.

Key facts

What matters

  • OpenRouter cut gpt-5.6-sol to $2.50 per million input tokens and $15 per million output tokens (50% off), verified on its model page August 19, 2026; Wayback snapshots show the same page at $5/$30 as recently as August 14.
  • The new rate exactly matches OpenAI's own Batch and Flex tier prices ($2.50 input, $0.25 cached input, $3.125 cache writes, $15 output, short context): the cut gives real-time traffic OpenAI's deferred-latency price.
  • OpenAI direct still charges $5.00 input, $0.50 cached input, $6.25 cache writes, and $30.00 output per 1M tokens for standard short-context gpt-5.6-sol calls, with long-context output at $45.
  • The model has a 1,050,000-token context window and a 128,000-token maximum output; it was released July 9, 2026 with a February 2026 knowledge cutoff.

What is GPT-5.6 Sol?

GPT-5.6 Sol is OpenAI's flagship frontier reasoning model, released July 9, 2026, and currently the model behind the daybreak-blue-latest alias on the OpenAI API. It accepts text, images, and PDF files and returns text, with a 1,050,000-token context window and a 128,000-token maximum output, per the OpenRouter model page.

OpenRouter describes it as designed for multi-step coding tasks and long-horizon problem solving. Its knowledge cutoff is February 2026. In OpenAI's lineup it sits above gpt-5.6-terra and gpt-5.6-luna on price and capability.

What exactly changed on OpenRouter this week?

OpenRouter cut gpt-5.6-sol's price by 50%. The model page now carries a 50% off badge and lists $2.50 per million input tokens and $15 per million output tokens. Wayback Machine snapshots of the same URL show $5/$30 per 1M tokens as recently as August 14, 2026, and the halved rate with the discount badge by August 17, 2026 UTC.

The change hit the Hacker News front page on August 17, drawing 616 points and 441 comments. The cut also halves cache pricing: cache reads now cost $0.25 per million tokens and cache writes $3.125 per million. Web Search stays at $10 per 1,000 calls.

OpenRouter vs OpenAI direct: the verified rates

All figures below were verified on the OpenRouter model page and the OpenAI pricing page on August 19, 2026. OpenAI's short-context standard tier is the fair comparison for normal real-time traffic; its Batch and Flex tiers are the discounted options.

The honest catch: this is OpenAI's Batch and Flex price

The number behind the badge deserves a second look. OpenRouter's new rate is exactly what OpenAI itself charges on the Batch and Flex tiers of the direct API: both list gpt-5.6-sol at $2.50 input, $0.25 cached input, $3.125 cache writes, and $15 output for short-context requests.

Batch is OpenAI's asynchronous tier for jobs that do not need a live response, and Flex is its flexible-latency tier; both trade response speed for the discount. OpenRouter sells the same rate for ordinary real-time API calls, which is the real story here: real-time gpt-5.6-sol traffic now runs at OpenAI's deferred-latency price. If your workload tolerates batch timing, OpenAI direct already charges you the same number through the Batch API.

Fast mode is the mirror image at the top: $10 input and $60 output per 1M tokens (short context), double the standard tier, for priority compute requested with service_tier "fast" (the old "priority" name was retired July 30, 2026). Regional processing for data residency adds a 10% uplift on OpenAI direct for models released on or after March 5, 2026.

Who should buy gpt-5.6-sol on OpenRouter vs OpenAI direct?

  • OpenRouter: you want real-time streaming at the batch price without managing batch jobs; you already route traffic through OpenRouter for multi-provider fallback; one API key and one bill across many models matters.
  • OpenAI direct: you run batch workloads (the Batch API already prices gpt-5.6-sol at $2.50/$15); you need Fast mode, first-party SLAs, data-residency regional endpoints, or an enterprise agreement.

A workable default: real-time traffic goes to OpenRouter at the cut price, and anything that can wait uses OpenAI's Batch API at the same number with first-party handling. If you are still choosing between Sol and the smaller family models, our Sol vs Luna breakdown covers that decision.

How Sol's new price fits the lineup

On OpenAI's rate card, the gpt-5.6 family ladder is luna at $0.20 input / $1.20 output, terra at $2 / $12, and sol at $5 / $30 (standard, short context). OpenRouter's cut narrows the gap between Sol and Terra for real-time traffic: Sol now costs $2.50 / $15 there, against Terra's $2 / $12 on OpenAI direct.

For cross-vendor context, Anthropic's pricing page lists Claude Opus 5 at $5 input / $25 output per million tokens and Claude Fable 5 at $10 / $50. Sol at OpenRouter's halved rate undercuts both on input price, and its 1.05M context window is the largest of the group.

Where the $2.50 rate is available right now

OpenRouter: select the openai/gpt-5.6-sol model on the platform; the $2.50 input / $15 output rate applies to real-time calls, with cache reads at $0.25 and cache writes at $3.125 per 1M tokens (verified August 19, 2026).

OpenAI direct: the identical numbers appear on the Batch and Flex tiers of the API rate card, both at $2.50 input / $15 output for short-context requests. Fast mode is the exception at the top: $10 input / $60 output, requested with service_tier "fast" (the retired name "priority" still works).

Context size changes the math. OpenAI's rate card splits short and long context, with long-context standard output at $45 per 1M tokens, while OpenRouter lists one flat rate for its full 1,050,000-token window. Workloads that push close to the context limit save the most on OpenRouter.

At a glance

RateOpenRouter (Aug 19)OpenAI direct standardOpenAI Batch / Flex
Input$2.50$5.00$2.50
Cached input$0.25$0.50$0.25
Cache writes$3.125$6.25$3.125
Output$15.00$30.00$15.00
Long-context output (1M window)$15.00 (flat rate)$45.00$22.50
Web search$10.00 per 1K callsNot listedNot listed

FAQ

Is the 50% OpenRouter discount permanent?

OpenRouter has not announced an end date. The 50% off badge and the $2.50/$15 rate were live on its model page as of August 19, 2026, after the page listed $5/$30 through August 14. Treat it as the current list price until OpenRouter says otherwise.

Is gpt-5.6-sol cheaper on OpenRouter than on the OpenAI API?

For standard real-time calls, yes: $2.50/$15 per 1M tokens vs $5/$30 direct (short context, verified August 19, 2026). It matches OpenAI's own Batch and Flex tier price, so workloads that tolerate batch latency can already pay the same rate directly via the Batch API.

Does gpt-5.6-sol support a 1M context window?

Yes: the OpenRouter listing shows a 1,050,000-token context window with a 128,000-token maximum output. On OpenAI direct, long-context requests carry higher rates (standard $10 input / $45 output per 1M tokens), while OpenRouter quotes one flat rate for the full window.

Related reading

GPT-5.6 Sol vs Luna, GPT-5.6 Sol Ultrafast pricing, OpenRouter alternatives, Claude review and pricing

Sources