Affiliate disclosure: ToolBistro may earn a commission from some links, at no extra cost to you. Facts come from official sources; we do not publish fabricated testing or ratings.

AI Tools

Ox Alpha vs DeepSeek: which coding model should you pay for?

Ox Alpha vs DeepSeek: Ox Alpha is Z.ai's GLM-5.3-Flash, confirmed on August 26, 2026 after the model was tested anonymously as "ox-alpha" on OpenCode and OpenRouter. Z.ai's API prices it at $0.15 per million input tokens ($0.50 output), undercutting DeepSeek V4 Flash's $0.22 and $0.66 off-peak cache-miss rates.

Key facts

What matters

  • Ox Alpha is Z.ai's GLM-5.3-Flash: confirmed August 26, 2026, after the model ran anonymously on OpenCode and OpenRouter for a week.
  • GLM-5.3-Flash API: $0.15 input and $0.50 output per 1M tokens, limited-time half price at $0.075 and $0.25.
  • DeepSeek V4 Flash API: $0.22 input (cache miss) and $0.66 output per 1M at off-peak hours; cached input drops to $0.007.
  • Both are open-weights models: GLM-5.3-Flash (MIT, 320B total, 18B active) and DeepSeek V4 are public on HuggingFace.
  • Free to try at oxalpha.com with no login; Z.ai's own benchmarks put GLM-5.3-Flash ahead of DeepSeek's vision model on DeepSWE (63.4 vs 59.3).

What is Ox Alpha?

Ox Alpha is the code name for GLM-5.3-Flash, a 320B-parameter reasoning model from Z.ai with only 18B active parameters and a 1M-token context window. It is the first natively multimodal model in the GLM-5 series, accepting text, image, and video input with up to 131K output tokens.

Z.ai's launch post on August 26, 2026 confirmed the identity directly: "Before release, we tested GLM-5.3-Flash anonymously as 'ox-alpha' on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week, with all of this traffic served on Chinese AI chips."

The architecture pairs sparse and linear attention to cut long-context serving cost, and the weights are public under an MIT license on HuggingFace. For self-hosting, Z.ai lists support for SGLang, vLLM, and TokenSpeed.

Ox Alpha vs DeepSeek: verified pricing

Z.ai's official pricing page (checked August 27, 2026) lists GLM-5.3-Flash at $0.15 per 1M input tokens and $0.50 per 1M output tokens, with a limited-time 50% discount to $0.075 and $0.25. Cached input is $0.03 per 1M, currently half price too. That undercuts DeepSeek V4 Flash, whose official API charges $0.22 per 1M input tokens (cache miss) and $0.66 per 1M output at off-peak hours.

DeepSeek's off-peak window covers every hour except 01:00-04:00 and 06:00-10:00 UTC on weekdays, and off-peak rates are exactly half of peak. Its cached-input price of $0.007 per 1M is the cheapest lever in this comparison for long agent sessions that reuse context.

Benchmarks: what Z.ai reports

Z.ai's launch post reports GLM-5.3-Flash at 63.4 on DeepSWE v1.1, versus 46.2 for GLM-5.2 and 59.3 for DeepSeek-V4-Vision-Exp, and 84.3 on Terminal Bench 2.1 versus 83.9 for the DeepSeek model. On Z.ai's internal Code Bench v1.0 at max effort it scores 29.0, close to Claude Opus 4.8's 29.5.

The oxalpha.com site adds an independent community benchmark: 8 of 10 real-world coding tasks solved, versus 65% for Fable 5, 62% for GLM-5.3, 62% for Grok 4.6, and 52% for GPT-5.6 Sol. The site itself flags this as third-party data with a small sample: directional, not definitive.

One cost-efficiency claim stands out: the Artificial Analysis Intelligence Index v4.1.1 scores GLM-5.3-Flash at 57 with a $0.045 cost per task (discounted), a price-performance point Z.ai says previously cost roughly 10x more. Treat the benchmark numbers as vendor-reported until independent suites replicate them.

Where to use Ox Alpha (GLM-5.3-Flash) right now

  • Free: oxalpha.com runs a no-login chat with the full 1M context window.
  • API: Z.ai's API at $0.15/$0.50 per 1M tokens, and OpenRouter lists z-ai/glm-5.3-flash at $0.075/$0.25 during the discount.
  • Subscription: the GLM Coding Plan starts at $18/month for Lite (10,000 credits per week), with Pro at $80/month and Max at $168/month; the launch post says Flash gets 3x the usable quota of GLM-5.3.
  • Agents: ZCode adds Browser Use and Computer Use on top of the multimodal input.
  • Self-host: MIT-licensed weights on HuggingFace (zai-org/GLM-5.3-Flash), with SGLang, vLLM, and TokenSpeed support.

Who should choose Ox Alpha over DeepSeek

Pick GLM-5.3-Flash when you want open weights with a permissive MIT license, native image and video input for UI coding, or the cheapest per-token multimodal frontier model. The free chat tier makes it the lowest-risk model to evaluate in a single afternoon.

Stick with DeepSeek when your workload is token-heavy and context-cached: V4 Flash cached input at $0.007 per 1M is about four times cheaper than GLM-5.3-Flash's $0.03, and off-peak batch work at $0.22/$0.66 is hard to beat. DeepSeek's API also offers thinking and non-thinking modes with a 384K max output, and its V4 weights are public as well.

Neither model has years of independent production evidence. Budget apps and agent prototypes fit both; mission-critical pipelines still warrant a paid tier with support, whichever you choose.

The honest catch

The limited-time 50% discount is doing real work in this comparison: at the standard $0.15/$0.50, GLM-5.3-Flash still undercuts DeepSeek V4 Flash, but the gap narrows. The launch was deliberately stealthy, so no one knew the vendor until Bloomberg and Z.ai confirmed the GLM-series identity on August 26. OpenRouter still serves a stealth/ox-alpha URL, but the model now appears in its catalog as z-ai/glm-5.3-flash.

Benchmark claims are Z.ai's own, and the flagship GLM-5.3 remains about 9x more expensive at $1.40/$4.40, so check which model a coding plan actually routes your traffic to. DeepSeek reserves the right to adjust prices, and its rate table already splits peak and off-peak, so re-verify both vendors before committing.

At a glance

Ox Alpha (GLM-5.3-Flash)DeepSeek V4 FlashDeepSeek V4 Pro
MakerZ.ai, GLM seriesDeepSeekDeepSeek
API input per 1M tokens$0.15 (limited-time $0.075)$0.22 off-peak, $0.44 peak (cache miss)$0.66 off-peak, $1.32 peak
API output per 1M tokens$0.50 (limited-time $0.25)$0.66 off-peak, $1.32 peak$1.98 off-peak, $3.96 peak
Cached input per 1M tokens$0.03 (limited-time $0.015)$0.007 off-peak, $0.014 peak$0.022 off-peak, $0.044 peak
Context window1M tokens1M tokens1M tokens
Max output131K tokens384K tokens384K tokens
Open weightsYes, MIT on HuggingFaceYes, on HuggingFaceYes, on HuggingFace
Free way to tryoxalpha.com free chatNo free API tierNo free API tier

FAQ

Is Ox Alpha free?

Yes for casual use: oxalpha.com offers a free, no-login chat with a 1M-token context. API access costs $0.15 per 1M input tokens (limited-time $0.075) plus $0.50 per 1M output, and the GLM Coding Plan starts at $18 per month.

Is Ox Alpha the same as GLM-5.3-Flash?

Yes. Z.ai confirmed on August 26, 2026 that Ox Alpha is GLM-5.3-Flash, a 320B-parameter GLM-5-series model with 18B active parameters, tested anonymously as 'ox-alpha' on OpenCode and OpenRouter before release.

Does Ox Alpha beat DeepSeek?

On Z.ai's own benchmarks, GLM-5.3-Flash edges DeepSeek's comparable model: 63.4 vs 59.3 on DeepSWE v1.1 and 84.3 vs 83.9 on Terminal Bench 2.1, at a lower per-token price. Those numbers are vendor-reported, and DeepSeek's cached-input pricing stays cheaper for long agent sessions.

Related reading

GLM-5.3 pricing, DeepSeek API pricing, Claude AI tool profile, radar index

Sources