Affiliate disclosure: ToolBistro may earn a commission from some links, at no extra cost to you. Facts come from official sources; we do not publish fabricated testing or ratings.

AI Tools

Qwen 3.8 27B on Cerebras vs Groq: Which Fast Inference API Costs Less?

Qwen 3.8 27B is now on Cerebras at $0.99 per million input and $1.49 per million output tokens, with a listed speed near 1,500 tokens per second, while Groq lists the same Alibaba model at 450 tokens per second for $0.80 in and $4.00 out. Both price pages were verified September 5, 2026.

Key facts

What matters

  • Cerebras added qwen-3.8-27b to public endpoints on September 3, 2026 at $0.99 per million input and $1.49 per million output tokens, listing roughly 1,500 tokens per second.
  • Groq lists the same Qwen3.8-27B at $0.80 input and $4.00 output per million tokens at 450 tokens per second, so Cerebras claims more than 3x the speed at under half the output price.
  • The weights are Apache-2.0 on Hugging Face (6.02 million downloads), so self-hosting is free; Cerebras' free trial caps Qwen 3.8 27B at 5 requests per minute and 1 million tokens per day.
  • Cerebras paid tiers give qwen-3.8-27b a 128K context window and 40K max output tokens; Groq lists about 131K context with a 16,384-token max completion.

What is Qwen 3.8 27B?

Qwen 3.8 27B is a 27-billion-parameter dense multimodal model from Alibaba's Qwen team, published on Hugging Face on August 5, 2026 under the Apache-2.0 license. The model card describes it as designed for agentic coding, tool use, research, and long-running workflows, with text and image input and text output, plus thinking mode that is on by default and adjustable per request.

Per the Hugging Face card, Qwen3.8-27B is a native vision-language model that understands images and video in its model form. Downloads passed 6.02 million and likes reached 13,988 as of September 5, 2026, which makes it one of the fastest-adopted open releases of the month. Alibaba also advertises a hosted Qwen Cloud version with a 1 million-token context, listed as coming soon on the model card.

The reason the model is on this page is commercial: open weights are free, but hosted inference is not, and the two fastest API hosts now serve it at very different prices.

Qwen 3.8 27B pricing on Cerebras

Cerebras added qwen-3.8-27b to its public endpoints on September 3, 2026, per the Cerebras inference change log. The model page, checked September 5, 2026, lists developer pricing of $0.99 per million input tokens and $1.49 per million output tokens, with a listed speed of about 1,500 tokens per second.

The same page splits limits by tier:

  • Free trial: 5 requests per minute, 30K uncached tokens per minute, 90K total tokens per minute, 1 million tokens per day, 64K context, 32K max output, and up to 2 images per request.
  • Developer: 300 requests per minute, 150K uncached tokens per minute, 450K total tokens per minute, no daily cap, 128K context, 40K max output, and up to 10 images per request.

Input images must be base64-encoded PNG or JPEG sent through Chat Completions; video is not supported on the Cerebras shared endpoint. Reasoning defaults to high and can be lowered or disabled with the reasoning_effort parameter, which matters for output-token cost. The Cerebras change log notes that Gemma 4 31B left public endpoints the same day, with Qwen 3.8 27B named as the migration target for shared-tier workloads.

Same model on Groq: price, speed, and limits

Groq's models documentation, checked September 5, 2026, lists the identical weights as qwen/qwen3.8-27b at $0.80 per million input tokens and $4.00 per million output tokens, with a listed speed of 450 tokens per second. On developer-plan limits Groq shows 250K tokens per minute and 1K requests per minute, a context window of 131,042 tokens, a 16,384-token max completion, and image files up to 20 MB.

The contrast is sharpest on output tokens: Groq prices output at 2.7x Cerebras' rate while listing one-third of Cerebras' speed. Groq's older Qwen3.6-27B sits nearby in the same table at $0.60 in and $3.00 out with 500 tokens per second, so the 3.8 generation costs more on Groq even as it lists slower.

Cerebras vs Groq for Qwen 3.8 27B: the price table

The table below pairs the official listings from both vendors for the same open model, all checked September 5, 2026. Speeds are vendor-reported figures, not independent benchmarks.

OpenRouter also lists qwen/qwen3.8-27b at $0.42 per million input and $3.00 per million output tokens with a 1 million-token context, per its public models API on the same date, though speed there depends on whichever upstream host answers a request. For teams that want no per-token bill at all, the Apache-2.0 weights run on any vLLM, SGLang, or Transformers deployment.

Who should pick Cerebras, Groq, OpenRouter, or self-host

  • Cerebras for latency-sensitive agents and chat products: it lists the highest speed at the lowest output price, and its paid tier keeps context at 128K with 40K max output.
  • Groq only if input-heavy traffic dominates and output volume is low, since input is 20% cheaper than Cerebras but output is 2.7x more expensive and capped at 16,384 completion tokens.
  • OpenRouter for multi-model flexibility and the cheapest listed input at $0.42 per million, accepting that endpoint choice and speed are out of your control.
  • Self-hosting for steady, high-volume workloads: the Apache-2.0 license means zero per-token cost on hardware you already run, at the price of ops and GPUs.

Free-tier evaluation is easiest on Cerebras, which gives 1 million tokens per day to try the model before paying anything. The Cerebras model page and the Groq models page are the two primary sources for every number in this comparison; both were reachable and current on September 5, 2026.

The honest catch: what the vendor pages do not say

Four caveats separate the marketing numbers from the bill.

  • Listed speeds are claims, not benchmarks. Cerebras says about 1,500 tokens per second and Groq says 450; neither page publishes the prompt length, hardware, or measurement method behind its figure, so treat the gap as directional.
  • Reasoning tokens are billable. Cerebras defaults Qwen 3.8 27B reasoning to high, and reasoning output counts toward the $1.49 per million rate. Set reasoning_effort to none for cheap, non-agentic calls.
  • Capability differs by host. The HF card advertises video understanding, but Cerebras' shared endpoint accepts only PNG or JPEG images, and Groq caps max completion at 16,384 tokens, below the long outputs the model can produce on Cerebras' 40K tier.
  • The lineup is shifting. Cerebras swapped Gemma 4 31B for Qwen 3.8 27B on September 3, 2026, and Qwen's own hosted cloud version with 1 million-token context is still marked coming soon, so host choice today may not be host choice next quarter.

At a glance

DetailCerebras (qwen-3.8-27b)Groq (qwen/qwen3.8-27b)
Listed speedAbout 1,500 tokens/s450 tokens/s
Input price per 1M tokens$0.99$0.80
Output price per 1M tokens$1.49$4.00
Context window64K free trial / 128K paid131,042 (as listed)
Max output tokens32K free trial / 40K paid16,384
Developer rate limits300 req/min, 450K total tokens/min250K tokens/min, 1K req/min
Image inputPNG/JPEG base64, up to 10 per requestFiles up to 20 MB

FAQ

Is Qwen 3.8 27B free to use?

The model weights are free under the Apache-2.0 license on Hugging Face, so self-hosting costs nothing in licensing. Hosted APIs charge per token: Cerebras lists $0.99 in and $1.49 out per million, Groq $0.80 and $4.00, and OpenRouter $0.42 and $3.00, all as of September 5, 2026.

Which is faster for Qwen 3.8 27B, Cerebras or Groq?

Per the vendors' own pages on September 5, 2026, Cerebras lists about 1,500 tokens per second and Groq lists 450. Both are vendor-reported figures without published benchmark methodology, so compare them as claims rather than measured results.

Does Qwen 3.8 27B support image and video input on these APIs?

The model itself is multimodal, but host support differs. Cerebras accepts base64-encoded PNG or JPEG images through Chat Completions (up to 10 per request on developer tiers) and does not support video on its shared endpoint. Groq accepts image files up to 20 MB. Video understanding is advertised on the Hugging Face model card, not on either API.

Related reading

Radar index, GPT-5.6 Sol on OpenRouter: pricing breakdown, OX Alpha vs DeepSeek: open model API pricing

Sources