Affiliate disclosure: ToolBistro may earn a commission from some links, at no extra cost to you. Facts come from official sources; we do not publish fabricated testing or ratings.

GPU Cloud

What Is the Cheapest GPU Cloud for LLM Inference in 2026?

The cheapest GPU cloud for LLM inference is Vast.ai at the floor: an H100 SXM 80GB is listed from $1.73/hr there, against $2.69/hr on RunPod's Community Cloud and $4.29/hr for a single-GPU H100 instance on Lambda. Median prices tell a different story, so here are the rates verified on each vendor's own pricing page on September 15, 2026.

Key facts

What matters

  • Vast.ai's $1.73/hr H100 SXM headline is the lowest live offer on its marketplace; the same page reports a 30-day daily median of $2.27/hr, which is the number to budget against.
  • Lambda cuts the per-GPU rate as the instance grows: H100 SXM runs $4.29/hr on a 1x instance and $3.99/hr on an 8x instance, per its GPU cloud page captured 2026-09-15.
  • RunPod prices the same H100 SXM at $3.49/hr on Secure Cloud and $2.69/hr on Community Cloud; its docs say Secure runs in T3/T4 data centers while Community pools vetted individual providers.
  • RunPod bills Pods by the minute with no ingress or egress fees, and its H100 serverless worker costs $4.79/hr against $3.49/hr for the equivalent Pod.
  • Vast bills per second with no minimum hours, which suits short experiments, but its per-SKU availability is limited: H100 SXM was marked 'Med' when captured.

Cheapest GPU cloud rates, verified on September 15, 2026

This question is live this month. The Economist's argument that Nvidia has become the central bank of AI held the top of Hacker News on September 12 with 580 points and 399 comments, and every version of that argument ends at the same buyer question: what does an hour of H100 actually cost? None of the three major rental clouds publishes a table comparing itself to the other two, so the one below was assembled from their own pricing pages.

Each figure is the list rate for the GPU alone. Bundled vCPU, RAM and SSD differ per vendor, so an hourly rate is not the whole unit price: Lambda bundles RAM and storage into the instance, RunPod bundles RAM into the Pod, and Vast lets hosts differentiate on the rest of the machine.

Two display details are worth knowing before you read the numbers. RunPod's pricing page defaults to Secure Cloud, the pricier of its two columns, so the cheaper Community rate is easy to miss. Lambda hides its price curve behind an 8x / 4x / 2x / 1x toggle, and the per-GPU rate changes with the selection.

The instance-size trap in Lambda's pricing

Lambda charges per GPU per hour, and the rate falls as the instance gets bigger. On the H100 SXM, the page shows $4.29 per GPU per hour at 1x, $4.19 at 2x, $4.09 at 4x and $3.99 at 8x. The B200 repeats the pattern: $6.99 at 1x down to $6.69 at 8x.

What you get for that money scales too. A 1x H100 SXM instance includes 26 vCPUs, 225 GiB of RAM and 2.75 TiB of SSD; the 8x version carries 208 vCPUs, 1800 GiB of RAM and 22 TiB of SSD. Lambda states pay-by-the-minute billing with no egress fees, and the rate table carries a footnote that prices exclude applicable sales tax, VAT or GST.

Read that as a small-deployment penalty. If you need exactly one H100 for an inference endpoint, Lambda is the most expensive of the three per GPU, and its discount only arrives at four or eight GPUs. The single-GPU table also shows an H100 PCIe at $3.29/hr and a GH200 at $2.29/hr, so the cheapest Lambda card for a given job is not always the headline model.

Vast.ai's $1.73 H100 is a floor, not a typical rate

Vast's H100 SXM page advertises GPUs 'from $1.73/hr' and, on the same page, reports a 30-day daily median of $2.27/hr. Both numbers come from the vendor, refresh hourly, and the distance between them is the whole experience of shopping a marketplace: the floor is one willing host, the median is what the platform usually clears at.

Vast describes the model as supply and demand pricing set by the market rather than by Vast, with per-second billing, no minimum hours, no rounding up, 68+ GPU types and no lock-in contracts. It also publishes availability per SKU, and H100 SXM was marked 'Med' at capture time.

For batch jobs that tolerate a restart, that is a genuine discount over fixed list prices. For an endpoint that has to stay up, plan against the median and treat the floor as best case. Vast's consumer-class cards move even further: RTX 4090 offers started at $0.14/hr against a $0.49/hr median, and the RTX 3090 ran $0.08/hr against a $0.16/hr median.

RunPod Community vs Secure: what the discount costs

RunPod splits its inventory, and its documentation is explicit about the difference. Secure Cloud operates in T3/T4 data centers aimed at enterprise and production workloads. Community Cloud connects individual compute providers to users through a vetted peer-to-peer system at lower prices.

The spread is real on the SKUs buyers actually rent. On the pricing page captured for this article, H100 SXM listed at $3.49/hr on Secure versus $2.69/hr on Community, H200 at $4.59/hr versus $3.59/hr, B200 at $6.79/hr versus $5.98/hr, and the RTX 4090 at $0.74/hr versus $0.34/hr. An A100 80GB SXM was $1.59/hr on Secure against $1.39/hr on Community. The same docs state Pods bill by the minute with no ingress or egress fees.

RunPod also rents serverless workers, and they are not the cheap option: its pricing page lists an H100 serverless worker at $4.79/hr and an A100 at $2.72/hr, against $3.49/hr for the equivalent H100 Pod. That premium buys autoscaling instead of an idle GPU you keep paying for.

What an H100 hour actually buys you

The H100 SXM is the SKU all three vendors price, so it is the fair unit of comparison. Vast's spec sheet describes the GH100 die: Hopper architecture, 5 nm process, 80 GB of HBM3 memory, 3.36 TB/s of memory bandwidth, 528 tensor cores, released March 21, 2023, with a 1590 MHz base clock and 1980 MHz boost clock.

Those 80 GB set the ceiling on what fits. Vast lists model-fit guidance per SKU, and its H100 SXM page names 30B-class multilingual and reasoning models such as Qwen3.8 27B, Gemma 4 31B IT and Muse Glimmer 30B as candidates, with the caveat that requirements vary by engine, precision and configuration.

Shared-model inference, the kind bought per token, is the other half of this decision. If your traffic is spiky or you would rather not run an endpoint, the per-token route is covered in our round-up of OpenRouter alternatives. Rent hourly when utilization is high and predictable; pay per token when it is not.

Which GPU cloud should you pick for LLM inference

  • One GPU, tightest possible budget: Vast.ai, budgeted at the 30-day median rather than the floor, or RunPod Community Cloud if you prefer a fixed list rate.
  • A production endpoint you answer for: RunPod Secure Cloud or Lambda, both of which place GPUs in managed data centers.
  • Multi-GPU fine-tuning or training: Lambda, where the per-GPU rate drops at 4x and 8x and the 8x H100 instance bundles 22 TiB of SSD.
  • Short, unattended experiments: Vast's per-second billing avoids paying for the tail of an hour; RunPod's serverless workers cost more per hour but remove idle time entirely.
  • Long-running agents that mostly wait: hourly GPU rental is the wrong shape. A plain VPS is cheaper for orchestration, which is why we compared VPS options for AI agents separately from GPU clouds.

The catches that never make the headline

  • A live minimum can vanish. Vast's floor price reflects currently available offers and updates hourly; the median is the budgeting figure.
  • The cheaper RunPod column is hidden by default. The page opens on Secure Cloud, so quote Community prices only after clicking through.
  • Lambda's rate depends on instance size. The $3.99 H100 figure applies at 8x, and every rate on that table excludes tax.
  • Serverless is a convenience premium. RunPod's H100 worker is $4.79/hr against a $3.49/hr Pod.
  • Availability is a real constraint. Vast marked H100 SXM availability as 'Med', and Lambda sells 1x through 8x instances on a self-serve, first-come basis.
  • Every rate here is a snapshot. All figures were captured from the vendors' pages on September 15, 2026 and marketplaces move hourly, so re-check before you commit.

At a glance

GPU (VRAM)RunPod CommunityRunPod SecureLambda CloudVast.ai (lowest live offer)
NVIDIA H100 SXM 80GB$2.69/hr$3.49/hr$4.29/hr at 1x, $3.99/hr at 8x$1.73/hr, 30-day median $2.27/hr
NVIDIA H200 141GB$3.59/hr$4.59/hrNot listed on the GPU cloud page$1.97/hr, 30-day median $4.00/hr
NVIDIA B200 180GB$5.98/hr$6.79/hr$6.99/hr at 1x, $6.69/hr at 8x$5.63/hr, 30-day median $7.75/hr
NVIDIA RTX 4090 24GB$0.34/hr$0.74/hrNot offered$0.14/hr, 30-day median $0.49/hr

FAQ

Is Vast.ai cheaper than RunPod for H100 GPUs?

At the floor, yes: Vast lists H100 SXM from $1.73/hr while RunPod lists $2.69/hr on Community Cloud and $3.49/hr on Secure Cloud. Vast's own 30-day daily median is $2.27/hr and its offers refresh hourly, whereas RunPod rates are fixed list prices billed by the minute.

Does Lambda charge per GPU or per instance?

Per GPU per hour, and the rate depends on instance size. Its GPU cloud page shows H100 SXM at $4.29 per GPU per hour on a 1x instance and $3.99 at 8x, with a footnote that listed prices exclude applicable sales tax, VAT or GST.

What is the difference between RunPod Community Cloud and Secure Cloud?

RunPod's docs say Secure Cloud runs in T3/T4 data centers for enterprise and production workloads, while Community Cloud pools vetted individual compute providers via a peer-to-peer system at lower prices. The gap is roughly 20 to 50 percent depending on the GPU, for example $3.49/hr versus $2.69/hr on the H100 SXM.

Related reading

VPS options for AI agents, OpenRouter alternatives, best VPS for local LLMs

Sources