BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/GPU Cost: The Same H100 at $0.35 or $14.90 an Hour
ToolsAugust 24, 2026
Read · 5 min
gpu cost · gpu cloud pricing

GPU Cost: The Same H100 at $0.35 or $14.90 an Hour

GPU cost varies more than 40 times for the same H100. The four things that explain the spread, and the utilisation where buying finally beats renting.

Key takeaways
  • The same H100 rents for as little as $0.35 an hour on spot capacity and as much as $14.90 an hour on demand at a hyperscaler, a spread of more than 40 times for identical silicon.
  • Four things explain almost all of it: the interconnect, the commitment model, whether storage and egress are billed on top, and whether the quote is per GPU or per instance.
  • The single most common misreading is a per instance price mistaken for a per GPU one. An eight GPU box at $50 an hour is not expensive, it is $6.25 a GPU.
  • Buying an H100 only beats renting above a level of use most teams never sustain. At list purchase prices near $25,000 a card, the break even needs close to full time utilisation.
  • Egress and idle time routinely cost more than the compute line, so the hourly rate you compared is rarely the bill you pay.

A technical founder pricing GPU capacity for the first time finds the same chip, the NVIDIA H100, quoted at $0.35 an hour in one place and $14.90 an hour in another. That is not a typo and it is not two different products. As reported in August 2026 by the price tracker GetDeploying, those are the low and high ends of the real market for one H100, a spread of more than 40 times for hardware that is physically identical. The founder's job is to understand what that spread is made of, because picking the wrong end of it can multiply an infrastructure bill by an order of magnitude for no benefit, and picking the cheap end blindly can cost you a training run that never finishes.

The spread is not noise. Each factor in it corresponds to something you are or are not buying, and once you can read the four factors you can place any quote on the map and know whether it is cheap for a reason you care about or expensive for one you do not.

It matters because GPUs are usually the largest single line in an AI budget, so a decision made once at the start compounds every hour the hardware runs. A founder who defaults to a hyperscaler because the brand feels safe can pay double what a specialist cloud charges for the identical card, month after month, and a founder who chases the cheapest spot rate for a production service can save money right up until the capacity is reclaimed in the middle of a demo. Neither the top nor the bottom of the spread is wrong, they are answers to different questions, and the cost of not knowing which question you are answering is measured in the same order of magnitude as the spread itself.

Why does the same chip cost $0.35 or $14.90 an hour?

Four factors, and together they explain nearly all of the gap. The interconnect is the first: a single GPU on a PCIe card is one thing, and a GPU wired into an NVLink and InfiniBand cluster for distributed training is another, because the second buys you the fabric that lets many GPUs act as one, and that fabric is expensive. The commitment model is the second: on demand is the anchor price, while reserved capacity and spot capacity undercut it heavily. The third is what is billed on top, since storage and especially data egress are metered by some providers and free at others. The fourth is the unit, per GPU or per instance, which is where the worst mistakes happen. The diagram below places the cheap and costly ends side by side, and the table after it turns each factor into something you can check on a quote.

A comparison diagram placing the cheap end of GPU pricing, spot capacity, specialist clouds and PCIe with no egress, against the costly end, on demand rates, hyperscalers and SXM cluster hardware
The two ends of the GPU price spread, and what sits at each.
What drives the spreadCheap endExpensive endHow to read a quote
InterconnectSingle PCIe GPUSXM with NVLink and InfiniBand clusterYou pay the fabric premium only if you train across many GPUs
CommitmentSpot, around $1.76 an hour on averageOn demand, the full anchorReserved and spot cut the anchor by 60 to 91 percent
Storage and egressIncluded, no egress chargeMetered per gigabyte movedA free hourly rate with paid egress can cost more overall
Billing unitPer GPUPer instance of 8 GPUsDivide an instance price by its GPU count before comparing

Two of those rows carry most of the confusion. On the commitment row, GetDeploying's August 2026 figures put the average spot rate at about $1.76 an hour against an on demand average near $3.86, and CloudZero notes committed and spot capacity can undercut the on demand anchor by 60 to 91 percent, which is why it warns that the on demand price is the anchor, not the answer. If your workload can tolerate interruption, spot is not a small discount, it is a different price class. On the provider row, CloudZero's own numbers show dedicated GPU clouds averaging about $4.17 an hour against hyperscalers near $7.89, an 89 percent premium for the same chip that buys you the enterprise wrapper of service levels, networking and compliance rather than any extra performance. Whether that wrapper is worth 89 percent depends entirely on what you are running.

Spot deserves the caveat that comes with the discount, because the price is low for a reason. Spot capacity is unused inventory a provider can reclaim with little warning, so a job running on it can be interrupted mid run. That is fine for work you have designed to checkpoint and resume, since a training run that saves state every few minutes loses only the minutes since its last save, and it is a disaster for a single long job with no checkpointing, which starts again from zero. The rule of thumb is that spot suits fault tolerant, resumable work while on demand suits the run you cannot afford to lose, and a lot of the apparent saving from spot is really the reward for having built your job to survive a restart.

The interconnect factor is the one founders most often overpay for by buying a capability they never use. A single H100 doing inference or fine tuning a modest model needs no cluster fabric at all, and paying the premium for NVLink and InfiniBand there is pure waste. That fabric earns its price only when you train a large model across many GPUs at once, where the speed at which the GPUs talk to each other becomes the bottleneck and a fast interconnect is the difference between a run that finishes and one that crawls. Match the interconnect to the job: a lone card for inference and small fine tunes, a tightly wired cluster only for genuine multi GPU training, and never the second when the first would do. For a merchant rather than a lab, none of this is the operative number anyway, since the median small business AI bill is measured in tens of dollars a month.

Is that price per GPU or per instance?

Ask this before any other question, because it is the error that makes a good deal look terrible and a terrible one look fine. Hyperscalers usually quote a full machine, an instance with eight GPUs, so a headline of roughly $55 an hour is not $55 a GPU, it is about $6.88 a GPU, which is exactly the AWS per GPU figure GetDeploying lists. A founder who compares a hyperscaler's per instance sticker against a specialist cloud's per GPU rate will conclude the hyperscaler is eight times more expensive when the real gap is far smaller. Normalise everything to a per GPU hour before you compare, every time, and half the apparent spread disappears. This is the same discipline that separates a real quote from a misleading one in model pricing, which we worked through in the piece on what a real feature costs in LLM API pricing.

Does renting or buying actually win?

Renting, for almost everyone, and the arithmetic shows why more honestly than a gut feeling does. Take the numbers Jarvislabs reported in April 2026 and treat the calculation as an illustration you should redo with your own inputs. A single H100 lists at roughly $25,000 to buy, before the $50,000 to $165,000 of power, cooling, networking and rack infrastructure a real deployment needs, and it draws up to 700 watts, which is about $60 a month in power per GPU. Rent the same card at around $2.99 an hour on demand and a full month of continuous use, 720 hours, is about $2,152. On paper the card pays for itself in roughly fourteen months of running flat out, which is the number that tempts founders into buying.

The catch is the phrase running flat out. That fourteen month break even assumes you use the card 24 hours a day, every day, and almost no team does. A GPU that sits idle overnight, over weekends, and between experiments might see a fraction of that utilisation, and at a realistic level the break even stretches past the card's useful life, at which point you have bought a depreciating asset to run it part time. Jarvislabs frames the threshold as sustained use above 500 hours a month before purchase even becomes a conversation, which is roughly two thirds of the hours in a month, running continuously, on one card. Below that line renting wins, and above it you are the rare workload that should model the purchase carefully. For most founders the honest answer is to rent, keep the capital, and revisit only when a single workload is genuinely saturating a card for months. If your goal is running a model rather than training one, the calculus shifts again, and our guide to running an LLM locally on hardware you already own covers the small end where no rented H100 is needed at all.

Which end of the spread should you pick?

Match the price class to the workload rather than chasing the lowest number. For inference, serving a model to users, you want a single GPU on demand or reserved, kept close to your users to hold egress and latency down, and you almost never need cluster hardware. For fine tuning, a short job on one or a few GPUs, spot capacity with checkpointing is often the cheapest sane choice, because an interruption costs you minutes rather than the run. For training a large model from scratch, you need the cluster fabric and you should be talking about reserved capacity, since the run is long, the interconnect is not optional, and spot interruptions on a multi day job are painful. The mistake is to read the $0.35 spot rate, assume it applies to your production inference service, and build a live product on capacity that can vanish underneath it. The cheap end is real, but it is cheap for workloads shaped to use it.

An illustration card listing the four parts of a real GPU bill beyond the headline hourly rate: compute per GPU hour, storage and data egress, idle and startup time, and any commitment discount left on the table

What does the total bill actually include?

The real GPU cost is more than the compute line, and the extras are where founders get surprised. The hourly rate you compared covers the GPU running. It does not cover the storage your dataset sits on, the data egress when you move results out, or the hours the machine is provisioned but idle while you debug, all of which are billed and none of which appear in the number that drew you to a provider. Egress is the quiet one, since GetDeploying notes some providers charge nothing for it while hyperscalers meter every gigabyte, and Jarvislabs puts egress at roughly $0.08 to $0.12 per gigabyte, which turns a large model checkpoint or dataset into a real line item every time it moves.

Put a number on the surprise to feel it, treating the figures as an illustration. Say you rent a single card at $2.99 an hour and run it 200 hours in a month, which is $598 of compute. Now add a terabyte of dataset and checkpoint egress at $0.10 a gigabyte, about $100 to move it out once, plus the storage it sat on, plus the 30 hours the instance was provisioned while you set up and debugged but ran nothing, another $90 of pure idle. The compute you compared was $598, and the bill is closer to $800 before storage, a third higher than the number that chose your provider. On a workload that moves large artefacts often, egress alone can rival the compute line, which is why a provider with a higher hourly rate and free egress can be the cheaper choice overall.

Idle time is the other quiet cost, and it is the one you control. A GPU billed by the hour while you read logs, fix a script, or wait for a colleague is pure waste, and on a per minute or per second billing model you can claw that back by shutting instances down between runs. The practical checklist is short. Count the compute, then add storage, then add egress for every gigabyte you will move out, then add the idle and startup hours you will realistically leave on the meter, and only then compare providers. A cheap hourly rate wrapped in paid egress and clumsy startup can lose to a dearer rate with neither. This is the same shape as the wider squeeze on AI infrastructure, where memory prices are now pushing server costs up, which we covered in the piece on the AI memory shortage and what it costs a business.

One caution on all of these numbers. GPU pricing moves fast, and every figure here carries the month it was read, August 2026 for the market ranges and April 2026 for the purchase math. GetDeploying itself notes the median on demand H100 rate rose about 12 percent in a year, from $3.02 to $3.38 an hour, so treat the exact rates as a snapshot and the four factors that explain the spread as the durable part.

There is a reason to expect the cheap end to stay cheap in relative terms even as the absolute numbers move. Specialist clouds compete almost entirely on price for raw capacity, while hyperscalers bundle the GPU into a wider platform and charge for that, so the gap between them is structural rather than a passing promotion. What shifts month to month is the level of the whole market, not the shape of the spread, which is why the four factors are worth learning once while the exact rates are worth re checking each time you buy. A founder who understands the factors can open any provider's pricing page, place it on the map in a minute, and read what the number is really quoting. That skill outlasts any single price, and it is the real deliverable of reading a market this volatile, not the cheapest quote today but the judgment to weigh the next one. The rates will drift. The reasons one quote is 40 times another will not. If you would rather not manage any of this and simply want predictable pricing for building a commerce site, our pricing for the MaShop builder runs on managed infrastructure so the GPU market is our problem to absorb, not yours to model.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building