- Published prices are per second of generated video and they span fourteen times, from 0.05 dollars a second for the cheapest tier to 0.70 for the most expensive, all read from vendor pricing pages on 4 August 2026.
- The per second price is not the cost. What you pay is the price multiplied by how many attempts you discard, and nobody publishes that number because it depends on your shot.
- Costed at a one in three acceptance rate, one minute of finished video runs from about 10 dollars on the cheapest tier to about 77 dollars on a standard tier with audio.
- Batch tiers halve the price in exchange for up to a day of latency, which is a straight win for anything not needed the same afternoon.
- Physical plausibility remains the hard limit. On the hard subset of the VideoPhy-2 benchmark the best evaluated model reached 22 percent joint adherence, and conservation of mass and momentum are named as specific weaknesses.
- An open weights video model reached the top of a public ranking for the first time on 3 August 2026, at 33 billion parameters, with a licence limiting commercial use to companies under 20 million dollars of revenue.
The demo reel problem is that a demo reel is edited. What you see is the take that worked, at the length the model is good at, on a subject chosen because it renders well. The question a marketer actually has is different and much duller: can I produce sixty usable seconds about my own product, this week, for a price that beats the alternative.
This is a costing exercise rather than a review. Every price below comes from a vendor pricing page read on 4 August 2026 and is stated per second of output. Everything derived from those prices is arithmetic you can redo.
What does a second of generated video cost?
Between 0.05 and 0.70 dollars, depending entirely on which tier you pick. OpenAI's published API pricing lists sora-2 at 0.10 dollars per second at 720p, with sora-2-pro at 0.30 at 720p, 0.50 at 1024p and 0.70 at 1080p. Batch pricing halves each of those, so the same 720p second falls to 0.05.
Google's Gemini API pricing page, last updated 30 July 2026, lists Veo 3.1 at 0.40 dollars per second for 720p or 1080p and 0.60 at 4K, with a Fast variant at 0.10, 0.12 and 0.30 for the same three resolutions, and a Lite variant at 0.05 for 720p and 0.08 for 1080p. Every Veo variant includes audio by default, which matters when you compare against a price that does not.
| Tier | Resolution priced | Per second | Ten seconds | Audio included |
|---|---|---|---|---|
| Veo 3.1 Lite | 720p | 0.05 dollars | 0.50 dollars | Yes |
| sora-2 batch | 720p | 0.05 dollars | 0.50 dollars | Not stated on the price row |
| sora-2 standard | 720p | 0.10 dollars | 1.00 dollars | Not stated on the price row |
| Veo 3.1 Fast | 1080p | 0.12 dollars | 1.20 dollars | Yes |
| Veo 3.1 standard | 1080p | 0.40 dollars | 4.00 dollars | Yes |
| Veo 3.1 standard | 4K | 0.60 dollars | 6.00 dollars | Yes |
| sora-2-pro | 1080p | 0.70 dollars | 7.00 dollars | Not stated on the price row |
The spread across that table is fourteen times for the same duration, which tells you the tier decision dominates every other decision you will make.
What does one usable minute really cost?
Multiply the honest number of seconds by an acceptance rate you assume out loud, because nobody publishes one.
Sixty finished seconds is not sixty generated seconds. Working in eight second clips, a minute needs eight of them, so 64 seconds of output before a single retry. Then the retries, and this is where every published comparison stops and every real project begins. A first attempt that is usable on a controlled shot might be one in two. On anything with hands, a specific product or on screen text, one in five is closer.
Take one in three as a working assumption, stated as an assumption. That is 24 generations of eight seconds, or 192 billed seconds for one finished minute.
| Tier | 64 seconds, one take each | 192 seconds at one in three |
|---|---|---|
| Veo 3.1 Lite 720p | 3.20 dollars | 9.60 dollars |
| sora-2 standard 720p | 6.40 dollars | 19.20 dollars |
| Veo 3.1 Fast 1080p | 7.68 dollars | 23.04 dollars |
| Veo 3.1 standard 1080p | 25.60 dollars | 76.80 dollars |
| sora-2-pro 1080p | 44.80 dollars | 134.40 dollars |
Change the acceptance rate and the whole column moves proportionally, which is why the rate is the number worth measuring on your own first project. Generate ten clips of the thing you actually sell, count how many you would ship, and you have a multiplier that makes every future estimate honest.
One more line belongs in the total and it is not billed by anyone. Someone has to cut eight clips together, fix the audio levels and decide the order. That is an hour of a person for a minute of video, every time, and at any realistic hourly rate it exceeds the generation cost on every tier except the most expensive one.
What does it still get wrong?
The visible failures are the ones everyone lists and they are getting rarer. The structural one is not.
Physics is the durable problem. VideoPhy-2, a benchmark of 200 actions built to test physical commonsense rather than visual quality, reports that even the best model it evaluated reached only 22 percent joint performance on the hard subset, where joint means the clip both matched the prompt and obeyed physical commonsense. The paper names conservation of mass and momentum as specific weak points. In practice that is why a poured liquid changes volume, why a dropped object lands wrong, and why anything involving weight looks subtly staged.
For commerce the other failures matter more, and three of them are worth stating plainly. Text rendered on screen is unreliable, so never generate a clip that has to contain your price or your brand name. Exact product fidelity is not achievable from a text prompt, since the model is producing a plausible object of that description rather than your object. And continuity across cuts is weak, so the same person or product in two clips will not quite be the same, which is the real reason multi shot sequences take so many attempts.
The same limitation shows up in still images, where it is cheaper to discover, and we set out what generated stills can and cannot honestly do in what a generated product image can honestly deliver. If your product must appear exactly as it is, the answer for video is currently the same as the answer for photography, which is covered in what holds up and what breaks in AI product photography.
Did anything actually change this month?
Yes, on the open side. On 3 August 2026 MiniMax released the weights for H3, and it became the first open model to lead a public video ranking, placing first in video editing, second in text to video and third in image to video on Artificial Analysis.
The details that decide whether it matters to you: 33 billion parameters, clips of four to 15 seconds with stereo sound, and a local ComfyUI path that tops out at 768p because the 2K module was held back from the open release. The prompt translation layer stays proprietary too. The licence restricts commercial use to companies with revenue under 20 million dollars, which covers almost every reader of this article and is worth reading rather than assuming.
What it changes: for the first time the price floor is your own hardware rather than a per second rate, and fine tuning on your own footage becomes possible. What it does not change: the physics, the text rendering or the continuity, which are properties of the generation approach rather than of the business model. The same day, ByteDance launched Seedance 2.5 as a closed model with longer clips, which we covered in the move past the thirty second barrier.
Is a batch tier worth the wait?
Almost always, and it is the single easiest saving on this page. OpenAI's batch rows are exactly half the standard rows at every resolution, so the same 720p second falls from 0.10 to 0.05 dollars and the 1080p pro second from 0.70 to 0.35.
What you give up is immediacy. The trade only hurts when someone is sitting watching the output appear, which is true during exploration and false during production. A sensible pattern is to explore at the standard rate with a handful of clips until the prompt is settled, then run the full batch of variations overnight at half price. On the one in three model above, that turns a 76.80 dollar minute into something closer to 40, for the cost of planning a day ahead.
The reason more people do not do this is workflow rather than economics. Same day generation fits how a small team works, and overnight generation requires deciding today what you want tomorrow. That is a habit worth acquiring, because it is the same habit that makes every other AI cost line predictable.
How does open weight video change the arithmetic?
It replaces a per second price with a fixed cost you already own or rent, and that flips which projects make sense.
Under per second billing, a hundred variations costs a hundred times one variation, so exploration is rationed. Under owned hardware the marginal clip is electricity and time, so exploration is nearly free and the constraint becomes throughput instead of budget. For anyone whose acceptance rate is poor, and the honest answer for product footage is that it is poor, that inversion matters more than any quality difference between the models.
The costs move rather than disappear. A 33 billion parameter video model is not a laptop proposition, the open release caps out below the resolution the hosted service offers, and fine tuning on your own footage is a project rather than an afternoon. Add the licence question, which for the current open leader means checking a revenue threshold before you ship anything commercial.
The reasonable position for a small business in August 2026 is to rent per second, measure the acceptance rate, and revisit owned inference only if the monthly bill passes the cost of the hardware. That is the same test we applied to text models when comparing the open weight families on licence and hardware, and the crossover sits further out than enthusiasm suggests in both cases.
Which tier should you actually pick?
Work down from what the clip has to survive.
Social feed, sound optional, watched once. The cheapest tier, generated in batch, at 720p. At around ten dollars a finished minute this competes with stock footage and wins on specificity. Nobody pauses a feed video to inspect a texture.
Paid advertising. Mid tier at 1080p, and budget for a low acceptance rate because an ad gets watched repeatedly and a flaw that reads as charming once reads as cheap on the fourth view. The audio inclusion matters here, since a separately produced soundtrack adds cost and sync work.
Anything showing the actual product. Reconsider. A phone, a window and twenty minutes gives you a video of the real item, which is what the customer wants to see, and no tier on the table can produce your product rather than something resembling it.
Background, texture, atmosphere. The strongest current use and the least discussed. Generated establishing shots, abstract motion behind text, ambient loops. No hands, no product, no continuity requirement, and a first attempt acceptance rate close to one, which changes the arithmetic completely.
Per second billing and per token billing have the same trap, which is that the unit price looks small and the retries are invisible until the invoice. Our pricing page shows how MaShop meters generation, and the habit worth carrying across every AI tool is to price the finished artefact rather than the unit.
What makes an acceptance rate good or terrible?
The subject, far more than the model. Four properties predict most of it, and knowing them lets you estimate before you spend anything.
Does a human appear? Faces and hands are where the remaining artefacts live. A clip with no person in it will be accepted several times more often than the same idea with someone holding the product.
Does anything have to be legible? Any text the model has to render, a label, a sign, a price, is a coin flip per attempt. Add it in the edit instead, where it is free and correct.
Does an object need to match a real one? If a viewer can compare the clip with the product page, the bar is exactness rather than plausibility, and no current text prompt clears it. This is the property that separates a usable atmosphere shot from an unusable product shot.
Does something have to move under its own weight? Pouring, dropping, bouncing, swinging. These are exactly the interactions the VideoPhy work found models handling worst, and they read as wrong to viewers who could not explain why.
Score your intended clip out of four. Zero of them is background footage and your acceptance rate will be high. Three or four is a product demonstration with a person, and you should film it instead. The middle is where the tier choice and the budget actually matter, and it is where this whole exercise earns its time.
Where does generated video fit against the alternatives?
Against stock footage, it wins on specificity and loses on certainty. Stock costs a known amount and arrives correct; generation costs an unknown amount and arrives specific. For anything generic, an establishing shot of a city or a hand on a keyboard, stock is usually cheaper once retries are counted. For something no stock library holds, generation is the only option that is not a shoot.
Against filming it yourself, the comparison depends entirely on whether you own the subject. A merchant with the product on the desk can film a usable clip on a phone in twenty minutes, and it will show the actual item. A merchant selling a service, or an idea, or something that photographs badly, has nothing to point a camera at, and that is where generation is genuinely the cheapest route to a moving image.
Against hiring, the honest framing is that the cost tables above are not competing with a videographer's day rate. They are competing with having no video at all, which is the real status quo for most small businesses. Judged against nothing, ten dollars a minute changes what is possible in a way that a comparison against professional production misses entirely.
The test worth running first
Before committing to a tier, spend five dollars on the cheapest one. Generate ten eight second clips of the thing you would actually publish, using the prompts you would actually write.
Count how many you would ship. That single number tells you your acceptance rate, and with it every price on this page becomes a real estimate rather than a headline. It also tells you something the tables cannot, which is whether your particular subject is one this technology handles at all. A cosmetics texture and a person demonstrating a tool are different problems and the second one is much harder.
Two rules survive whatever launches next. Price per finished minute, never per second, because the per second number is marketing and the per minute number is your budget. And never let a generated clip carry a fact, whether a price, a claim or a specification, because the model is generating something plausible and plausible is not the same as true. The same discipline applies to the copy that goes around the video, which is where we started in what a real feature costs once you count every call.