- A generated clip fails differently from a generated still. A still can be wrong once. A video is wrong progressively, and the last frame is often a different product from the first.
- YouTube requires disclosure when AI generates a realistic scene that did not occur, and states plainly that disclosing does not limit a video's audience or its ability to earn.
- Minor aesthetic edits, beauty filters, colour adjustment and voice repair are named on that same page as things you do not have to disclose.
- TikTok has offered a creator label for content completely generated or significantly edited by AI since September 2023, with automatic labelling tested alongside it.
- Amazon takes shoppable video between 1 and 12 minutes, up to 5 GB, in .mov or .mp4, four per ASIN, and puts it in the main media block only when the listing has fewer than six images.
- The jobs generated video does well are the ones where your product is not the subject. The moment the customer is judging the item itself, shoot it.
Video on a product page is not a nice extra any more, and every seller knows it. What has changed in the last year is that making one no longer requires a camera, a tripod or a Saturday. What has not changed is that most of the video a small shop needs is a close look at an object, and a close look at an object is the exact thing generative video is worst at.
This piece is about that gap. Not whether AI video is impressive, which it obviously is, but which specific jobs on a listing it can take over this quarter without costing you a return or a strike.
Where does a product video actually earn its place?
In three distinct places that behave nothing alike, and lumping them together is why so much of this advice is useless. There is the listing page video, whose job is to answer the question the photographs left open. There is the social clip, whose job is to stop a thumb. And there is the paid ad, whose job is to survive an automated review and then convert.
The listing video is the demanding one. A shopper who plays it has already decided they are interested and is now looking for a reason to stop worrying: how big is it really, what is the finish like, does the lid close properly. Every one of those questions is answered by the physical truth of your object, which means the video has to show your object. Amazon's own guidance for sellers pushes in exactly this direction, recommending that you show the product honestly and in everyday settings with clear visibility from a few different angles.
The social clip has different physics. Nobody there is auditing your stitching. They are deciding in half a second whether to keep watching, and a stylised, obviously constructed scene is not just acceptable, it often performs better than a faithful one. That is the corner of the job where generation earns its keep today.
What breaks when a model animates your product?
Identity, and it breaks gradually rather than all at once, which is why it slips past a quick review. The first frame usually looks superb, because a first frame is just an image and image models are now good at holding a reference. Then the camera moves, and the model has to invent what the unseen side of your product looks like. It invents something plausible. Plausible is not yours.
Watch a generated clip of a specific object frame by frame and you can usually name the moment it stops being your item. Proportions creep first, because the model has no fixed geometry, only a statistical sense of what objects of that kind look like. Printed text goes next, dissolving into a texture that reads as writing at a glance and as noise on pause. Then anything a hand touches deforms, because contact is where physics has to be simulated rather than sampled.
The same identity problem exists in stills, and we walked through its current state in the note on what ChatGPT Images 2.5 changes for product photos. Video multiplies it by the frame count. A still gives you one chance to be wrong. Thirty seconds at 24 frames gives you seven hundred.
Pause a candidate clip at four random frames and screenshot them. If you would not have published any one of those four as a product photograph, the clip is not ready, however good it looks at full speed.
Which video jobs can AI do today?
Here is the division we would apply, sorted by how directly the customer is judging your object. The rule underneath it is the same one that governs stills: generation is safe where it supplies context and unsafe where it supplies evidence.
| Video job | Is your product the subject? | Main risk | Verdict today |
|---|---|---|---|
| Animated background behind a real clip of the item | No, the item is filmed | Minimal | Use it |
| Text and motion graphics over real footage | No | Minimal | Use it |
| Abstract social hook that never shows the product close up | No | Brand fit only | Use it |
| Voiceover for a real product clip | No | Disclosure if you clone another person's voice | Use it with care |
| Lifestyle scene with your product held or worn | Yes, at distance | Drift on contact and proportion | Review every frame |
| Close inspection of finish, stitching or mechanism | Yes, entirely | Invented detail the buyer will not receive | Film it |
| Demonstration of how the product works | Yes, entirely | Physics the model guesses wrong | Film it |
| Size and scale reference | Yes, entirely | Reads as deception when wrong | Film it |
Four of those eight are usable now with essentially no risk, which is more than most sellers assume. The trick is that all four are jobs where the generated pixels are around your product rather than of it.
Do you have to label an AI product video?
Sometimes, and the rule is narrower than the panic suggests. YouTube's help page on disclosing altered or synthetic content asks for disclosure in three situations: making a real person appear to say or do something they did not, altering footage of a real event or place, and generating a realistic scene that never occurred. It then lists what does not need disclosing, and the list is generous: non realistic content, beauty filters, colour adjustment, voice repair, cloning your own voice for a voiceover, captions and thumbnails.
Read those two lists against the table above and a useful line appears. A generated animated background behind your real footage is a minor aesthetic edit. A fully generated scene of a person using your product in a kitchen that does not exist is a realistic scene that did not occur, and it wants the label. YouTube also states that disclosing does not limit a video's audience or its eligibility to earn, which removes the only rational reason anyone had for hiding it.
TikTok has run a creator applied label for content completely generated or significantly edited by AI since it introduced the disclosure tool in September 2023, and said at the time it would test automatic labelling for content it detects as AI made. The practical consequence for a seller is that the label may arrive whether or not you apply it, so applying it yourself is simply the version where you control the framing. The same logic applies on Meta's surfaces, which we covered in the piece on what shops actually have to declare on Instagram.
In the European Union there is a further layer, and it is worth reading carefully rather than assuming the worst. Article 50 of the AI Act, applicable since 2 August 2026, puts machine readable marking on the provider of the generating system. The deployer duty in paragraph 4 attaches to content that constitutes a deep fake. A generated clip of your own ceramic mug is not one. A generated clip of a recognisable person endorsing that mug is a different conversation entirely, and not one to have without permission.
What do the platforms require technically?
The specifications are boring and they are also where most uploads die, so they are worth having in one place. Amazon accepts shoppable video in .mov or .mp4, up to 5 GB, at up to 1080p, with a duration between 1 and 12 minutes. You get four videos per ASIN. Placement in the main media block on the detail page is conditional: the listing has to have fewer than six images. Review takes a few days, and the most recently uploaded video takes priority in the ordering.
That last condition catches people. A seller who has diligently uploaded seven photographs has locked their video out of the position they wanted it in. Whether six images beat five images plus a video is an empirical question about your product, and it is one you can test in an afternoon.
What does an AI voiceover change?
More than the pictures do, for most small sellers, and with a cleaner rule attached. YouTube's list of things that do not require disclosure includes cloning your own voice to create voice overs. That single clause is a gift to anyone who hates the sound of themselves recording twelve product explainers, and it draws the line exactly where it should sit: your voice is yours to synthesise, somebody else's is not.
The second use is bigger and less discussed. A product video with a spoken track is a product video in one language. Dubbing it opens the listing to markets you already ship to, and the quality bar for a thirty second product explainer is far lower than for drama, which is the case we made in the breakdown of what actually decides dubbing quality. If you already run translated listings, and the practicalities of that are in the guide to running a multilingual store on machine translation, the audio track is the piece most sellers forget to translate.
One caution that costs people. Background music generated by a model sits in the YouTube list as content that does require disclosure, since AI generated music is named there explicitly. That is a strange asymmetry when a synthesised voice reading your own script does not, and it is the sort of detail that only matters when an automated system flags a video you thought was uncontroversial.
How does this fit a catalogue rather than one hero product?
Badly, if you approach it product by product, which is how almost everyone starts. Twelve minutes of allowed duration and four videos per ASIN sounds like an invitation to make a lot of video. In practice a shop with 200 listings that tries to make 200 bespoke videos makes about nine and stops.
The version that survives contact with a real week is a template. One structure, one music bed, one caption style, one closing frame, and the only variable is the thirty seconds of real footage in the middle. Generation is what makes the fixed parts cheap to produce once and reuse two hundred times. It is not what makes each video special, and treating it as a personalisation engine is how the budget disappears.
Aspect ratio is the other tax nobody budgets for. The listing wants one shape, the social feed wants a taller one, and the ad platform often wants a square as well. Shooting the real footage wide enough to crop three ways costs nothing at capture time and cannot be fixed afterwards, so decide it before you press record rather than after.
Plan around the review delay too. Amazon says approved videos typically appear within three days, which means video is not a launch day asset unless you uploaded it the previous week. Sellers who discover this on the morning of a product drop tend to discover it once.
A workflow that will not embarrass you
- Film thirty seconds of the real object once. A phone on a stack of books, a window, one rotation. This becomes the truth layer for everything else you make this year.
- Generate around it, never through it. Backgrounds, transitions, text, atmosphere. Keep the frames where the customer is judging the item as real footage.
- Screenshot four random frames before publishing. If any of them would not pass as a product photograph, the clip goes back.
- Apply the platform label when the scene is realistic and did not happen. It costs you nothing on YouTube by their own statement, and it removes the enforcement question entirely.
- Keep the source files and a note of what was generated. When a marketplace asks, and they do ask, an answer beats an argument.
When is a real camera still cheaper?
More often than the tooling suggests, and the reason is iteration count rather than shooting time. A generated clip that is nearly right is not fixable in the way a filmed clip is. You cannot ask a model to keep everything and change one strap. You regenerate, and you get a new set of small wrongnesses, and the fourth attempt is often further from usable than the second.
Filmed footage has the opposite economics. It is expensive to acquire and nearly free to adjust. Cut it differently, speed it up, change the music, add generated text over the top. One honest thirty second rotation of your product supports a year of social clips, which is why the workflow above starts there rather than ending there.
There is also a quieter argument. The reason your listing needs a video at all is that photographs leave a doubt, and the doubt is about whether the real thing matches the pictures. Answering that doubt with footage that is itself generated is a strange thing to do, and shoppers are getting better at spotting it. The same tension runs through ad creative, where we found that testing generated variants works best when the underlying asset is real.
What we would do with a hundred pound budget
Spend nothing on a video model this month. Spend an hour filming every product you sell on a white surface next to a window, one slow rotation each, and store the files somewhere you will find them again. Then use generation for everything that surrounds that footage, which is where it is genuinely good and where no disclosure question arises.
If you are still deciding where these videos will live, the storefront matters as much as the asset: our ecommerce website builder generates product pages you own outright, which means the video embed and its markup are yours to fix rather than a platform's to change. That turns out to matter the first time a marketplace deprecates a video field and you need to re-upload two hundred files.
The summary a busy seller can act on: generated video is ready for the frames where your product is not the subject, and it is not ready for the frames where it is. That line will move. It has not moved yet.