I have generated somewhere north of four hundred clips in the last three months, across two brands I actually run and pay for. Most of them were unusable. The ones that worked were not the ones I spent the most time on — and the pattern behind which is which turned out to be boringly consistent.
This is what an AI video maker is genuinely good at in 2026, where it still wastes your money, and the working order I now use so that a clip is either useful in two attempts or abandoned.
What "AI video maker" means in 2026
The phrase covers four categories that behave nothing alike, and choosing wrong is how people conclude the whole field is a toy.
- Slideshow and template tools. Feed them product photos, get a beat-cut montage. Reliable, instant, and immediately readable as a template. Fine for a catalogue reel, weak for cold traffic.
- Avatar and talking-head tools. A synthetic presenter reads your script. This is the category most people mean, and the one that improved most in the last year.
- Generative video models. Footage built from a prompt or a start image. Most control, most variance — and the one where prompt quality dominates the result.
- Editing assistants. They cut, caption and reframe footage you already have. Not generators at all, but often what someone actually needed.
If you're starting from zero, the full workflow guide covers the ground before this one. If you're comparing specific tools, the generator comparison is the better starting point.

The number that changed how I work
Last week I ran two ads for the same product, same budget, same day, same audience. One was cinematic — slow dolly, controlled lighting, the kind of frame you screenshot. It returned 0.9. The other looked like a phone held at arm's length in a hallway. It returned 3.3.
The organic side told me the same thing twice over. On our video channel, one clip has 796 views. Everything published before it sits between 1 and 6. The difference was not production value — that clip was worse looking. It opened with a specific, surprising number instead of a product name.
On the other brand I run, the same pattern: a clip opening with a problem ("rain never ruins my outfit anymore") got 643 views against 8–22 for clips titled with the product name. Same camera, same pipeline, same day of the week.
So the honest summary of four hundred generations is this: the generator is not the variable. The first two seconds are. An AI video maker gets you to a usable clip cheaply, and then almost all remaining leverage sits in the writing.
The working order that stopped wasting my credits
I burn far fewer generations since I stopped starting with the prompt.
- Write the spoken line first. Fifteen seconds is 35–40 words. Write them, read them out loud, and delete anything you would never say to a friend. If the line is weak, no amount of rendering rescues it.
- Decide the single proof. One. Pockets, machine washable, fits in a carry-on. Listing five features is the most common reason a clip flatlines — five proofs means none is remembered.
- Lock the character before the motion. Generate the still first, approve the face, then animate from it. Fixing identity drift after the fact costs more than getting the frame right. The method is in the guide to consistent AI video characters.
- Generate short. Ask for 5–8 seconds per shot and stitch. Long single generations drift, and a drifted 15-second clip is a total loss where a drifted 5-second shot costs you one retry.
- Caption last, lower third, plain white. Most feeds autoplay muted. Moving captions from mid-frame to the lower third improved watch-through for me before any other change did.
Prompting for "real" instead of "beautiful"
This is the counter-intuitive part, and it is worth more than any tool choice. Generative models are perfectly capable of producing footage that reads as ordinary — you just have to ask, and the default is the opposite.
Prompt for cinematic, dramatic lighting, shallow depth of field and you get something a viewer instantly classifies as an advert. Prompt for handheld phone camera, slightly uneven framing, ordinary indoor light, no music and you get something they read as a person. On a feed made almost entirely of people talking to phones, the second one earns its two seconds and the first one doesn't.
Four things I now name explicitly in every prompt: the camera ("handheld phone"), the light ("overcast window light"), the motion ("slight natural sway"), and what I don't want ("no text, no logos, no music"). Vague prompts produce the model's average, and the model's average is a commercial.
What it costs, realistically
The pricing question everyone asks has an unsatisfying answer: the render is rarely the expensive part. My last month of generation cost less than a single day of ad spend on the campaign it fed.
What costs money is testing badly. I once ran a campaign split across four countries for thirty days and spent 894 euros for two conversions — not because the videos were bad, but because splitting a small budget four ways meant nothing ever got enough data to be judged. The clips were fine. The structure around them was the waste.
So budget the way the leverage actually sits: cheap on generation, generous on the one thing you're testing, ruthless about killing what has spent 3–4× the product price without a sale.

Where AI video makers still fail
- Hands and on-screen text. Still the tell. Cut away before a hand does anything complicated, and never let a model render a price or a label — burn text on afterwards.
- Continuity across shots. Faces, garments and rooms drift between generations. Reference images help; assuming it will "just match" does not.
- Facts. A model will invent a statistic without hesitation. Everything factual in a script has to come from you.
- Your offer. No generator fixes a weak offer. A better video just gets more people to discover the price is mid and shipping is slow.
- Disclosure. Meta and TikTok both expect AI content to be labelled. Use the platform's own toggle — it costs nothing and skipping it risks the account.
FAQ
What is an AI video maker?
A tool that produces finished video from text, images or a script — generating or assembling the footage, adding captions and audio, and exporting in the aspect ratios each platform expects.
Is an AI video maker good enough for real ads?
Yes, when the output is built to look native rather than cinematic. In my own testing the low-fi, single-proof, talking-head structure beat the polished version by roughly 3× on return at identical spend.
How long should an AI-generated clip be?
Fifteen seconds for cold traffic, generated as two or three short shots and stitched. Long single generations drift and cost more to retry.
Do I need a paid tool to start?
No. Start free, learn what a usable clip looks like, and only pay once you know which step you're repeating most. The unlimited generation guide covers the free routes.
Why do my generations look obviously AI?
Usually because the prompt asked for beauty. Name the camera, the light and the imperfection you want. "Handheld phone camera, ordinary indoor light" produces something far more believable than "cinematic".
How many versions should I test?
Three, differing only in the first two seconds. Keep product, proof and close identical so the opening is the only variable you're measuring.
Where to go next
Take one product you already sell. Find the single detail that makes someone buy it, write 40 words that mention it once, and generate three different openings for those same 40 words. Run them two days and keep the cheapest click. That loop is the entire job, and it works regardless of which tool you picked.
If you'd rather work through it alongside people running the same tests, I run a free Skool community where we post the hooks, prompts and numbers that actually landed — winners and losers both. Come say hi.
Or, if you'd rather hand the whole thing over, that's something I offer as a service.


Share:
AI Lip Sync Video: How to Make a Character Actually Talk (2026)
Seedance AI: The Prompt Structure That Fixed My Hit Rate (2026)