You watch the clip back and something is off. The lighting is beautiful, the model is beautiful, the dress moves, and within a second and a half everyone who sees it knows a machine made it. Nobody can quite say why.
I've rendered a few hundred vertical clips this year for a clothing store and for content channels, roughly three a day, and I've learned that the tells are almost never what people assume. It isn't resolution, and it usually isn't the model. Six things give an AI clip away, and five of them are decisions made in the prompt before a single frame is generated.
The short answer
- Over-saturation and glossy skin — the default look of every model. Ask for muted tones and real skin texture, every time.
- Impossibly smooth camera motion — real cameras wobble. Say so in the prompt.
- Cuts — the moment a generated clip cuts, continuity breaks. One shot beats five.
- Hands, text and small logos — keep them out of frame instead of fixing them.
- Audio — a scene with no ambient sound reads as synthetic before you've looked at it.
- Grade — the last 5% is colour, and it happens after the render, not inside it.
1. The look is too clean
Every video model ships with the same house style: high contrast, warm golden light, saturated colour, skin with no pores. It's optimised to look impressive in a demo grid, and it is the single loudest signal that nobody held a camera.
Real footage is duller than you remember. Overcast light is grey. Indoor light is green under fluorescents and orange under lamps. Skin has texture and uneven tone. If your clip has none of that, it reads as a render regardless of how good the composition is.
What I put at the end of every prompt, near enough word for word: muted natural tones, real skin texture, overcast daylight, no CGI look, no oversaturation, subtle film grain. It is boring and it works. The failure mode is asking for "cinematic" — every model interprets that as more contrast and more orange, which is the opposite of what you want.
2. The camera moves like nothing in the real world
Generated camera motion is perfect: a flawless dolly, a glide with no weight, a slow push at a constant speed. No human operator produces that, and no gimbal quite does either. Your eye has watched a hundred thousand hours of slightly imperfect footage and it flags perfect motion instantly.
The fix is to name the imperfection. "Handheld, slight natural vibration, small drift as the operator adjusts" produces something recognisably human. So does naming a real rig — "shot on a shoulder rig walking behind her" gives the model a physical constraint to imitate.
The related mistake is asking for too much movement in fifteen seconds. Real shots are slower than you think. One move per shot — a push, or a pan, not both.

3. Every cut is a chance to break continuity
This is the expensive one, and it took me months to accept.
A five-cut sequence sounds better than a single shot. It isn't, because the model regenerates the world at every cut. The jacket loses a button. The street changes. The woman's hair parts on the other side. Viewers don't consciously catalogue those errors; they just distrust the clip.
A single continuous shot — one camera moving through the whole fifteen seconds — has no seams to break. It is also harder to write, because you have to specify a camera path and an action timeline rather than a list of pretty frames. I now write prompts as a timed sequence: what happens at 0–4s, at 4–9s, at 9–15s, with one camera and one location locked for all of it.
When you genuinely need multiple shots, lock the elements rather than re-describing them. Reference images and named character handles keep a person and a product consistent across generations far better than prose does — that mechanism is the whole subject of consistent AI video characters, and for a clip that must match existing footage, motion control transfers real movement onto your own still.
4. Hands, text and logos — don't fix them, avoid them
Models have improved enormously on faces and barely at all on fingers doing a specific task. Fastening a button, holding scissors, counting money, pouring into a glass: these fail often, and they fail in the uncanny way rather than the funny way.
Text is worse and shows no sign of improving. Any lettering the model renders — a sign, a label, a price on a card — comes out as plausible nonsense. The same goes for logos, which additionally is not a place you want a model improvising.
So frame them out. Hands below the crop line, products held still rather than manipulated, and every word added afterwards in the editor. Burned-in text is a post-production job, and doing it properly is covered in AI video text overlay. One rule I hold to hard: never burn a price into a frame. Pixels can't be edited after publishing and prices move.
5. Silence is the tell nobody mentions
Put a clip with no ambient track next to one with room tone, and the silent one reads as synthetic before the viewer has examined a single frame. We are far more sensitive to wrong audio than to wrong pixels.
Two routes. Either let the model generate native audio and name the sounds you want — "wind through the street, distant traffic, footsteps on wet stone, no music" — or strip the audio and lay your own bed underneath: a quiet ambient loop at low level plus music if the brand calls for it. What kills a clip is the third option, which is nothing at all.
If there's a voice, the mix matters more than the voice does. Voice on top, ambient well underneath, and music that doesn't fight the first three words — that's where the hook lives, and hooks are the subject of AI video hooks.
6. The last five per cent happens after the render
Even a good generation comes out slightly too contrasty and slightly too saturated for a feed. A light grade afterwards — pull the saturation down a few per cent, lift the blacks slightly, add a touch of grain — closes most of the remaining distance. It takes under a minute per clip and it is the step people skip because the render already looked fine on its own.
Do it consistently across a batch and something else happens: the clips start to look like they came from the same camera, which is its own kind of credibility.

What none of this fixes
Two things, and I'd rather say them than sell around them.
Close-up human performance is still not there. A face delivering an emotional line, sustained eye contact, a genuine laugh — you can get away with it for two seconds, not for ten. If your concept depends on it, film a person or use a disclosed avatar.
And realism is not the same as effectiveness. My best-performing clip in the last 30 days did 11,175 views out of 38,488 total across 90 clips; several technically better clips did under fifty. Looking real is table stakes — it stops people scrolling past for the wrong reason. It doesn't decide whether they stop at all, and neither does the render: on the same 90 clips, the posting slot moved the median by roughly twenty times, which is more than any craft decision on this page.
One thing worth checking before you spend another evening tuning a prompt: most AI video models do not have a negative prompt field at all. Wan 2.2 and Kling 3.0 do, and there it is a real second conditioning input. Veo 3.1 and Runway do not, so anything you type as an exclusion is just more positive prompt. That distinction matters because a negative prompt can only push away from a concept the model already represents strongly — "blurry" works, "no extra fingers" essentially does not. Five rendering-quality terms outperform a forty-term list badly enough that the long version came out worse than using none, which is what you would expect once you see that each extra term dilutes the rest. Which models have the field, and the five terms that earn their place, in negative prompts for AI video.
FAQ
Why do AI videos look fake even at high resolution?
Because the tells are lighting, motion and continuity, not pixels. An over-saturated 4K clip with impossibly smooth camera movement looks more synthetic than a grainy 720p one shot handheld.
What prompt wording makes AI video look more realistic?
Name the imperfections: muted natural tones, overcast light, real skin texture, handheld with slight vibration, no CGI look, no oversaturation. Avoid the word "cinematic" — models read it as more contrast and more orange.
Is one long shot better than several cuts?
Usually yes. The model regenerates the scene at every cut, so details drift between them. A single continuous camera move over fifteen seconds has no seams to break.
How do I stop AI videos from garbling text?
Don't ask for text at all. Keep signs, labels and prices out of frame and add every word afterwards in an editor, where you can also fix it later.
Does adding audio really matter that much?
Yes. A clip with no ambient sound reads as synthetic before a viewer has looked closely. Either generate native audio and name the sounds, or lay a quiet ambient bed under the clip yourself.
Which model looks the most realistic right now?
They are closer to each other than to the gap made by the fixes above. A well-prompted, well-graded clip on a mid-tier model beats a lazy prompt on the best one, and the price difference per clip is real — see AI video cost per video.
Build this with us
Short on time? We also make custom AI videos for you.
Looking for something else? Browse all our AI video guides in one list.
In our Skool community I share the exact prompts, the grade settings and the renders that failed, so you can skip the expensive lessons. Join the AI Video Generator community on Skool.


Share:
AI Video Generator for Hotels and Short-Term Rentals: What to Automate, What to Shoot
AI Video API Pricing: What a Finished Clip Actually Costs