The first time image-to-video actually clicked for me, it was an accident. I'd been fighting a text prompt for twenty minutes trying to get a specific shot — a jacket on a hook, morning light, slow push in — and getting a different jacket every time. Then I generated the still first, fed it in, and got the shot on the second try.

That's the whole argument for image-to-video. You stop describing and start animating. The look is already decided, so the model only has one job left.

Why start from an image at all

Text-to-video asks the model to invent two things at once: what everything looks like, and how it moves. Every generation re-rolls both. That's why you get a good composition with bad motion, then good motion with the wrong subject.

Starting from a still locks the first variable. Framing, colour, wardrobe, product, lighting — all fixed before motion is ever considered. What's left is a much narrower question, and narrow questions get answered correctly more often.

In practice this is the single biggest credit saver available to you. Fewer re-rolls, and the re-rolls you do run are about motion alone.

The workflow, start to finish

  1. Make the still first. Any decent image generator works. Get the framing and lighting genuinely right — every flaw here gets animated and amplified.
  2. Upload it as the start frame. This is the anchor for everything that follows.
  3. Describe only motion. This is where most people go wrong. Don't re-describe the scene the image already shows. Describe what moves: the subject, the fabric, the camera.
  4. Add a negative prompt. Cheap insurance. Warping, morphing, extra limbs, text artefacts — name them.
  5. Generate low, then commit. Get the motion right at the cheapest settings. Only re-run the winner at full resolution.

Writing a motion prompt that works

Here's the shift that matters. A text-to-video prompt is a description. An image-to-video prompt is a set of instructions.

Weak: "A woman in a blue coat on a city street, cinematic, golden hour." The image already says all of that. You've told the model nothing new, so it invents — and inventing is what you were trying to stop.

Strong: "She turns her head slowly to the left. Coat hem moves in a light breeze. Camera pushes in slowly, ending at chest height. Everything else stays still."

Note the parts doing the work: one subject action, one secondary detail, one camera instruction, and an explicit hold on everything else. Speed and direction are both stated. "Camera movement" on its own produces randomness; "slow push in, ending at chest height" produces a shot.

That last clause — everything else stays still — is worth adding every single time. Unrequested motion is where most of the drift comes from.

AI Video Generator Skool Community Banner

Motion brush and end frames

Two features are worth learning properly, because they solve problems prompting can't.

Motion brush lets you paint over the part of the image you want animated. Photo of a lake with mountains behind it: paint the water, leave the mountains, and only the water ripples. One rule people get wrong — use separate strokes for separate elements. Dress moving left and hair moving right is two strokes, not one. Combine them and you get mush.

Flat vector illustration of a landscape where only the water is highlighted, showing motion brush animating one element

Start and end frames let you supply both the first and last image, and the model interpolates between them. This is the closest thing to real direction available here. If you know exactly where a shot begins and ends, you're no longer hoping — you're specifying. It's also the cleanest way to get a deliberate transition rather than a lucky one.

Flat vector diagram of a start frame and an end frame with interpolated frames between them

Camera controls — pan, tilt, zoom, custom paths — are worth a pass too, but they're blunter than the brush. Reach for them when the whole frame should move, not part of it.

What still goes wrong

  • Hands. Still the weak spot in 2026. Visible hands doing something specific will cost you extra generations. Frame them out when you can.
  • Faces over long clips. Identity drifts as duration grows. Keep it short. Our guide to keeping characters consistent goes deeper on this.
  • Low-quality source stills. Soft or noisy inputs produce worse video than clean ones. The model amplifies whatever you feed it.
  • Over-painting with the brush. Selecting most of the frame defeats the point. Be stingy.
  • Busy backgrounds. The more objects behind your subject, the more things there are to rearrange themselves.

Where this fits with everything else

Image-to-video isn't Kling-specific — it's the default professional workflow across every current tool, and the technique transfers. If you want the broader version, our image-to-video AI guide covers the approach tool-agnostically, and the Kling costs and limits breakdown covers what each of these generations will actually charge you.

For sharpening the motion instructions themselves, the prompt guide is the one to read next.

The short version

Generate the still. Lock it as your start frame. Describe motion and nothing else. Name a speed and a direction. Hold everything you didn't ask to move. Test cheap, commit expensive.

Do that and your hit rate goes up immediately — not because you got better at prompting in the abstract, but because you stopped asking the model to guess at things you already knew.

If you'd rather not work it out alone: I run a free Skool community where we trade the motion prompts that landed, the brush setups that held, and the settings that stopped us burning credits on tests. It's free, and it's where the working versions get posted.

Join the free AI Video Generator community on Skool →

Latest Stories

This section doesn’t currently include any content. Add content to this section using the sidebar.