KitanaAI Photo & VideoStudio
video maker · 4 min read

AI Photo to Video: How the Technology Works

What happens when a still becomes a clip: how the model predicts motion, why some photos animate well and others warp, and what the current limits actually are.

By the Kitana team
APVIDEO MAKER

Photo to video looks like the photo starting to move. What actually happens is stranger and worth understanding, because it explains precisely which photos work and which produce the warping everyone has seen.

What the model is predicting

Give it a still and it treats that image as the first frame of a clip that does not exist. Its job is to generate the frames that follow.

It does that by having learned, from an enormous quantity of video, what tends to happen next. Not in any particular scene — it has never seen your room — but in general. Hair in still air drifts slightly. Fabric settles. A person standing does not stay perfectly rigid. Light through a window shifts.

Two constraints govern the result, and they pull against each other:

Temporal coherence. Every frame must be consistent with the ones before it. The face in frame 60 has to be the same face as in frame 1.

Plausible motion. Something has to actually move, or you have produced a still image played 24 times a second.

Almost every failure is these two constraints losing to each other. Too much weight on motion and identity drifts — the warping. Too much on coherence and you get an image that barely moves.

Why some photos animate well

The pattern is consistent, and it is about how much the model has to invent.

Animates well:

  • A clear subject, large in frame, sharply focused
  • Depth behind the subject — a background that is separate from them
  • Front-facing or slightly turned, not profile
  • Even lighting with visible direction
  • Room around the subject for the camera to move into

Animates badly:

  • Small or distant subjects — too little facial structure to keep coherent
  • Busy, cluttered scenes — many things the model must decide whether to move
  • Hands prominent in frame, still the weakest area across all these models
  • Text or logos, which smear as soon as anything moves
  • Tight crops with no space for camera movement
  • Multiple people, where identities can swap between them

The through-line: the more the model has to invent, the more it drifts. A sharp, simple, well-separated subject leaves little to invent.

Why the clips are short

Four seconds feels arbitrary. It is not.

Cost scales with frames, and video frames are expensive — roughly an order of magnitude more compute per second than a still image. But the harder limit is coherence. Identity drift compounds: a face that is 99 percent right at one second is noticeably off at eight, because each frame is generated relative to the last and small errors accumulate.

Four seconds, vertical, is roughly where current consumer tools stay reliably coherent. It is also, not coincidentally, about the length that loops cleanly on short-form platforms.

What to give it

The practical version:

  1. A sharp photo. Soft in, soft and wobbling out.
  2. A clear subject, large in frame. Head and shoulders or half body, not a distant figure.
  3. Separation from the background. A metre or two of space behind the subject.
  4. Room in the frame. If you want a push in, leave space to push into.
  5. A simple motion description. "A slow push in." "Hair moving gently." "Light shifting across the face." One idea, not three.

If your source is soft or small, upscale it first as a separate step. The video model has nothing to sharpen — it can only propagate what is already there, including the softness.

What it still cannot do

Worth stating plainly:

  • Hands. Unreliable in motion across every current tool.
  • Text. Smears immediately.
  • Complex interaction. Two people doing something together, an object being picked up. The model does not understand causality.
  • Specific choreography. You can direct it; you cannot script it.
  • Length. Anything beyond a handful of seconds drifts.

None of this is a settings problem. These are properties of how the models work, and they will move over time, but not by you trying harder today.

Where it is genuinely useful

Short vertical clips for social, where four seconds is the native format rather than a compromise. Bringing a still portrait to life. Giving a static product shot some movement. Turning an archive photo into something that holds attention for a moment longer than a still would.

For the practical side of planning and shooting for this, see preparing for AI video, and for what it is doing to short-form platforms generally, how short-form AI video is reshaping TikTok and Reels.

Photo to Video runs in the Kitana apps rather than the browser studio, as part of Kitana Pro — ten clips a month at four seconds and 768 by 1344. The video maker page has the details.

Frequently asked questions

Does the model know what happens next in my photo?
No. It predicts motion that is consistent with what the image shows, based on having seen enormous numbers of videos. A photo of someone mid-stride produces walking because that is what such images usually precede, not because the model knows where that person went.
Why do faces sometimes warp during the clip?
Because the model has to keep a face coherent across every frame while also moving it, and identity is fragile under motion. Warping is most common when the face is small in frame, partly turned, or partly obscured. A large, sharp, front-facing subject warps far less.
Why are AI videos so short?
Cost and coherence, both of which scale badly. Every additional frame is compute, and every additional second is another opportunity for the subject to drift from what it looked like at the start. Four seconds is roughly where current consumer tools hold together reliably.
Is the audio generated too?
In most photo-to-video tools, no — the output is silent. That suits short-form platforms, where the majority of clips are watched with the platform's own audio over them anyway.
What resolution does the output have?
Kitana produces vertical clips at 768 by 1344 and four seconds, as part of Kitana Pro. Vertical because that is the shape short-form platforms display without letterboxing.
Can I control the motion, or is it random?
You can steer it with a description — a slow push in, hair moving, light shifting — and the model will generally follow. What you cannot do is specify it precisely, frame by frame. Treat the prompt as direction rather than instruction.

Ready to put this into practice?

Create with Kitana using the tool that fits this guide.

Turn a photo into video

Ready to try it yourself?

Download Kitana and create your first AI photo in under a minute.