KitanaAI Photo & VideoStudio
video maker · 4 min read

Preparing for AI Video: Tips for Content Creators

How to shoot source photos that animate well, plan a clip that survives four seconds, and build a repeatable process instead of generating and hoping.

By the Kitana team
PFVIDEO MAKER

Most people approach AI video by generating first and evaluating afterwards, which turns a production process into a slot machine. The clips that look effortless were almost always planned, and the planning happens before any model is involved.

Here is what that looks like in practice.

Decide the shot before you shoot the still

A four-second clip holds exactly one idea. Not a sequence, not a story — one movement.

So decide it first: a slow push in on a face, hair moving in wind, light shifting across a room, a subject turning slightly toward camera. Write it down in a sentence. That sentence becomes both your framing guide and your motion prompt.

The reason to do this first is practical. The framing that supports a push in — headroom, space at the edges — is not the framing you would use for a static portrait. Decide the motion and the framing follows; decide the framing and you may have made the motion impossible.

Shoot for animation specifically

The requirements differ from ordinary portrait photography in ways that matter:

  • Frame loose. Leave space around the subject. Camera movement needs somewhere to go, and a tight crop means the motion pushes your subject out of frame or the model refuses to move at all.
  • Separate the subject from the background. A metre or two of air. This gives depth for the model to work with and is the single biggest factor in whether motion looks three-dimensional or like a flat image being warped.
  • Shoot vertical. Output is 9:16. Shooting horizontal and cropping loses most of your frame.
  • Keep it sharp. Softness propagates. The model cannot sharpen; it can only carry existing detail forward.
  • Simplify the scene. Every additional element is something the model must decide whether and how to move. Clutter produces incoherent motion.
  • Avoid hands in frame where you can. Still the least reliable element in motion.

Ten minutes of shooting with these in mind produces a better batch than an hour of searching an existing library.

Plan for a low hit rate

This is the expectation that most needs resetting. You will not get a usable clip every time.

Budget roughly one keeper from every two or three attempts, and prefer varying the source over regenerating the same one — a photo that warped once will usually warp again, because the cause is in the image rather than in the run.

That means preparing more source stills than you need clips. If you want five clips for a week of posting, start with fifteen to twenty candidate stills.

Keep a consistent look

Consistency across a set does not come from the model, which varies every run. It comes from fixing everything upstream:

  • Same lighting setup and time of day
  • Same framing and distance
  • Same wardrobe, or a deliberate rotation
  • Same motion description, or a small fixed set of them
  • Same source-photo treatment before generation

Vary one thing at a time when you want variety, rather than letting everything drift.

Build a short pipeline

A workable loop:

  1. Write the motion sentence. One idea.
  2. Shoot fifteen to twenty stills for it, vertical, loose, separated, sharp.
  3. Cull to the best five or six on sharpness and subject clarity.
  4. Upscale any that need it, as a separate step.
  5. Generate, two or three attempts per still.
  6. Select — this takes longer than you expect.
  7. Post, with the platform's own audio over the top.

Step 6 is the one people underestimate. The generating is fast. The judging is the work.

What not to attempt yet

Be honest about the limits so you do not burn a day discovering them:

  • Anything with legible text
  • Two people interacting
  • Hands doing something specific
  • Precise choreography
  • Anything longer than a handful of seconds without cuts

The mechanics behind these limits are in how photo to video works, and it is worth reading before you plan a shoot around something the model cannot do.

Where this is heading

Short-form platforms already reward volume and consistency over production value, and this lowers the cost of both. The analysis is in how short-form AI video is reshaping TikTok and Reels.

Photo to Video runs in the Kitana apps as part of Pro — ten clips a month at four seconds and 768 by 1344. Details on the video maker page, and the thirteen photo tools that feed it are on the photo maker page.

Frequently asked questions

How many source photos should I prepare for a batch of clips?
More than the number of clips you need, by a wide margin. Hit rates are not high — plan for roughly one usable clip from every two or three attempts, and each attempt is better served by a different source than by regenerating the same one. Twenty good stills is a sensible afternoon's worth.
Should I shoot specifically for animation, or reuse existing photos?
Shoot for it if you can. The requirements are specific enough that photos taken for other purposes often miss them — particularly headroom for camera movement and separation from the background. Ten minutes of shooting with animation in mind beats an hour of sorting an existing library.
How do I keep a consistent look across several clips?
Fix the variables you control before you generate anything: the same lighting setup, the same framing, the same wardrobe, and the same motion description. Consistency comes from the inputs, because the generation itself varies every run.
What aspect ratio should I shoot in?
Vertical, and frame loosely. Short-form platforms are 9:16, and clips arrive in that shape. Shooting loose gives you room to crop rather than discovering the motion pushed your subject out of frame.
Is it worth generating the same prompt multiple times?
Yes, within reason. Generation is stochastic, so two runs from identical inputs give different results and one is often clearly better. Three attempts is a reasonable ceiling before the problem is the source photo rather than luck.
How should I budget time for this?
Generation is minutes; selection is the real cost. Expect to spend longer choosing between attempts than waiting for them. Build that into the plan rather than treating each clip as a single quick step.

Ready to put this into practice?

Create with Kitana using the tool that fits this guide.

Create an AI video

Ready to try it yourself?

Download Kitana and create your first AI photo in under a minute.