The Rise of AI Video Generation: What to Expect
Where the capability boundary actually sits, which limits are engineering problems and which are structural, and what that implies for what improves next.
Writing about where this technology is going tends to produce either breathless prediction or a list of model releases that is stale by publication. Neither helps. What is more durable is knowing which limits are engineering problems and which are structural, because that tells you what will move and what will not.
The constraint everything else follows from
A video model has to satisfy two demands at once: produce plausible content, and keep every frame consistent with its neighbours.
Those pull against each other. Weight motion too heavily and identity drifts — the warping everyone has seen. Weight consistency too heavily and you get a still image played repeatedly.
Every current limitation is downstream of that tension, which is why it is worth knowing before reading anything about the future.
Which limits are structural
These will improve but not disappear, because they follow from how the generation works.
Error accumulation over time. Each frame depends on the last, so small errors compound. Longer clips are not merely more expensive; they are harder in a way that scales badly. This is why four seconds is a common ceiling and why extending it requires changing the approach rather than raising a budget.
Plausible is not accurate. The model produces something consistent with its training, never something verified. No amount of scale changes that, and it is why generated text, counts and specific facts remain unreliable.
Cost per second. Video frames cannot be generated independently — the model attends across time as well as space — so a second of video will keep costing far more than a still. The gap narrows with efficiency work; it does not close.
Which limits are engineering
These are the ones genuinely moving.
Hands. Steadily improving. Hard because they are articulated, self-occluding and high-degree-of-freedom, and because an error in motion is far more visible than in a still. Not structural — just difficult.
Resolution and frame rate. Straightforwardly a function of compute and efficiency.
Prompt adherence. How closely output follows a description has improved markedly and continues to. Largely a training and conditioning problem.
Control. The ability to specify camera movement, keep a character consistent across shots, or edit a generated clip after the fact. This is where the most visible product progress is happening, because it is what separates a toy from a tool.
What has actually changed so far
Worth being concrete, because the headline is usually about quality when the real change is elsewhere.
Cost collapsed. Producing a few seconds of usable motion used to require a camera, a location and a person. It now requires a photograph and a sentence. That is a change in who can produce, not in what the best output looks like.
The format found its shape. Short, vertical, silent clips are the native output, and that happens to match what short-form platforms display. This is why photo-to-video found a use before long-form generation did.
Provenance became infrastructure. Platforms now label AI content using signals including C2PA Content Credentials, which attach signed metadata at creation. The direction is toward provenance declared at source rather than detection after the fact — detection being unreliable in both directions.
What that implies for using it
Use it where four seconds is the native format, not as a compromise. Openers, cutaways, bringing a still to life.
Do not plan around capabilities that have not arrived. If your idea needs hands doing something specific, or legible text, or a consistent character across eight shots, it is not a settings problem.
Expect a low hit rate and budget for selection rather than generation. The waiting is fast; the judging is the work. See preparing for AI video.
Evaluate tools on difficult inputs. Every showreel is a selected result. Give a new tool a busy background, a partly turned face and visible hands, and look at the floor rather than the ceiling.
What gets more valuable
When execution becomes cheap, execution stops distinguishing anyone. What remains scarce is specificity — something only you could have made — judgement about which of twenty attempts to publish, and consistency over time.
That argument, and what it means for short-form platforms specifically, is in how short-form AI video is reshaping TikTok and Reels. The mechanics underlying all of it are in how photo to video works and why video is harder than images.
Photo to Video runs in the Kitana apps as part of Pro: ten clips a month at four seconds and 768 by 1344.
Frequently asked questions
- Why are clips still only a few seconds long?
- Because errors compound. Each frame is generated relative to the previous one, so a subject that is slightly off at one second is visibly wrong by eight. That is a property of generating sequentially under a consistency constraint, not a compute budget that will simply be raised.
- Will clip length keep increasing?
- Gradually, and not by brute force. The approaches that extend length meaningfully change how consistency is maintained — keeping a longer memory of the subject, or generating keyframes first and filling between them — rather than simply generating more frames the same way.
- Are hands going to be fixed?
- They keep improving and they remain the most common tell in motion. Hands are hard because they are high-degree-of-freedom articulated objects that self-occlude constantly, and a moving error is far more visible than a static one.
- Is generated video going to replace filming?
- Not for anything where the specific thing being filmed matters. It is displacing stock footage, simple cutaways and concept work — the categories where the footage was already generic. What only you can film is the part that gets more valuable, not less.
- How should I evaluate a new tool?
- Give it a difficult input — a subject with hands visible, a busy background, a face partly turned — and look at what it does. Every showreel is a selected result from many attempts, so the demo tells you about the ceiling and nothing about the floor.
- What is the biggest practical change so far?
- Cost, not quality. Producing a few seconds of usable motion went from requiring a camera, a location and a person to requiring a photograph and a sentence. That changes who can produce, which changes the competitive landscape more than any quality improvement does.
Ready to put this into practice?
Create with Kitana using the tool that fits this guide.
Try image generationReady to try it yourself?
Download Kitana and create your first AI photo in under a minute.