Nine AI Photo Tools Worth Knowing
Not apps — the nine kinds of edit these models can do, what each is reliably good at, and which one to reach for.
Lists of AI photo apps go stale within a year. The operations underneath them do not. Learn these nine and you will know what to search for regardless of what the products are called this year.
1. Background replacement
What it does: separates subject from background and replaces everything behind them.
Reliable at: clean subjects with space behind them, plain replacement backdrops.
Fails at: hair against a close or similar-coloured background, glass, anything partly transparent.
The highest-value single edit and the safest to learn, because the subject is untouched. Guide.
2. Inpainting — object removal
What it does: fills a masked region with generated content consistent with its surroundings.
Reliable at: distractions against simple backgrounds — sky, pavement, grass, plain walls.
Fails at: anything over text, faces, or overlapping your subject. It invents rather than recovers. How it works and when it fails.
3. Upscaling
What it does: predicts detail to enlarge an image.
Reliable at: soft photos, heavy crops, small originals, scanned prints — two to four times up.
Fails at: genuinely blurred faces, motion blur, text, anything evidential. What it can and cannot recover.
4. Style transfer
What it does: keeps composition, replaces how the image is rendered.
Reliable at: strong simple shapes — landscapes, architecture, a single subject.
Fails at: fine repeating texture, busy scenes, small faces. Fine detail is style, not content, so it does not survive. How style filters work.
5. Portrait relighting and headshot passes
What it does: evens out lighting on a face, usually alongside a background change.
Reliable at: fixing harsh or uneven light on a decent source.
Fails at: anything asked as a judgement rather than a change. "Professional" is how faces drift. Guide.
6. Avatars and stylised portraits
What it does: redraws a face according to a style's conventions.
Reliable at: producing a recognisable stylised version when hair, face shape and accessories are visible.
Fails at: small faces, extreme angles, preserving fine features — which it replaces by design. Preparing photos.
7. Generation from text
What it does: produces an image from a description, with no source photo.
Reliable at: scenes, objects, concepts, backgrounds.
Fails at: specific real people or places, legible text, counting, exact spatial relations. How it works.
8. Restoration
What it does: repairs damage — scratches, tears, fading — and colourises.
Reliable at: scanned prints with even damage.
Fails at: the same caveat as upscaling. Colour added to a black-and-white photo is invented, not recovered. Keep the original scan.
9. Photo to video
What it does: animates a still into a short clip.
Reliable at: sharp images, clear subjects, separation from the background, one simple motion.
Fails at: hands, text, multiple people, anything beyond a few seconds. How it works.
Why these are converging into one model
Worth knowing, because it changes how you should think about the list.
These began as genuinely separate systems — a segmentation model for backgrounds, an inpainting model for removal, a super-resolution network for upscaling. Each was built and trained for its own job.
They are increasingly the same class of model with different framings around it. A single generative model can be conditioned to fill a masked region, to enlarge, to restyle or to relight, depending on what it is given and what it is asked. Which is why capabilities now tend to arrive across all of these at once rather than one at a time.
The categories stay useful anyway, for a reason that outlives the architecture: they describe what you are asking for, and therefore predict how it will fail. An inpainting request over text fails the same way whether it is served by a dedicated model or a general one, because the failure comes from the request being unanswerable rather than from the implementation.
So learn them as descriptions of problems, not as a list of products.
Choosing between them
Ask what you want to change:
| You want to change | Reach for |
|---|---|
| Everything behind the subject | Background replacement |
| One object | Inpainting |
| Not enough pixels | Upscaling |
| How it is drawn | Style transfer |
| The light on a face | Headshot pass |
| Your face into a style | Avatar |
| Nothing exists yet | Generation |
| Damage on an old print | Restoration |
| A still into motion | Photo to video |
The rule that runs through all nine
Every one of these produces something plausible, never something verified. That is the property to carry between them: excellent when the surrounding context makes the answer inferable, confidently wrong when it does not.
The five errors that follow from forgetting it are in common mistakes.
Four of these run in Kitana's browser studio; all thirteen photo tools and Photo to Video are in the apps.
Frequently asked questions
- Why list kinds of tool rather than named apps?
- Because apps change, merge and disappear on a timescale of months while the operations stay the same. Knowing that your problem is an inpainting problem tells you what to search for and what to expect; knowing last year's app names does not.
- Which one should a beginner learn first?
- Background replacement. It is the highest-value single edit, it never touches the subject's face, and it fails visibly rather than subtly — which makes it the safest place to learn what these tools actually do.
- Which are the least reliable?
- Anything involving fine structure the model must invent: text repair, patterned fabric, jewellery, and anything partly transparent. Generation from text is reliable in itself but cannot produce a specific real person or place.
- Do these run on my phone or on a server?
- Mostly on a server, because the models are too large to run well on a phone. Some simple operations are done locally. On-device processing is better for privacy and meaningfully more limited in capability.
- How do I know which tool my problem needs?
- Ask what you want to change. Everything behind the subject — background replacement. One object — inpainting. Not enough pixels — upscaling. How it is drawn — style transfer. Nothing exists yet — generation.
- Are these all separate tools or one model?
- Increasingly one class of model with different framings around it. The categories remain useful because they describe what you are asking for and predict how it will fail, which is the part that helps.
Ready to try it yourself?
Download Kitana and create your first AI photo in under a minute.