KitanaAI Photo & VideoStudio
photo maker · 4 min read

What Is Text-to-Image AI and How Does It Work?

Diffusion explained without the maths: what the model learned, why the same prompt never gives the same picture, and what that means for writing prompts.

By the Kitana team
WIPHOTO MAKER

Text-to-image is the one part of this field where the popular mental model is most wrong, and the wrong model leads directly to bad prompts. So it is worth getting right.

It is not searching for anything

The most common assumption is that the model has a vast library of images and stitches together the ones matching your words. It does not. There is no library.

What it has is a set of learned associations between language and visual structure, built by looking at enormous numbers of image-and-caption pairs. It learned that "golden retriever" tends to co-occur with a particular arrangement of shapes, fur texture and colour; that "at dusk" tends to co-occur with long shadows and warm low light.

When you prompt it, it is not retrieving. It is generating pixels that satisfy those learned associations. That is why it can draw something that has never been photographed, and why you cannot ask it to find a picture you remember.

How the picture actually appears

The technique is called diffusion, and the intuition is simpler than the name suggests.

During training, the model was repeatedly shown an image with a bit of random noise added, then more, then more, until the image was pure static. It learned to run that backwards: given a noisy image, predict what the slightly less noisy version looked like.

To generate, it starts from pure random noise and applies that learned step over and over — twenty, thirty, fifty times — each time removing a little noise and moving toward something coherent. Your prompt steers every step, nudging the result toward the associations your words evoke.

An image emerges the way a shape emerges from fog: not drawn stroke by stroke, but resolved all at once, gradually.

Two things follow immediately.

The same prompt gives different images. Different starting noise, different path, different destination. This is not a bug and there is no "correct" output being approximated.

Nothing is retrieved, so nothing is guaranteed. The model produces something plausible for your words. Plausible is not accurate, which is why generated text is nonsense and generated hands are unreliable.

What this means for prompts

If the model is matching learned associations between words and visual structure, then a good prompt is one made of words with strong visual associations.

Works — concrete, visible things:

  • "A glass greenhouse beside a lake, early morning mist, soft light"
  • "Overhead view of a wooden desk, scattered brass instruments, warm lamp"
  • "Close-up of weathered rope on a stone wall, shallow depth of field"

Does not work — judgements with no visual referent:

  • "Stunning masterpiece, 8k, award-winning, best quality"
  • "Beautiful and professional"

The second set feels like it should help, and it is where most people start. But "award-winning" has no consistent visual form — it appeared in captions across every genre and style, so it steers nothing. Worse, it occupies attention that a describable detail could have used.

The single most useful edit to a weak prompt is to delete every adjective that is not describing something you could point at in the picture.

The parts that are still unreliable

  • Text. Letterforms are visual patterns to the model, not language. Short words sometimes land; sentences do not.
  • Hands and fingers. Improving, still the most common tell.
  • Counting. "Five birds" reliably produces somewhere between three and eight.
  • Spatial relations. "The cup to the left of the book" is followed loosely.
  • Anything factual. It has no facts, only associations.

These are not settings you have got wrong. They are consequences of how the model works, and no amount of prompt engineering fully resolves them.

Where it fits alongside editing tools

Text-to-image creates something from nothing. The other tools take a photo you already have and change it. Different jobs, and it is worth being clear about which you need — restyling a photo keeps your subject, while generating from text does not.

If you want an image of you in a different style, that is restyling or an avatar, not text-to-image. A prompt cannot produce your face, because your face is not one of the associations it learned.

Text to image is one of the four tools in Kitana's browser studio — the one that needs no photo at all, which makes it the easiest place to see how prompts behave. The other thirteen photo tools are in the apps, and the beginner's guide covers which edit to reach for when.

Frequently asked questions

Is the model searching for images and combining them?
No. There is no image database to search. The model holds learned patterns — statistical relationships between words and visual structure — and generates new pixels each time. That is why you cannot ask it to find a specific photo, and why it can produce things that have never existed.
Why do I get a different picture from the same prompt?
Generation starts from random noise and refines it. A different starting noise takes a different path to a different plausible image. Most tools expose this as a seed; fixing the seed and the prompt gives you a repeatable result.
Why is text in generated images always wrong?
The model learned what text looks like, not what it says. Letterforms are visual patterns to it, and it reproduces the texture of writing without the underlying rules. Newer models are better at short words and still unreliable at sentences.
Do longer prompts work better?
Up to a point. Specific visual detail helps; piling on adjectives does not. A prompt is competing for a fixed amount of attention, so every word that is not describing something visible is diluting the words that are.
Who owns what I generate?
Unsettled and jurisdiction-dependent. The US Copyright Office has held that works without human authorship are not copyrightable, which puts purely prompt-generated images on uncertain ground. Check the terms of the tool you used and do not assume exclusivity.
Can it copy a specific artist's style?
It can approximate styles it saw during training, and many tools now restrict naming living artists. Whether that is acceptable is contested and varies by jurisdiction; the practical advice is to describe the visual qualities you want rather than borrowing a person's name.

Ready to put this into practice?

Create with Kitana using the tool that fits this guide.

Generate an image from text

Ready to try it yourself?

Download Kitana and create your first AI photo in under a minute.