Treadstone Associates
Article · 8 min read

How AI image generation works

Typing a sentence and getting a picture back looks like magic. Mechanically, it is closer to sculpting an image out of random noise, one small correction at a time, guided by what the words in the prompt mean to the model.

Treadstone Associates · Updated 2026

Key takeaways

  • • The prompt text is converted into a numerical representation that captures how the model learned words and images relate to each other.
  • • Most current image generators start from random noise and refine it step by step until it matches that representation — they do not retrieve an existing image.
  • • The same prompt can produce a different image each run, because the starting noise and the refinement steps both involve randomness.
  • • Provenance tools like Content Credentials and Google's SynthID try to mark AI-made images afterward, but neither is required by Canadian law, and neither catches every case.

Turning words into a shape the model can use

Canada's Cyber Centre lists this as one of the established uses of the technology today: generative AI can be used to analyze, alter and create visual content for personal or business use. Before any of that happens, the prompt itself has to become something the model can work with mathematically. The text is converted into a numerical representation — an embedding — that captures the relationships between words the model learned during training, so that “a red bicycle leaning against a brick wall” sits, in that mathematical space, close to other things the model has learned are visually associated with those words.

This embedding step is also why a prompt's precise wording matters more than it might seem to. Two phrasings that mean the same thing to a person — “a red bicycle by a brick wall” and “a crimson bike beside a brick building” — can sit in noticeably different places in that mathematical space if the model's training saw those particular words in different visual contexts, which is part of why small wording changes can shift a generated image more than expected.

Starting from noise, not a blank canvas

Most mainstream image generators do not draw the way a person does, building up a scene from scratch. They start from a field of random noise and repeatedly refine it: at each of many small steps, the model nudges the noise slightly closer to a pattern that matches the prompt's embedding, until a coherent image gradually emerges from what began as static. This step-by-step denoising approach is the mechanism behind most current text-to-image tools, whichever product name sits on top of it.

Each individual step makes a comparatively small correction — removing a bit of the noise, shifting a region slightly closer to something recognizable — and it is the accumulation of many such small corrections, not one large leap, that turns static into a finished picture. This is also why generating an image typically takes a noticeable moment rather than returning instantly the way retrieving an existing file would: the system is doing real iterative work at inference time, not looking anything up.

Why the same prompt gives a different image each time

The starting noise is randomly generated, and the refinement process itself has randomness built into several of its steps. The same words can therefore land on genuinely different finished images from one run to the next — a structural property of how the images are produced, not an inconsistency in the model's understanding of the prompt.

Some tools let a user fix that randomness deliberately, by supplying the same starting seed twice, which reproduces a very similar image on both runs. Without deliberately fixing it, variation between runs on an identical prompt is the expected, normal behaviour of the mechanism, not a sign anything went wrong.

Knowing whether an image was AI-made, after the fact

Because the image itself carries no obvious signature of how it was made, a separate set of tools has been built to mark that afterward. The Coalition for Content Provenance and Authenticity describes its standard, Content Credentials, as functioning like a nutrition label for digital content, giving a peek at the content's history available for anyone to access, at any time. Content Credentials' own framing is candid about the limit of that approach: our goal is to provide good actors with a way to demonstrate the authenticity of their content — it proves what a good-faith creator disclosed, and does not, on its own, stop someone determined to strip that information out.

Google's SynthID: the same idea, applied inside the pixels

Google DeepMind takes a different technical route to a similar goal: SynthID embeds digital watermarks directly into AI-generated images, audio, text or video. The watermarks are embedded across Google's generative AI consumer products, and are imperceptible to humans – but can be detected by SynthID's technology. The company is specific about the boundary of that coverage: the watermark applies across Google's generative AI consumer products. It says nothing about an image made with a different company's generator, which is the practical limit of any single vendor's provenance system.

The watermark itself is designed to survive ordinary handling of the image — resizing, a modest crop, a filter — without requiring anything visible to the viewer. That resilience is the point of embedding it in the image data rather than adding a visible logo a person could simply crop out, but it is still one company's system covering that company's own products, not a universal marker every generator applies.

Why this matters beyond curiosity

A business generating marketing images, product mockups or illustrative content at any volume will eventually face a question about disclosure — to a client, a regulator, or simply an audience that wants to know what it is looking at. Understanding that provenance tools prove what a good-faith creator disclosed, rather than catching every case of undisclosed AI content, is what keeps that disclosure honest: relying on the absence of a watermark as proof an image is not AI-generated would be relying on a guarantee neither C2PA nor SynthID actually makes.

See also how AI turns speech into text and generative AI vs agentic AI.

Common questions

Can you tell an AI image apart from a photo just by looking at it?

Increasingly, not reliably by eye alone, which is exactly why provenance tools like Content Credentials and SynthID exist — to carry information a viewer cannot otherwise recover just by looking.

Is a watermark like SynthID required by Canadian law?

No. Neither Content Credentials nor SynthID is required by any Canadian statute. The closest federal hook is a voluntary code of conduct that asks its own signatories to label AI output; it does not bind businesses generally.

If an image has no visible watermark, does that mean it was not AI-generated?

No. Many tools embed no provenance marker at all, and where one exists it can be stripped during editing. Absence of a mark is not evidence of authenticity either way.

Does the same denoising mechanism apply to AI-generated video?

The underlying principle is closely related — most current video generators extend the same noise-to-image refinement across a sequence of frames, with the added requirement of keeping the result consistent from one frame to the next, which is a harder problem than a single still image.

Can a business remove or add a provenance marker itself?

Adding one generally requires the generation tool itself to embed it at creation time; removing one, deliberately or through ordinary editing, is possible for some formats, which is exactly the limit both C2PA and SynthID are candid about above.

Where this goes next

Once a business is generating images at any volume, the practical question becomes how to build that into a repeatable content process. For that, see building an AI content engine from one idea a week.