An image sent to a multimodal AI model is not read by a second, separate system. It is sliced up and converted into the same kind of unit the model already reasons over. What that mechanism is, and what it changes.
Key takeaways
“Multimodal” sounds like a model gained a second brain for pictures. Mechanically, what actually happens is closer to the opposite: the image is converted into the same kind of unit the model already reasons over — a token — and dropped into the same sequence as the surrounding text. Anthropic’s own developer documentation for its Claude models describes exactly how this works for its vision feature: (Anthropic, Claude API documentation — vision) “Claude views images in patches instead of pixels. Each patch is a 28×28-pixel block of the image, referred to as a visual token. An image, therefore, costs ⌈width / 28⌉ × ⌈height / 28⌉ visual tokens.” A picture, in other words, is chopped into a grid of small squares, and each square becomes one more token sitting in the same context window as your words — not a separate channel the model interprets on the side.
This is also why image size has a direct, calculable cost in the same currency as everything else the model reads: a larger image simply means more visual tokens competing for space in the same context window your text and conversation history are also using.
The 28×28-patch figure above is specific to Anthropic’s implementation, and it is worth naming as such rather than treating it as a universal constant — a different provider’s vision system tokenizes images with its own scheme and its own numbers. What does generalize across providers is the underlying idea: a genuinely multimodal model is not switching between separate systems per input type. It is extending one sequence-based architecture so that images (and, in other products, audio and video) get encoded into the same kind of representation as text, so a single model can reason over all of them together in one pass.
Google DeepMind’s description of its own watermarking tool, SynthID, incidentally confirms this holds across media types rather than being unique to still images: (Google DeepMind, SynthID) the company describes watermarking “images, audio, text or video” through the same family of models, embedding a signal “across Google’s generative AI consumer products” regardless of which of those formats is being generated. Different media, handled by variations of the same underlying mechanism — not four unrelated tools stitched together.
Even inside one sequence, order is not irrelevant. Anthropic’s vision documentation notes a real, measurable effect: (Anthropic — vision documentation) “Claude works best when images come before text,” and when a request includes several images, the documentation recommends labelling each one in the surrounding text (“Image 1:”, “Image 2:”) “so you can refer to them by name in your prompt and in follow-up turns.” None of that would be necessary if the model had a genuinely separate visual-understanding subsystem consulting an index of “the images in this conversation” — the labelling is doing real work precisely because, mechanically, the images are simply more tokens in a sequence, and the model benefits from being told in plain language which token block corresponds to which label.
The same documentation also flags a cost implication worth knowing before building anything at scale: (Anthropic — vision documentation) sending images as inline base64 data means “the full image bytes are included in the payload on every turn” of a multi-turn conversation, since — as covered in why AI forgets what you told it — a model re-reads the whole conversation on every turn by default. An uploaded-once, referenced-by-id approach avoids resending the same image repeatedly, which matters once a conversation runs long.
It does not mean the model gained a sense — no camera, no microphone, no lived perception. It means one more format has been converted into the numerical representation the model already works with. That distinction matters for expectations: a model can describe what is depicted in an image with genuine, useful precision while still being subject to exactly the same failure mode covered in how ChatGPT generates an answer — a fluent, confident description is not proof the underlying interpretation is correct, because the generation mechanism producing the description is the same probability-driven process either way, just now reading visual tokens instead of only text ones.
The same tokenized approach runs in the opposite direction too, when a model generates an image or video rather than describing one — and that direction has produced its own, separate governance question: how does anyone tell that a given image was machine-generated at all? That is a different problem from anything covered above; it is not about whether the model interpreted an image correctly, but about labelling what the model itself produced. The Coalition for Content Provenance and Authenticity addresses exactly this, independent of which company’s model did the generating: (C2PA) it “provides an open technical standard for publishers, creators and consumers to establish the origin and edits of digital content”, through what it calls Content Credentials — described on the same page as functioning “like a nutrition label for digital content”, giving anyone “a peek at the content’s history”. Google DeepMind’s SynthID applies a related idea directly inside the generation step itself, across more than one modality: (Google DeepMind, SynthID) it “embeds digital watermarks directly into AI-generated images, audio, text or video”, watermarks that are “imperceptible to humans” but detectable by the tool built to read them.
Neither mechanism helps a model understand an image better on the way in — they exist to mark what a model produced on the way out, which is a completely separate use of “multimodal” from the input-side mechanism this article is mainly about, and worth not conflating with it.
Related, within this hub: how large language models work and what open-weight AI models change. For connecting a multimodal workflow to systems you already run, see AI integration and automation.
Not in the way that phrase suggests. Anthropic's own documentation describes images being cut into fixed-size patches and converted into the same kind of token the model already uses for text, then placed in the same sequence — one mechanism, extended to a new input type, not a second system bolted on.
Yes. Anthropic's documentation gives an explicit formula for the number of visual tokens an image costs, based on its dimensions, and every one of those tokens counts against the same context window budget as the rest of the conversation.
No reason to assume so. The description is still produced by the same token-by-token generation mechanism described elsewhere in this hub, which is fluent by design and not independently fact-checked as it writes.
A short call is enough to map your inputs against what the mechanism can realistically carry.