Multimodal AI (Vision + Language)
One model, multiple modalities
A multimodal model processes and reasons over more than one type of input — most commonly text and images, though audio and video are increasingly common too. Modern flagship models like GPT-4V/GPT-4o, Claude, and Gemini are all vision-language models (VLMs): they can take an image and a text prompt together and produce a text response that draws on both, whether that's describing a photo, reading a chart, debugging a screenshot of code, or answering questions about a document's layout.
The core idea that makes this possible: turn the image into something that looks, structurally, like a sequence of tokens — the same kind of input a transformer already knows how to process for text — and let the same attention mechanism that relates words to each other also relate image content to words.
Turning an image into tokens: patches
Text is naturally sequential — a sentence is already a sequence of discrete units. An image is a 2D grid of pixels, not a sequence, so it needs to be converted into a sequence-like form before a transformer can process it. The standard approach, introduced by the Vision Transformer (ViT), is to slice the image into a grid of fixed-size patches — commonly 16x16 pixel squares — and treat each patch as one "token."
Each patch is flattened into a vector and passed through a linear projection to produce a patch embedding, exactly analogous to how a word gets converted into a word embedding. Positional information is added (just as with text) so the model knows where in the image each patch came from. The result is a sequence of patch embeddings that a transformer can process with the same self-attention machinery used for text — a 224x224 pixel image split into 16x16 patches becomes a sequence of 196 "tokens."
Cross-modal attention: connecting words and pixels
Once an image is represented as a sequence of patch embeddings, the model needs a way to relate those visual tokens to the text tokens in the prompt — to let a question like "what color is the car on the left?" actually attend to the specific patches showing that car. This is cross-modal attention: the same query-key-value attention mechanism used within text, extended so text tokens can attend to image tokens and vice versa.
Architecturally, this is commonly implemented by first encoding the image with a vision encoder (often a ViT, sometimes contrastively pretrained via a method like CLIP, which learns to align images and their text descriptions in a shared embedding space), then projecting the resulting visual tokens into the same embedding space the language model uses for text tokens. From that point, the image tokens and text tokens are concatenated into a single combined sequence and processed together by the transformer's ordinary self-attention layers — the model doesn't need a fundamentally different mechanism for cross-modal reasoning, just a shared representational space that lets attention operate across both.
Note
What VLMs can and can't reliably do
Vision-language models handle a wide range of tasks well: describing scenes, reading and interpreting charts and diagrams, transcribing text in images (OCR-adjacent tasks), answering questions about document layout, comparing multiple images, and reasoning about spatial relationships in reasonably simple scenes.
They still struggle with certain categories: precise counting of many similar objects, exact pixel-level measurements or coordinates, fine-grained detail in small or low-resolution image regions (since detail smaller than a patch is compressed away during patch tokenization), and reading dense, small text in cluttered images. It's also worth being clear about what a VLM is *not*: it doesn't have true visual perception the way a human does — it's converting pixels into a token sequence and pattern-matching over that sequence the same way it pattern-matches over word tokens, with all the same strengths and hallucination risks that implies.
Text tokens vs. image patch tokens
| Aspect | Text tokens | Image patch tokens |
|---|---|---|
| Source unit | Subword pieces from a vocabulary | Fixed-size pixel patches (e.g. 16x16) |
| Vocabulary | Fixed, discrete vocabulary (e.g. ~100k tokens) | Continuous embeddings, no fixed vocabulary |
| Positional meaning | Sequential order in a sentence | 2D spatial location in the image |
| Granularity tradeoff | Rare words split into multiple tokens | Smaller patches = more detail but more tokens/cost |
| Processed by | Standard transformer self-attention | Same self-attention, after projecting to shared space |
See attention connect tokens
This attention visualization illustrates how tokens relate to each other via query-key-value scoring — the exact same mechanism, extended across modalities, that lets a VLM connect a question's text tokens to the specific image patches that answer it.
🧠 Attention Visualizer
InteractiveNote
What's next
Multimodal models extend the transformer's attention mechanism across modalities using a shared embedding space — the same core idea behind text embeddings, just generalized. If you want the foundational version of that idea, the Embeddings lesson covers how it works for text, which is the simpler case this builds on.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.