AI & MLDeep
Advanced

Multimodal AI (Vision + Language)

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Advanced
Requires advanced math/coding
Interactive
Live Simulator
Contains interactive viz

One model, multiple modalities

A multimodal model processes and reasons over more than one type of input — most commonly text and images, though audio and video are increasingly common too. Modern flagship models like GPT-4V/GPT-4o, Claude, and Gemini are all vision-language models (VLMs): they can take an image and a text prompt together and produce a text response that draws on both, whether that's describing a photo, reading a chart, debugging a screenshot of code, or answering questions about a document's layout.

The core idea that makes this possible: turn the image into something that looks, structurally, like a sequence of tokens — the same kind of input a transformer already knows how to process for text — and let the same attention mechanism that relates words to each other also relate image content to words.

Turning an image into tokens: patches

Text is naturally sequential — a sentence is already a sequence of discrete units. An image is a 2D grid of pixels, not a sequence, so it needs to be converted into a sequence-like form before a transformer can process it. The standard approach, introduced by the Vision Transformer (ViT), is to slice the image into a grid of fixed-size patches — commonly 16x16 pixel squares — and treat each patch as one "token."

Each patch is flattened into a vector and passed through a linear projection to produce a patch embedding, exactly analogous to how a word gets converted into a word embedding. Positional information is added (just as with text) so the model knows where in the image each patch came from. The result is a sequence of patch embeddings that a transformer can process with the same self-attention machinery used for text — a 224x224 pixel image split into 16x16 patches becomes a sequence of 196 "tokens."

text

Cross-modal attention: connecting words and pixels

Once an image is represented as a sequence of patch embeddings, the model needs a way to relate those visual tokens to the text tokens in the prompt — to let a question like "what color is the car on the left?" actually attend to the specific patches showing that car. This is cross-modal attention: the same query-key-value attention mechanism used within text, extended so text tokens can attend to image tokens and vice versa.

Architecturally, this is commonly implemented by first encoding the image with a vision encoder (often a ViT, sometimes contrastively pretrained via a method like CLIP, which learns to align images and their text descriptions in a shared embedding space), then projecting the resulting visual tokens into the same embedding space the language model uses for text tokens. From that point, the image tokens and text tokens are concatenated into a single combined sequence and processed together by the transformer's ordinary self-attention layers — the model doesn't need a fundamentally different mechanism for cross-modal reasoning, just a shared representational space that lets attention operate across both.

Note

The elegance of the modern VLM approach is that it doesn't require inventing new attention mechanisms for vision. It requires getting images and text into the *same* vector space, so that the existing transformer machinery — already good at finding relationships between tokens — just works across modalities without modification. Most of the engineering effort in building a strong VLM goes into training the vision encoder and the projection between the two spaces well, not into changing the transformer itself.

What VLMs can and can't reliably do

Vision-language models handle a wide range of tasks well: describing scenes, reading and interpreting charts and diagrams, transcribing text in images (OCR-adjacent tasks), answering questions about document layout, comparing multiple images, and reasoning about spatial relationships in reasonably simple scenes.

They still struggle with certain categories: precise counting of many similar objects, exact pixel-level measurements or coordinates, fine-grained detail in small or low-resolution image regions (since detail smaller than a patch is compressed away during patch tokenization), and reading dense, small text in cluttered images. It's also worth being clear about what a VLM is *not*: it doesn't have true visual perception the way a human does — it's converting pixels into a token sequence and pattern-matching over that sequence the same way it pattern-matches over word tokens, with all the same strengths and hallucination risks that implies.

Text tokens vs. image patch tokens

AspectText tokensImage patch tokens
Source unitSubword pieces from a vocabularyFixed-size pixel patches (e.g. 16x16)
VocabularyFixed, discrete vocabulary (e.g. ~100k tokens)Continuous embeddings, no fixed vocabulary
Positional meaningSequential order in a sentence2D spatial location in the image
Granularity tradeoffRare words split into multiple tokensSmaller patches = more detail but more tokens/cost
Processed byStandard transformer self-attentionSame self-attention, after projecting to shared space

See attention connect tokens

This attention visualization illustrates how tokens relate to each other via query-key-value scoring — the exact same mechanism, extended across modalities, that lets a VLM connect a question's text tokens to the specific image patches that answer it.

🧠 Attention Visualizer

Interactive
The
cat
sat
on
the
mat
The
0.35
0.14
0.14
0.13
0.11
0.12
cat
0.15
0.34
0.14
0.13
0.13
0.11
sat
0.12
0.15
0.34
0.13
0.13
0.13
on
0.11
0.13
0.15
0.35
0.14
0.13
the
0.13
0.11
0.13
0.15
0.35
0.14
mat
0.12
0.13
0.11
0.13
0.16
0.35
Hover over cells to view exact attention weights.
Low
High
Each row shows how much that token attends to every other token. Rows sum to 1.

Note

High-resolution images get split into more patches, which means more tokens consumed from the context window and higher processing cost — this is why vision-capable APIs typically charge per image based on resolution, and why downsizing images before sending them (when fine detail isn't needed) is a practical cost-optimization technique in production applications.

What's next

Multimodal models extend the transformer's attention mechanism across modalities using a shared embedding space — the same core idea behind text embeddings, just generalized. If you want the foundational version of that idea, the Embeddings lesson covers how it works for text, which is the simpler case this builds on.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.