AI, decoded

What is multimodal AI?

Multimodal AI is a model that takes in and reasons across more than one kind of data — text, images, audio, video — in a single system. Instead of a separate model for each, one model can read a chart and answer questions about it, transcribe speech and act on it, or describe a video, and some multimodal models can now generate images as well. A hard part isn't handling each modality; it's alignment — getting the model to connect what it sees, hears, and reads into one coherent understanding.

· Updated · Chain of Thought

Level 1: How AI works · 1.5 From chatbot to agent

Multimodal AIModel Architecture

Text, image, audio and video streams converging into a single model, beside the line: one model, many inputs.

One model, many kinds of input

A text-only model reads and writes words. A multimodal model takes in several kinds of data — words, pixels, sound, frames of video — and reasons over them together. That’s what lets a single system answer a question about a photo, pull the key figure out of a chart, or follow a spoken instruction without a separate pipeline bolted on for each input type. Some multimodal models can now generate images as well as read them.

Alignment is a hard part

Accepting an image and accepting text is the easy half. The difficulty is cross-modal alignment: making the model relate the thing it sees to the thing it reads so they form one understanding rather than two disconnected ones. When alignment is weak, the model describes an image fluently but gets the details wrong, or answers about what it expected to see instead of what’s actually there.

Update, October 3, 2026: Combining modalities also raises a system design problem. At 3:39 in episode 70, Tormod Ree describes cameras and microphones supplying signals to help interpret who is present, who is talking and how people interact. Ree identifies combining signals across multiple devices as a hard problem. At 6:27, Ree locates the differentiation in the harness and specialized data. These passages describe a system combining sensor signals; they do not establish that it uses one general multimodal model.

Why it matters, and where it breaks

Multimodal models open up tasks that pure-text systems can’t touch — visual question answering, document understanding, video analysis, voice interfaces. They also widen the surface for failure: a model can hallucinate about an image as confidently as about text, and evaluating “did it get the picture right” is harder than grading words. The capability is real; so is the need to measure it carefully.

Multimodal AI: from a separate pipeline per input type to one model that reasons across all of them Before: a separate system for each kind of input, a text model, a vision model and a speech model, each its own pipeline; Logan Kilpatrick recalls that one domain-specific computer vision problem took on the order of nine to twelve months to solve in a successful case, collecting data, training end to end and putting it in production. After: text, images, audio and video flow into one model that reasons over them together, and he notes such models now do image segmentation, bounding boxes and object detection. Where the inputs meet is a hard part, alignment: connecting what the model sees to what it reads into one understanding. Underneath, where it breaks: weak alignment gives a fluent description with the details wrong, a model can hallucinate about an image as confidently as about text, and checking whether it got the picture right is harder than grading words, so measure it. Many pipelines become one model Words, pixels, sound and video frames, reasoned over together instead of bolted on one at a time. BEFORE: ONE SYSTEM PER INPUT text → text model image → vision model speech → speech model Logan Kilpatrick: one domain-specific vision problem took nine to twelve months, in a successful case. NOW: ONE MODEL, MANY INPUTS text image audio video one model answers about a photo, follows a spoken instruction, segments, boxes, detects alignment A hard part: connecting what it sees to what it reads, so it's one understanding, not two. WHERE IT BREAKS A fluent description of the picture, with the details wrong. A model can hallucinate about an image as confidently as about text, and checking "did it get the picture right" is harder than grading words. The capability is real; measure it.
Many pipelines become one model, and alignment is where it can break. The nine-to-twelve-month figure and the vision capabilities are Logan Kilpatrick’s, from episode 12. Download the image

Hear it from the guest

“The amount of time it would take to solve one of these, like, domain specific computer vision problems was on the order of nine to twelve months in a successful case.”
Logan Kilpatrick, Google DeepMind EP 12 · around 19:10 · read the transcript
“The models can do image segmentation now, and they can do bounding boxes and object detection.”
Logan Kilpatrick, Google DeepMind EP 12 · around 19:53 · read the transcript
“I think the real differentiation isn't so much in the model for us.”

Quotes lightly edited to remove filler words.

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Multimodal AIGenerative AI