What is multimodal AI?
Multimodal AI is a model that takes in and reasons across more than one kind of data — text, images, audio, video — in a single system. Instead of a separate model for each, one model can read a chart and answer questions about it, transcribe speech and act on it, or describe a video, and some multimodal models can now generate images as well. A hard part isn't handling each modality; it's alignment — getting the model to connect what it sees, hears, and reads into one coherent understanding.
Level 1: How AI works · 1.5 From chatbot to agent
Multimodal AIModel Architecture
One model, many kinds of input
A text-only model reads and writes words. A multimodal model takes in several kinds of data — words, pixels, sound, frames of video — and reasons over them together. That’s what lets a single system answer a question about a photo, pull the key figure out of a chart, or follow a spoken instruction without a separate pipeline bolted on for each input type. Some multimodal models can now generate images as well as read them.
Alignment is a hard part
Accepting an image and accepting text is the easy half. The difficulty is cross-modal alignment: making the model relate the thing it sees to the thing it reads so they form one understanding rather than two disconnected ones. When alignment is weak, the model describes an image fluently but gets the details wrong, or answers about what it expected to see instead of what’s actually there.
Update, October 3, 2026: Combining modalities also raises a system design problem. At 3:39 in episode 70, Tormod Ree describes cameras and microphones supplying signals to help interpret who is present, who is talking and how people interact. Ree identifies combining signals across multiple devices as a hard problem. At 6:27, Ree locates the differentiation in the harness and specialized data. These passages describe a system combining sensor signals; they do not establish that it uses one general multimodal model.
Why it matters, and where it breaks
Multimodal models open up tasks that pure-text systems can’t touch — visual question answering, document understanding, video analysis, voice interfaces. They also widen the surface for failure: a model can hallucinate about an image as confidently as about text, and evaluating “did it get the picture right” is harder than grading words. The capability is real; so is the need to measure it carefully.
Hear it from the guest
“The amount of time it would take to solve one of these, like, domain specific computer vision problems was on the order of nine to twelve months in a successful case.”
“The models can do image segmentation now, and they can do bounding boxes and object detection.”
“I think the real differentiation isn't so much in the model for us.”
Quotes lightly edited to remove filler words.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.