What it is
A multimodal model handles several types of information, called modalities, in one system. Text is one modality. Images, audio and video are others. A text-only model reads and writes words. A multimodal model might look at a photo and describe it, listen to a recording and summarize it, or answer a question about a chart.
Think of the difference between a person who can only read letters and one who can also look at pictures and listen to a call. The second one understands more of what is going on.
How it works
Everything a model sees has to become numbers. For text, a tokenizer splits words into tokens. A multimodal model uses similar encoders to turn image patches, slices of audio or video frames into tokens too, so one network can process them together. That shared representation lets it connect a spoken question to what is in a picture.
Models differ in what they output. Some accept images and audio but reply only in text. Others also generate images or speech. Check which direction a given model supports, because "multimodal" often describes input only.
Why it matters to you
- More tasks, one tool. You can send a screenshot of an error, a photo of a receipt, or a spreadsheet chart and ask a question about it, with no separate tool.
- Cost. Images, audio and video turn into many tokens, so they use more of the context window and cost more than the same request in text.
- Limits. Models can misread small text in images, miss details in long video, or describe things that are not there, so check important results.
- Open versions. Some multimodal models are open-weight models you can run yourself, which matters for privacy-sensitive material.
Example
You photograph a whiteboard after a meeting and ask the model to turn it into a task list with owners. A text-only model cannot see the board. A multimodal model reads the handwriting, groups the items and drafts the list, and you correct any misread names.
Many multimodal systems use a mixture of experts design to keep running costs down, and the extra inference work for large images or video is a real part of the bill.