Multimodal AI
The assumption that AI just reads and writes text is already a few years out of date — and most people's habits haven't caught up to that yet.
An outdated mental model
The habit of only ever typing into a chat box comes from an earlier generation of AI tools, where text really was the only input available. Most capable models today can take a screenshot, a photo, an audio clip, a video, or a full document in the same conversation as a text message — and reason across all of it together, not as separate, disconnected steps. The gap between what these tools can actually do and what most people habitually ask of them is largely a leftover mental model, not a technical limitation anymore.
What "multimodal" actually means, technically
A genuinely multimodal model processes different input types within one unified system, rather than running a separate image tool, transcribing audio elsewhere, and stitching the text results together afterward. That distinction matters: a truly multimodal system can reason about how the image and the text relate to each other, not just process each one in isolation and hope the connection comes through.
What each modality is actually good for
Text
Precise instructions, nuance, and anything requiring exact wording.
Images
Layouts, diagrams, visual errors — anything easier to show than describe.
Audio
Tone, emphasis, and spoken content where the words alone lose something.
Video
Sequences and motion — a process happening over time, not a single frame.
Documents & PDFs
Long-form structured content, preserving layout and context together.
Screenshots
Exact current state of something on screen — often faster than describing it.
Why combining beats using one at a time
A screenshot of an error message plus a sentence describing what you were doing when it happened gives a model far more to work with than either alone — the image supplies the exact wording and visual context, the text supplies intent and sequence. Splitting these into two separate exchanges forces the model to work from a text description of an image instead of the image itself, quietly losing detail in the retelling.
Practical techniques worth using
Where this still falls short
Fine visual detail, small text in a low-resolution image, subtle audio cues, and precise spatial reasoning ("is this exactly aligned") remain harder than they look — a multimodal system can process an image without perceiving it with the same precision a person glancing at the same picture would. Treat these edge cases as worth a second, careful check, not as a reason to avoid multimodal input generally.
The habit worth breaking
Typing out a long description of a chart, screenshot, or document that could simply be attached is the most common missed opportunity here — not a mistake exactly, just leftover habit from tools that couldn't take anything but text. The fix is almost embarrassingly simple: if you have the actual file, send the actual file.
The short version
Modern AI tools reason across text, images, audio, video, and documents together, not as bolted-together separate features — and the biggest gap left isn't the technology, it's the habit of defaulting to text when a screenshot, a document, or a photo would give the model far more to work with. Send the actual thing when you have it, combine modalities in one request instead of splitting them, and stay a little more careful with fine visual or audio detail than with plain text.