Multimodal AI

The assumption that AI just reads and writes text is already a few years out of date — and most people's habits haven't caught up to that yet.

Academy · AI Basics · Advanced Guides

An outdated mental model

The habit of only ever typing into a chat box comes from an earlier generation of AI tools, where text really was the only input available. Most capable models today can take a screenshot, a photo, an audio clip, a video, or a full document in the same conversation as a text message — and reason across all of it together, not as separate, disconnected steps. The gap between what these tools can actually do and what most people habitually ask of them is largely a leftover mental model, not a technical limitation anymore.

What "multimodal" actually means, technically

A genuinely multimodal model processes different input types within one unified system, rather than running a separate image tool, transcribing audio elsewhere, and stitching the text results together afterward. That distinction matters: a truly multimodal system can reason about how the image and the text relate to each other, not just process each one in isolation and hope the connection comes through.

What each modality is actually good for

Text

Precise instructions, nuance, and anything requiring exact wording.

Images

Layouts, diagrams, visual errors — anything easier to show than describe.

Audio

Tone, emphasis, and spoken content where the words alone lose something.

Video

Sequences and motion — a process happening over time, not a single frame.

Documents & PDFs

Long-form structured content, preserving layout and context together.

Screenshots

Exact current state of something on screen — often faster than describing it.

Why combining beats using one at a time

A screenshot of an error message plus a sentence describing what you were doing when it happened gives a model far more to work with than either alone — the image supplies the exact wording and visual context, the text supplies intent and sequence. Splitting these into two separate exchanges forces the model to work from a text description of an image instead of the image itself, quietly losing detail in the retelling.

Practical techniques worth using

Attach the actual file instead of describing it in words when both are available
Combine modalities in one message rather than sending them as separate follow-ups
Be specific about what to focus on within an image or document — "check the total in this invoice," not just "look at this"

Where this still falls short

Fine visual detail, small text in a low-resolution image, subtle audio cues, and precise spatial reasoning ("is this exactly aligned") remain harder than they look — a multimodal system can process an image without perceiving it with the same precision a person glancing at the same picture would. Treat these edge cases as worth a second, careful check, not as a reason to avoid multimodal input generally.

The habit worth breaking

Typing out a long description of a chart, screenshot, or document that could simply be attached is the most common missed opportunity here — not a mistake exactly, just leftover habit from tools that couldn't take anything but text. The fix is almost embarrassingly simple: if you have the actual file, send the actual file.

The short version

Modern AI tools reason across text, images, audio, video, and documents together, not as bolted-together separate features — and the biggest gap left isn't the technology, it's the habit of defaulting to text when a screenshot, a document, or a photo would give the model far more to work with. Send the actual thing when you have it, combine modalities in one request instead of splitting them, and stay a little more careful with fine visual or audio detail than with plain text.