Glossary — Images, audio and video

Multimodal

A multimodal model can work with more than one kind of input — text, images and audio at the same time, for instance.

3 tokensMultimodal

Simply put#

Older models handled text only. Today's take images, audio and video, and can produce them too.

Example#

You can screenshot an error message and ask "what is wrong here" — the model reads the text and the context straight off the picture.

Why it matters to you#

It changes what is worth asking at all. Photographing a document and asking about it is often faster than retyping it.

Related terms