Glossary — Images, audio and video
Multimodal
A multimodal model can work with more than one kind of input — text, images and audio at the same time, for instance.
3 tokensMultimodal
Simply put#
Older models handled text only. Today's take images, audio and video, and can produce them too.
Example#
You can screenshot an error message and ask "what is wrong here" — the model reads the text and the context straight off the picture.
Why it matters to you#
It changes what is worth asking at all. Photographing a document and asking about it is often faster than retyping it.