What is Multimodal?
Multimodal means an AI can handle more than one kind of input: not just text, but also images, sound or video.
A pure text model understands only text. A multimodal model can also look at a photo, read a chart or listen to sound. You could upload a photo of a form and ask what it says. The newest tools are often multimodal.
It makes AI more useful in practice: you do not have to type everything out, but can also have it look at an image or a recording.
Put precisely
Being able to look is not the same as looking well. A table or handwriting in a photo regularly trips it up.
Common misconception
That the model understands an image the way you do. It recognizes patterns in the image.
Related words
Understand AI, calmly
This word is a start. The inklaretaal Learning curve explains AI step by step in plain language, tuned to your field. The first month is free.