In more detail
A multimodal model works in more than one medium — it can look at pictures, listen to audio, watch video and read text, often producing several of those too. Earlier AI was single-sense: text in, text out. Multimodal models are closer to how people actually experience the world.
It matters because most of real life isn’t typed. You can snap a photo of a rash, a broken appliance or a homework problem and just ask about it. Tools like ChatGPT, Claude and Gemini are all multimodal now — you talk to them in whatever form is handiest.
Goes with
