LLMs GuruLLMsGuru
Find your AI
AI ModelsAI ToolsLearnGlossaryAI NewsFree ToolsFind your AISpeakRightAstrology
Explore All Tools
← The AI Dictionary
Plain-English definition

Multimodal

In one breath

One model, many senses: text, images, audio and video — in and out.

In more detail

A multimodal model works in more than one medium — it can look at pictures, listen to audio, watch video and read text, often producing several of those too. Earlier AI was single-sense: text in, text out. Multimodal models are closer to how people actually experience the world.

It matters because most of real life isn’t typed. You can snap a photo of a rash, a broken appliance or a homework problem and just ask about it. Tools like ChatGPT, Claude and Gemini are all multimodal now — you talk to them in whatever form is handiest.

📌 See it in action

You photograph the inside of your fridge and ask, ‘What can I cook tonight?’ The model recognizes the eggs, spinach and leftover rice in the image, then writes you a recipe — reading a picture and answering in words, two different senses in a single exchange.

Goes with
Text-to-VideoDiffusion ModelLLM

Now put the vocabulary to work.

Meet the models →Which AI is mine?