Multimodal AI
AI that can work with more than one kind of input or output — text, images, audio, and video — rather than just one. A multimodal model might describe a photo, answer a spoken question, or generate a picture from a written description.
Early AI systems tended to specialise in a single medium: one model read text, another recognised images, another handled speech. Multimodal AI brings these together, so a single model can take in and reason across different kinds of data at once. Show it a photograph and ask a question in words, and it can answer using both.
The appeal is that the real world is not neatly divided into separate channels. A doctor reads a scan and the notes beside it; a shopper looks at a product and its description. A multimodal model can mirror that, combining what it sees, hears, and reads to produce a more complete response. The same underlying approach also works in reverse, generating an image from a sentence or a voice from written text.
Most of the best-known general-purpose models are now multimodal to some degree, and the term comes up whenever an AI product handles pictures, sound, or video rather than plain text alone.