Multimodal AI

From ALT-TEXT
Revision as of 10:36, 7 September 2026 by imported>ALT-TEXT (Import: AI terminology and people glossary)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

Multimodal AI

An AI system that can process and generate more than one type of data, such as text, images, audio, and video, within the same model. Modern chatbots that can describe an uploaded photo, transcribe a voice note, or generate video from a text prompt are all examples of multimodal AI. (See also: Generative AI, Large language model)