Multimodal AI
Jump to navigation
Jump to search
Multimodal AI
An AI system that can process and generate more than one type of data, such as text, images, audio, and video, within the same model. Modern chatbots that can describe an uploaded photo, transcribe a voice note, or generate video from a text prompt are all examples of multimodal AI. (See also: Generative AI, Large language model)