Multimodal AI

From ALT-TEXT
Jump to navigation Jump to search

Multimodal AI

An AI system that can process and generate more than one type of data, such as text, images, audio, and video, within the same model. Modern chatbots that can describe an uploaded photo, transcribe a voice note, or generate video from a text prompt are all examples of multimodal AI. (See also: Generative AI, Large language model)