AI systems capable of processing and generating multiple types of data — such as text, images, and audio — together.
Multimodal models can, for example, answer questions about an uploaded image or generate a description from audio, expanding what AI applications can be built to do.