METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㄴTechnical words in the news

Native Multimodal Model

An AI model that learns to understand text, images, sound, and video all together from the start, within a single model.

In plain words

A native multimodal model is an AI that learns text, images, sound, and video all at once, from the very beginning, inside one single model. Here, "modal" refers to the form information comes in (whether it's text, an image, or sound), and "multimodal" means handling several of these forms together.

The older approach was like relaying information through a chain of different translators. A model good at reading images would first generate a caption like "there's a dog in this photo," and that text would then be passed to a language model to produce the final answer. Along the way, subtle details—like a facial expression in a photo or the tone of a voice—were easily lost.

A native multimodal model is closer to someone who grew up learning multiple languages and senses at the same time. Instead of passing images, text, and voice through separate steps, it connects and understands them directly within one model, so less information gets lost and it can answer questions that span multiple formats more naturally.

Try it yourself

Try uploading a photo along with a voice memo (or a short video) to an AI chatbot, and ask it to connect the two in its answer. Example prompt: "Check whether the object in this photo matches what's described in the voice memo I just uploaded, and point out any differences." If the model smoothly explains both pieces of information together without relying on separate passes through different models, it's likely using a native multimodal approach.

See also

Stories using this term

Browse every entry