METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅁTechnical words in the news

Multimodal agent

An AI agent that carries out real tasks by exchanging not just text but multiple formats of material such as images, audio, video, and 3D models

In plain words

A multimodal agent is like a highly capable assistant that doesn't just trade words back and forth — it can look at photos, listen to sounds, produce video, and even work with 3D models. While older chatbots only answered questions with sentences, this kind of assistant uses several senses at once: identifying a bird species from its call, spotting abnormalities in a heart ultrasound image, or sketching a furniture layout from a photo of a kitchen.

The catch is that these outputs don't have a single correct answer. You can't score them the way you'd check a math problem's number or run code to see if it passes or fails. So to evaluate this kind of ability, people sometimes pit outputs from different agents against each other and have a human or another agent judge which one is better.

In the end, a multimodal agent is AI built to handle real-world work that mixes multiple types of material — tasks that text alone can't solve.

How it shows up in the news

Meta AI introduced a new grading framework, describing it as an attempt to "take a step further in measuring the practical value that multimodal agents actually deliver." What's easy to misread here is that this isn't an announcement of a new model — it's the release of the grading standard itself for measuring how good multimodal agents are.

Try it yourself

  1. In the task viewer on Meta AI's developer page, look through the list of tasks such as bird sound analysis, heart ultrasound reading, and 3D modeling.
  2. Pick one task and check how its three components — the instruction, the input material (image, audio, YAML), and the grading rubric — are structured.
  3. Feed the same instruction into two or three different AI chatbots and compare their outputs (image descriptions, tables, summaries, etc.) side by side. Judging for yourself which one followed the instructions better will give you a feel for why evaluation without a single correct answer is so hard.

See also

Stories using this term

Browse every entry