AI GlossaryㅁTechnical words in the news
Multimodal agent
An AI agent that carries out real tasks by exchanging not just text but multiple formats of material such as images, audio, video, and 3D models
In plain words
A multimodal agent is like a highly capable assistant that doesn't just trade words back and forth — it can look at photos, listen to sounds, produce video, and even work with 3D models. While older chatbots only answered questions with sentences, this kind of assistant uses several senses at once: identifying a bird species from its call, spotting abnormalities in a heart ultrasound image, or sketching a furniture layout from a photo of a kitchen.
The catch is that these outputs don't have a single correct answer. You can't score them the way you'd check a math problem's number or run code to see if it passes or fails. So to evaluate this kind of ability, people sometimes pit outputs from different agents against each other and have a human or another agent judge which one is better.
In the end, a multimodal agent is AI built to handle real-world work that mixes multiple types of material — tasks that text alone can't solve.
How it shows up in the news
Meta AI introduced a new grading framework, describing it as an attempt to "take a step further in measuring the practical value that multimodal agents actually deliver." What's easy to misread here is that this isn't an announcement of a new model — it's the release of the grading standard itself for measuring how good multimodal agents are.
Try it yourself
- In the task viewer on Meta AI's developer page, look through the list of tasks such as bird sound analysis, heart ultrasound reading, and 3D modeling.
- Pick one task and check how its three components — the instruction, the input material (image, audio, YAML), and the grading rubric — are structured.
- Feed the same instruction into two or three different AI chatbots and compare their outputs (image descriptions, tables, summaries, etc.) side by side. Judging for yourself which one followed the instructions better will give you a feel for why evaluation without a single correct answer is so hard.
See also
Stories using this term
- Google Research unveils AI that prioritizes depression biomarker candidates from wearable dataAI · 2026.08.22
- Meta Unveils First 10 Tasks in WildArtifactBench, a Benchmark for AI AgentsAI · 2026.08.21
- Meta releases 30B-parameter open model for local agentsAI · 2026.08.10
- Meta's Muse Spark 1.2 ties for third with Intelligence Index score of 54AI · 2026.08.09
- Grok 4.6 unveiled: top-tier performance at half the priceAI · 2026.08.13
- Meta's paid agent Hatch to launch alongside new Watermelon model in OctoberAI · 2026.08.26
