AI GlossaryㅁTechnical words in the news
Multimodal Media Embedding
A technology that converts different forms of content, such as video, sound, and text, into a single set of numerical coordinates so they can be compared and searched by meaning.
In plain words
Multimodal media embedding is a technology that turns different kinds of content, like video, sound, and text, into a single set of numbers a computer can compare against each other.
Imagine a broadcaster's storage room piled with tapes, where the only record of what's on each tape is a handwritten notebook kept by one staff member. If that person quits, nobody knows what's inside anymore. This technology takes that person's place: it scans through the scenes in a video, the words people say, and even the background sounds, then compresses the features of each moment into a single coordinate, like a fingerprint. Photos, videos, and audio all end up placed on the same coordinate system despite their different forms, which makes it possible to calculate how similar they are to one another.
Once these coordinates exist, content can be searched without anyone tagging it in advance. If someone types a sentence like "stairs" or "a shot taken on a rainy day," that sentence is converted into a coordinate the same way, and the system finds the moments whose coordinates are closest to it. This coordinate conversion is exactly what makes searching by meaning, rather than by tags, possible.
How it shows up in the news
The article explains that Marengo, Twelve Labs' multimodal media embedding model, powers this meaning-based search. One easy point of confusion is that the technology itself is not a search bar or a service. It's the groundwork that converts content into coordinates; the actual search is carried out by a separate system that compares those coordinates.
Try it yourself
When using a service with video search features, try typing a full sentence describing a situation instead of tags or keywords. For example, enter something specific like "a player in a white shirt shooting from left to right on a rainy day," and see whether results come back even without any human-added tags. That's a good way to feel what this technology actually does.
See also
Stories using this term
- Twelve Labs and Mimir Integrate AI Search for Video ArchivesAI · 2026.08.05
- Runway unveils Solaris, which redraws the screen with every clickAI · 2026.09.01
- Perplexity's local agent beats Hermes, Pi in benchmarksAI · 2026.08.26
- Modly: Open-Source App Converts Photos to 3D Models Using Only a GPUAI · 2026.08.21
- FLUX 3 Video launches to general availability, adds sound at 20 seconds, but no Korean supportAI · 2026.08.05
- Runway's Ruby converts AI video to broadcast HDR standardsCreative · 2026.08.25
