UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval
arXiv:2608.185042026-08-20
A universal search AI learns to compare candidates side by side to explain why one is the right match
Multimodal retrieval systems usually look at a query and each candidate separately and just measure similarity, which makes it hard to tell a correct match from a deceptively similar wrong one. UMER instead makes the model reason over a query and a candidate together, explicitly comparing what matches and what doesn't, while jointly training fast vector-based embedding search and precise pairwise ranking inside one model so the two can reinforce each other. On the 78-task MMEB-V2 benchmark, this approach outperformed prior methods while letting users trade off speed and accuracy as needed.
What they did
- Problem identified: existing chain-of-thought (CoT) methods reason about the query and each candidate independently, so they never produce evidence explaining why a correct answer beats a visually or semantically similar wrong one (a hard negative).
- Solution: UMER trains 'Pair-Aware Discriminative Reasoning,' where the model is shown a query alongside a positive candidate and a confusable hard-negative candidate together, and must explain the matching and discrepancy evidence between them.
- Architecture: a single multimodal large language model (MLLM) jointly learns a fast embedding function for large-scale vector search and a discriminative ranking function that scores query-candidate pairs directly, linked through 'Complementary Mutual Distillation (CMD)' that transfers knowledge between the two only when the source is reliable.
- Data construction: Qwen3.5-9B generates 'query intent' and 'target observations' reasoning for each positive/negative pair, and a small verifier model (Qwen3.5-0.8B) checks whether the reasoning alone is enough to infer the correct label, filtering out low-quality traces.
- Results: on the 78-task MMEB-V2 benchmark, UMER's hybrid mode reached an overall score of 65.5, the best among compared methods, and its embedding-only mode was up to 118.6x faster than a prior reasoning-based method (UME-R1) at query time.
Table 1: Main results on MMEB-V2. The best and second-best scores in each column are in bold and underlined, respectively.| Model | Image | Video | VisDoc | All |
|---|
| CLS | QA | RET | GD | Overall | CLS | QA | RET | MRET | Overall | VDRv1 | VDRv2 | VR | OOD | Overall | |
| # of Datasets | 10 | 10 | 12 | 4 | 36 | 5 | 5 | 5 | 3 | 18 | 10 | 4 | 6 | 4 | 24 | 78 |
| Baseline Models |
| GME | 54.4 | 29.9 | 66.9 | 55.5 | 51.9 | 34.9 | 42.0 | 25.6 | 32.4 | 33.9 | 86.1 | 54.0 | 82.5 | 43.1 | 72.7 | 54.1 |
| VLM2Vec | 58.7 | 49.3 | 65.0 | 72.9 | 59.7 | 33.4 | 30.5 | 20.6 | 33.0 | 29.0 | 49.8 | 13.5 | 51.8 | 33.5 | 41.6 | 47.0 |
| VLM2Vec-V2 | 62.9 | 56.3 | 69.5 | 77.3 | 64.9 | 39.3 | 34.3 | 28.8 | 38.5 | 34.9 | 75.5 | 44.9 | 79.4 | 39.4 | 65.4 | 58.0 |
| DUME | 59.3 | 55.0 | 66.3 | 78.0 | 62.5 | 37.7 | 46.6 | 17.1 | 30.0 | 33.2 | 67.6 | 43.3 | 47.1 | 33.8 | 52.8 | 52.7 |
| BToks | 64.3 | 59.8 | 68.8 | 77.4 | 66.0 | 43.7 | 47.0 | 33.0 | 33.6 | 39.9 | 71.1 | 38.6 | 81.3 | 38.1 | 62.7 | 59.0 |
| UME-R1 | 64.8 | 62.8 | 67.6 | 77.2 | 66.6 | 44.3 | 51.2 | 32.9 | 39.7 | 42.2 | 72.4 | 46.2 | 79.2 | 37.2 | 63.9 | 60.1 |
| PLUME | 66.5 | 59.2 | 67.6 | 79.7 | 66.3 | 45.0 | 52.3 | 33.5 | 46.7 | 44.1 | 72.1 | 49.8 | 78.1 | 57.4 | 67.5 | 61.6 |
| RIME | 67.9 | 64.4 | 69.8 | 82.1 | 69.1 | 48.0 | 52.1 | 33.6 | 39.2 | 43.7 | 76.4 | 51.4 | 81.7 | 63.9 | 71.4 | 64.1 |
| Ours |
| UMER-E | 64.4 | 64.9 | 70.2 | 78.4 | 68.0 | 43.0 | 50.6 | 34.7 | 32.8 | 41.1 | 76.1 | 50.1 | 82.8 | 68.3 | 72.2 | 63.1 |
| UMER-R | 64.9 | 68.4 | 69.7 | 74.1 | 68.5 | 51.4 | 51.6 | 35.3 | 33.3 | 44.0 | 75.1 | 48.3 | 84.6 | 67.9 | 71.8 | 63.9 |
| UMER-H | 66.1 | 68.9 | 71.7 | 80.8 | 70.4 | 50.6 | 52.8 | 37.7 | 35.0 | 45.0 | 78.0 | 50.5 | 85.0 | 68.8 | 73.6 | 65.5 |
Table 2: Controlled ablations of pair-aware CoT, ranking supervision and complementary mutual distillation (CMD). The best and second-best scores in each column are in bold and underlined, respectively.| Configuration | Embedding mode (UMER-E) | Ranking mode (UMER-R) | Hybrid mode (UMER-H) |
|---|
| Image | Video | VisDoc | All | Image | Video | VisDoc | All | Image | Video | VisDoc | All |
| Embedding only | 66.5 | 37.8 | 69.2 | 60.7 | – | – | – | – | – | – | – | – |
| + pair-aware CoT, w/o ranking losses | 68.2 | 41.1 | 71.2 | 62.9 | 44.4 | 31.8 | 57.1 | 45.4 | 68.6 | 42.2 | 71.3 | 63.3 |
| + ranking losses, w/o CoT | 68.0 | 40.2 | 69.3 | 62.0 | 68.6 | 43.1 | 69.4 | 63.0 | 70.8 | 45.1 | 71.2 | 65.0 |
| + pair-aware CoT and ranking losses | 67.9 | 39.6 | 71.1 | 62.4 | 67.2 | 44.2 | 72.1 | 63.4 | 70.3 | 44.9 | 73.4 | 65.4 |
| Pair-aware model, w/o CMD | 67.9 | 39.6 | 71.1 | 62.4 | 67.2 | 44.2 | 72.1 | 63.4 | 70.3 | 44.9 | 73.4 | 65.4 |
| + ranking → embedding only | 67.9 | 40.4 | 71.5 | 62.7 | 68.5 | 43.3 | 70.9 | 63.4 | 70.6 | 44.7 | 72.7 | 65.3 |
| + embedding → ranking only | 67.8 | 40.4 | 70.1 | 62.2 | 69.2 | 43.9 | 71.1 | 63.9 | 70.8 | 45.1 | 72.6 | 65.5 |
| + always-on bidirectional CMD | 68.0 | 40.2 | 71.1 | 62.6 | 68.5 | 43.6 | 71.0 | 63.6 | 70.7 | 45.3 | 73.0 | 65.5 |
| + selective bidirectional CMD (full) | 68.0 | 41.1 | 72.2 | 63.1 | 68.5 | 44.0 | 71.8 | 63.9 | 70.4 | 45.0 | 73.6 | 65.5 |
Table 3: Accuracy–efficiency on MMEB-V2 using one NVIDIA A100 GPU (batch size=1). UMER-H reuses the UMER-E index; 0. denotes no extra indexing cost.| Model | K | Score ↑ | Reasoning Tokens | Query Latency | Indexing Time |
|---|
| (/query) | (s/query) ↓ | (s/candidate) ↓ |
| UME-R1 | – | 60.1 | 352 | 9.963 | 11.755 |
| PLUME | – | 61.6 | 8 | 0.329 | 0.366 |
| UMER-E | – | 63.1 | 0 | 0.084 | 0.118 |
| UMER-H | 3 | 64.8 | 410 | 13.566 | 0. |
| 5 | 65.5 | 665 | 19.880 |
| 10 | 66.0 | 1305 | 42.764 |
| 20 | 66.1 | 2508 | 81.726 |
Table 1: Implementation and evaluation settings for UMER.| Setting | Value |
|---|
| Model and optimization |
| Initialization | Qwen2-VL-2B-Instruct. |
| Embedding tokens | Four learnable tokens per input (M=4). |
| Embedding temperature | 0.02. |
| Objective weights | λcot=0.2, λbce=0.1, λmargin=0.2, and λcmd=0.2. |
| Optimizer and schedule | AdamW; per-device batch size 64; one accumulation step; linear schedule with peak rate 5×10−5, 100 warmup steps and 10,000 maximum steps. |
| Adaptation | LoRA rank 16, scaling 64 and dropout 0.1 on attention/MLP projections and the language-model head; visual encoder frozen. |
| Precision and preprocessing | BF16 with FlashAttention-2; maximum image pixels 2,359,296. |
| Inference |
| Candidate selection | Retrieve with the embedding branch and rerank the top K=5 candidates. |
| Ranking generation | At most 256 newly generated tokens per candidate. |
| Hybrid fusion | Separately z-score normalize embedding and ranking scores over the five candidates. |
| Ranking-score weight | α=2.0. |
| Evaluation and reporting |
| Benchmark | MMEB-V2: 78 datasets across image, video and visual-document modalities. |
| Benchmark versions | Corrected ViDoSeek-page and MMLongBench-page versions. |
| Metrics | Hit@1 for image and video tasks; NDCG@5 for visual-document tasks. |
| Aggregation | Unweighted macro averages over datasets, including all 78 tasks for All. |
Table 4: Ranking-score-weight sensitivity of UMER-H at K=5, with candidates, decoding configuration, and normalization fixed.| Ranking-score weight | 0.5 | 0.75 | 1 | 1.5 | 2 | 3 | 5 |
|---|
| UMER-H | 64.7 | 65.0 | 65.2 | 65.5 | 65.5 | 65.4 | 65.3 |
Table 5: Decision overlap of embedding retrieval (E) and pair-aware ranking (R) over 83,530 MMEB-V2 queries. Both, E only, R only, and Neither indicate which branch produces the correct top decision. Each cell is a within-family percentage.| Family | Both | E only | R only | Neither |
|---|
| Reasoning / semantic | 48.8 | 9.0 | 12.1 | 30.1 |
| Content matching | 45.4 | 8.9 | 8.3 | 37.4 |
| All | 47.2 | 9.0 | 10.3 | 33.5 |
Why it matters
This offers a practical approach for real-world search or recommendation systems that need to quickly filter huge pools of images, videos, or documents while still accurately distinguishing tricky near-miss candidates. The ability to adjust the speed-accuracy tradeoff also makes it directly useful for production system design.
Terms in this paper
- Embedding · A numeric vector representation of text or images placed so that similar items end up close together
- Hard Negative · A wrong candidate that looks or reads very similarly to the correct answer, making it easy to confuse
- Chain-of-Thought (CoT) · A technique where a model writes out intermediate reasoning steps before giving a final answer
- Reranking · A second-stage process that re-orders a small set of quickly retrieved candidates using a more precise but slower model
- Distillation · A method where one model or function's judgments are used to train another so it learns similar behavior
Figures we cannot republish
- Figure 1: Two overlooked issues in universal multimodal retrieval. (a) Item-wise self-reflective CoT lacks pair-aware evidence for hard-negative discrimination. (b) Different meta- tasks require different capabilities, with embedding and ranking offering complementary strengths.
- Figure 2: Overview of UMER, a unified multimodal embedding and ranking framework. (a) Query and candidate inputs are first encoded independently through masked branch attention to extract their embeddings, and are then jointly used for pair-aware autoregressive reasoning followed by a ranking token. (b) Multi-task heads support metric embedding learning, generative CoT learning and discriminative ranking learning. (c) Unified co-optimization combines multi-task supervision and complementary mutual distillation to jointly improve embedding and ranking.
- Figure 3: Embedding separation under item-wise and pair-aware supervision on 3,600 queries from the 36 image tasks of MMEB-V2. (a) Bootstrap distribution of the macro-averaged positive–hard-negative margin. (b) Fraction of queries exceeding margin thresholds.
- Figure 4: Capability specialization and transfer via CMD. (a) Embedding is stronger for content matching, while ranking is stronger for reasoning-intensive relevance judgment. (b) CMD transfers these complementary strengths between the two functions.
See the figures in the original paper →Original abstract (English)
Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically derive representations from hidden states, while Chain-of-Thought (CoT) reasoning is emerging as a promising strategy for embedding enhancement by encoding intermediate semantic evidence into the representation space. However, exist
Authors · Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang, Zhenyu Tang
Read on arXiv