One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

arXiv:2608.185042026-08-20

A universal search AI learns to compare candidates side by side to explain why one is the right match

Multimodal retrieval systems usually look at a query and each candidate separately and just measure similarity, which makes it hard to tell a correct match from a deceptively similar wrong one. UMER instead makes the model reason over a query and a candidate together, explicitly comparing what matches and what doesn't, while jointly training fast vector-based embedding search and precise pairwise ranking inside one model so the two can reinforce each other. On the 78-task MMEB-V2 benchmark, this approach outperformed prior methods while letting users trade off speed and accuracy as needed.

What they did

  1. Problem identified: existing chain-of-thought (CoT) methods reason about the query and each candidate independently, so they never produce evidence explaining why a correct answer beats a visually or semantically similar wrong one (a hard negative).
  2. Solution: UMER trains 'Pair-Aware Discriminative Reasoning,' where the model is shown a query alongside a positive candidate and a confusable hard-negative candidate together, and must explain the matching and discrepancy evidence between them.
  3. Architecture: a single multimodal large language model (MLLM) jointly learns a fast embedding function for large-scale vector search and a discriminative ranking function that scores query-candidate pairs directly, linked through 'Complementary Mutual Distillation (CMD)' that transfers knowledge between the two only when the source is reliable.
  4. Data construction: Qwen3.5-9B generates 'query intent' and 'target observations' reasoning for each positive/negative pair, and a small verifier model (Qwen3.5-0.8B) checks whether the reasoning alone is enough to infer the correct label, filtering out low-quality traces.
  5. Results: on the 78-task MMEB-V2 benchmark, UMER's hybrid mode reached an overall score of 65.5, the best among compared methods, and its embedding-only mode was up to 118.6x faster than a prior reasoning-based method (UME-R1) at query time.
Table 1: Main results on MMEB-V2. The best and second-best scores in each column are in bold and underlined, respectively.
ModelImageVideoVisDocAll
CLSQARETGDOverallCLSQARETMRETOverallVDRv1VDRv2VROODOverall
# of Datasets101012436555318104642478
Baseline Models
GME54.429.966.955.551.934.942.025.632.433.986.154.082.543.172.754.1
VLM2Vec58.749.365.072.959.733.430.520.633.029.049.813.551.833.541.647.0
VLM2Vec-V262.956.369.577.364.939.334.328.838.534.975.544.979.439.465.458.0
DUME59.355.066.378.062.537.746.617.130.033.267.643.347.133.852.852.7
BToks64.359.868.877.466.043.747.033.033.639.971.138.681.338.162.759.0
UME-R164.862.867.677.266.644.351.232.939.742.272.446.279.237.263.960.1
PLUME66.559.267.679.766.345.052.333.546.744.172.149.878.157.467.561.6
RIME67.964.469.882.169.148.052.133.639.243.776.451.481.763.971.464.1
Ours
UMER-E64.464.970.278.468.043.050.634.732.841.176.150.182.868.372.263.1
UMER-R64.968.469.774.168.551.451.635.333.344.075.148.384.667.971.863.9
UMER-H66.168.971.780.870.450.652.837.735.045.078.050.585.068.873.665.5
Table 2: Controlled ablations of pair-aware CoT, ranking supervision and complementary mutual distillation (CMD). The best and second-best scores in each column are in bold and underlined, respectively.
ConfigurationEmbedding mode (UMER-E)Ranking mode (UMER-R)Hybrid mode (UMER-H)
ImageVideoVisDocAllImageVideoVisDocAllImageVideoVisDocAll
Embedding only66.537.869.260.7
+ pair-aware CoT, w/o ranking losses68.241.171.262.944.431.857.145.468.642.271.363.3
+ ranking losses, w/o CoT68.040.269.362.068.643.169.463.070.845.171.265.0
+ pair-aware CoT and ranking losses67.939.671.162.467.244.272.163.470.344.973.465.4
Pair-aware model, w/o CMD67.939.671.162.467.244.272.163.470.344.973.465.4
+ ranking → embedding only67.940.471.562.768.543.370.963.470.644.772.765.3
+ embedding → ranking only67.840.470.162.269.243.971.163.970.845.172.665.5
+ always-on bidirectional CMD68.040.271.162.668.543.671.063.670.745.373.065.5
+ selective bidirectional CMD (full)68.041.172.263.168.544.071.863.970.445.073.665.5
Table 3: Accuracy–efficiency on MMEB-V2 using one NVIDIA A100 GPU (batch size=1). UMER-H reuses the UMER-E index; 0. denotes no extra indexing cost.
ModelKScore ↑Reasoning TokensQuery LatencyIndexing Time
(/query)(s/query) ↓(s/candidate) ↓
UME-R160.13529.96311.755
PLUME61.680.3290.366
UMER-E63.100.0840.118
UMER-H364.841013.5660.
565.566519.880
1066.0130542.764
2066.1250881.726
Table 1: Implementation and evaluation settings for UMER.
SettingValue
Model and optimization
InitializationQwen2-VL-2B-Instruct.
Embedding tokensFour learnable tokens per input (M=4).
Embedding temperature0.02.
Objective weightsλcot=0.2, λbce=0.1, λmargin=0.2, and λcmd=0.2.
Optimizer and scheduleAdamW; per-device batch size 64; one accumulation step; linear schedule with peak rate 5×10−5, 100 warmup steps and 10,000 maximum steps.
AdaptationLoRA rank 16, scaling 64 and dropout 0.1 on attention/MLP projections and the language-model head; visual encoder frozen.
Precision and preprocessingBF16 with FlashAttention-2; maximum image pixels 2,359,296.
Inference
Candidate selectionRetrieve with the embedding branch and rerank the top K=5 candidates.
Ranking generationAt most 256 newly generated tokens per candidate.
Hybrid fusionSeparately z-score normalize embedding and ranking scores over the five candidates.
Ranking-score weightα=2.0.
Evaluation and reporting
BenchmarkMMEB-V2: 78 datasets across image, video and visual-document modalities.
Benchmark versionsCorrected ViDoSeek-page and MMLongBench-page versions.
MetricsHit@1 for image and video tasks; NDCG@5 for visual-document tasks.
AggregationUnweighted macro averages over datasets, including all 78 tasks for All.
Table 4: Ranking-score-weight sensitivity of UMER-H at K=5, with candidates, decoding configuration, and normalization fixed.
Ranking-score weight0.50.7511.5235
UMER-H64.765.065.265.565.565.465.3
Table 5: Decision overlap of embedding retrieval (E) and pair-aware ranking (R) over 83,530 MMEB-V2 queries. Both, E only, R only, and Neither indicate which branch produces the correct top decision. Each cell is a within-family percentage.
FamilyBothE onlyR onlyNeither
Reasoning / semantic48.89.012.130.1
Content matching45.48.98.337.4
All47.29.010.333.5

Why it matters

This offers a practical approach for real-world search or recommendation systems that need to quickly filter huge pools of images, videos, or documents while still accurately distinguishing tricky near-miss candidates. The ability to adjust the speed-accuracy tradeoff also makes it directly useful for production system design.

Terms in this paper

  • Embedding · A numeric vector representation of text or images placed so that similar items end up close together
  • Hard Negative · A wrong candidate that looks or reads very similarly to the correct answer, making it easy to confuse
  • Chain-of-Thought (CoT) · A technique where a model writes out intermediate reasoning steps before giving a final answer
  • Reranking · A second-stage process that re-orders a small set of quickly retrieved candidates using a more precise but slower model
  • Distillation · A method where one model or function's judgments are used to train another so it learns similar behavior

Figures we cannot republish

  • Figure 1: Two overlooked issues in universal multimodal retrieval. (a) Item-wise self-reflective CoT lacks pair-aware evidence for hard-negative discrimination. (b) Different meta- tasks require different capabilities, with embedding and ranking offering complementary strengths.
  • Figure 2: Overview of UMER, a unified multimodal embedding and ranking framework. (a) Query and candidate inputs are first encoded independently through masked branch attention to extract their embeddings, and are then jointly used for pair-aware autoregressive reasoning followed by a ranking token. (b) Multi-task heads support metric embedding learning, generative CoT learning and discriminative ranking learning. (c) Unified co-optimization combines multi-task supervision and complementary mutual distillation to jointly improve embedding and ranking.
  • Figure 3: Embedding separation under item-wise and pair-aware supervision on 3,600 queries from the 36 image tasks of MMEB-V2. (a) Bootstrap distribution of the macro-averaged positive–hard-negative margin. (b) Fraction of queries exceeding margin thresholds.
  • Figure 4: Capability specialization and transfer via CMD. (a) Embedding is stronger for content matching, while ranking is stronger for reasoning-intensive relevance judgment. (b) CMD transfers these complementary strengths between the two functions.
See the figures in the original paper →

Original abstract (English)

Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically derive representations from hidden states, while Chain-of-Thought (CoT) reasoning is emerging as a promising strategy for embedding enhancement by encoding intermediate semantic evidence into the representation space. However, exist

Authors · Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang, Zhenyu Tang

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB