UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval
arXiv:2608.185042026-08-20
이미지·영상·문서를 아우르는 검색 AI가 후보를 비교하며 '왜 맞는지' 스스로 따져보게 만들었다
멀티모달 검색 AI는 보통 질문과 후보를 각각 따로 살펴본 뒤 비슷한 정도만 계산해서 순위를 매긴다. UMER은 질문과 후보 쌍을 나란히 놓고 어디가 맞고 어디가 다른지 비교 추론을 시킨 뒤, 빠른 임베딩 검색과 정밀한 순위 판정을 한 모델 안에서 함께 학습시켜 서로의 약점을 보완하게 했다. 그 결과 MMEB-V2라는 78개 과제 벤치마크에서 기존 방법들보다 높은 점수를 내면서도, 필요에 따라 속도와 정확도를 조절할 수 있었다.
무엇을 했나
- 문제의식: 기존 방식은 질문과 후보 이미지를 각각 따로 설명하는 추론(아이템별 CoT)만 하기 때문에, 비슷하게 생긴 오답(하드 네거티브)과 정답을 구별할 근거를 만들지 못한다는 한계를 지적했다.
- 해결책: 정답 후보와 헷갈리는 오답 후보를 한 쌍으로 묶어 모델에게 같이 보여주고, 무엇이 일치하고 무엇이 어긋나는지 비교하며 설명하는 '쌍 기반 추론(Pair-Aware Discriminative Reasoning)'을 학습시켰다.
- 구조: 하나의 멀티모달 대형언어모델(MLLM) 안에서 빠른 벡터 검색용 임베딩 학습과, 쌍을 직접 비교해 관련도를 점수로 매기는 순위 학습을 동시에 진행했고, 서로의 판단이 믿을 만할 때만 지식을 주고받는 '상호 증류(CMD)'로 두 기능을 맞물려 개선했다.
- 데이터: Qwen3.5-9B로 정답·오답 쌍마다 '질문 의도'와 '후보 관찰 내용'을 담은 추론 문장을 만들고, 작은 검증 모델(Qwen3.5-0.8B)이 그 설명만 보고 정답/오답을 맞힐 수 있는지 걸러내 품질을 확보했다.
- 성능: 78개 과제로 구성된 MMEB-V2 벤치마크에서 전체 평균 65.5점으로 최고 수준을 기록했고, 순수 임베딩 검색만 쓸 때도 기존 추론 기반 방법(UME-R1) 대비 최대 118.6배 빠른 속도를 보였다.
Table 1: Main results on MMEB-V2. The best and second-best scores in each column are in bold and underlined, respectively.| Model | Image | Video | VisDoc | All |
|---|
| CLS | QA | RET | GD | Overall | CLS | QA | RET | MRET | Overall | VDRv1 | VDRv2 | VR | OOD | Overall | |
| # of Datasets | 10 | 10 | 12 | 4 | 36 | 5 | 5 | 5 | 3 | 18 | 10 | 4 | 6 | 4 | 24 | 78 |
| Baseline Models |
| GME | 54.4 | 29.9 | 66.9 | 55.5 | 51.9 | 34.9 | 42.0 | 25.6 | 32.4 | 33.9 | 86.1 | 54.0 | 82.5 | 43.1 | 72.7 | 54.1 |
| VLM2Vec | 58.7 | 49.3 | 65.0 | 72.9 | 59.7 | 33.4 | 30.5 | 20.6 | 33.0 | 29.0 | 49.8 | 13.5 | 51.8 | 33.5 | 41.6 | 47.0 |
| VLM2Vec-V2 | 62.9 | 56.3 | 69.5 | 77.3 | 64.9 | 39.3 | 34.3 | 28.8 | 38.5 | 34.9 | 75.5 | 44.9 | 79.4 | 39.4 | 65.4 | 58.0 |
| DUME | 59.3 | 55.0 | 66.3 | 78.0 | 62.5 | 37.7 | 46.6 | 17.1 | 30.0 | 33.2 | 67.6 | 43.3 | 47.1 | 33.8 | 52.8 | 52.7 |
| BToks | 64.3 | 59.8 | 68.8 | 77.4 | 66.0 | 43.7 | 47.0 | 33.0 | 33.6 | 39.9 | 71.1 | 38.6 | 81.3 | 38.1 | 62.7 | 59.0 |
| UME-R1 | 64.8 | 62.8 | 67.6 | 77.2 | 66.6 | 44.3 | 51.2 | 32.9 | 39.7 | 42.2 | 72.4 | 46.2 | 79.2 | 37.2 | 63.9 | 60.1 |
| PLUME | 66.5 | 59.2 | 67.6 | 79.7 | 66.3 | 45.0 | 52.3 | 33.5 | 46.7 | 44.1 | 72.1 | 49.8 | 78.1 | 57.4 | 67.5 | 61.6 |
| RIME | 67.9 | 64.4 | 69.8 | 82.1 | 69.1 | 48.0 | 52.1 | 33.6 | 39.2 | 43.7 | 76.4 | 51.4 | 81.7 | 63.9 | 71.4 | 64.1 |
| Ours |
| UMER-E | 64.4 | 64.9 | 70.2 | 78.4 | 68.0 | 43.0 | 50.6 | 34.7 | 32.8 | 41.1 | 76.1 | 50.1 | 82.8 | 68.3 | 72.2 | 63.1 |
| UMER-R | 64.9 | 68.4 | 69.7 | 74.1 | 68.5 | 51.4 | 51.6 | 35.3 | 33.3 | 44.0 | 75.1 | 48.3 | 84.6 | 67.9 | 71.8 | 63.9 |
| UMER-H | 66.1 | 68.9 | 71.7 | 80.8 | 70.4 | 50.6 | 52.8 | 37.7 | 35.0 | 45.0 | 78.0 | 50.5 | 85.0 | 68.8 | 73.6 | 65.5 |
Table 2: Controlled ablations of pair-aware CoT, ranking supervision and complementary mutual distillation (CMD). The best and second-best scores in each column are in bold and underlined, respectively.| Configuration | Embedding mode (UMER-E) | Ranking mode (UMER-R) | Hybrid mode (UMER-H) |
|---|
| Image | Video | VisDoc | All | Image | Video | VisDoc | All | Image | Video | VisDoc | All |
| Embedding only | 66.5 | 37.8 | 69.2 | 60.7 | – | – | – | – | – | – | – | – |
| + pair-aware CoT, w/o ranking losses | 68.2 | 41.1 | 71.2 | 62.9 | 44.4 | 31.8 | 57.1 | 45.4 | 68.6 | 42.2 | 71.3 | 63.3 |
| + ranking losses, w/o CoT | 68.0 | 40.2 | 69.3 | 62.0 | 68.6 | 43.1 | 69.4 | 63.0 | 70.8 | 45.1 | 71.2 | 65.0 |
| + pair-aware CoT and ranking losses | 67.9 | 39.6 | 71.1 | 62.4 | 67.2 | 44.2 | 72.1 | 63.4 | 70.3 | 44.9 | 73.4 | 65.4 |
| Pair-aware model, w/o CMD | 67.9 | 39.6 | 71.1 | 62.4 | 67.2 | 44.2 | 72.1 | 63.4 | 70.3 | 44.9 | 73.4 | 65.4 |
| + ranking → embedding only | 67.9 | 40.4 | 71.5 | 62.7 | 68.5 | 43.3 | 70.9 | 63.4 | 70.6 | 44.7 | 72.7 | 65.3 |
| + embedding → ranking only | 67.8 | 40.4 | 70.1 | 62.2 | 69.2 | 43.9 | 71.1 | 63.9 | 70.8 | 45.1 | 72.6 | 65.5 |
| + always-on bidirectional CMD | 68.0 | 40.2 | 71.1 | 62.6 | 68.5 | 43.6 | 71.0 | 63.6 | 70.7 | 45.3 | 73.0 | 65.5 |
| + selective bidirectional CMD (full) | 68.0 | 41.1 | 72.2 | 63.1 | 68.5 | 44.0 | 71.8 | 63.9 | 70.4 | 45.0 | 73.6 | 65.5 |
Table 3: Accuracy–efficiency on MMEB-V2 using one NVIDIA A100 GPU (batch size=1). UMER-H reuses the UMER-E index; 0. denotes no extra indexing cost.| Model | K | Score ↑ | Reasoning Tokens | Query Latency | Indexing Time |
|---|
| (/query) | (s/query) ↓ | (s/candidate) ↓ |
| UME-R1 | – | 60.1 | 352 | 9.963 | 11.755 |
| PLUME | – | 61.6 | 8 | 0.329 | 0.366 |
| UMER-E | – | 63.1 | 0 | 0.084 | 0.118 |
| UMER-H | 3 | 64.8 | 410 | 13.566 | 0. |
| 5 | 65.5 | 665 | 19.880 |
| 10 | 66.0 | 1305 | 42.764 |
| 20 | 66.1 | 2508 | 81.726 |
Table 1: Implementation and evaluation settings for UMER.| Setting | Value |
|---|
| Model and optimization |
| Initialization | Qwen2-VL-2B-Instruct. |
| Embedding tokens | Four learnable tokens per input (M=4). |
| Embedding temperature | 0.02. |
| Objective weights | λcot=0.2, λbce=0.1, λmargin=0.2, and λcmd=0.2. |
| Optimizer and schedule | AdamW; per-device batch size 64; one accumulation step; linear schedule with peak rate 5×10−5, 100 warmup steps and 10,000 maximum steps. |
| Adaptation | LoRA rank 16, scaling 64 and dropout 0.1 on attention/MLP projections and the language-model head; visual encoder frozen. |
| Precision and preprocessing | BF16 with FlashAttention-2; maximum image pixels 2,359,296. |
| Inference |
| Candidate selection | Retrieve with the embedding branch and rerank the top K=5 candidates. |
| Ranking generation | At most 256 newly generated tokens per candidate. |
| Hybrid fusion | Separately z-score normalize embedding and ranking scores over the five candidates. |
| Ranking-score weight | α=2.0. |
| Evaluation and reporting |
| Benchmark | MMEB-V2: 78 datasets across image, video and visual-document modalities. |
| Benchmark versions | Corrected ViDoSeek-page and MMLongBench-page versions. |
| Metrics | Hit@1 for image and video tasks; NDCG@5 for visual-document tasks. |
| Aggregation | Unweighted macro averages over datasets, including all 78 tasks for All. |
Table 4: Ranking-score-weight sensitivity of UMER-H at K=5, with candidates, decoding configuration, and normalization fixed.| Ranking-score weight | 0.5 | 0.75 | 1 | 1.5 | 2 | 3 | 5 |
|---|
| UMER-H | 64.7 | 65.0 | 65.2 | 65.5 | 65.5 | 65.4 | 65.3 |
Table 5: Decision overlap of embedding retrieval (E) and pair-aware ranking (R) over 83,530 MMEB-V2 queries. Both, E only, R only, and Neither indicate which branch produces the correct top decision. Each cell is a within-family percentage.| Family | Both | E only | R only | Neither |
|---|
| Reasoning / semantic | 48.8 | 9.0 | 12.1 | 30.1 |
| Content matching | 45.4 | 8.9 | 8.3 | 37.4 |
| All | 47.2 | 9.0 | 10.3 | 33.5 |
왜 중요한가
검색 서비스나 추천 시스템처럼 대량의 이미지·영상·문서 중에서 빠르게 후보를 걸러내면서도 헷갈리는 오답을 정확히 걸러내야 하는 실무 환경에 바로 적용할 수 있는 방법을 제시했다. 속도와 정확도 중 상황에 맞게 조절할 수 있다는 점도 실제 서비스 설계에 유용하다.
이 논문의 용어
- 임베딩(Embedding) · 텍스트나 이미지를 숫자 벡터로 바꿔 비슷한 것끼리 가깝게 배치하는 표현 방식
- 하드 네거티브(Hard Negative) · 정답과 겉모습이나 주제가 비슷해서 헷갈리기 쉬운 오답 후보
- CoT(Chain-of-Thought) · 결론을 내기 전에 중간 추론 과정을 글로 풀어써서 모델이 단계적으로 생각하게 하는 기법
- 리랭킹(Reranking) · 1차로 빠르게 뽑은 후보들을 더 정밀한 모델로 다시 순위를 매기는 2단계 검색 방식
- 지식 증류(Distillation) · 한 모델(또는 기능)의 판단을 다른 모델(또는 기능)이 학습해서 닮아가게 하는 방법
본문에 싣지 못한 그림
- Figure 1: Two overlooked issues in universal multimodal retrieval. (a) Item-wise self-reflective CoT lacks pair-aware evidence for hard-negative discrimination. (b) Different meta- tasks require different capabilities, with embedding and ranking offering complementary strengths.
- Figure 2: Overview of UMER, a unified multimodal embedding and ranking framework. (a) Query and candidate inputs are first encoded independently through masked branch attention to extract their embeddings, and are then jointly used for pair-aware autoregressive reasoning followed by a ranking token. (b) Multi-task heads support metric embedding learning, generative CoT learning and discriminative ranking learning. (c) Unified co-optimization combines multi-task supervision and complementary mutual distillation to jointly improve embedding and ranking.
- Figure 3: Embedding separation under item-wise and pair-aware supervision on 3,600 queries from the 36 image tasks of MMEB-V2. (a) Bootstrap distribution of the macro-averaged positive–hard-negative margin. (b) Fraction of queries exceeding margin thresholds.
- Figure 4: Capability specialization and transfer via CMD. (a) Embedding is stronger for content matching, while ranking is stronger for reasoning-intensive relevance judgment. (b) CMD transfers these complementary strengths between the two functions.
원문에서 그림 보기 →논문 원문 초록 (영문)
Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically derive representations from hidden states, while Chain-of-Thought (CoT) reasoning is emerging as a promising strategy for embedding enhancement by encoding intermediate semantic evidence into the representation space. However, exist
저자 · Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang, Zhenyu Tang
arXiv에서 원문 보기