매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

arXiv:2608.185042026-08-20

이미지·영상·문서를 아우르는 검색 AI가 후보를 비교하며 '왜 맞는지' 스스로 따져보게 만들었다

멀티모달 검색 AI는 보통 질문과 후보를 각각 따로 살펴본 뒤 비슷한 정도만 계산해서 순위를 매긴다. UMER은 질문과 후보 쌍을 나란히 놓고 어디가 맞고 어디가 다른지 비교 추론을 시킨 뒤, 빠른 임베딩 검색과 정밀한 순위 판정을 한 모델 안에서 함께 학습시켜 서로의 약점을 보완하게 했다. 그 결과 MMEB-V2라는 78개 과제 벤치마크에서 기존 방법들보다 높은 점수를 내면서도, 필요에 따라 속도와 정확도를 조절할 수 있었다.

무엇을 했나

  1. 문제의식: 기존 방식은 질문과 후보 이미지를 각각 따로 설명하는 추론(아이템별 CoT)만 하기 때문에, 비슷하게 생긴 오답(하드 네거티브)과 정답을 구별할 근거를 만들지 못한다는 한계를 지적했다.
  2. 해결책: 정답 후보와 헷갈리는 오답 후보를 한 쌍으로 묶어 모델에게 같이 보여주고, 무엇이 일치하고 무엇이 어긋나는지 비교하며 설명하는 '쌍 기반 추론(Pair-Aware Discriminative Reasoning)'을 학습시켰다.
  3. 구조: 하나의 멀티모달 대형언어모델(MLLM) 안에서 빠른 벡터 검색용 임베딩 학습과, 쌍을 직접 비교해 관련도를 점수로 매기는 순위 학습을 동시에 진행했고, 서로의 판단이 믿을 만할 때만 지식을 주고받는 '상호 증류(CMD)'로 두 기능을 맞물려 개선했다.
  4. 데이터: Qwen3.5-9B로 정답·오답 쌍마다 '질문 의도'와 '후보 관찰 내용'을 담은 추론 문장을 만들고, 작은 검증 모델(Qwen3.5-0.8B)이 그 설명만 보고 정답/오답을 맞힐 수 있는지 걸러내 품질을 확보했다.
  5. 성능: 78개 과제로 구성된 MMEB-V2 벤치마크에서 전체 평균 65.5점으로 최고 수준을 기록했고, 순수 임베딩 검색만 쓸 때도 기존 추론 기반 방법(UME-R1) 대비 최대 118.6배 빠른 속도를 보였다.
Table 1: Main results on MMEB-V2. The best and second-best scores in each column are in bold and underlined, respectively.
ModelImageVideoVisDocAll
CLSQARETGDOverallCLSQARETMRETOverallVDRv1VDRv2VROODOverall
# of Datasets101012436555318104642478
Baseline Models
GME54.429.966.955.551.934.942.025.632.433.986.154.082.543.172.754.1
VLM2Vec58.749.365.072.959.733.430.520.633.029.049.813.551.833.541.647.0
VLM2Vec-V262.956.369.577.364.939.334.328.838.534.975.544.979.439.465.458.0
DUME59.355.066.378.062.537.746.617.130.033.267.643.347.133.852.852.7
BToks64.359.868.877.466.043.747.033.033.639.971.138.681.338.162.759.0
UME-R164.862.867.677.266.644.351.232.939.742.272.446.279.237.263.960.1
PLUME66.559.267.679.766.345.052.333.546.744.172.149.878.157.467.561.6
RIME67.964.469.882.169.148.052.133.639.243.776.451.481.763.971.464.1
Ours
UMER-E64.464.970.278.468.043.050.634.732.841.176.150.182.868.372.263.1
UMER-R64.968.469.774.168.551.451.635.333.344.075.148.384.667.971.863.9
UMER-H66.168.971.780.870.450.652.837.735.045.078.050.585.068.873.665.5
Table 2: Controlled ablations of pair-aware CoT, ranking supervision and complementary mutual distillation (CMD). The best and second-best scores in each column are in bold and underlined, respectively.
ConfigurationEmbedding mode (UMER-E)Ranking mode (UMER-R)Hybrid mode (UMER-H)
ImageVideoVisDocAllImageVideoVisDocAllImageVideoVisDocAll
Embedding only66.537.869.260.7
+ pair-aware CoT, w/o ranking losses68.241.171.262.944.431.857.145.468.642.271.363.3
+ ranking losses, w/o CoT68.040.269.362.068.643.169.463.070.845.171.265.0
+ pair-aware CoT and ranking losses67.939.671.162.467.244.272.163.470.344.973.465.4
Pair-aware model, w/o CMD67.939.671.162.467.244.272.163.470.344.973.465.4
+ ranking → embedding only67.940.471.562.768.543.370.963.470.644.772.765.3
+ embedding → ranking only67.840.470.162.269.243.971.163.970.845.172.665.5
+ always-on bidirectional CMD68.040.271.162.668.543.671.063.670.745.373.065.5
+ selective bidirectional CMD (full)68.041.172.263.168.544.071.863.970.445.073.665.5
Table 3: Accuracy–efficiency on MMEB-V2 using one NVIDIA A100 GPU (batch size=1). UMER-H reuses the UMER-E index; 0. denotes no extra indexing cost.
ModelKScore ↑Reasoning TokensQuery LatencyIndexing Time
(/query)(s/query) ↓(s/candidate) ↓
UME-R160.13529.96311.755
PLUME61.680.3290.366
UMER-E63.100.0840.118
UMER-H364.841013.5660.
565.566519.880
1066.0130542.764
2066.1250881.726
Table 1: Implementation and evaluation settings for UMER.
SettingValue
Model and optimization
InitializationQwen2-VL-2B-Instruct.
Embedding tokensFour learnable tokens per input (M=4).
Embedding temperature0.02.
Objective weightsλcot=0.2, λbce=0.1, λmargin=0.2, and λcmd=0.2.
Optimizer and scheduleAdamW; per-device batch size 64; one accumulation step; linear schedule with peak rate 5×10−5, 100 warmup steps and 10,000 maximum steps.
AdaptationLoRA rank 16, scaling 64 and dropout 0.1 on attention/MLP projections and the language-model head; visual encoder frozen.
Precision and preprocessingBF16 with FlashAttention-2; maximum image pixels 2,359,296.
Inference
Candidate selectionRetrieve with the embedding branch and rerank the top K=5 candidates.
Ranking generationAt most 256 newly generated tokens per candidate.
Hybrid fusionSeparately z-score normalize embedding and ranking scores over the five candidates.
Ranking-score weightα=2.0.
Evaluation and reporting
BenchmarkMMEB-V2: 78 datasets across image, video and visual-document modalities.
Benchmark versionsCorrected ViDoSeek-page and MMLongBench-page versions.
MetricsHit@1 for image and video tasks; NDCG@5 for visual-document tasks.
AggregationUnweighted macro averages over datasets, including all 78 tasks for All.
Table 4: Ranking-score-weight sensitivity of UMER-H at K=5, with candidates, decoding configuration, and normalization fixed.
Ranking-score weight0.50.7511.5235
UMER-H64.765.065.265.565.565.465.3
Table 5: Decision overlap of embedding retrieval (E) and pair-aware ranking (R) over 83,530 MMEB-V2 queries. Both, E only, R only, and Neither indicate which branch produces the correct top decision. Each cell is a within-family percentage.
FamilyBothE onlyR onlyNeither
Reasoning / semantic48.89.012.130.1
Content matching45.48.98.337.4
All47.29.010.333.5

왜 중요한가

검색 서비스나 추천 시스템처럼 대량의 이미지·영상·문서 중에서 빠르게 후보를 걸러내면서도 헷갈리는 오답을 정확히 걸러내야 하는 실무 환경에 바로 적용할 수 있는 방법을 제시했다. 속도와 정확도 중 상황에 맞게 조절할 수 있다는 점도 실제 서비스 설계에 유용하다.

이 논문의 용어

  • 임베딩(Embedding) · 텍스트나 이미지를 숫자 벡터로 바꿔 비슷한 것끼리 가깝게 배치하는 표현 방식
  • 하드 네거티브(Hard Negative) · 정답과 겉모습이나 주제가 비슷해서 헷갈리기 쉬운 오답 후보
  • CoT(Chain-of-Thought) · 결론을 내기 전에 중간 추론 과정을 글로 풀어써서 모델이 단계적으로 생각하게 하는 기법
  • 리랭킹(Reranking) · 1차로 빠르게 뽑은 후보들을 더 정밀한 모델로 다시 순위를 매기는 2단계 검색 방식
  • 지식 증류(Distillation) · 한 모델(또는 기능)의 판단을 다른 모델(또는 기능)이 학습해서 닮아가게 하는 방법

본문에 싣지 못한 그림

  • Figure 1: Two overlooked issues in universal multimodal retrieval. (a) Item-wise self-reflective CoT lacks pair-aware evidence for hard-negative discrimination. (b) Different meta- tasks require different capabilities, with embedding and ranking offering complementary strengths.
  • Figure 2: Overview of UMER, a unified multimodal embedding and ranking framework. (a) Query and candidate inputs are first encoded independently through masked branch attention to extract their embeddings, and are then jointly used for pair-aware autoregressive reasoning followed by a ranking token. (b) Multi-task heads support metric embedding learning, generative CoT learning and discriminative ranking learning. (c) Unified co-optimization combines multi-task supervision and complementary mutual distillation to jointly improve embedding and ranking.
  • Figure 3: Embedding separation under item-wise and pair-aware supervision on 3,600 queries from the 36 image tasks of MMEB-V2. (a) Bootstrap distribution of the macro-averaged positive–hard-negative margin. (b) Fraction of queries exceeding margin thresholds.
  • Figure 4: Capability specialization and transfer via CMD. (a) Embedding is stronger for content matching, while ranking is stronger for reasoning-intensive relevance judgment. (b) CMD transfers these complementary strengths between the two functions.
원문에서 그림 보기 →

논문 원문 초록 (영문)

Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically derive representations from hidden states, while Chain-of-Thought (CoT) reasoning is emerging as a promising strategy for embedding enhancement by encoding intermediate semantic evidence into the representation space. However, exist

저자 · Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang, Zhenyu Tang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사