Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
arXiv:2608.185212026-08-20
图文检索AI误以为自己已经全学会了,结果学不会区分那些几乎一样的描述句子
把图片和长篇详细描述句子进行匹配的AI,在训练中使用的损失函数往往在第一轮训练内就迅速跌到接近零,导致模型几乎收不到信号去学习那些最容易混淆、最难区分的错误配对。作者提出HN-CLIP,利用文本编码器自身对句子相似度的判断,自动给更相似(更难区分)的负样本设置更大的惩罚间隔,且不需要额外数据或模型结构。在四个长文本检索基准上,该方法的准确率比此前最强方法高出2.4到4.3个百分点,训练速度也快2.4到5.4倍。
他们做了什么
- 与以往引入分割图、边缘图、LLM筛选文本等复杂机制的方法不同,本文直接针对训练用的损失函数本身的缺陷下手,损失函数负责告诉模型哪里错了、该往哪个方向调整。
- 在长描述文本数据集中,很多句子内容几乎重复;第一轮训练内就有80%的批次损失值降到接近零,47%的测量中梯度(告诉模型如何更新参数的信号)精确为零,说明模型很早就停止从难题中学习了。
- 解决方法HN-CLIP会计算一批数据中文本与文本之间的相似度矩阵,将其从反向传播中剥离(不参与梯度更新),再加到正确配对与错误配对的得分差上,使得越相似(越难区分)的负样本需要更大的得分差距才能被判定为已学会。
- 这一方法不增加额外数据、不新增模型部件、推理阶段零额外开销,仅需一次矩阵乘法和加法,却在四个基准的全部八个检索方向上都取得了最佳的第一名命中率,且仅用20%的训练数据就超过了此前最强方法用100%数据训练的效果。
- 将该损失函数直接替换进六种不同的现有微调框架(Long-CLIP、FineLIP、GOAL、StructXLIP、LoRA、DoRA)中,每一种在同域测试上都获得了提升,说明这一改进具有普适性,不依赖特定模型结构。
Table 1: Cross-modal retrieval performance of CLIP-based fine-tuning methods on four dense-caption benchmarks. We report Recall@K (%) on both Text→Image and Image→Text settings. All fine-tuned methods start from the same Long-CLIP-L backbone with an identical training budget; Long-CLIP denotes the released checkpoint. Best results in bold; second best underlined. Δ denotes the margin over the best competitor per column, with gain in ↑ green. | DOCCI | DCI |
|---|
| Method | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |
| Long-CLIP[ECCV’24] | 78.78 | 95.24 | 98.02 | 66.75 | 91.92 | 96.31 | 67.83 | 83.19 | 87.69 | 64.13 | 84.84 | 89.74 |
| FineLIP[CVPR’25] | 77.51 | 96.02 | 98.41 | 69.90 | 93.43 | 97.45 | 72.69 | 87.14 | 90.65 | 65.48 | 86.84 | 91.00 |
| GOAL[CVPR’25] | 81.53 | 97.02 | 98.80 | 80.86 | 96.24 | 98.63 | 77.29 | 90.25 | 93.30 | 74.84 | 89.94 | 93.25 |
| StructXLIP[CVPR’26] | 84.73 | 97.69 | 99.00 | 82.61 | 97.08 | 98.71 | 75.84 | 89.94 | 93.65 | 74.49 | 90.05 | 93.40 |
| HN-CLIP | 88.25 | 98.45 | 99.43 | 86.24 | 98.12 | 99.22 | 80.69 | 92.40 | 95.10 | 78.84 | 91.90 | 94.65 |
| Δ | ↑ 3.52 | ↑ 0.76 | ↑ 0.43 | ↑ 3.63 | ↑ 1.04 | ↑ 0.51 | ↑ 3.40 | ↑ 2.15 | ↑ 1.45 | ↑ 4.00 | ↑ 1.85 | ↑ 1.25 |
Table 2: Plug-and-play enhancement of our ℒHN on CLIP-based fine-tuning. Results on DOCCI and Long-DCI for Text→Image and Image→Text retrieval. Upper: full-parameter fine-tuning; lower: parameter-efficient tuning. Our method consistently boosts diverse CLIP variants, and the gains grow with caption length. Best in bold, with gain in ↑ green. | DOCCI | Long-DCI |
|---|
| Method | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |
| Long-CLIP[ECCV’24] | 86.45 | 98.00 | 99.31 | 84.10 | 97.84 | 99.04 | 70.24 | 89.62 | 94.00 | 68.44 | 89.35 | 94.67 |
| + our ℒHN | 88.31 | 98.75 | 99.45 | 86.39 | 98.08 | 99.20 | 78.43 | 92.52 | 95.27 | 76.02 | 91.78 | 94.76 |
| Δ | ↑ 1.86 | ↑ 0.75 | ↑ 0.14 | ↑ 2.29 | ↑ 0.24 | ↑ 0.16 | ↑ 8.19 | ↑ 2.90 | ↑ 1.27 | ↑ 7.58 | ↑ 2.43 | ↑ 0.09 |
| FineLIP[CVPR’25] | 77.51 | 96.02 | 98.41 | 69.90 | 93.43 | 97.45 | 59.24 | 77.86 | 83.19 | 49.52 | 75.08 | 82.39 |
| + our ℒHN | 85.47 | 97.71 | 99.04 | 83.88 | 97.41 | 98.82 | 74.92 | 90.52 | 93.62 | 73.39 | 89.87 | 93.23 |
| Δ | ↑ 7.96 | ↑ 1.69 | ↑ 0.63 | ↑ 13.98 | ↑ 3.98 | ↑ 1.37 | ↑ 15.68 | ↑ 12.66 | ↑ 10.43 | ↑ 23.87 | ↑ 14.79 | ↑ 10.84 |
| GOAL[CVPR’25] | 81.53 | 97.02 | 98.80 | 80.86 | 96.24 | 98.63 | 74.29 | 92.77 | 95.58 | 73.31 | 92.14 | 95.77 |
| + our ℒHN | 86.24 | 98.14 | 99.35 | 84.86 | 97.69 | 99.20 | 84.27 | 94.99 | 96.57 | 82.28 | 94.26 | 96.18 |
| Δ | ↑ 4.71 | ↑ 1.12 | ↑ 0.55 | ↑ 4.00 | ↑ 1.45 | ↑ 0.57 | ↑ 9.98 | ↑ 2.22 | ↑ 0.99 | ↑ 8.97 | ↑ 2.12 | ↑ 0.41 |
| StructXLIP[CVPR’26] | 84.73 | 97.69 | 99.00 | 82.61 | 97.08 | 98.71 | 75.34 | 93.15 | 95.84 | 72.30 | 92.79 | 95.78 |
| + our ℒHN | 85.45 | 97.92 | 99.12 | 83.31 | 97.45 | 98.92 | 82.04 | 94.18 | 96.23 | 78.98 | 93.22 | 95.65 |
| Δ | ↑ 0.72 | ↑ 0.23 | ↑ 0.12 | ↑ 0.70 | ↑ 0.37 | ↑ 0.21 | ↑ 6.70 | ↑ 1.03 | ↑ 0.39 | ↑ 6.68 | ↑ 0.43 | ↓ 0.13 |
| LoRA[ICLR’22] | 80.49 | 96.16 | 98.55 | 77.96 | 95.80 | 98.04 | 60.08 | 80.68 | 86.86 | 58.97 | 80.02 | 86.53 |
| + our ℒHN | 82.82 | 97.14 | 98.86 | 80.84 | 96.49 | 98.63 | 64.87 | 83.09 | 88.35 | 62.79 | 81.77 | 87.31 |
| Δ | ↑ 2.33 | ↑ 0.98 | ↑ 0.31 | ↑ 2.88 | ↑ 0.69 | ↑ 0.59 | ↑ 4.79 | ↑ 2.41 | ↑ 1.49 | ↑ 3.82 | ↑ 1.75 | ↑ 0.78 |
| DoRA[ICML’24] | 80.76 | 96.25 | 98.61 | 78.20 | 95.90 | 98.10 | 60.76 | 81.19 | 87.48 | 59.65 | 80.31 | 87.13 |
| + our ℒHN | 83.41 | 97.33 | 99.00 | 81.41 | 96.65 | 98.67 | 65.46 | 83.57 | 88.61 | 63.38 | 82.15 | 87.56 |
| Δ | ↑ 2.65 | ↑ 1.08 | ↑ 0.39 | ↑ 3.21 | ↑ 0.75 | ↑ 0.57 | ↑ 4.70 | ↑ 2.38 | ↑ 1.13 | ↑ 3.73 | ↑ 1.84 | ↑ 0.43 |
Table 3: Ablations. (a) Sensitivity to the boost strength γ (Recall@1): every γ∈[0.25,1] beats γ=0 in-domain, while under transfer (Urban-1K) milder boosts generalize better. (b) Loss components under an identical recipe; only the loss changes. (c) G¯ recomputed from the live encoder (default) vs. frozen at initialization; R@1/5/10 in Table A2. Best per column in bold. | DOCCI | DCI | Long-DCI | Urban-1K | |
|---|
| γ | R@1 | R@1 | R@1 | R@1 | R@1 | R@1 | R@1 | R@1 | Avg |
| γ=0 | 86.45 | 84.29 | 79.24 | 77.54 | 70.92 | 68.00 | 91.30 | 93.00 | 81.34 |
| γ=0.25 | 87.82 | 86.63 | 80.74 | 79.29 | 76.38 | 74.80 | 91.70 | 93.00 | 83.80 |
| γ=0.5 (default) | 88.25 | 86.24 | 80.69 | 78.84 | 79.03 | 76.44 | 91.10 | 90.30 | 83.86 |
| γ=0.75 | 88.20 | 85.73 | 81.04 | 78.59 | 79.96 | 76.72 | 89.10 | 89.50 | 83.61 |
| γ=1.0 | 87.98 | 85.02 | 80.34 | 78.19 | 79.94 | 76.59 | 87.90 | 88.40 | 83.04 |
Table A1: Benchmark statistics. Mean caption length is measured on the evaluation split in words; similarities use the pre-trained Long-CLIP-L text encoder.| Benchmark | #train | #test | words | pairwise sim | hardest sim |
|---|
| DOCCI | 9450 | 5100 | 123 | 0.84 | 0.93 |
| DCI | 5445 | 1999 | 133 | 0.85 | 0.92 |
| Long-DCI | 5445 (=DCI) | 7444 | 134 | 0.85 | 0.93 |
| Urban-1K | 14579 (VG) | 1000 | 107 | 0.88 | 0.94 |
Table A2: Recomputed vs. frozen G¯ (Recall@K, Text→Image and Image→Text). Building the boost from a frozen copy of the pre-trained text encoder makes the margins exactly stationary but removes their implicit annealing: accuracy peaks at epoch 1 on DCI and Urban-1K and then declines, and the default wins 22 of 24 columns. Best per column in bold. Upper band: DOCCI and DCI; lower band: Long-DCI and the transfer setting Urban-1K. | DOCCI | DCI |
|---|
| G¯ source | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |
| Recomputed G¯ (default) | 88.25 | 98.45 | 99.43 | 86.24 | 98.12 | 99.22 | 80.69 | 92.40 | 95.10 | 78.84 | 91.90 | 94.65 |
| Frozen G¯ | 85.86 | 97.61 | 99.18 | 82.43 | 96.82 | 98.73 | 77.79 | 90.80 | 93.35 | 75.84 | 89.49 | 92.95 |
| Δ (default − frozen) | ↑ 2.39 | ↑ 0.84 | ↑ 0.25 | ↑ 3.81 | ↑ 1.30 | ↑ 0.49 | ↑ 2.90 | ↑ 1.60 | ↑ 1.75 | ↑ 3.00 | ↑ 2.41 | ↑ 1.70 |
| best epoch (default / frozen) | 6 / 5 | 5 / 1 |
| Long-DCI | Urban-1K |
| G¯ source | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |
| Recomputed G¯ (default) | 79.03 | 93.76 | 96.34 | 76.44 | 92.79 | 95.91 | 91.10 | 98.10 | 99.40 | 90.30 | 98.20 | 99.30 |
| Frozen G¯ | 79.93 | 90.84 | 93.82 | 76.75 | 89.54 | 92.62 | 87.30 | 97.70 | 98.90 | 83.50 | 95.90 | 97.90 |
| Δ (default − frozen) | ↓ 0.90 | ↑ 2.92 | ↑ 2.52 | ↓ 0.31 | ↑ 3.25 | ↑ 3.29 | ↑ 3.80 | ↑ 0.40 | ↑ 0.50 | ↑ 6.80 | ↑ 2.30 | ↑ 1.40 |
| best epoch (default / frozen) | 10 / 10 | 4 / 1 |
Table A3: Training-efficiency comparison on DCI fine-tuning (10 epochs, 5.4k images, single Ascend 910B, effective batch 128). Wall-clock for GOAL/StructXLIP is measured from same-device sequential runs minus evaluation overhead; FineLIP ran in parallel and is not attributable. Offline preprocessing time (segmentation, edge extraction, LLM filtering) is not included in the wall-clock.| Method | Aux. training inputs | Offline prep. | Wall-clock | Throughput |
|---|
| Long-CLIP (plain FT) | none | none | ≈16 min | 55.5 img/s |
| FineLIP | none | none | — | — |
| GOAL | SAM segments | required | ≈41 min | ≈22 img/s |
| StructXLIP | edges + LLM lexicon | required | ≈92 min | ≈10 img/s |
| HN-CLIP (ours) | none | none | 17 min | 53 img/s |
Table A4: Plug-and-play enhancement of ℒHN on CLIP-based fine-tuning, all four benchmarks. Results on DOCCI, DCI, Long-DCI, and Urban-1K for Text→Image and Image→Text retrieval; per framework we report the official baseline, the same recipe with ℒHN replacing its global InfoNCE term, and the per-column difference Δ (gain in ↑ green, drop in ↓ gray; differences within ±0.15 are shown as ≈0). Bold marks the +ℒHN value where it is the better of the pair. Upper band: DOCCI and DCI; lower band: Long-DCI and the transfer setting Urban-1K. | DOCCI | DCI |
|---|
| Method | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |
| Long-CLIP | 86.45 | 98.00 | 99.31 | 84.10 | 97.84 | 99.04 | 79.39 | 91.60 | 94.65 | 78.34 | 92.85 | 95.25 |
| +our ℒHN | 88.31 | 98.75 | 99.45 | 86.39 | 98.08 | 99.20 | 81.04 | 92.20 | 95.00 | 78.54 | 91.95 | 94.60 |
| Δ | ↑ 1.86 | ↑ 0.75 | ≈0 | ↑ 2.29 | ↑ 0.24 | ↑ 0.16 | ↑ 1.65 | ↑ 0.60 | ↑ 0.35 | ↑ 0.20 | ↓ 0.90 | ↓ 0.65 |
| FineLIP† | 77.12 | 95.94 | 98.29 | 70.16 | 93.14 | 97.29 | 72.69 | 87.14 | 90.65 | 65.48 | 86.84 | 91.00 |
| +our ℒHN | 85.47 | 97.71 | 99.04 | 83.88 | 97.41 | 98.82 | 80.74 | 91.90 | 94.55 | 79.99 | 91.80 | 94.35 |
| Δ | ↑ 8.35 | ↑ 1.77 | ↑ 0.75 | ↑ 13.72 | ↑ 4.27 | ↑ 1.53 | ↑ 8.05 | ↑ 4.76 | ↑ 3.90 | ↑ 14.51 | ↑ 4.96 | ↑ 3.35 |
| GOAL† | 81.96 | 96.94 | 98.78 | 80.84 | 96.33 | 98.61 | 77.29 | 90.25 | 93.30 | 74.84 | 89.94 | 93.25 |
| +our ℒHN | 86.24 | 98.14 | 99.35 | 84.86 | 97.69 | 99.20 | 79.59 | 90.80 | 93.30 | 77.09 | 90.10 | 93.00 |
| Δ | ↑ 4.28 | ↑ 1.20 | ↑ 0.57 | ↑ 4.02 | ↑ 1.36 | ↑ 0.59 | ↑ 2.30 | ↑ 0.55 | ≈0 | ↑ 2.25 | ↑ 0.16 | ↓ 0.25 |
| StructXLIP | 84.73 | 97.69 | 99.00 | 82.61 | 97.08 | 98.71 | 75.84 | 89.94 | 93.65 | 74.49 | 90.05 | 93.40 |
| +our ℒHN | 85.45 | 97.92 | 99.12 | 83.31 | 97.45 | 98.92 | 77.34 | 90.75 | 93.10 | 74.14 | 89.89 | 92.30 |
| Δ | ↑ 0.72 | ↑ 0.23 | ≈0 | ↑ 0.70 | ↑ 0.37 | ↑ 0.21 | ↑ 1.50 | ↑ 0.81 | ↓ 0.55 | ↓ 0.35 | ↓ 0.16 | ↓ 1.10 |
| LoRA | 80.49 | 96.16 | 98.55 | 77.96 | 95.80 | 98.04 | 73.84 | 89.44 | 92.85 | 71.89 | 88.74 | 92.85 |
| +our ℒHN | 82.82 | 97.14 | 98.86 | 80.84 | 96.49 | 98.63 | 75.84 | 90.20 | 93.00 | 73.44 | 88.54 | 92.25 |
| Δ | ↑ 2.33 | ↑ 0.98 | ↑ 0.31 | ↑ 2.88 | ↑ 0.69 | ↑ 0.59 | ↑ 2.00 | ↑ 0.76 | ↑ 0.15 | ↑ 1.55 | ↓ 0.20 | ↓ 0.60 |
| DoRA | 80.76 | 96.25 | 98.61 | 78.20 | 95.90 | 98.10 | 74.14 | 89.74 | 93.00 | 72.44 | 89.24 | 92.95 |
| +our ℒHN | 83.41 | 97.33 | 99.00 | 81.41 | 96.65 | 98.67 | 76.24 | 90.45 | 93.15 | 73.89 | 88.59 | 92.30 |
| Δ | ↑ 2.65 | ↑ 1.08 | ↑ 0.39 | ↑ 3.21 | ↑ 0.75 | ↑ 0.57 | ↑ 2.10 | ↑ 0.71 | ↑ 0.15 | ↑ 1.45 | ↓ 0.65 | ↓ 0.65 |
| Long-DCI | Urban-1K |
| Method | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |
| Long-CLIP | 70.24 | 89.62 | 94.00 | 68.44 | 89.35 | 94.67 | 91.70 | 99.10 | 99.50 | 93.20 | 99.00 | 99.50 |
| +our ℒHN | 78.43 | 92.52 | 95.27 | 76.02 | 91.78 | 94.76 | 90.60 | 98.30 | 99.40 | 90.20 | 98.30 | 99.40 |
| Δ | ↑ 8.19 | ↑ 2.90 | ↑ 1.27 | ↑ 7.58 | ↑ 2.43 | ≈0 | ↓ 1.10 | ↓ 0.80 | ≈0 | ↓ 3.00 | ↓ 0.70 | ≈0 |
| FineLIP† | 59.24 | 77.86 | 83.19 | 49.52 | 75.08 | 82.39 | 81.10 | 94.50 | 97.40 | 77.50 | 94.80 | 97.50 |
| +our ℒHN | 74.92 | 90.52 | 93.62 | 73.39 | 89.87 | 93.23 | 91.80 | 98.10 | 99.30 | 93.20 | 98.80 | 99.30 |
| Δ | ↑ 15.68 | ↑ 12.66 | ↑ 10.43 | ↑ 23.87 | ↑ 14.79 | ↑ 10.84 | ↑ 10.70 | ↑ 3.60 | ↑ 1.90 | ↑ 15.70 | ↑ 4.00 | ↑ 1.80 |
| GOAL† | 74.34 | 92.89 | 95.77 | 72.03 | 92.37 | 95.69 | 86.30 | 97.00 | 98.90 | 86.10 | 97.40 | 98.90 |
| +our ℒHN | 84.27 | 94.99 | 96.57 | 82.28 | 94.26 | 96.18 | 85.70 | 96.70 | 98.20 | 86.80 | 96.60 | 98.00 |
| Δ | ↑ 9.93 | ↑ 2.10 | ↑ 0.80 | ↑ 10.25 | ↑ 1.89 | ↑ 0.49 | ↓ 0.60 | ↓ 0.30 | ↓ 0.70 | ↑ 0.70 | ↓ 0.80 | ↓ 0.90 |
Table A5: Loss ablation, full resolution (all four benchmarks, all ranks). Best per column in bold. | | DOCCI | DCI |
|---|
| ℒtok | ℒHN | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 |
| ✗ | ✗ | 86.45 | 98.00 | 99.31 | 99.86 | 99.98 | 84.10 | 97.84 | 99.04 | 99.78 | 99.98 | 79.39 | 91.60 | 94.65 | 97.25 | 97.95 | 78.34 | 92.85 | 95.25 |
| ✓ | ✗ | 86.45 | 97.98 | 99.43 | 99.86 | 99.98 | 84.29 | 97.90 | 98.94 | 99.78 | 99.94 | 79.24 | 91.55 | 94.25 | 97.00 | 98.00 | 77.54 | 92.95 | 95.35 |
| ✗ | ✓ | 88.31 | 98.75 | 99.45 | 99.84 | 99.98 | 86.39 | 98.08 | 99.20 | 99.82 | 99.96 | 81.04 | 92.20 | 95.00 | 97.40 | 98.10 | 78.54 | 91.95 | 94.60 |
| ✓ | ✓ | 88.25 | 98.45 | 99.43 | 99.86 | 99.98 | 86.24 | 98.12 | 99.22 | 99.84 | 99.94 | 80.69 | 92.40 | 95.10 | 97.35 | 98.15 | 78.84 | 91.90 | 94.65 |
| | Long-DCI | Urban-1K |
| ℒtok | ℒHN | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 |
| ✗ | ✗ | 70.24 | 89.62 | 94.00 | 97.66 | 98.79 | 68.44 | 89.35 | 94.67 | 98.11 | 98.95 | 91.70 | 99.10 | 99.50 | 99.80 | 99.90 | 93.20 | 99.00 | 99.50 |
| ✓ | ✗ | 70.92 | 89.51 | 93.90 | 97.27 | 98.64 | 68.00 | 89.12 | 94.17 | 97.84 | 98.93 | 91.30 | 99.00 | 99.30 | 99.80 | 99.90 | 93.00 | 99.10 | 99.50 |
| ✗ | ✓ | 78.43 | 92.52 | 95.27 | 97.21 | 98.47 | 76.02 | 91.78 | 94.76 | 97.31 | 98.29 | 90.60 | 98.30 | 99.40 | 99.60 | 99.80 | 90.20 | 98.30 | 99.40 |
| ✓ | ✓† | 78.79 | 92.32 | 95.12 | 97.19 | 98.27 | 76.18 | 91.99 | 95.15 | 97.26 | 98.40 | 91.00 | 98.50 | 99.50 | 99.70 | 99.90 | 90.50 | 98.50 | 99.30 |
Table A6: γ sweep, full resolution (all four benchmarks, all ranks). Best per column in bold. | DOCCI | DCI |
|---|
| γ | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 |
| γ=0 | 86.45 | 97.98 | 99.43 | 99.86 | 99.98 | 84.29 | 97.90 | 98.94 | 99.78 | 99.94 | 79.24 | 91.55 | 94.25 | 97.00 | 98.00 | 77.54 | 92.95 | 95.35 | 97.55 |
| γ=0.25 | 87.82 | 98.63 | 99.45 | 99.88 | 99.98 | 86.63 | 98.04 | 99.16 | 99.84 | 99.96 | 80.74 | 92.40 | 95.35 | 97.55 | 98.30 | 79.29 | 92.85 | 95.40 | 97.70 |
| γ=0.5 (default) | 88.25 | 98.45 | 99.43 | 99.86 | 99.98 | 86.24 | 98.12 | 99.22 | 99.84 | 99.94 | 80.69 | 92.40 | 95.10 | 97.35 | 98.15 | 78.84 | 91.90 | 94.65 | 97.40 |
| γ=0.75 | 88.20 | 98.67 | 99.49 | 99.82 | 99.96 | 85.73 | 98.00 | 99.25 | 99.78 | 99.98 | 81.04 | 92.20 | 94.70 | 97.20 | 98.25 | 78.59 | 91.05 | 93.90 | 96.90 |
| γ=1.0 | 87.98 | 98.47 | 99.51 | 99.82 | 99.96 | 85.02 | 97.69 | 99.02 | 99.76 | 99.96 | 80.34 | 92.00 | 94.40 | 97.10 | 98.30 | 78.19 | 90.80 | 93.70 | 96.70 |
| Long-DCI | Urban-1K |
| γ | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 |
| γ=0 | 70.92 | 89.51 | 93.90 | 97.27 | 98.64 | 68.00 | 89.12 | 94.17 | 97.84 | 98.93 | 91.30 | 99.00 | 99.30 | 99.80 | 99.90 | 93.00 | 99.10 | 99.50 | 99.60 |
| γ=0.25 | 76.38 | 92.49 | 95.34 | 97.73 | 98.68 | 74.80 | 91.94 | 95.51 | 97.94 | 98.68 | 91.70 | 98.60 | 99.50 | 99.70 | 99.90 | 93.00 | 98.90 | 99.60 | 99.70 |
| γ=0.5 (default)† | 78.79 | 92.32 | 95.12 | 97.19 | 98.27 | 76.18 | 91.99 | 95.15 | 97.26 | 98.40 | 91.00 | 98.50 | 99.50 | 99.70 | 99.90 | 90.50 | 98.50 | 99.30 | 99.60 |
| γ=0.75 | 79.96 | 92.45 | 95.06 | 97.06 | 98.19 | 76.72 | 91.35 | 94.43 | 96.76 | 97.97 | 89.10 | 98.50 | 99.10 | 99.60 | 99.60 | 89.50 | 97.90 | 99.00 | 99.80 |
| γ=1.0 | 79.94 | 92.17 | 94.79 | 96.76 | 97.89 | 76.59 | 90.78 | 93.91 | 96.49 | 97.70 | 87.90 | 98.00 | 99.10 | 99.50 | 99.60 | 88.40 | 97.50 | 98.70 | 99.80 |
Table A7: Cross-domain transfer DOCCI→Long-DCI. Best per column in bold. | DOCCI → Long-DCI |
|---|
| Method | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 |
| Long-CLIP (zero-shot) | 54.61 | 72.80 | 78.33 | 85.29 | 89.15 | 47.35 | 73.04 | 80.10 | 86.66 | 90.60 |
| FineLIP† | 54.65 | 73.04 | 79.46 | 85.87 | 89.90 | 46.70 | 73.01 | 80.51 | 87.10 | 90.64 |
| GOAL† | 55.71 | 74.93 | 80.66 | 86.93 | 90.66 | 56.61 | 75.46 | 81.19 | 87.55 | 91.28 |
| StructXLIP | 57.70 | 76.03 | 81.68 | 88.06 | 91.60 | 57.90 | 76.36 | 82.09 | 87.71 | 91.03 |
| HN-CLIP | 62.60 | 79.16 | 84.19 | 89.88 | 93.05 | 63.77 | 79.89 | 84.22 | 89.05 | 92.14 |
Table A8: Seed replication on Long-DCI. max|Δ| is the largest absolute difference across the six metrics.| Method | Seed | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 | max|Δ| |
|---|
| HN-CLIP (ours) | seed 42 | 79.03 | 93.76 | 96.34 | 76.44 | 92.79 | 95.91 | |
| seed 43 | 78.79 | 92.32 | 95.12 | 76.18 | 91.99 | 95.15 | 1.44 |
| GOAL | seed 42 | 74.29 | 92.77 | 95.58 | 73.31 | 92.14 | 95.77 | |
| seed 43 | 74.34 | 92.89 | 95.77 | 72.03 | 92.37 | 95.69 | 1.28 |
| StructXLIP | seed 42 | 75.34 | 93.15 | 95.84 | 72.30 | 92.79 | 95.78 | |
| seed 43 | 75.58 | 93.12 | 96.20 | 72.39 | 93.03 | 95.80 | 0.36 |
Table A9: Sample efficiency, full resolution. Best per pair in bold; Δ is the average-R@1 margin of HN-CLIP over GOAL at that fraction.| Data | Method | R@1 | R@5 | R@10 | R@25 | R@50 | R@1 | R@5 | R@10 | R@25 | R@50 | Δ |
|---|
| DOCCI |
| 5% | GOAL | 77.88 | 95.35 | 97.88 | 99.59 | 99.84 | 75.45 | 94.65 | 97.76 | 99.35 | 99.84 | |
| HN-CLIP | 82.73 | 96.94 | 98.69 | 99.69 | 99.84 | 80.08 | 96.22 | 98.35 | 99.57 | 99.98 | ↑ 4.74 |
| 20% | GOAL | 78.22 | 95.65 | 98.22 | 99.55 | 99.86 | 76.47 | 94.88 | 97.96 | 99.39 | 99.84 | |
| HN-CLIP | 85.43 | 97.86 | 99.20 | 99.80 | 99.96 | 83.75 | 97.25 | 98.80 | 99.76 | 99.98 | ↑ 7.25 |
| 50% | GOAL | 80.20 | 96.47 | 98.33 | 99.63 | 99.88 | 78.18 | 95.37 | 97.96 | 99.43 | 99.92 | |
| HN-CLIP | 87.16 | 98.39 | 99.43 | 99.84 | 99.94 | 85.37 | 97.69 | 99.18 | 99.80 | 99.94 | ↑ 7.07 |
| 100% | GOAL† | 81.96 | 96.94 | 98.78 | 99.78 | 99.92 | 80.84 | 96.33 | 98.61 | 99.67 | 99.92 | |
| HN-CLIP | 88.25 | 98.45 | 99.43 | 99.86 | 99.98 | 86.24 | 98.12 | 99.22 | 99.84 | 99.94 | ↑ 5.85 |
| DCI |
| 5% | GOAL | 70.94 | 87.19 | 90.95 | 94.95 | 96.95 | 69.03 | 86.89 | 90.80 | 94.75 | 96.85 | |
| HN-CLIP | 73.59 | 87.39 | 92.30 | 95.10 | 96.95 | 74.59 | 89.19 | 93.20 | 96.00 | 97.15 | ↑ 4.11 |
| 20% | GOAL | 73.69 | 87.99 | 91.40 | 95.40 | 97.10 | 71.39 | 87.24 | 91.55 | 95.45 | 97.30 | |
| HN-CLIP | 78.09 | 90.60 | 93.70 | 96.30 | 97.65 | 76.79 | 90.40 | 93.55 | 96.55 | 97.70 | ↑ 4.90 |
| 50% | GOAL | 75.79 | 89.19 | 93.00 | 95.55 | 96.80 | 73.79 | 88.54 | 92.65 | 95.90 | 97.30 | |
| HN-CLIP | 79.24 | 91.60 | 94.55 | 97.05 | 97.95 | 78.44 | 91.35 | 94.25 | 96.70 | 98.00 | ↑ 4.05 |
| 100% | GOAL | 77.29 | 90.25 | 93.30 | 96.10 | 97.40 | 74.84 | 89.94 | 93.25 | 96.50 | 98.15 | |
| HN-CLIP | 80.69 | 92.40 | 95.10 | 97.35 | 98.15 | 78.84 | 91.90 | 94.65 | 97.40 | 98.30 | ↑ 3.70 |
Table A10: Cross-domain generalization between DCI and DOCCI. Train on one dataset and test on another. Values are Recall@K (%), using Text→Image and Image→Text retrieval. In-domain best in italic bold, cross-domain best in bold.| Setting | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |
|---|
| Train on DCI → Test on DCI vs. DOCCI |
| Long-CLIP (DCI→DCI) | 67.83 | 83.19 | 87.69 | 64.13 | 84.84 | 89.74 |
| Long-CLIP (DCI→DOCCI) | 78.78 | 95.24 | 98.02 | 66.75 | 91.92 | 96.31 |
| FineLIP (DCI→DCI) | 72.69 | 87.14 | 90.65 | 65.48 | 86.84 | 91.00 |
| FineLIP (DCI→DOCCI) | 80.69 | 96.47 | 98.53 | 65.63 | 91.51 | 96.31 |
| GOAL (DCI→DCI) | 77.29 | 90.25 | 93.30 | 74.84 | 89.94 | 93.25 |
| GOAL (DCI→DOCCI) | 80.14 | 96.25 | 98.39 | 77.24 | 95.22 | 97.90 |
| StructXLIP (DCI→DCI) | 75.84 | 89.94 | 93.65 | 74.49 | 90.05 | 93.40 |
| StructXLIP (DCI→DOCCI) | 76.94 | 95.06 | 97.76 | 74.24 | 94.12 | 97.51 |
| HN-CLIP (DCI→DCI) | 80.69 | 92.40 | 95.10 | 78.84 | 91.90 | 94.65 |
| HN-CLIP (DCI→DOCCI) | 85.00 | 97.90 | 99.25 | 83.14 | 96.98 | 98.75 |
| Train on DOCCI → Test on DOCCI vs. DCI |
| Long-CLIP (DOCCI→DOCCI) | 78.78 | 95.24 | 98.02 | 66.75 | 91.92 | 96.31 |
| Long-CLIP (DOCCI→DCI) | 67.83 | 83.19 | 87.69 | 64.13 | 84.84 | 89.74 |
| FineLIP (DOCCI→DOCCI) | 77.51 | 96.02 | 98.41 | 69.90 | 93.43 | 97.45 |
| FineLIP (DOCCI→DCI) | 68.23 | 83.84 | 88.74 | 63.48 | 85.79 | 89.84 |
| GOAL (DOCCI→DOCCI) | 81.53 | 97.02 | 98.80 | 80.86 | 96.24 | 98.63 |
| GOAL (DOCCI→DCI) | 69.58 | 85.24 | 89.39 | 69.88 | 85.34 | 89.79 |
| StructXLIP (DOCCI→DOCCI) | 84.73 | 97.69 | 99.00 | 82.61 | 97.08 | 98.71 |
| StructXLIP (DOCCI→DCI) | 69.93 | 86.74 | 91.05 | 71.39 | 86.99 | 90.60 |
| HN-CLIP (DOCCI→DOCCI) | 88.25 | 98.45 | 99.43 | 86.24 | 98.12 | 99.22 |
| HN-CLIP (DOCCI→DCI) | 74.84 | 88.54 | 92.40 | 75.64 | 88.04 | 91.35 |
为什么重要
仅通过修正训练目标函数就能大幅提升图文检索效果,无需增加预处理或模型复杂度,为实际检索系统的改进提供了低成本、高效率的思路。同时,该研究揭示了看似收敛的损失值可能掩盖模型早已停止学习关键难例的问题,这对其他基于对比学习的AI项目也具有参考价值。
本文术语
- InfoNCE · 一种对比学习损失函数,让匹配的图文对靠近、不匹配的对拉远
- 近重复(near-duplicate)描述 · 内容极其相似、难以区分的描述句子
- 梯度(gradient) · 指示模型参数该往哪个方向、调整多少的信号
- detach/停止梯度 · 让某个计算结果只作为参考值使用,不参与反向传播更新
- R@1(Recall@1) · 检索结果中正确答案排在第一位的比例
无法转载的图表
- Figure 1: Dense-caption benchmarks are dominated by hard negatives. Distributions of (a) all pairwise caption-caption cosine similarities and (b) each caption’s hardest-negative similarity, measured with the pre-trained Long-CLIP-L text encoder on the test sets. Dotted lines mark the means. The pairs in (b) receive the strongest boosts.
- Figure 2: Overview of HN-CLIP. For a batch of B image–text pairs, HN-CLIP encodes images and captions with the dual encoder being fine-tuned, and forms the image–text similarity matrix S=vt⊤ and a detached, diagonal-masked caption-similarity matrix G¯ from G=tt⊤. Their combination yields boosted logits s(S+γG¯), so near-twin negatives receive larger margins. The token-level term of Eq. 3 acts alongside this objective, improving supervision on hard negatives during training while keeping inference unchanged.
- Figure 3: Illustration of HN-CLIP. A real DOCCI query and its hardest in-batch negative (cos 0.89). The strongest baseline ranks the ground truth 52nd (its top-1 (pink) is the near-twin’s own image); HN-CLIP ranks it first.
- Figure 4: Empirical gradient analysis (Long-CLIP-L, DOCCI, 10 epochs; losses and full-parameter gradients measured every 5 optimizer steps, 148 measurements). Standard InfoNCE declares 80% of batches solved within the first epoch and its gradient is exactly zero in fp32 in 47% of measurements (median zero at five epochs; plotted clamped to 10−10). The boosted loss keeps at least 26% of batches active in every epoch and, where the standard gradient is nonzero, exceeds it by 102–106× in per-epoch median, with interquartile bands disjoint at nine epochs, and the retrieval error it buys keeps falling.
- Figure 5: Convergence comparison. Average R@1 (T→I, I→T) per fine-tuning epoch. HN-CLIP’s first epoch matches or exceeds most baselines’ final accuracy on all four benchmarks.
- Figure 6: Sample efficiency (avg R@1, identical subsets). HN-CLIP at 20% of the data already clears the strongest baseline trained on 100% (dotted line, from Table 1); full numbers in Table A9.
- Figure A1: Near-duplicate caption pairs on DOCCI. Each row: a test image, its caption, the three nearest other captions’ images (cosine printed underneath), and the caption texts with shared content words highlighted.
- Figure A2: Near-duplicate caption pairs on DCI.
- Figure A3: Near-duplicate caption pairs on Long-DCI (full-length captions).
- Figure A4: Near-duplicate caption pairs on Urban-1K.
- Figure A5: One real training batch per benchmark. Text–text similarity matrices G¯ (pre-trained encoder; diagonal masked, max off-diagonal entry annotated). Uniformly dark = every negative is hard; this matrix, detached and scaled by γ, is the entire mechanism of HN-CLIP.
- Figure A6: Statistical analysis of the caption geometry across the four benchmarks (pre-trained Long-CLIP-L text encoder, full test splits). (A) Mean pairwise caption similarity. (B) Mean similarity of each caption’s hardest companion. (C) Composition of in-batch negatives by hardness bucket.
- Figure A8: Gradient dynamics across data scales. Rows: fraction of the training set; columns: dataset. Per cell: (a) raw loss traces (thin) with running medians (thick), (b) gradient norms and their ratio (dashed), (c) gradient cosine. Log axes in (a,b); values clamped at 10−10 (exact zeros at fp32).
- Figure A9: Per-direction convergence. Recall@1 per epoch; top: Text→Image, bottom: Image→Text.
- Figure A10: Qualitative T→I retrieval on DOCCI (green = ground truth).
在原文中查看图表 →论文原文摘要(英文)
Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, these methods largely inherit the same InfoNCE objective, whose optimization can prematurely saturate under a strong pre-trained initialization: on dense captions, the loss falls below 10^{-3} on 80% of batches within the first epoch, while its gradient reaches exact zero in
作者 · Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoying Tang
在 arXiv 阅读