Figure 1: The proposed SPK framework, a proactive OoD hallucination mitigation framework, further reduces OoD-induced hallucinations beyond previous state-of-the-art methods (48), achieving additional improvements in challenging high-performance regimes.
Table 1: Comparison of OoD detection performance using FPR95 across different detector architectures trained on PASCAL-VOC and BDD-100K. Lower is better.
Method
YOLO
Faster R-CNN
RT-DETR
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
MSP
67.48
67.18
72.93
74.12
68.36
78.69
77.46
74.41
70.44
67.81
77.46
74.41
EBO
90.49
90.84
87.22
87.06
60.62
56.21
94.83
94.37
98.03
96.69
94.83
94.37
MLS
89.88
90.08
86.47
87.06
59.62
57.89
92.55
91.08
92.84
89.75
92.55
91.08
SCALE
80.67
80.92
77.44
70.59
92.34
80.70
86.35
84.98
81.12
76.92
86.35
84.98
MDS
57.67
69.47
68.42
82.35
49.96
56.38
78.90
79.34
48.65
52.90
78.90
79.34
BAM
45.36
43.72
49.63
52.18
65.44
42.16
65.73
61.34
75.61
65.27
75.48
68.44
KNN
48.20
39.50
41.95
45.24
61.95
37.53
50.54
49.76
77.10
63.00
62.10
58.00
iForest
70.27
67.82
60.42
65.23
75.43
52.38
63.25
62.78
79.52
65.92
68.23
62.30
SPK-MDS
14.99
17.28
23.35
6.46
21.03
23.50
42.42
42.73
18.43
17.83
41.77
16.48
SPK-BAM
21.96
21.73
16.75
3.07
18.98
18.17
9.07
6.18
23.12
25.19
18.76
16.84
SPK-KNN
19.64
17.99
13.47
1.17
15.50
13.69
4.70
3.09
18.26
20.43
16.97
13.74
SPK-iForest
14.25
11.84
9.86
0.70
13.92
10.52
2.31
1.52
15.48
17.32
11.42
9.25
Figure 2: Overview of the proposed SPK framework. SPK elicits semantic, geometric, and contextual priors from a pretrained object detector and organizes them into a compact five-dimensional representation for OoD hallucination detection.
Table 2: OoD detection counts (Near-OoD/Far-OoD) across different detector architectures. Lower is better.
Model
Method
VOC (N/F)
BDD (N/F)
YOLO
Original
946 / 440
701 / 666
Proximal-OoD
134 / 60
80 / 47
SPK
135 / 52
69 / 5
Faster R-CNN
Original
2150 / 1335
2576 / 1634
Proximal-OoD
710 / 253
207 / 167
SPK
299 / 140
60 / 25
RT-DETR
Original
2311 / 1589
3145 / 1220
Proximal-OoD
386 / 470
525 / 240
SPK
358 / 275
359 / 113
Figure 3: Automated part-level annotation pipeline. Given an RoI crop and its object category, OWLv2 grounds the corresponding GPT-5-generated concept vocabulary into part bounding boxes. Each box prompts SAM 2 to produce a refined pixel-level part mask. Masks satisfying the object-mask coverage threshold are projected into detector-RoI coordinates and rasterized as binary 7×7 targets. Concepts without a retained mask receive an all-zero target.
Table 3: Ablation study of the SPK loss components on YOLO trained on PASCAL-VOC and BDD-100K. Results are reported as Near-OoD / Far-OoD FPR95. The average is computed over all four results. Lower is better.
ℒdice
ℒsuppress
ℒgroup
VOC
BDD
Average
Near / Far
Near / Far
FPR95 ↓
✗
✓
✓
25.80 / 21.30
17.85 / 18.26
20.80
✓
✓
✗
22.10 / 16.40
15.29 / 15.97
17.44
✓
✗
✓
20.50 / 15.80
14.18 / 12.93
15.85
✓
✓
✓
14.25 / 11.84
9.86 / 0.70
9.16
Figure 4: Part-Level Concept Annotation Example. We use a bird RoI crop to illustrate the step-by-step annotation process for the concepts wing, head, beak, and torso. Step 1: OWLv2 grounds each concept within the detector RoI and generates a bounding-box proposal (a). Step 2: SAM 2 refines each proposal into a pixel-level part mask, which is retained only if at least 70% of its pixels overlap with the corresponding SAM 2 object mask (b). Step 3: Each retained mask is projected into detector-RoI coordinates and converted into a binary 7×7 supervision target (c). Step 4: The binary target is overlaid on the RoI crop for visualization (d).
Table 4: Ablation study of different prior components on YOLO trained on PASCAL-VOC and BDD-100K. Each dataset column reports Near-OoD / Far-OoD FPR95. The average is computed over all four results. Lower is better.
Prior components
VOC
BDD
Average
Near / Far
Near / Far
FPR95 ↓
Semantic
15.43 / 13.28
30.37 / 25.58
21.17
Semantic + Geometric
13.23 / 11.50
20.88 / 3.84
12.36
Semantic + Geometric + Contextual
14.25 / 11.84
9.86 / 0.70
9.16
Figure 5: Qualitative visualization of part-level semantic responses. The learned concept maps are well aligned with the corresponding regions, demonstrating that the semantic elicitation head successfully decodes spatially grounded semantic evidence from detector RoI features.
Table 5: Quality of the automated part-level annotations. We report the overall concept-mIoU, Recall@0.5, number of covered concepts, and per-class concept-mIoU. All values except concept coverage are percentages. Best results are shown in bold.
Method
Overall
Per-class concept-mIoU
concept-mIoU
Recall@0.5
# Concepts
bird
bus
car
cat
cow
dog
horse
Ours
40.8
38.4
25
46.5
37.9
43.9
40.4
37.4
41.8
38.2
VLPart (37)
37.0
36.8
25
35.6
32.8
32.5
38.1
36.8
44.1
36.5
Grounded SAM (31)
32.3
30.7
25
32.1
31.7
31.3
33.0
29.5
38.3
29.4
(b) Prediction: wing
Table 6: RoI feature sources and dimensions for each detector. YOLO and RT-DETR concatenate RoI-aligned features from three feature scales, whereas Faster R-CNN uses the Detectron2 box_pooler.
Detector
Feature source
Channels
RoI feature 𝐅i
YOLO
Detect neck
128+256+512
896×7×7
RT-DETR
Hybrid-encoder neck
256+256+256
768×7×7
Faster R-CNN
ResNet-FPN
256
256×7×7
(c) Prediction: torso
Table 7: Hyperparameters of the Semantic Elicitation Head.
Semantic Elicitation Head
Hyperparameter
Value
Architecture
Hidden channels
256
Dropout
0.1
Normalization
GroupNorm
Activation
GELU
Residual blocks
2
Output head
1×1 conv → num_concepts
Inference activation
Sigmoid
RoI spatial size
7×7
Inference pooling
LogSumExp, τ=0.5
Training
Training epochs
80
Batch size
2000
Optimizer
AdamW
Learning rate
2×10−4
Weight decay
5×10−4
Random seed
42
Validation split
10%
Training sampler
WeightedRandomSampler
Suppress loss weight
0.25
Group loss weight
0.75
Early-stopping patience
10 epochs
(d) Prediction: foot
Table 8: Image-level embedding sources and dimensions for each detector. We globally pool each selected feature map by its spatial mean and standard deviation, then concatenate the resulting statistics. YOLO and RT-DETR use their deepest selected backbone stage, whereas Faster R-CNN aggregates all four ResNet-FPN levels.
Detector
Feature source
Channels
Embedding dimension
YOLO
Backbone L6 (stride 16)
256
mean+std:512
RT-DETR
HGBlock L9 backbone (stride 32)
2048
mean+std:4096
Faster R-CNN
ResNet-FPN P2–P5
4×256
mean+std:2048
Figure 6: Distributions of the learned semantic group responses. Samples from different data sources predominantly activate their corresponding semantic groups, validating the effectiveness of the proposed group objective.
Table 9: Semantic-head group-classification accuracy. Computed by applying argmax to concept activations for YOLO on PASCAL-VOC.
Data split
Accuracy (%)
ID training set
96.5
Proximal OoD
81.0
Background
77.0
Figure 7: UMAP visualization of representation spaces for detections from the dog, sheep, and cat categories. The visualizations are obtained using YOLO trained on PASCAL-VOC. We compare detector classification logits (left) with the proposed SPK representations (right), using the same ID-validation, Near-OoD, and Far-OoD samples. SPK yields a more structured representation space with clearer distributional differences between ID and OoD samples.
Table 10: AUROC comparison across detector architectures on PASCAL-VOC and BDD-100K. Higher is better.
Method
YOLO
Faster R-CNN
RT-DETR
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
MSP
81.24
79.47
77.63
75.12
78.71
73.84
72.48
76.39
79.16
78.72
74.31
75.06
EBO
60.73
57.92
65.41
62.86
82.58
86.14
52.07
50.83
40.12
43.31
49.68
52.74
MLS
59.41
61.26
63.78
65.02
84.76
83.58
54.82
59.46
56.37
59.74
54.91
57.43
SCALE
69.84
71.73
72.46
79.31
57.49
69.72
64.18
67.82
69.37
75.16
66.42
65.73
MDS
85.62
77.81
80.34
68.46
88.39
84.41
73.68
70.94
87.43
85.71
73.27
71.18
BAM
90.14
89.38
89.57
86.94
80.96
91.56
80.75
83.65
75.75
81.65
76.50
80.45
KNN
89.52
90.31
91.26
88.47
81.73
92.58
88.43
86.72
75.12
81.46
83.68
85.29
iForest
77.38
80.71
82.46
80.19
76.31
85.64
81.37
83.29
70.82
81.94
78.46
83.57
SPK-MDS
96.12
93.55
95.80
98.65
96.10
94.20
92.10
91.80
94.35
95.54
89.10
99.21
SPK-BAM
94.55
92.50
96.90
99.35
96.45
95.40
98.30
98.85
93.75
93.90
93.45
99.18
SPK-KNN
94.91
93.31
97.40
99.71
97.07
96.46
99.06
99.36
94.40
94.87
93.94
99.36
SPK-iForest
96.31
95.43
98.10
99.81
97.38
97.25
99.50
99.65
95.23
95.67
95.78
99.60
(b) sheep
Table 11: Ablation of SPK loss components. Evaluated across Faster R-CNN and RT-DETR on PASCAL-VOC and BDD-100K. Lower FPR95 is better.
ℒdice
ℒsuppress
ℒgroup
Faster R-CNN
RT-DETR
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
✗
✓
✓
27.87
19.98
18.36
18.74
32.31
27.08
25.35
29.14
✓
✓
✗
24.42
15.42
14.15
16.11
27.86
22.53
20.71
24.36
✓
✗
✓
22.19
14.73
11.34
13.27
24.86
21.64
18.15
20.81
✓
✓
✓
13.92
10.52
2.31
1.52
15.48
17.32
11.42
9.25
(c) cat
Table 12: Ablation of SPK prior components. Evaluated across Faster R-CNN and RT-DETR on PASCAL-VOC and BDD-100K. Lower FPR95 is better.
Prior components
Faster R-CNN
RT-DETR
PASCAL-VOC
BDD-100K
PASCAL-VOC
BDD-100K
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Near-OoD
Far-OoD
Semantic
18.17
18.91
23.92
25.21
18.34
29.17
33.36
34.75
Semantic + Geometric
15.78
16.43
14.52
15.71
15.85
25.35
26.51
26.82
Semantic + Geometric + Contextual
13.92
10.52
2.31
1.52
15.48
17.32
11.42
9.25
Table 13: Comparison with competitive OoD detection methods for Deformable-DETR. Results are reported on PASCAL-VOC and BDD-100K as ID datasets, with MS-COCO and OpenImages as OoD datasets. Higher AUROC and lower FPR95 indicate better OoD detection performance. Methods marked with † employ an external DINO ViT encoder to extract additional visual representations for OoD detection, rather than relying solely on Deformable-DETR features. SPK (DINO ViT) is included to provide a fair comparison with UNO-Adapter, as both methods operate under this setting. The best results are highlighted in bold.
Method
ID: PASCAL-VOC
ID: BDD-100K
OoD: MS-COCO
OoD: OpenImages
OoD: MS-COCO
OoD: OpenImages
FPR95↓
AUROC↑
FPR95↓
AUROC↑
FPR95↓
AUROC↑
FPR95↓
AUROC↑
MDS (18)
97.39
50.28
97.88
49.08
70.86
76.83
71.43
77.98
Gram matrices (32)
94.16
43.97
95.29
38.81
73.81
60.13
71.56
57.14
KNN (38)
91.80
62.15
91.36
59.64
64.75
80.90
61.13
79.64
CSI (39)
84.00
55.07
79.16
51.37
70.27
77.93
71.30
76.42
VOS (8)
97.46
54.40
97.07
52.77
76.44
77.33
72.58
76.62
OW-DETR (11)
93.09
55.70
93.82
57.80
80.78
70.29
77.37
73.78
DisMax (25)
82.05
75.21
76.37
70.66
77.62
72.14
81.23
67.18
SIREN-vMF (7)
75.49
76.10
78.36
71.05
67.54
80.06
66.31
79.77
SIREN-KNN (7)
64.77
78.23
65.99
74.93
53.97
86.56
47.28
89.00
SAFE (42)
48.88
78.88
8.99
96.73
39.18
85.95
21.10
94.31
InfoBound (52)
44.88
89.76
43.89
88.00
44.88
89.76
43.89
88.00
UNO-Adapter† (28)
32.61
91.68
19.90
95.40
9.88
97.61
3.80
99.04
SPK
52.32
75.84
24.38
90.20
1.68
99.42
0.37
99.93
SPK (DINO ViT)†
28.55
92.38
14.85
96.25
0.00
99.80
0.00
99.97
Table 14: Comparison with competitive OoD detection methods for Faster R-CNN. Results are reported on PASCAL-VOC as the ID dataset, with MS COCO and OpenImages as OoD datasets. Higher AUROC and lower FPR95 indicate better OoD detection performance. Methods marked with † employ an external DINO ViT encoder to extract additional visual representations for OoD detection, rather than relying solely on Faster R-CNN features. SPK (DINO ViT) is included to provide a fair comparison with UNO-Adapter, as both methods operate under this setting. The best results are highlighted in bold.
Method
MS-COCO
OpenImages
AUROC↑
FPR95↓
AUROC↑
FPR95↓
CSI (39)
82.95
57.41
81.83
59.91
GAN-Synthesis (17)
82.67
59.97
83.67
60.93
VOS (8)
85.23
51.33
88.70
47.53
SIREN (7)
85.36
64.68
82.78
68.53
TIB (44)
90.36
41.55
88.09
47.19
DFDD (43)
90.79
41.34
88.65
44.52
WFS (45)
89.01
40.05
90.35
39.17
UNO-Adapter† (28)
91.25
38.73
92.40
35.74
SPK
91.58
41.20
95.50
25.61
SPK (DINO ViT)†
95.48
26.28
98.16
10.97
Table 15: Per-image runtime of the complete SPK inference pipeline. Reported for a YOLO model pretrained on PASCAL-VOC. The complete SPK pipeline introduces an additional 2.72 ms latency per image, corresponding to a 26.8% runtime overhead over the original detector inference. *SPK inference obtains detection outputs, image-level contextual embeddings, and RoI features within the same forward pass.
Object detectors often produce over-confident predictions for objects outside their training categories, leading to so-called out-of-distribution (OoD) hallucinations. Existing approaches for detecting or mitigating such hallucinations typically either construct scoring functions directly over learned object detector representations or modify the object detector itself to suppress hallucination emergence. However, the latent priors implicitly encoded in these representations remain largely unexplored and have not been explicitly decoded for OoD detection. To uncover and exploit these latent priors, we propose Structured Prior Knowledge (SPK), a hallucination-oriented framework that explicitly elicits OoD-relevant priors from pretrained object detectors. Specifically, SPK leverages in-distribution data and hallucination-inducing samples as diagnostic supervision to elicit part-level semantic concepts underlying object detector decision-making, rather than using them merely for rejection or object detector adaptation. The elicited semantic priors are further integrated with geometric and contextual priors to form a compact five-dimensional SPK representation for OoD detection. Extensive experiments across diverse object detector architectures and multiple OoD benchmarks demonstrate that SPK achieves state-of-the-art OoD detection. Our findings reveal that pretrained object detectors already encode substantially richer latent knowledge than is typically exploited for OoD detection. More importantly, this knowledge can be explicitly elicited and organized into a compact, structured, and interpretable knowledge space for prediction reliability analysis. This suggests a promising proactive route for improving object detector reliability by explicitly uncovering and leveraging latent priors. Code and data are available at: https://gricad-gitlab.univ-grenoble-alpes.fr/dnn-safety/spk