Figure 1: Referential dangling with a missing bridge. Independent scoring retains the query subject and the answer string but removes the fact that Tim DuBois was born in Southwest City, leaving the inference chain incomplete.
Table 1: Referential dangling and complete evidence retention under Beaver at r=0.30.
Dataset
Hops
ρd (%)
ρe (%)
HotpotQA
2
34.2
61.0
2WikiMultiHopQA
2
53.5
30.7
MuSiQue
2 to 4
54.2
27.0
Figure 2: Referential dangling under Beaver. Panel (a) reports ρd across compression ratios on HotpotQA (n=269 to 300 per point, including partial paragraph retention). Panel (b) reports ρd by annotated hop count on HotpotQA (n=234), 2WikiMultiHopQA (n=241), and MuSiQue (n=286). Panel (c) reports dangling rates for 4,649 reference pairs in LongBench-v2 Single-Document QA by sentence distance from first mention to later reference. Error bars are bootstrap 95% confidence intervals.
Table 2: Pairwise Jaccard similarities between dangling case sets on the shared HotpotQA bridge set (n=184) at compression ratio 0.30. The first row reports the dangling rate of each compressor. Abbreviations match Figure 3.
BEAVER
PartPr.
Sel.-Ctx
LLML-2
DAC
LongLL
Dangling rate (%)
32.1
47.8
51.6
56.0
58.7
59.8
BEAVER
N/A
0.36
0.23
0.30
0.29
0.32
PartPr.
0.36
N/A
0.36
0.44
0.44
0.37
Sel.-Ctx
0.23
0.36
N/A
0.39
0.35
0.51
LLML-2
0.30
0.44
0.39
N/A
0.47
0.45
DAC
0.29
0.44
0.35
0.47
N/A
0.48
LongLL
0.32
0.37
0.51
0.45
0.48
N/A
Figure 3: Dangling rates for six compressors on the shared HotpotQA bridge set (n=184) at compression ratio 0.30. PartPr. denotes PartPrompt, Sel.-Ctx denotes Selective-Context, LLML-2 denotes LLMLingua-2, and LongLL denotes LongLLMLingua. All outputs are evaluated using the content-word overlap criterion with threshold 0.5. Light bars denote methods that use the query, and darker bars denote methods that do not.
Table 3: Answer accuracy with base contexts produced by Beaver at target compression ratio 0.30. Panel 1 uses the dangling subsets of HotpotQA (n=80), 2WikiMultiHopQA (n=72), and MuSiQue (n=102), with McNemar p values comparing Base and Reselected. Panel 2 uses a separate set of 200 HotpotQA examples, with McNemar p values comparing Base and Full support.
Panel 1: dangling subsets evaluated with Qwen3-8B
Dataset
Downstream LLM
Base
Reselected
Full support
McNemar p
HotpotQA
Qwen3-8B
0.287
0.575
0.600
1.6×10−6
2WikiMultiHopQA
Qwen3-8B
0.097
0.403
0.444
1.1×10−5
MuSiQue
Qwen3-8B
0.147
0.490
0.471
3.1×10−8
Panel 2: a 200 example HotpotQA evaluation set with four downstream LLMs
Dataset
Downstream LLM
Base
Full support
McNemar p
HotpotQA
Qwen3-8B
0.535
0.615
0.001
Qwen3-4B
0.500
0.585
0.002
Llama-3.1-8B
0.575
0.660
0.004
Mistral-7B
0.455
0.545
0.0005
Figure 4: Mean salience percentiles (%) for answer and definition sentences among all sentences in 180 bridge examples. Beaver similarity is query-aware; self-information is not.
Table 4: Answer accuracy of proprietary models under Base and Full support, with base contexts produced by Beaver at target compression ratio 0.30. HotpotQA uses the full shared bridge set, while MuSiQue uses the dangling subset. GLM-5.2 returned answers for 95 of the 102 MuSiQue contexts because of API timeouts.
Model
Dataset
n
Base
Full support
McNemar p
GPT-5.5
HotpotQA
184
0.913
0.913
1.0
GPT-5.5
MuSiQue
102
0.775
0.863
0.011
GLM-5.2
MuSiQue
95
0.695
0.937
2.4×10−7
Figure 5: Dangling rate across content-word-overlap thresholds for 184 bridge examples (Figure 3; ratio 0.30).
Table 5: Changes in answer accuracy, in percentage points relative to Base, for candidate sources with a fixed classifier and Qwen3-8B (K=3). Hybrid augments first-mention candidates with embedding retrieval, and the final row includes the annotated supporting sentence in the candidate set.
Candidate source
HotpotQA
2WikiMultiHopQA
First mention
+4.7 (p=0.022)
+0.5 (not significant)
All mentions
+4.5 (p=0.15)
+4.0 (p=0.20)
Hybrid
+4.5 (p=0.12)
+5.5 (p=0.063)
Annotated support included
+8.0 (p=0.008)
N/A
Figure 6: Referential dangling examples from HotpotQA, 2WikiMultiHopQA, MuSiQue, and LongBench-v2 Single-Document QA. Each panel shows the original context and the compressed output.
Table 6: Dangling rate (%) across content-word overlap retention thresholds on the same 184 bridge examples as Figure 3 at compression ratio 0.30. The 0.5 column matches Figure 3.
Overlap threshold
0.3
0.4
0.5
0.6
0.7
LLMLingua-2 (token)
28.3
43.5
56.0
57.6
36.4
Beaver (chunk)
19.6
25.0
32.1
36.4
40.2
Figure 7: Accuracy gains from full support and first-mention automatic restoration on HotpotQA (Beaver at ratio 0.30, K=3). Full support uses 200 examples; restoration uses 300 for Qwen3-8B and Llama-3.1-8B and 200 for Mistral-7B.
Table 7: Robustness of the dangling diagnostic to its three main free choices (HotpotQA, n=300, ratio 0.30 unless swept). Embedding shifts are measured in percentage points relative to the released Qwen3-0.6B embedding setup.
Check
Variation
Outcome
Embedding scorer
Qwen3-0.6B embeddings → GPT-2
+0.9 points
Overlap threshold
0.3 to 0.7
substantial throughout
Compression ratio
0.70 to 0.20
monotonic increase
Table 8: Referential dangling on LongBench-v2 Single-Document QA by subdomain (Beaver, ratio 0.30, n=80 documents). “Mean rate” is the per-document average fraction of retained sentences that are dangling, macro-averaged over documents. “Affected docs” is the fraction of documents with at least one dangling reference.
Subdomain
n
Mean rate
Affected docs
Academic
13
36.5%
100%
Literary
12
34.4%
100%
Financial
12
32.3%
100%
Legal
8
29.1%
100%
Detective
15
27.3%
100%
Event ordering
11
27.0%
100%
Governmental
9
25.1%
100%
All
80
30.5%
𝟏𝟎𝟎%
Table 9: Official checkpoint and API identifiers. Display names are the shorthand used in the paper; exact identifiers are shown for reproducibility.
Table 10: Automatic restoration results with the classifier fixed at K=3. The evaluation uses 300 HotpotQA examples, except for Mistral-7B, which uses 200. Base is Beaver at compression ratio 0.30, and Restored has an average ratio of 0.31. The reported p values use paired McNemar tests.
Downstream LLM
Candidate source
Base
Restored
p
Qwen3-8B
First mention
0.567
0.613
0.022
Mistral-7B
First mention
0.455
0.520
0.012
Llama-3.1-8B
First mention
0.587
0.600
0.60
Llama-3.1-8B
Hybrid
0.587
0.610
0.17
Table 11: Restoration statistics when the Beaver baseline was incorrect (HotpotQA, n=300, compression ratio 0.30; downstream Qwen3-8B). SD denotes standard deviation.
Feature
Fixed (23)
Failed (107)
Sentences added, mean ± SD
2.13 ± 1.08
1.79 ± 1.17
Sentences added, median
3.0
2.0
McNemar: 23 fixes, 9 breaks, p=0.022
Table 12: Matched addition control on HotpotQA with Qwen3-8B (n=300). Random insertion and targeted restoration add the same number of sentences per example (m: mean 1.81, median 2, interquartile range [1,3]; approximately 40 tokens; compression ratio 0.30 to 0.31; K=3). Brackets report bootstrap 95% confidence intervals.
Condition
Accuracy [95% CI]
Δ
Base compressor
0.567 [.51,.62]
N/A
Random insertion, m sentences
0.587 [.53,.64]
+2.0
Targeted restoration, m sentences
0.613 [.55,.67]
+4.7
Table 13: Transfer of one restoration configuration across four compressor outputs on HotpotQA with downstream Qwen3-8B (n≈150 to 300).
Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and retaining the highest-scoring units under a budget. We identify a structural failure in this procedure: independent selection can split dependent evidence pairs, retaining one member while deleting the other. When retained text contains an answer but deleted text defines the entity needed to interpret it, we call the result referential dangling. At a compression ratio of 0.30, Beaver, which ranks coherent chunks using Qwen3-0.6B embeddings, leaves the answer path incomplete in 34-54% of bridge examples across three multi-hop question answering datasets. On a shared HotpotQA bridge set, all six hard compressors we test exhibit dangling at rates up to 60%, and every document in LongBench-v2 Single-Document QA contains at least one dangling reference. On dangling examples evaluated with Qwen3-8B, reinserting the missing supporting paragraph while removing nonsupporting paragraphs to maintain the token budget improves accuracy by 29-34 percentage points (p < 0.0001), recovering at least 88% of the gap to contexts retaining both supporting paragraphs. Stronger answer models do not absorb the loss: on MuSiQue, GPT-5.5 is 8.8 points less accurate on compressed contexts than on contexts retaining both supporting paragraphs. Finally, we train a compact classifier to rank omitted sentences by whether they are needed to interpret retained text and reinsert the top-ranked candidates without support annotations at inference. On HotpotQA with Qwen3-8B, this automatic restoration improves accuracy by 4.7 points while changing the compression ratio only from 0.30 to 0.31. Hard compressors should optimize both relevance and referential completeness.