Figure 1: FlashPrefill V2 evolves FlashPrefill [8] from an algorithmic prototype toward practical long-context serving along three dimensions: (1) a mean correction term that preserves model accuracy under extreme sparsity; (2) an FA3/4-aligned sparse attention kernel with PackGQA memory access, warp specialization, pingpong pipelining, and FP8 support; and (3) native compatibility with paged KV cache and continuous batching for integration as an attention backend in modern serving frameworks such as SGLang [42].
Table 1: RULER scores (left) and attention operator speedups over FlashAttention-2 (right) at sequence lengths from 4K to 128K. The speedups are measured end-to-end on needle-in-a-haystack inputs, including pattern discovery, thresholding, and mean correction. Results marked with ∗ (in gray) are measured with direct online FP8 quantization without corrected weights.
Method
RULER Score
Operator Speedup over FA2
4K
8K
16K
32K
64K
128K
Avg.
4K
8K
16K
32K
64K
128K
Llama-3.1-8B-Instruct
Full
95.98
94.74
93.07
89.06
86.27
73.82
88.82
1.00×
1.00×
1.00×
1.00×
1.00×
1.00×
MInference
95.06
94.23
92.18
87.68
83.01
71.61
87.30
0.12×
0.18×
0.46×
0.83×
1.34×
2.45×
FlexPrefill
95.18
94.09
92.35
86.79
84.08
72.01
87.42
0.12×
0.33×
0.98×
2.21×
4.16×
5.18×
XAttention
95.24
94.37
91.61
87.03
83.65
72.33
87.37
0.79×
1.29×
1.83×
2.34×
3.19×
3.48×
FlashPrefill
95.16
94.29
92.11
87.49
84.18
70.91
87.36
1.21×
2.38×
4.31×
7.46×
13.62×
22.67×
FlashPrefill V2
95.52
94.65
92.95
88.48
83.07
72.08
87.79
1.66×
2.62×
4.28×
7.63×
14.39×
27.21×
FlashPrefill V2-FP8∗
95.28
94.48
92.78
88.80
80.28
67.78
86.57
2.67×
4.56×
7.79×
13.89×
25.76×
47.33×
Qwen3-4B-Instruct-2507
Full
94.82
92.98
91.68
88.86
82.81
71.22
87.06
1.00×
1.00×
1.00×
1.00×
1.00×
1.00×
MInference
94.06
92.18
91.09
88.07
82.08
68.01
85.92
0.11×
0.15×
0.42×
0.83×
1.32×
2.53×
FlexPrefill
93.81
90.96
91.41
87.62
81.37
67.24
85.40
0.12×
0.33×
0.99×
2.28×
4.01×
5.27×
XAttention
94.42
91.49
91.42
87.31
81.56
68.24
85.74
0.81×
1.35×
1.92×
2.38×
3.07×
3.43×
FlashPrefill
94.71
92.04
91.28
87.29
81.08
69.02
85.90
1.21×
2.40×
4.36×
6.91×
11.54×
19.26×
FlashPrefill V2
94.49
92.42
91.56
87.48
81.67
69.76
86.23
1.72×
2.82×
4.63×
7.92×
13.93×
27.45×
FlashPrefill V2-FP8
94.13
92.37
91.36
87.47
80.59
68.97
85.82
2.75×
4.87×
8.40×
14.44×
25.06×
47.59×
Qwen3-30B-A3B-Instruct-2507
Full
95.78
95.23
94.36
92.72
90.11
84.12
92.05
1.00×
1.00×
1.00×
1.00×
1.00×
1.00×
MInference
95.71
94.33
93.81
92.01
89.02
82.09
91.16
0.11×
0.17×
0.46×
0.83×
1.43×
2.76×
FlexPrefill
95.32
94.71
93.52
91.81
88.09
83.11
91.09
0.12×
0.34×
0.98×
2.31×
3.92×
5.43×
XAttention
95.43
94.12
92.78
91.06
88.39
82.23
90.67
0.86×
1.44×
2.01×
2.42×
2.95×
3.42×
FlashPrefill
95.62
94.47
93.28
92.17
88.14
82.96
91.11
1.22×
2.43×
4.36×
6.92×
11.45×
18.67×
FlashPrefill V2
95.35
94.85
94.19
92.48
89.88
83.83
91.76
1.75×
2.95×
4.62×
8.21×
14.11×
27.19×
FlashPrefill V2-FP8
95.28
94.53
94.12
92.34
89.54
82.55
91.39
2.78×
5.12×
8.41×
14.93×
25.36×
47.26×
Figure 2: Speedup of various attention operators relative to FlashAttention-2 [5] on NVIDIA H20 GPUs, including an FA3/4-aligned dense baseline [22, 38]. All results are measured with a batch size of 4. FlashPrefill V2 exhibits a dominant advantage, particularly in long-context scenarios, and its FP8 variant further amplifies the gains.
Table 2: Performance comparison on all 21 tasks of LongBench [2], grouped by task category. Results marked with ∗ (in gray) are measured with direct online FP8 quantization without corrected weights.
Method
Single-Document QA
Multi-Document QA
Summarization
Few-shot Learning
Synthetic
Code
Avg.
NarrQA
Qasper
MF-en
MF-zh
Hotpot
2Wiki
MuSiQue
DuRead
GovRep
QMSum
MNews
VCSum
TREC
TQA
SAMSum
LSHT
Count
PR-en
PR-zh
LCC
Llama-3.1-8B-Instruct
Full
29.39
44.61
56.34
63.63
57.48
48.08
32.28
34.43
34.51
25.29
26.80
17.37
72.50
91.48
43.59
46.50
11.10
100.00
90.45
MInference
28.13
42.36
53.67
62.08
56.28
46.72
32.01
33.41
33.28
25.16
27.02
16.84
69.50
90.83
42.87
43.00
6.00
96.00
84.00
FlexPrefill
27.62
42.67
52.98
62.34
57.14
45.32
30.97
33.72
33.87
24.58
26.34
17.02
67.50
90.36
44.72
43.00
6.00
94.00
83.33
XAttention
28.52
43.12
54.01
60.19
56.73
45.98
32.17
33.15
34.55
25.37
27.21
17.41
71.00
91.37
44.05
44.00
6.88
95.50
85.68
FlashPrefill
28.31
42.89
54.23
61.97
57.01
45.76
31.96
33.60
34.02
24.83
26.47
17.13
71.00
91.12
44.21
46.50
7.00
95.50
83.45
FlashPrefill V2
30.21
44.56
55.97
62.37
58.02
48.84
32.57
33.78
34.90
25.24
27.13
17.37
72.50
91.45
44.76
45.50
7.39
97.00
89.47
FlashPrefill V2-FP8∗
29.08
45.51
54.70
62.59
55.41
46.40
32.03
32.97
34.70
24.88
26.97
17.61
70.00
91.16
44.80
45.00
6.88
96.50
85.68
Qwen3-4B-Instruct-2507
Full
27.97
44.98
51.12
64.59
59.39
44.18
24.85
26.17
30.68
22.62
23.93
11.73
74.00
87.74
45.55
44.00
0.75
100.00
98.04
MInference
25.89
41.69
48.26
62.16
54.17
42.76
22.16
24.89
29.12
21.98
23.15
11.28
72.00
87.62
44.32
42.50
0.00
91.00
94.67
FlexPrefill
25.26
41.48
48.02
61.98
54.96
43.21
22.71
25.43
30.94
22.15
24.47
11.47
70.50
86.18
45.87
41.00
0.50
94.00
94.20
XAttention
24.31
42.67
49.01
63.02
54.14
43.14
22.56
25.12
29.63
22.41
23.68
11.56
72.50
86.47
45.13
41.50
0.00
96.00
93.89
FlashPrefill
25.06
42.78
49.22
62.31
55.07
42.97
23.42
25.66
29.80
22.09
23.32
11.62
72.00
87.95
45.42
42.00
0.00
92.00
94.00
FlashPrefill V2
27.56
43.73
51.07
63.72
56.39
43.11
23.18
26.36
30.48
22.67
23.88
11.76
74.00
87.74
45.78
41.25
0.50
96.00
97.17
FlashPrefill V2-FP8
26.55
42.70
50.28
61.71
55.21
41.95
24.50
26.03
30.74
22.47
24.07
11.48
73.50
88.16
45.89
41.25
0.50
93.00
94.67
Qwen3-30B-A3B-Instruct-2507
Full
31.29
44.09
54.79
66.48
63.86
56.56
32.29
24.64
31.10
21.90
23.48
11.09
76.50
91.01
45.79
51.00
16.50
100.00
100.00
MInference
28.62
41.08
53.06
64.28
61.21
51.01
31.71
23.52
30.53
21.07
22.79
10.72
75.00
90.16
44.96
46.00
5.00
95.00
97.67
FlexPrefill
28.07
40.96
52.41
64.07
60.72
50.91
31.65
23.97
30.02
21.62
23.86
10.88
75.50
89.37
46.38
48.50
4.50
97.00
96.00
XAttention
27.98
42.86
52.83
63.82
60.93
53.16
33.01
23.81
31.31
21.28
22.93
10.97
73.00
91.14
45.92
47.50
3.00
97.00
99.00
FlashPrefill
28.35
41.11
53.34
64.81
61.27
51.09
30.08
24.11
30.55
21.74
23.21
11.02
77.00
90.68
46.12
46.50
5.00
96.50
99.00
FlashPrefill V2
29.74
41.58
53.82
66.80
63.66
52.62
34.11
24.94
31.00
22.04
23.45
11.13
77.00
90.99
46.59
49.00
6.50
99.50
99.50
FlashPrefill V2-FP8
30.03
41.37
55.28
66.31
61.40
51.83
30.63
24.58
31.19
21.62
23.56
10.96
77.50
89.61
46.08
47.50
8.50
98.00
99.00
Figure 3: The FlashPrefill V2 prefill pipeline. Stage 1 packs queries with PackGQA, pools block-mean statistics, and produces a CSR index of selected blocks in a single fused pass; Stage 2 runs the warp-specialized sparse attention kernel over the selected blocks (exact path) while the unselected blocks are compensated through the mean correction path (Sec. 3.2), and the two streams merge inside the online softmax.
Table 3: Attention density of FlashPrefill V2 measured on needle-in-a-haystack inputs at each sequence length.
Model
4K
8K
16K
32K
64K
128K
Llama-3.1-8B
76.0%
54.2%
33.6%
18.7%
9.0%
4.6%
Qwen3-4B
72.0%
48.5%
29.8%
17.6%
9.2%
4.9%
Qwen3-30B-A3B
70.4%
46.0%
29.6%
16.2%
9.4%
4.9%
Figure 4: Mean correction. Selected blocks (red) are computed exactly, while each pruned block is pooled into its mean statistics (k¯J,v¯J) and contributes a surrogate term |ℬJ|es¯J(v¯J,1) to the softmax numerator and denominator, recovering the discarded probability mass without per-token computation.
Table 4: End-to-end time-to-first-token (TTFT, seconds) of Llama-3.1-8B-Instruct, Qwen3-4B-Instruct-2507, and Qwen3-30B-A3B-Instruct-2507 served by SGLang on the needle-in-a-haystack workload. Each entry is averaged over 30 repeated measurements. Parenthesized numbers are speedups over the FA3/4 backend.
BSZ
FA3/4
FlashPrefill V2
FlashPrefill V2-FP8
4K
8K
16K
32K
64K
128K
4K
8K
16K
32K
64K
128K
4K
8K
16K
32K
64K
128K
Llama-3.1-8B-Instruct
1
0.15
0.29
0.63
1.49
3.95
11.81
0.15 (1.02×)
0.29 (1.02×)
0.58 (1.09×)
1.17 (1.28×)
2.46 (1.60×)
5.49 (2.15×)
0.10 (1.53×)
0.18 (1.63×)
0.35 (1.81×)
0.71 (2.09×)
1.48 (2.66×)
3.23 (3.66×)
4
0.58
1.17
2.54
5.94
12.30
32.94
0.57 (1.01×)
1.14 (1.03×)
2.32 (1.09×)
4.79 (1.24×)
7.69 (1.60×)
15.39 (2.14×)
0.35 (1.66×)
0.68 (1.73×)
1.36 (1.87×)
2.80 (2.12×)
4.49 (2.74×)
8.84 (3.73×)
8
1.07
2.21
4.87
9.83
21.23
56.67
1.06 (1.01×)
2.14 (1.03×)
4.39 (1.11×)
7.77 (1.27×)
13.18 (1.61×)
26.13 (2.17×)
0.67 (1.58×)
1.39 (1.59×)
2.81 (1.73×)
4.51 (2.18×)
7.66 (2.77×)
14.93 (3.80×)
16
2.60
3.85
8.14
16.10
37.12
102.84
2.58 (1.01×)
3.59 (1.07×)
7.31 (1.11×)
12.56 (1.28×)
22.93 (1.62×)
46.96 (2.19×)
1.44 (1.80×)
2.97 (1.30×)
4.62 (1.76×)
7.28 (2.21×)
13.25 (2.80×)
26.82 (3.83×)
Qwen3-4B-Instruct-2507
1
0.09
0.19
0.43
1.13
3.35
11.10
0.09 (1.02×)
0.18 (1.07×)
0.36 (1.19×)
0.77 (1.47×)
1.71 (1.96×)
4.12 (2.70×)
0.07 (1.42×)
0.12 (1.61×)
0.23 (1.87×)
0.49 (2.33×)
1.08 (3.09×)
2.44 (4.55×)
4
0.35
0.74
1.69
4.35
10.61
31.47
0.35 (1.01×)
0.69 (1.06×)
1.43 (1.18×)
3.06 (1.42×)
5.38 (1.97×)
11.55 (2.72×)
0.22 (1.59×)
0.43 (1.72×)
0.88 (1.93×)
1.86 (2.34×)
3.26 (3.25×)
6.73 (4.67×)
8
0.64
1.38
3.25
7.40
18.25
53.84
0.63 (1.02×)
1.29 (1.07×)
2.70 (1.20×)
5.05 (1.47×)
9.19 (1.99×)
19.35 (2.78×)
0.43 (1.49×)
0.87 (1.59×)
1.78 (1.82×)
3.06 (2.42×)
5.51 (3.31×)
11.28 (4.77×)
16
1.07
2.31
5.53
12.20
31.73
97.26
1.05 (1.02×)
2.15 (1.08×)
4.56 (1.21×)
8.18 (1.49×)
15.84 (2.00×)
34.48 (2.82×)
0.90 (1.18×)
1.83 (1.26×)
2.95 (1.87×)
4.92 (2.48×)
9.43 (3.37×)
20.13 (4.83×)
Qwen3-30B-A3B-Instruct-2507
1
0.10
0.19
0.45
1.26
4.00
13.78
0.09 (1.02×)
0.18 (1.09×)
0.36 (1.26×)
0.78 (1.61×)
1.75 (2.29×)
4.17 (3.30×)
0.09 (1.06×)
0.15 (1.24×)
0.30 (1.50×)
0.61 (2.07×)
1.32 (3.03×)
3.00 (4.59×)
4
0.26
0.73
1.76
4.86
13.05
40.44
0.24 (1.09×)
0.67 (1.10×)
1.42 (1.24×)
3.11 (1.56×)
5.70 (2.29×)
12.16 (3.33×)
0.22 (1.18×)
0.57 (1.29×)
1.15 (1.53×)
2.41 (2.01×)
4.21 (3.10×)
8.59 (4.71×)
8
0.64
1.35
3.36
8.47
22.37
68.56
0.62 (1.03×)
1.22 (1.11×)
2.65 (1.27×)
5.25 (1.61×)
9.69 (2.31×)
20.35 (3.37×)
0.55 (1.15×)
1.02 (1.32×)
2.11 (1.59×)
4.03 (2.10×)
7.13 (3.14×)
14.35 (4.78×)
16
1.16
3.14
5.90
13.98
38.50
123.23
1.15 (1.01×)
2.92 (1.08×)
4.60 (1.28×)
8.49 (1.65×)
16.55 (2.33×)
36.21 (3.40×)
1.04 (1.12×)
1.73 (1.81×)
3.66 (1.61×)
6.47 (2.16×)
12.16 (3.17×)
25.51 (4.83×)
Figure 5: Block-sparse attention latency at 64K sequence length in FP8. HPC-BSA does not support BF16, so only FP8 is compared. Both sparse operators track the theoretical bound; FlashPrefill V2 is 6 to 7% faster than HPC-BSA at every sparsity level.
Table 5: Open-loop serving results on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 under Poisson arrivals with mixed prompt lengths (4K–128K RULER needle-in-a-haystack, 100 requests per cell, 64 output tokens each). Parenthesized numbers are speedups over the FA3/4 backend.
Rate (req/s)
Backend
TTFT (s)
TPOT (ms)
Throughput
P50
P99
P50
P99
(tok/s)
(req/s)
Qwen3-4B-Instruct-2507
1
FA3/4
76.73
154.64
419
1170
24.0
0.37
FlashPrefill V2
17.23 (4.45x)
48.73 (3.17x)
258 (1.62x)
601 (1.94x)
45.5 (1.90x)
0.71
FlashPrefill V2-FP8
2.79 (27.47x)
17.27 (8.95x)
47 (8.84x)
487 (2.40x)
58.4 (2.44x)
0.91
4
FA3/4
79.05
161.54
513
1102
23.8
0.37
FlashPrefill V2
42.28 (1.87x)
80.34 (2.01x)
258 (1.99x)
509 (2.16x)
47.3 (1.99x)
0.74
FlashPrefill V2-FP8
24.47 (3.23x)
46.05 (3.51x)
165 (3.10x)
331 (3.33x)
82.3 (3.46x)
1.29
8
FA3/4
79.55
161.01
463
1110
23.9
0.37
FlashPrefill V2
43.83 (1.81x)
80.12 (2.01x)
236 (1.96x)
523 (2.12x)
48.3 (2.03x)
0.76
FlashPrefill V2-FP8
25.83 (3.08x)
46.24 (3.48x)
152 (3.05x)
302 (3.68x)
82.2 (3.45x)
1.28
16
FA3/4
81.91
160.99
481
1110
23.8
0.37
FlashPrefill V2
43.22 (1.90x)
78.68 (2.05x)
263 (1.83x)
501 (2.21x)
48.9 (2.05x)
0.76
FlashPrefill V2-FP8
25.14 (3.26x)
45.16 (3.56x)
152 (3.17x)
285 (3.89x)
85.5 (3.59x)
1.34
Qwen3-30B-A3B-Instruct-2507
1
FA3/4
90.69
195.43
582
1335
19.9
0.31
FlashPrefill V2
17.63 (5.14x)
52.15 (3.75x)
241 (2.42x)
604 (2.21x)
44.5 (2.24x)
0.70
FlashPrefill V2-FP8
7.93 (11.44x)
21.92 (8.91x)
102 (5.69x)
637 (2.10x)
56.1 (2.82x)
0.88
4
FA3/4
101.44
195.33
462
1312
19.9
0.31
FlashPrefill V2
46.26 (2.19x)
85.57 (2.28x)
255 (1.81x)
520 (2.52x)
45.5 (2.29x)
0.71
FlashPrefill V2-FP8
30.62 (3.31x)
57.43 (3.40x)
198 (2.33x)
407 (3.22x)
64.6 (3.25x)
1.01
8
FA3/4
105.52
195.34
481
1333
19.9
0.31
FlashPrefill V2
45.73 (2.31x)
83.04 (2.35x)
234 (2.05x)
529 (2.52x)
47.0 (2.36x)
0.73
FlashPrefill V2-FP8
33.17 (3.18x)
58.46 (3.34x)
165 (2.92x)
369 (3.62x)
66.9 (3.36x)
1.05
16
FA3/4
99.68
195.20
547
1347
19.9
0.31
FlashPrefill V2
45.20 (2.21x)
81.95 (2.38x)
266 (2.06x)
514 (2.62x)
47.4 (2.38x)
0.74
FlashPrefill V2-FP8
32.47 (3.07x)
57.33 (3.40x)
185 (2.95x)
359 (3.75x)
67.7 (3.41x)
1.06
Figure 6: Overhead of mean correction at 64K sequence length. The corrected variants (dashed) stay close to the uncorrected ones (solid) in both BF16 and FP8, and the absolute overhead shrinks as sparsity increases.
Table 6: Open-loop serving at 16 req/s with chunked prefill at chunk sizes of 8K and 16K (same workload as Tab. 5, where chunking is disabled). Parenthesized numbers are speedups over the FA3/4 backend at the same chunk size.
Chunk size
Backend
TTFT (s)
TPOT (ms)
Throughput
P50
P99
P50
P99
(tok/s)
(req/s)
Qwen3-4B-Instruct-2507
8K
FA3/4
80.98
165.14
597
1228
22.5
0.35
FlashPrefill V2
50.15 (1.61x)
95.23 (1.73x)
360 (1.66x)
684 (1.80x)
38.7 (1.72x)
0.60
FlashPrefill V2-FP8
33.47 (2.42x)
61.20 (2.70x)
259 (2.30x)
529 (2.32x)
60.7 (2.70x)
0.95
16K
FA3/4
79.42
161.96
580
1203
22.9
0.36
FlashPrefill V2
45.09 (1.76x)
85.47 (1.89x)
315 (1.84x)
615 (1.96x)
43.9 (1.91x)
0.69
FlashPrefill V2-FP8
26.80 (2.96x)
50.34 (3.22x)
193 (3.00x)
382 (3.15x)
73.3 (3.20x)
1.15
Qwen3-30B-A3B-Instruct-2507
8K
FA3/4
96.61
199.39
693
1494
18.7
0.29
FlashPrefill V2
79.81 (1.21x)
149.52 (1.33x)
574 (1.21x)
1200 (1.24x)
24.6 (1.32x)
0.38
FlashPrefill V2-FP8
63.23 (1.53x)
121.53 (1.64x)
505 (1.37x)
1009 (1.48x)
29.2 (1.56x)
0.46
16K
FA3/4
94.78
195.84
672
1466
19.0
0.30
FlashPrefill V2
46.47 (2.04x)
88.15 (2.22x)
323 (2.08x)
652 (2.25x)
42.1 (2.21x)
0.66
FlashPrefill V2-FP8
33.72 (2.81x)
62.19 (3.15x)
250 (2.69x)
464 (3.16x)
57.1 (3.00x)
0.89
Table 7: Design comparison with the HPC-Ops block-sparse attention kernel. Index memory is evaluated at batch size 1 with 32 query heads and 4 KV heads: the dense mask stores one byte per (query head, 128-token Q tile, 128-token KV tile) regardless of density, whereas the CSR index stores only the selected 64-token tiles as int32 at the measured needle-in-a-haystack densities (Tab. 3), is organized per KV head under PackGQA, and matches the SGLang index format. The fused correction path has no counterpart in HPC-Ops BSA.
HPC-Ops BSA
FlashPrefill V2
Sparse index format
dense mask, padded, per Q head
CSR of selected blocks, per KV head
Index memory @ 64K / 128K
8 MB / 32 MB (any density)
3 MB / 6 MB (9% / 5% density)
GEMM–softmax overlap
serial
intra-warpgroup pingpong
Paged KV addressing
per-page TMA (page size ∣ 128)
cp.async, arbitrary page size
KV splitting
none
count-balanced CSR splits
Mean correction
none
fused into the mainloop
Precision
FP8 only
BF16 + FP8
Table 8: Ablation of the mean correction term on RULER with Qwen3-4B-Instruct-2507, in both BF16 and FP8. Discarding unselected blocks without compensation degrades accuracy increasingly toward longer contexts, while the zero-order correction stays close to full attention at all lengths in both precisions.
Method
4K
8K
16K
32K
64K
128K
Avg.
Full Attention
94.82
92.98
91.68
88.86
82.81
71.22
87.06
V2 w/o correction
94.38
92.06
91.18
87.21
80.92
68.86
85.77
FlashPrefill V2
94.49
92.42
91.56
87.48
81.67
69.76
86.23
V2-FP8 w/o correction
93.78
91.82
90.68
85.04
76.86
62.78
83.49
FlashPrefill V2-FP8
94.13
92.37
91.36
87.47
80.59
68.97
85.82
Table 9: Ablation of the selection threshold α on RULER at 64K with Qwen3-4B-Instruct-2507 in FP8. Density is the fraction of selected attention blocks. Red numbers are the score drops of the uncorrected variant relative to the corrected pipeline at the same α. Full attention scores 82.81.
Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic prototype that is still distant from production deployment. In this paper, we present FlashPrefill V2, which evolves FlashPrefill from a prototype toward practical long-context serving along three dimensions. First, we introduce a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements. Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing integration as an attention backend in modern inference frameworks such as SGLang. Extensive evaluations on NVIDIA H20 GPUs---among the most widely deployed inference accelerators---demonstrate that FlashPrefill V2 delivers up to 47.26x and 27.19x speedups over FlashAttention-2 at 128K context length under FP8 and BF16 precision, respectively, and, in FP8, still achieves a 30.49x speedup against an FA3/4-aligned dense baseline.
作者 · Qihang Fan, Huaibo Huang, Zhiying Wu, Bingning Wang, Ran He