Figure 1: Overview of LeapTalk. Given audio and a reference image, our method enables open-ended streaming talking-head generation with consistent identity. It achieves 1-step inference per chunk at up to 200 FPS, delivering up to 𝟏𝟓𝟎𝟎𝟎× speedup while maintaining strong lip-sync accuracy.
Table 1: Comparison on HDTF and CelebV-HQ datasets.
HDTF
CelebV-HQ
Model
NFE
FID↓
FVD↓
Sync-C↑
Sync-D↓
IQA↑
ASE↑
FPS↑
FID↓
FVD↓
Sync-C↑
Sync-D↓
IQA↑
ASE↑
FPS↑
StableAvatar
50
176
329
8.11
8.05
6.51
2.69
0.42
318
492
4.73
8.61
5.42
3.30
0.42
Echomimic
30
722
981
5.32
10.02
6.16
2.29
0.52
642
1885
1.09
11.02
0.15
3.19
0.52
Hallo3
51
871
972
7.14
9.23
6.24
2.60
0.30
842
1104
2.49
8.85
5.23
3.22
0.29
FantasyTalking
30
459
884
6.33
9.41
6.13
2.45
0.15
544
1429
1.84
8.61
5.21
3.18
0.15
OmniAvatar
50
168
623
3.10
12.36
6.32
2.68
0.18
351
923
1.25
10.03
5.39
3.27
0.16
Soulx-Flashhead
4
30
452
8.07
8.24
6.60
2.95
14.42
71
642
4.77
8.23
5.57
3.35
14.42
OURS (Lite)
1
38
285
8.14
7.89
6.22
2.70
200
47
456
4.80
8.22
5.51
3.33
200
OURS (Pro)
1
21
197
8.38
7.69
6.53
2.74
55
42
370
5.05
8.21
5.58
3.34
55
Figure 2: Comparison between conventional noise-to-data flow matching and our Bridge Forcing paradigm. The identity consistency is preserved along bridge process, while error accumulates in the forward diffusion process.
Table 2: Ablation study results.
Method
FID
Sync-C
Sync-D
LeapTalk
21
8.38
7.69
w/o Brownian Bridge
217
7.16
11.05
w/o Time Transformation
378
7.84
8.13
w/o Audio-Driven CFG
162
4.34
10.21
Figure 3: Distillation pipeline of our approach. The one-step student generates frames conditioned on a static reference. To enable stable distillation across heterogeneous processes, we apply a time transformation t=Φ(τ) to align noise levels between the flow-matching teacher and the bridge-based student, allowing consistent score supervision and a well-defined DMD objective.
Table 3: Perceptual loss weight sensitivity. Bold: best.
Metric
λperc=0
λperc=1.0
λperc=2.0
λperc=4.0
λperc=8.0
PSNR↑
18.62
19.18
19.30
19.70
17.31
SSIM↑
0.625
0.683
0.692
0.704
0.349
LPIPS↓
0.203
0.197
0.194
0.183
0.562
Figure 4: Qualitative comparison on long-video streaming generation. Frames are sampled from streaming rollouts at increasing time indices. The red boxes mark identity inconsistency and drift, and the blue boxes mark slight or incorrect lip-sync. LeapTalk maintains stable identity and lip motion as generation proceeds.
Table 4: User study results across different evaluation criteria. The table reports the percentage of participants who prefer our method to another method. A higher percentage indicates stronger user preference and better perceived performance.
Method
Identity Cons.
Lip-sync Acc.
Visual Quality
Overall Pref.
StableAvatar
92.40%
93.14%
91.25%
91.37%
Echomimic
91.26%
96.14%
94.36%
95.11%
SoulX-FlashHead
95.23%
92.16%
91.89%
95.32%
Hallo3
94.81%
95.48%
96.29%
96.21%
FantasyTalking
96.89%
93.32%
95.36%
97.90%
OmniAvatar
95.95%
94.31%
97.43%
95.75%
Figure 5: DINO similarity over video time. We plot the similarity between each generated frame and the reference image as streaming generation progresses. A flatter and higher curve indicates less identity drift over long-duration generation.
Table 5: Comparison of autoencoder backbones used in Lite and Pro variants. Encoding and decoding speeds are measured on video clips of 81 frames under BF16 precision.
Autoenc.
Arch.
Enc. Speed
Dec. Speed
Enc. Mem.
Dec. Mem.
WanVAE
Causal Conv3D
4.17s
5.26s
8.495GB
10.128GB
TAEHV
Conv2D
0.39s
0.24s
0.008GB
0.411GB
Figure 6: Ablation results demonstrating the contribution of each component in our framework.
Table 6: Motion diversity vs. baselines on the HDTF samples. Bold: best; uline: 2nd best.
Method
Yaw Std↑
Pitch Std↑
Roll Std↑
Avg Std↑
BAS↑
EchoMimic
1.734
1.867
0.771
1.457
0.650
OmniAvatar
4.332
5.888
1.909
4.043
0.652
SoulX-Flashhead
3.594
2.918
1.959
2.824
0.684
Ours
6.177
5.632
2.146
4.652
0.696
Figure 7: Visualization of ablation effects of audio-driven CFG.
Table 7: Audio CFG scale ablation on HDTF.
CFG Scale
Yaw Std↑
Pitch Std↑
Roll Std↑
Avg Std↑
BAS↑
1.0
1.553
2.653
0.757
1.655
0.723
3.0
4.021
4.493
1.668
3.394
0.658
5.0
6.840
5.764
2.396
5.000
0.696
7.0
8.533
7.635
2.801
6.323
0.650
Figure 8: Sensitivity analysis of audio-driven CFG.
Table 8: Chunk size ablation at 512×512, 1-step inference, single A100 GPU. Tgen/Tchunk<1 indicates real-time. Bold represents chosen default, which achieves the best ratio.
Chunk Size
Tchunk (s)
Tgen (s)
Tgen/Tchunk ↓
FPS ↑
9
0.36
0.088
0.24
45.4
17
0.68
0.146
0.22
82.0
33
1.32
0.268
0.20
104.7
49
1.96
0.432
0.22
101.9
65
2.60
0.629
0.24
95.3
Figure 9: Visual comparison across different VAEs.
Table 9: Inference speed and average chunk generation latency under different resolutions on a single A100 GPU (1-step inference).
Resolution
256×256
384×384
𝟓𝟏𝟐×𝟓𝟏𝟐
768×768
1024×1024
FPS ↑
476
191
104
35
14
Latency (s) ↓
0.059
0.146
0.270
0.793
1.908
Figure 10: Generated results under diverse and challenging input scenarios.
Figure 11: More qualitative results on challenging conditions including non-human faces (e.g., animals), artistic portraits (e.g., paintings and stylized illustrations), side-view photos, low-light conditions, partial occlusions, and even non-photorealistic objects such as sculptures.
Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos. At the heart of our approach lies a single-step