4DAnyone: Create Anyone in 4D from a Casual Monocular Video
arXiv:2608.203352026-08-19
스마트폰으로 찍은 영상 한 편으로 어느 각도에서나 볼 수 있는 4D 사람 모델 만들기
4DAnyone은 사람을 찍은 평범한 단안 영상 한 개만으로, 그 사람을 여러 방향의 가상 카메라로 촬영한 것처럼 일관된 영상 수십 개를 생성한 뒤 이를 4D Gaussian Splatting(4DGS) 모델로 재구성해 어느 시점에서든 자유롭게 렌더링할 수 있게 한다. 기존 카메라 제어 비디오 생성 모델은 몇 개의 새로운 시점 영상은 그럴듯하게 만들지만, 실제 3D 복원에 필요한 수십 개 시점으로 늘리면 일관성이 깨지는 문제가 있었는데, 저자들은 Reference Context Packing과 Target Context Routing이라는 두 가지 장치와 자체 제작한 합성 데이터셋 MVGameHuman으로 이를 해결했다고 밝혔다. 두 개의 벤치마크 데이터셋에서 기존 방법들보다 나은 결과를 보였다.
무엇을 했나
문제 정의: 기존 카메라 제어 영상 생성 모델은 소수의 새 시점에는 그럴듯한 영상을 만들지만, 4DGS 복원에 필요한 16개 이상의 시점으로 확장하면 모델이 한 번에 처리할 수 있는 '주의 문맥'의 한계 때문에 시점 간 일관성이 무너진다.
Reference Context Packing(RCP)은 지금까지 생성된 참조 시점 영상들을 계속 늘어나게 두지 않고 고정된 길이의 혼합 해상도 문맥으로 압축해, 참조 영상이 아무리 많아져도 계산 비용이 일정하게 유지되도록 한다.
Target Context Routing(TCR)은 노이즈 제거 과정에서 시점 그룹을 회전시킨다. 노이즈가 많은 초반 단계에서는 그룹을 섞어 전체 몸 구조에 대한 합의를 이루게 하고, 노이즈가 적은 후반 단계에서는 인접한 시점끼리 고정된 그룹으로 묶어 세부 디테일을 안정화한다.
3D 인지 스켈레톤 조건화: 신뢰하기 어려운 조밀한 깊이 지도 대신, 정확도가 높은 성긴 3D 뼈대 정보를 z-버퍼(깊이 순서를 반영한 렌더링)로 만들어 2D 스켈레톤이 갖는 앞뒤 모호성 문제를 해결했다.
DNA-Rendering과 DyMVHumans 두 테스트 데이터셋에서 4DAnyone은 MV-Performer, TrajectoryCrafter, 그리고 동일 조건으로 미세조정한 ReCamMaster 대비 생성 영상 품질, 시점 간 일관성, 최종 4D 복원 품질 모두에서 앞섰고 실제 촬영된 다양한 영상에서도 잘 일반화됐다.
Figure 1. Given a casually captured monocular video, typically with mild camera motion and unknown camera intrinsics and poses, 4DAnyone generates multiview-consistent human videos, enabling 4DGS reconstruction rendered from free viewpoints. See the project page and Fig. 7 for more results.One continuous 3D scene: a smartphone at the lower left films a bass-playing subject, a ring of twenty-four generated camera views surrounds the stage, and the reconstructed subject stands at the center on a turntable, rendered from a free viewpoint.
Table 1. Training dataset statistics.
Dataset
Videos
Cameras
Actors
Resolution
Type
MVGameHuman
38k
24
318
2560×1440
Multi-view
SynCamVideo
34k
10
66
1280×1280
Multi-view
DNA-Rendering
51k
48
548
2048×2048
Multi-view
TedTalk
42k
1
413
2160×1620
Monocular
Pexels
20k
1
1,411
3840×2160
Monocular
Figure 2. Overview of 4DAnyone. Given a source video, an HMR model (GVHMR (33)) estimates a 3D skeleton sequence, which is rendered into depth-buffered skeleton videos for v target views. A 3D-aware skeleton encoder produces residual skeleton tokens that are added to the noisy target latents. The source tokens, RCP reference tokens, and skeleton-conditioned target tokens are concatenated along the view dimension and processed by the DiT with video attention, multiview attention, and text cross-attention. Generated target videos are packed back into the RCP context for subsequent generation rounds, while TCR improves consistency during grouped target-view generation. The final multi-view videos are used to train a 4DGS model with FreeTimeGS (39).Pipeline diagram. A source video is converted into skeleton-conditioned target-view tokens. Reference Context Packing supplies compact reference tokens, Target Context Routing exchanges information among target-view groups, and the generated multiview videos are reconstructed as a 4D Gaussian Splatting model.
Table 2. Quantitative comparison on DNA-Rendering and DyMVHumans. We evaluate 4DGS reconstruction, generated video consistency, and generated video reconstruction. Best results are in bold. † denotes our fine-tuned version.
Method
DNA-Rendering
DyMVHumans
PSNR↑
SSIM↑
LPIPS↓
PSNR↑
SSIM↑
LPIPS↓
Gen. Video Consistency
MV-Performer (59)
21.25
0.799
0.222
19.98
0.797
0.182
TrajectoryCrafter (52)
13.56
0.641
0.331
15.19
0.769
0.204
ReCamMaster† (2)
21.47
0.806
0.210
21.94
0.833
0.142
Ours
24.33
0.862
0.163
24.48
0.862
0.109
4DGS Reconstruction
MV-Performer (59)
20.38
0.826
0.191
18.69
0.787
0.158
TrajectoryCrafter (52)
14.81
0.718
0.331
15.11
0.748
0.217
ReCamMaster† (2)
20.55
0.807
0.214
19.86
0.795
0.159
Ours
24.15
0.863
0.159
23.28
0.846
0.117
Gen. Video Reconstruction
MV-Performer (59)
19.33
0.803
0.204
14.36
0.731
0.230
TrajectoryCrafter (52)
13.68
0.698
0.358
14.11
0.725
0.247
ReCamMaster† (2)
20.74
0.809
0.204
19.18
0.778
0.168
Ours
23.69
0.850
0.165
21.03
0.808
0.143
Figure 3. Progressive inference with Reference Context Packing and Target Context Routing. We first generate reference videos and pack them with the source video into a fixed RCP context. During target-view generation, this fixed RCP context is shared by all groups. TCR partitions v target views into m four-view groups, cyclically regroups them at high noise levels to propagate global structure, and fixes adjacent groups at low noise levels for stable detail refinement.Three-stage inference diagram. Initial views become a fixed multiscale reference context; target views are cyclically regrouped at high noise to share global structure and arranged into fixed adjacent groups at low noise to refine local details.
Table 3. Ablation study. Following the Gen. Video Consistency setting, we evaluate the consistency among generated videos.
Configuration
PSNR↑
SSIM↑
LPIPS↓
w/o TCR & RCP
21.09
0.766
0.216
w/o RCP
22.03
0.780
0.203
w/o TCR
22.21
0.788
0.196
Full (Random)
22.20
0.788
0.197
Full (Strided)
22.06
0.786
0.198
Full (Sliding)
22.63
0.796
0.191
Figure 4. Qualitative comparison with baselines. We show target-view generated videos (Gen.) and their corresponding 4DGS renderings (Rend.) across diverse human-centric videos. 4DAnyone produces geometrically accurate and visually detailed results across viewpoints, while the baselines suffer from inaccurate camera control (our fine-tuned ReCamMaster†) or geometric distortions (MV-Performer). See the project page for dynamic results.A grid compares generated target views and reconstructed 4D renderings from several methods on diverse human videos. The proposed method preserves body geometry, clothing appearance, and camera viewpoint more consistently than ReCamMaster and MV-Performer.
Table 4. Training data sampling. Source cameras are uniformly sampled from the listed options.
Dataset
No. Tgt Cam
No. Src Cam
No. Frame
Weight
MVGameHuman
6
1 / 4 / 8
41
0.4
MVGameHuman
4
1 / 4 / 8
61
0.4
MVGameHuman
1
1 / 4 / 8
121
0.2
SynCamVideo
4
1 / 4
61
0.8
SynCamVideo
1
1 / 4 / 8
81
0.2
DNA-Rendering
6
1 / 4 / 8
41
0.4
DNA-Rendering
4
1 / 4 / 8
61
0.4
DNA-Rendering
1
1 / 4 / 8
121
0.2
Pexels
1
1
121
1.0
TedTalk
1
1
121
1.0
Figure 5. Qualitative ablation results. Left: 3D-aware skeleton conditioning resolves the inherent ambiguity of 2D skeletons, guiding the model to generate geometrically correct content. Right: RCP and TCR maintain an effective cross-view context, leading to consistent appearance across generated views.Side-by-side ablation grids compare outputs with and without 3D-aware skeleton conditioning, Reference Context Packing, and Target Context Routing. The complete model has more accurate body geometry and more consistent appearance across views.
Table 5. Stage-specific training settings. “Indep. Src Prob” denotes the probability of independently sampling the source frame range. Skeleton: B=body, H=hands, F=feet, Fi=fingers.
Stage
Dataset
Bg. Removal
Indep. Src Prob
Skeleton
1
DNA-Rendering
✓
0.2
B+H+F+Fi
2
+ MVGameHuman, SynCam.
–
0.0
B+H+F+Fi
3
+ Pexels, TedTalk
–
0.0
B+H+F
Figure 6. Challenging cases. 4DAnyone robustly handles back-view source videos (left), complex subject appearance (right), and complex human motions.Two challenging examples show source frames and generated viewpoints for a person initially seen from the back and a person with complex clothing and appearance.
Table 6. Multi-GPU inference configurations. Configurations are shown for different camera setups.
No. Layers
No. Cam / Layer
No. GPUs
No. Tgt Cam / GPU
1
16
4
4
2
16
8
4
3
16
8
6
Figure 7. Robust generalization to diverse in-the-wild human videos. For each example, we show the source video (left), generated target-view videos (middle, 4 of 16 views), and a 4DGS novel-view rendering (right). See the project page for dynamic results.Sixteen in-the-wild examples are arranged in two columns. Each example shows a source human video frame, four generated target viewpoints, and a novel-view rendering of the reconstructed 4D Gaussian avatar.
Table 7. TCR switching-time sweep. Gen. Video Consistency when varying the number of sliding denoising steps.
ts/T
Sliding steps
PSNR↑
SSIM↑
LPIPS↓
1.00
0
22.2079
0.7880
0.1964
0.75
5
22.3030
0.7898
0.1947
0.50
10
22.4575
0.7925
0.1933
0.25
15
22.6093
0.7955
0.1912
0.20
16
22.6294
0.7963
0.1906
0.15
17
22.6023
0.7961
0.1908
0.10
18
22.6217
0.7967
0.1905
0.05
19
22.6367
0.7972
0.1903
0.00
20
22.6414
0.7971
0.1903
Figure 8. Representative MVGameHuman samples. Each row shows one frame from four evenly spaced cameras in a synchronized 24-camera sequence captured with our in-house game data engine. MVGameHuman provides diverse actors, clothing, motions, lighting conditions, virtual scenes, and backgrounds for training multi-view human video generation.Eight rows of synthetic human scenes, each viewed simultaneously from four evenly spaced virtual cameras. The samples vary in actor, clothing, pose, lighting, environment, and background.
왜 중요한가
이 방법은 비싼 다중 카메라 스튜디오 장비 없이도 스마트폰 영상 한 편만으로 자유 시점 3D 인간 아바타를 만들 수 있는 길을 보여준다. AR/VR 콘텐츠, 가상 제작, 디지털 휴먼을 다루는 사람들에게는 특수 장비 없이 생성 후 복원하는 실용적인 파이프라인을 제시한다는 점에서 의미가 있다.
Figure 9. Single image to 4D avatar. Given a single input image, we first generate a source video via pose-driven video generation (Wan-Animate), then produce multi-view target videos with 4DAnyone, and finally reconstruct a 4DGS avatar via FreeTimeGS.A left-to-right pipeline turns one portrait into a pose-driven source video, generates synchronized target-view videos with 4DAnyone, and reconstructs an animatable 4D Gaussian avatar.
이 논문의 용어
4D Gaussian Splatting(4DGS) · 움직이는 3D 장면을 작은 가우시안 덩어리들의 집합으로 표현해 실시간으로 어느 시점, 어느 순간에서든 렌더링할 수 있게 하는 기법
DiT(Diffusion Transformer) · 노이즈를 점진적으로 제거해 이미지나 영상을 생성하는 확산 모델을, 기존 U-Net 대신 트랜스포머 구조로 구현한 것
주의 문맥(attention context) · 모델이 한 번의 연산에서 서로 비교할 수 있는 정보 토큰들의 범위. 넓을수록 더 많은 시점을 한꺼번에 맞춰볼 수 있지만 메모리와 연산량이 늘어난다
HMR(Human Mesh Recovery) · 평범한 영상에서 사람의 3D 몸 형태와 자세(뼈대·메시)를 추정하는 AI 기술
z-버퍼 렌더링 · 카메라에 더 가까운 표면이 먼 표면을 올바르게 가리도록 각 화소별 거리 정보를 기록해 렌더링하는 기법
Figure 10. Failure cases. Left: skeleton guidance is uninformative for the large flowing fabric, which is generated inconsistently across views and yields a degraded reconstruction (red box). Right: HMR mis-estimates the en-pointe pose (green circle in the source view) as flat feet, and all generated views inherit the pose error (red box).Two failure examples. Red boxes highlight inconsistent large flowing fabric in one reconstruction and incorrect feet in generated views caused by a source-pose estimation error marked with a green circle.
논문 원문 초록 (영문)
We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existing camera-controlled video diffusion models synthesize plausible novel-view videos but fail to maintain consistency when scaled to the tens of target views required for 4DGS reconstruction. We identify this failure as a bounded-attention-context problem: when target views exceed the capacity of a single DiT forward pass, they must be split into groups, exposing two coupled bottlenecks. On the reference-context side, conditioning on all previously generated views grows as O(N), weakening cross-view appearance guidance. On the target-context side, disjoint groups cannot directly exchange information, causing global structural drift. 4DAnyone addresses both bottlenecks with two complementary designs: Reference Context Packing (RCP) compresses growing reference views into a fixed-length mixed-resolution context with O(1) reference-context complexity, while Target Context Routing (TCR) rotates target-view groupings during denoising to share context across groups at high-noise steps and stabilize details at low-noise steps. We further build the MVGameHuman dataset using our in-house game engine and combine it with light-stage and in-the-wild video datasets for training. Experiments on DNA-Rendering and DyMVHumans show that 4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization. See our project page for video results and source code: https://4danyone.github.io.