AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox

METAL LAB

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

arXiv:2608.027112026-08-02

Tencent's Hunyuan team merges 3D understanding, generation, and editing into one model called Buffalo 1.0

Hunyuan3D-Buffalo 1.0 handles understanding 3D objects, generating them from text, editing them via instructions, and generating individual parts, all within a single model. To make this possible, the team built an 87-million-sample 3D multimodal training corpus, tackling the scarcity of 3D editing data with a new automated pipeline called Nano3D-v2. Experiments show the model outperforms prior methods on text-to-3D generation and editing benchmarks, and reveal that stronger generation ability also boosts editing ability.

METAL LAB explanatory visual

How Hunyuan3D-Buffalo 1.0 is structured

Evidence statusMeasured results reported

  1. Data engineAutomated pipelines build 25M understanding, 50M text-to-3D, and 12M editing samples, totaling 87M
  2. Nano3D-v2Five stages: anchor-view selection, edit-region prediction, voxel editing, geometry/texture refinement, VLM verification, used to auto-generate editing training pairs
  3. Hunyuan3D-VLMReads 3D point clouds to understand an object's meaning, structure, and part locations, producing conditions for generation
  4. Hunyuan3D DiTA diffusion model that takes the VLM's conditions plus the source object representation to generate or edit the 3D shape
  5. Benchmark validationOn Edit3D-Bench, achieves 86.7% lower Chamfer Distance and 2.39x higher F1 than Steer3D
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.

What they did

  1. The team assembled an 87M-sample 3D multimodal corpus made of 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs, all built through fully automated pipelines.
  2. A 3D vision-language model, Hunyuan3D-VLM, understands an object's semantics, structure, and spatial layout and produces conditioning signals for a diffusion model called Hunyuan3D DiT, which actually generates the 3D shape.
  3. To solve the scarcity of editing data, the team built Nano3D-v2, an agent-based pipeline with five stages: anchor-view selection, a learned model that localizes the 3D edit region at voxel level, voxel editing, fine-grained geometry/texture refinement, and VLM-based verification.
  4. On the Edit3D-Bench benchmark, compared to the previous strongest baseline Steer3D, the model reduced Chamfer Distance (a shape-error metric) from 0.0684 to 0.0091 (an 86.7% relative reduction) and raised average F1 from 0.2729 to 0.6515 (a 2.39x improvement).
  5. Adding just 1,000 extra text-to-3D training samples of chicken heads (with no new editing data at all) gave the model the new ability to edit that body part, showing that scaling up generation data also strengthens editing.
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing figure 0
Table 1: Overview of the full training corpus across 3D understanding, text-to-3D generation, and instruction-guided 3D editing.
CapabilitySubset#Samples
3D understandingText / image / 3D instruction data∼25M
Text-to-3DText–asset pairs∼50M
3D editingHuman edits∼7M
Object edits∼3M
Part generation∼2M
Subtotal∼12M
Figure 1: Hunyuan3D-Buffalo 1.0 is an unified 3D multimodal framework that combines autoregressive modeling with diffusion-based generation, enabling 3D understanding, text-to-3D generation, 3D editing, and text-grounded part generation within a single architecture.
Figure 1: Hunyuan3D-Buffalo 1.0 is an unified 3D multimodal framework that combines autoregressive modeling with diffusion-based generation, enabling 3D understanding, text-to-3D generation, 3D editing, and text-grounded part generation within a single architecture.
Table 2: Six caption tiers produced per asset.
TierLengthEmphasis
Detailed4–6 sentences (≤120 w)subject, parts, pose, features
Main24–30 tokensstructure + key parts
Simplified15–20 tokensmain parts and fit
Paraphrase15–20 tokensreworded Simplified
Short6–10 tokenssubject + ≤1 attribute
Tags≤8 keywordsdisentangled keywords
Figure 2: Pipeline of constructing text-to-3D training corpus.
Figure 2: Pipeline of constructing text-to-3D training corpus.
Table 3: Comparison of part-level question answering and object-level captioning on UniPart-Bench [103]. The two tasks assess complementary aspects of 3D understanding, including localized part-aware reasoning and holistic object-level semantic description.
ModelPart Understanding Q&AOverall 3D Object Captioning
SBERTSimCSEBLEU-1ROUGE-LMETEORSBERTSimCSEBLEU-1ROUGE-LMETEOR
GPT4Point [63]48.3245.1715.1622.5516.1925.6027.0011.5012.0012.70
PointLLM-7B [93]61.3058.4821.7829.2622.4542.7942.4411.5814.3916.90
PointLLM-13B [93]56.3651.4721.4029.1621.8043.5143.1213.5415.7417.45
ShapeLLM-13B [62]61.1957.2623.3232.5624.4525.1527.1411.7712.1412.84
ShapeLLM-Omni-7B [103]57.3551.1622.7729.5723.2431.1831.9317.7919.0414.30
Part-X-MLLM [78]78.9884.2540.5442.2634.2453.8251.9736.0438.1130.71
UniVerse3D [101]83.1187.1646.7943.9442.0565.1866.2542.7544.1741.11
Hunyuan3D-VLM (Ours)85.4789.0649.9545.0145.7972.9473.6050.9352.8450.47
Figure 3: Pipeline of constructing 3D editing training corpus (Nano3D-v2).
Figure 3: Pipeline of constructing 3D editing training corpus (Nano3D-v2).
Table 4: Detailed all-task evaluation of Hunyuan3D-VLM on UniPart-Bench [103]. The benchmark covers pure box listing, multi-part grounding, single-part grounding, box-to-text generation, and part-level question answering.
TaskNameIoUSBERTSimCSEBLEU-1ROUGE-LMETEOR
0Pure box listing0.864-----
1Multi-Part Grounding (Q1)0.88068.0068.5552.0652.0926.35
2Multi-Part Grounding (Q2)0.84470.9269.4740.0441.8638.19
3Single-Part Grounding (Q1)0.62678.9577.9245.7447.4744.07
4Single-Part Grounding (Q2)0.525-----
5Box-to-Text (Q1)-67.6468.2749.8950.0025.45
6Box-to-Text (Q2)-74.1372.9942.0144.2840.96
7Part QA0.63385.4789.0649.9545.0145.79
Figure 4: Examples of editing pairs in the training corpus created by Nano3D-v2.
Figure 4: Examples of editing pairs in the training corpus created by Nano3D-v2.
Table 5: Results of the user study for text-to-3D generation. We report the preference rate (%), i.e., the percentage of comparisons in which each model is selected as the best among the four candidates. In each comparison, the four results are shown side by side in randomized order, and participants pick the single best result for each criterion (a “tie” option is also allowed, so columns may sum to slightly below 100%). Higher is better, with a random-choice baseline of 25%. Best results are shown in bold.
ModelText alignmentGeometry qualityOverall preference
Universe3D [101]8.27.48.3
TRELLIS [92]14.912.414.4
Omni123 [100]17.521.018.4
Hunyuan3D-Buffalo 1.0 (Ours)55.257.156.6
Figure 5: Examples of multi-round editing by Nano3D-v2.
Figure 5: Examples of multi-round editing by Nano3D-v2.
Table 6: Ablation on the quantity of training data. We report the preference rate (%), i.e., the percentage of comparisons in which a model is chosen as the best among the three variants. In each comparison, the three results are shown side by side in randomized order, and participants select the best result for each criterion (a “tie” option is also allowed, so columns may sum to slightly below 100%). Higher is better, and the random-choice baseline is 33.3%. Best results are shown in bold.
Num. of samplesText alignmentGeometry qualityOverall preference
300w9.98.88.4
1500w28.829.028.6
5000w54.557.457.5
Figure 6: Hunyuan3D-Buffalo 1.0 pipeline. The framework unifies 3D QA and grounding, text-to-3D generation, and 3D editing through a shared Hunyuan3D-VLM backbone, which connects language, 3D representations, and generative Hunyuan3D DiT modules for multimodal understanding, generation, and editing.
Figure 6: Hunyuan3D-Buffalo 1.0 pipeline. The framework unifies 3D QA and grounding, text-to-3D generation, and 3D editing through a shared Hunyuan3D-VLM backbone, which connects language, 3D representations, and generative Hunyuan3D DiT modules for multimodal understanding, generation, and editing.
Table 7: Quantitative comparison of language-guided 3D shape editing on Edit3D-Bench [85]. Both CLIP-conditioned and 3D-VLM-conditioned variants of Hunyuan3D-Buffalo 1.0 are included to analyze the effect of stronger 3D instruction understanding.
MethodAddRemoveAvg
CD ↓F1 ↑CD ↓F1 ↑CD ↓F1 ↑
ShapeLLM-Omni [103]0.25460.08770.22370.11660.23920.1022
3DEditFormer [90]0.16760.19550.13420.18360.15090.1896
Tailor3D [64]0.16610.12170.17550.13520.17080.1285
Steer3D [85]0.14040.24140.09760.30440.11900.2729
Omni123 [100]0.07360.17430.06320.22590.06840.2001
Hunyuan3D-Buffalo 1.0 w/ CLIP (Ours)0.01540.56570.01620.70150.01580.6336
Hunyuan3D-Buffalo 1.0 w/ 3D-VLM (Ours)0.01270.56100.00540.74200.00910.6515
Figure 8: Qualitative text to 3D results.
Figure 8: Qualitative text to 3D results.

Findings

  • On Edit3D-Bench, Hunyuan3D-Buffalo 1.0 (3D-VLM conditioned) reduced average Chamfer Distance from 0.0684 to 0.0091 (86.7% relative reduction) and improved average F1 from 0.2729 to 0.6515 (2.39x) compared to the strongest prior baseline Steer3D.
  • On addition and removal editing tasks respectively, the model achieved CD/F1 of 0.0127/0.5610 and 0.0054/0.7420, showing accurate localized editing while preserving the rest of the geometry.
  • Switching the model's conditioning from CLIP embeddings to the 3D-VLM improved average CD from 0.0158 to 0.0091 and average F1 from 0.6336 to 0.6515, confirming that stronger 3D understanding improves editing.
  • Adding only 1,000 extra chicken-related text-to-3D training samples, with no new editing data, gave the model a new ability to edit that specific body part.
  • In qualitative comparisons, Hunyuan3D-Buffalo 1.0 showed a better balance between following edit instructions and preserving the original shape than baselines Omni123 and Steer3D.
Figure 9: Qualitative shape editing results. Our method significantly outperforms all baselines in both geometric consistency before and after editing, and responsiveness to editing instructions.
Figure 9: Qualitative shape editing results. Our method significantly outperforms all baselines in both geometric consistency before and after editing, and responsiveness to editing instructions.

Where it can be used

  • Tools for game and animation asset creation that let artists modify 3D characters or props via text instructions (e.g., adding glasses, removing wings)
  • Pipelines that decompose and reassemble 3D objects part-by-part to build reusable 3D asset libraries
  • 3D search and annotation tools that answer questions about or localize (ground) parts of a 3D asset
Figure 10: Scaling up text-to-3d data facilitates 3D editing. Model A is our base model; when instructed to edit an object by replacing its head with a chicken head, it fails to produce a satisfactory result. Model B is built upon Model A by adding only 1,000 additional chicken samples for the text-to-3D task during the Omni pre-training stage—crucially, without introducing any new editing data. After incorporating this text-to-3D data, the model can successfully replace the head with a chicken head. This suggests a clear direction: to improve 3D editing, the text-to-3D generation capability should be maximized as much as possible. Since constructing text-to-3D data is far less costly than constructing 3D editing data, scaling up text-to-3D data is a relatively more feasible path toward stronger 3D editing.
Figure 10: Scaling up text-to-3d data facilitates 3D editing. Model A is our base model; when instructed to edit an object by replacing its head with a chicken head, it fails to produce a satisfactory result. Model B is built upon Model A by adding only 1,000 additional chicken samples for the text-to-3D task during the Omni pre-training stage—crucially, without introducing any new editing data. After incorporating this text-to-3D data, the model can successfully replace the head with a chicken head. This suggests a clear direction: to improve 3D editing, the text-to-3D generation capability should be maximized as much as possible. Since constructing text-to-3D data is far less costly than constructing 3D editing data, scaling up text-to-3D data is a relatively more feasible path toward stronger 3D editing.

Limits and open work

  • The model currently focuses on geometry editing and does not yet handle texture editing, so fully complete edits involving color or material remain future work.
  • The editing data construction pipeline cannot guarantee consistency for non-edited regions inside the editing mask, which can introduce noise into training.
  • Text-to-3D captions still rely on multimodal language models like Gemini, which produce ambiguous descriptions and add noise to the training pairs.
  • The multi-stage diffusion pipeline built on TRELLIS-style architectures requires multi-stage editing to reach high quality, which fundamentally limits scalability.
  • The authors state that both the volume and quality of the 3D data have not yet reached an ideal scale, indicating further scaling is still needed.
Figure 11: Qualitative part generation results.
Figure 11: Qualitative part generation results.

Why it matters

3D understanding, generation, and editing models have traditionally been built as separate systems, but this work shows concrete evidence that combining them in one model lets the tasks reinforce each other. For anyone building tools for game assets, animation, or product design, this points to a way to unify understanding, creation, and editing without maintaining separate pipelines.

Figure 12: Qualitative shape editing results.
Figure 12: Qualitative shape editing results.

Terms in this paper

  • Nano3D-v2 · An agent-based pipeline that automatically generates edited 3D objects from a source object plus an editing instruction, used to build training data
  • Hunyuan3D-VLM · A 3D-specific vision-language model that reads point-cloud data and understands an object's meaning, structure, and part locations
  • DiT (Diffusion Transformer) · A diffusion-based generative model that starts from noise and gradually produces a target output, here a 3D shape
  • Chamfer Distance · A metric measuring average surface distance between two 3D shapes; lower means the edited result is geometrically closer to the intended target
  • Voxel · A 3D grid unit, the 3D equivalent of a 2D pixel, used to represent and edit shapes locally

Original abstract (English)

Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part gene

Authors · Junliang Ye

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB

Figures: Junliang Ye et al., arXiv:2608.02711, arxiv-nonexclusive