Source-linked AI summary
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Xin Huang, Zhuo Chen, Chunchao Guo
TL;DR
통합 3D 모델링은 희소하고 기하학적으로 일관된 멀티모달 데이터의 부족으로 제약되며, 특히 편집에서 그 제약이 크다. Hunyuan3D-Buffalo 1.0은 이해, 생성, 편집, 파트 생성을 아우르는 통합 프레임워크와 대규모 멀티모달 코퍼스로 이를 해결하고, 생성 및 편집 벤치마크 전반에서 최첨단 또는 선도적 성능을 달성하면서 강력한 이해 및 파트 생성 능력을 보인다.
문제
통합 3D 모델링에는 대규모의 기하학적으로 일관된 멀티모달 데이터가 부족해, 이해·생성·편집이 서로 다른 시스템에 걸쳐 대체로 분절되어 있다.
방법
Hunyuan3D-Buffalo 1.0은 autoregressive 이해와 diffusion 기반 3D 합성을 결합하고, 이해·text-to-3D·편집·파트 생성을 아우르는 87M개 샘플 코퍼스를 사용한다.
결과
Hunyuan3D-Buffalo 1.0은 text-to-3D 생성 및 3D 편집 벤치마크에서 최첨단 또는 선도적 성능을 달성하며, 생성과 이해가 편집 성능을 향상시킨다.
시사점 및 한계
통합 3D 멀티모달 학습은 과제 간 능력 전이를 지원하며, 특히 생성과 이해에서 편집으로의 전이를 가능하게 한다.
시사점 및 한계
Nano3D-v2 편집 데이터 파이프라인은 편집 마스크 내부에 불일치를 유발해 편집 품질을 저하시킬 수 있다.
Abstract
from arXiv · showhide
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
1 서론
Hunyuan3D-Buffalo 1.0은 이해, 생성, 편집, 파트 생성을 지원하는 통합 프레임워크로 대규모의 기하학적으로 일관된 3D 편집 데이터 부족 문제를 해결한다. 대규모 multimodal corpus와 Hunyuan3D-VLM을 diffusion 기반 합성과 결합해 benchmark에서 선도적 성능을 달성하고 task 간 상호 시너지 효과를 보여준다.
- 1 서론: 대규모의 기하학적으로 일관된 3D 편집 데이터는 3D asset을 수집·주석화하고 identity, 구조, 편집되지 않은 영역을 보존하면서 편집하기 어렵기 때문에 여전히 부족하다.이 병목은 3D understanding, generation, editing 모델의 발전을 제한한다 [63] [64] [93] [40] [47] [92] [109] [12] [43] [104].
- 1 서론: 87M-sample 3D multimodal corpus는 25M understanding samples, 50M text-to-3D pairs, 12M editing pairs로 구성된다.Nano3D-v2는 agent 기반 데이터 구축을 통해 고품질의 기하학적으로 일관된 editing pairs를 대규모로 생성한다.
- 1 서론: Hunyuan3D-VLM은 3D grounding, part-level reasoning, edit-aware understanding을 위해 세밀한 semantic, structural, spatial understanding을 제공한다.captioning, part-level question answering, edit-instruction synthesis, edit-outcome reasoning을 지원하도록 기하학적 구조와 외관 단서를 인코딩한다.
- 1 서론: 이 통합 프레임워크는 autoregressive modeling과 diffusion 기반 3D generation을 결합해 하나의 architecture에서 understanding, text-to-3D generation, 3D editing, text-grounded part generation을 지원한다.Hunyuan3D-VLM은 통합 conditional interface를 통해 synthesis와 editing을 유도하는 multimodal reasoning을 제공한다.
- 1 서론: Hunyuan3D-Buffalo 1.0은 text-to-3D generation 및 3D editing benchmark에서 state-of-the-art 또는 선도적 성능을 달성하며, 강력한 3D understanding 및 part-generation 능력을 보인다.통합 학습은 더 강력한 generation과 더 강력한 understanding이 3D editing을 향상시킨다는 점을 추가로 보여준다.
2 관련 연구
관련 연구는 optimization 기반 및 diffusion 기반 3D 생성, text-guided editing, 구조화된 part generation, unified multimodal modeling을 포괄한다. 이러한 패러다임은 서로 다른 설계를 통해 scalable synthesis, 3D-consistent modification, semantic decomposition, joint image-text understanding and generation을 다룬다.
- 3D Content Generation: 3D generation은 DreamFusion [62]로 대표되는 pretrained 2D diffusion prior의 optimization 기반 distillation에서 출발해, 3DShape2VecSet [108]과 TRELLIS [92]를 포함하여 더 높은 효율성과 multi-view consistency를 지향하는 방법으로 발전했다.초기 방법은 visual prior를 3D representation으로 distill하여 large-scale 3D supervision 없이 text-to-3D generation을 가능하게 했다.
- Text-Guided 3D Editing: Text-guided 3D editing은 natural-language instruction에 따라 기존 asset을 수정하면서 unedited region과 3D consistency를 보존하며, instance별 SDS 또는 2D diffusion optimization을 사용하거나 edited-view fusion 또는 reconstruction을 사용한다.이 부분은 optimization 기반 방법과 rendered 2D view를 먼저 편집한 뒤 edited 3D asset을 생성하는 pipeline을 모두 제시한다.
- 3D Part Generation: 3D part generation은 semantic component마다 distinct mesh를 갖는 구조화된 asset을 추구하며, fixed taxonomy [60] [95]와 view-inconsistent 2D-segmentation pipeline에서 coordinated multi-diffusion branch [35] [48] [49] [58] [97] [98] [117]로 발전했다.최근 3D-native 방법은 coordinated diffusion path를 통해 part를 합성하며, 이 부분에서는 part segmentation을 위한 continuous feature field도 소개한다.
- Unified Multimodal Modeling: Unified image understanding and generation은 Chameleon [67] [74] [82] [89]과 같은 token-based autoregressive design 및 MetaQuery [9] [10] [61] [75] [87] [88] [112]로 대표되는 보다 decoupled된 multimodal architecture를 따른다.Chameleon은 이미지를 discrete VQVAE visual token으로 변환하여 joint text-image next-token modeling을 수행하며, MetaQuery는 별도의 decoupled 방향을 나타낸다.
3 데이터 큐레이션
저자들은 부족한 기하학적으로 일관된 멀티모달 학습 데이터를 보완하기 위해 상호 보완적인 이해, text-to-3D, editing corpus를 생성하는 통합 3D 데이터 엔진을 구축한다. 자동화된 pipeline은 semantic coverage, geometric fidelity, 그리고 편집 및 뷰 간 일관성을 유지하면서 이러한 dataset을 생성한다.
- Corpus 개요: 이 data engine은 3D understanding, text-to-3D generation, instruction-guided 3D editing을 위한 세 가지 상호 보완적 corpus를 구축한다.이 설계는 세 가지 capability를 모두 포괄하는 대규모·고품질·기하학적으로 일관된 data의 부족을 직접 겨냥한다.
- 3D understanding corpus: Understanding corpus는 captioning, question answering, grounding, editing 관련 task를 다루는 point-cloud dialogue와 text 및 image–text data를 결합한다.이 dialogue는 geometry, structure, part를 설명하고 spatial 및 attribute question에 답하며 region을 localize하고 editing operation을 그 기하학적 결과와 연결한다.
- Text-to-3D corpus: Text-to-3D corpus는 compositional prompt를 합성하고, asset을 생성·rendering하며, caption을 작성하고, geometry quality를 평가한 뒤 결과를 filtering하는 fully automated five-stage pipeline을 사용한다.Prompt는 hierarchical taxonomy를 통해 구성되며, asset은 canonical multi-view render와 sampled surface point cloud와 함께 저장된다.
- Editing corpus: Nano3D-v2는 natural-language edit를 실행하면서 non-target geometry, identity, multi-view consistency를 보존해 scalable하고 고품질인 editing pair를 생성한다.Quality control은 edit 전후 image를 비교하고, detected difference 바깥의 unchanged region을 검증하며, 비현실적으로 큰 edit mask를 거부한다.
- Part-generation data: Part-level instruction data는 specialized geometry module이나 bounding-box supervision에 의존하지 않고 language-guided part localization and decomposition으로 unified model의 범위를 확장한다.Semantic decomposition은 wheel segmenting이나 handle removing과 같은 query를 지원하며, raw mesh-level part에 포함된 제한적인 semantic information 문제를 다룬다.
4 방법
Hunyuan3D-Buffalo 1.0은 Hunyuan3D-VLM backbone과 Hunyuan3D DiT를 연결해 3D 이해, 생성, grounding, editing을 통합한다. 이 multimodal architecture는 구조·외관 3D representation, semantic conditioning, source-object conditioning을 결합해 구조적으로 일관된 editing과 part generation을 수행한다.
- Architecture: 통합 pipeline은 MLP-Connector를 통해 Hunyuan3D-VLM과 Hunyuan3D DiT를 연결하며, VLM은 multimodal reasoning을, DiT는 3D synthesis를 담당한다.3D-DiT는 Hunyuan3D-2.1에서 초기화되며, connector는 VLM hidden states를 DiT conditioning space에 정렬한다.
- 3D-aware Vision Language Model: Hunyuan3D-VLM은 geometric pathway와 RGB pathway를 통해 colored point cloud를 encode한 뒤, latent token을 512-token sequences로 압축해 효율적인 multimodal fusion을 수행한다.Geometric input에는 XYZ coordinates와 surface normals가 포함되며, RGB cue는 기하학적으로 유사하지만 시각적으로 다른 part를 구분하는 데 도움을 준다.
- 3D Editing and Part Generation: Editing과 part generation에서는 diffusion process가 VLM semantic embedding과 source object representation을 결합하며, source object representation은 DiT self-attention에서 noisy latent와 concatenate된다.이를 통해 denoiser가 generation 중 원래 geometry에 직접 접근할 수 있다.
- Training Pipeline: Training은 3D-VLM pre-training, text-to-3D pre-training, unified omni pre-training, task-specific continued pre-training의 four stages로 진행된다.마지막 stage는 editing, text-to-3D, part-generation path로 분기되며, generative stage에서는 flow matching을 사용해 target 3D latent를 향하는 velocity field를 예측한다.
- Training Pipeline: Omni pre-training에서는 text-to-3D와 editing plus part generation을 1:1 sampling ratio로 균형화하며, continued training에서는 task-specific data mixing을 통해 generation을 유지한다.Editing과 part-generation path는 text-to-3D data를 절반 유지하는 반면, text-to-3D path는 text-to-3D data만 사용한다.
5 실험
실험 결과, Hunyuan3D-Buffalo 1.0은 강력한 통합 3D 이해, text-to-3D 선호도, instruction-guided 편집, open-vocabulary 파트 생성을 제공한다. 3D-VLM conditioning과 확장 가능한 text-to-3D 학습은 의미 정확도, 기하학적 품질, 국소적 구조 보존 편집을 향상한다.
- 3D 이해: Hunyuan3D-VLM은 UniPart-Bench의 파트 수준 Q&A와 객체 캡셔닝에서 보고된 성능 중 최고를 달성하며, 파트 이해에서 85.47 SBERT, 89.06 SimCSE, 49.95 BLEU-1, 45.79 METEOR를 기록한다.이 벤치마크는 상호보완적인 국소 파트 추론과 전체 객체 수준의 설명 능력을 평가한다.
- 3D 이해: Hunyuan3D-VLM은 pure box listing에서 0.864 IoU를 얻으며, multi-part grounding, single-part grounding, box-to-text generation, part QA 전반에서도 강한 성능을 보인다.이 과제들은 localization, region-conditioned description, part-aware reasoning을 종합적으로 평가한다.
- Text-to-3D 생성: Hunyuan3D-Buffalo 1.0은 text alignment, geometry quality, overall preference에서 각각 55.2%, 57.1%, 56.6%의 선호도를 얻었으며, Omni123 [101]은 각각 17.5%, 21.0%, 18.4%를 기록했다.인간 평가는 100개의 다양한 text prompt와 four-way comparison group을 사용했으며, 보고된 모든 격차는 25%의 random-choice 수준을 초과한다.
- Instruction-Guided 3D 편집: Edit3D-Bench [86]에서 3D-VLM 버전은 Omni123 [101] 대비 평균 CD를 0.0684에서 0.0091로 낮춰 86.7% relative reduction을 달성하며, CD와 F1 모두에서 기존 방법을 능가한다.CLIP-conditioned variant와 비교하면 3D-VLM conditioning은 평균 CD를 0.0158에서 0.0091로 낮추고 평균 F1을 0.6336에서 0.6515로 높인다.
- Instruction-Guided 3D 편집: 정성적 결과는 입력 shape의 전체 구조, pose, 세밀한 디테일, 관련 없는 영역을 보존하면서 국소적으로 추가하고 제거하는 편집을 보여준다.이 방법은 편집되지 않은 geometry를 유지하면서 추가된 component를 올바른 semantic location에 배치한다.
- Text-Grounded 파트 생성: 이 방법은 다양한 객체에 대해 open-vocabulary, text-grounded 파트 생성을 지원하며, 입력 shape에 높은 충실도로 질의된 geometry를 추출한다.이 기능은 구조적으로 복잡한 객체와 기하학적으로 단순한 객체 모두에 적용되며, 여러 추출 파트를 결합할 수도 있다.
6 결론 및 향후 연구
통합 multimodal 3D model은 understanding, generation, editing을 공동으로 다루며 세 과제 모두에서 state-of-the-art 성능을 달성한다. 향후 연구는 representation, data quality, texture editing, editing-data robustness, architecture scaling을 대상으로 한다.
- 결론: 통합 model은 3D understanding, generation, editing에서 state-of-the-art 성능을 달성하며, generation과 editing에서 기존 방법보다 크게 향상된다.이는 하나의 framework 안에서 여러 과제를 처리하는 유망한 방향으로서 unified 3D modeling을 뒷받침한다.
- 향후 연구: TRELLIS [92]와 같은 접근법에서 요구되는 multi-stage pipeline 없이 고품질 geometry를 제공하는 scalable single-stage representation을 개발해야 한다.Multi-stage 설계는 고품질 unified editing을 확장하기 어렵게 만든다.
- 향후 연구: captioning quality를 개선하고 3D data를 volume과 quality 양方面에서 확장하는 것은 noisy text-to-3D training pair를 줄이고 성능을 향상하는 데 중요하다.Gemini와 같은 multimodal model에서 생성된 현재 caption은 여전히 모호하며, 3D data도 이상적인 scale에 도달하지 못했다.
- 향후 연구: End-to-end texture editing은 아직 탐구되지 않았으며, 적절한 data와 geometry 및 texture representation을 공동으로 model링하는 방법이 필요할 수 있다.고품질 geometry와 texture를 위한 unified single-stage representation은 여전히 미해결 문제다.
- 향후 연구: replacement mask 내부의 non-edited region이 일관성을 잃고 end-to-end editing quality를 저하시킬 수 있으므로 더욱 robust한 editing-data construction이 필요하다.Nano3D-v2 기반 pipeline은 mask 외부의 일관성을 보존하지만 mask 내부의 non-edited content를 처리하는 데 어려움을 겪는다.
- 향후 연구: Transfusion-style architecture를 탐구하면 현재의 cascaded AR + DiT framework를 넘어 modality 간 정보를 깊이 융합하여 3D generation을 향상할 수 있다.이러한 architecture는 image와 video generation에서 효과를 보였으며, 3D를 위한 유망한 다음 단계로 제시된다.
7 저자 목록
이 논문은 저자 16명을 밝히고, 3명의 프로젝트 리더와 여러 핵심 기여자를 식별하며, 주요 연구 영역별로 기여자를 나열한다.
- 저자: 저자 목록은 Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Xin Huang, Zhuo Chen, Chunchao Guo로 구성된다.Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang은 핵심 기여자로 표시된다.
- 저자 역할: Yang Li, Zhuo Chen, Chunchao Guo는 프로젝트 리더로 식별되며, 별표는 핵심 기여자를 나타낸다.Yang Li, Zhuo Chen, Chunchao Guo에는 프로젝트 리더 지정이 부여되어 있으며, 원문에서는 핵심 기여자를 별표로 표시한다.
- 연구 기여: 기여자는 3D editing, text-to-3D, 3D understanding에 걸쳐 배정되며, Yang Li는 세 영역 모두에 참여한다.3D-editing 기여자로는 Junliang Ye, Guocun Wang, Yansong Qu, Yang Li, Chunshi Wang, Kenkun Liu가 나열되고, text-to-3D에는 Kenkun Liu, Junliang Ye, Yang Li가, 3D understanding에는 Guocun Wang, Junliang Ye, Kenkun Liu, Yang Li가 나열된다.
A 추가 결과
이 절에서는 shape editing과 text-to-3D generation에 대한 정성적 결과를 제시한다. 그림들은 두 과제 전반에서 모델의 성능을 보여준다.
- Fig. 12에 정성적 shape editing 결과를 제시한다.
- 추가 결과에서는 shape editing과 text-to-3D generation을 모두 정성적으로 다룬다.
- Fig. 13에 정성적 text-to-3D 결과를 제시한다.