월~금 오전 7시, 일요일 오전 8시 — AI 뉴스와 용어를 보내드립니다메일로 받아보기

METAL LAB

AI 에이전트들이 팀을 이뤄 글 한 줄로 걸어다닐 수 있는 3D 오픈월드 전체를 만든다

arXiv:2608.052482026-08-07

WorldClaw: Agentic 3D Open-World Generation at Scale

AI 에이전트들이 팀을 이뤄 글 한 줄로 걸어다닐 수 있는 3D 오픈월드 전체를 만든다

WorldClaw는 텍스트 한 문장을 받아 지형, 건물, 나무, 배 같은 요소가 갖춰진 대규모 3D 세계를 만드는 에이전트 시스템이다. 계획 담당 AI가 장면을 지역·지형·물체로 나눈 설계도를 먼저 만들고, 그 다음 전체 지형을 짠 뒤, 세부가 필요한 구역만 골라 물체를 생성해 배치한다. 렌더링 결과를 직접 보고 스스로 검토·수정하는 에이전트들이 이 과정을 반복해 자연스러운 장면으로 다듬는다.

METAL LAB 해설 도표

WorldClaw의 3단계 전역-지역 생성 흐름

증거 상태측정 결과와 예정된 검증이 함께 있음

  1. 1. 의도 분석·계획사용자의 짧은 문장을 지역·지형·물체·재질·공간관계가 담긴 구조화된 장면 설계도로 변환
  2. 2. 전역 지형 생성시맨틱 레이아웃 맵과 지형 파라미터로 높이 필드를 만들고 재사용 소품·재질을 배치해 전체 지형 완성
  3. 3. 지역별 물체 생성·배치세부가 필요한 구역의 지형을 렌더링해 이미지 편집 AI로 물체를 그려 넣고, 이미지-투-3D로 개별 오브젝트를 복원해 지형에 배치
  4. 4. 렌더 기반 정제 루프렌더링 결과를 에이전트가 검사해 지형 이음새, 재질 스케일, 물체 자세·크기, 지면 접촉 문제를 반복 수정
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 사용자의 짧은 문장을 '의도 분석 에이전트'와 '장면 계획 에이전트'가 받아, 지역 구성·지형 종류·물체 종류·재질·공간 관계까지 담긴 구조화된 설계도로 바꾼다.
  2. 이 설계도를 바탕으로 색으로 구역을 표시한 2D 배치도를 만들고, 여기서 높낮이 지형(높이 필드)을 계산하며, 재사용 가능한 바위·식물 같은 3D 소품과 재질(질감)을 생성해 전체 지형을 완성한다.
  3. 세부 묘사가 필요한 구역만 골라 그 지형을 2D로 렌더링한 뒤 이미지 편집 AI로 건물·배·인물 같은 물체를 그려 넣고, 이를 다시 개별 3D 오브젝트로 잘라내 복원한 다음 지형 위 정확한 위치에 배치한다.
  4. 렌더링 기반 에이전트가 결과 화면을 직접 검사해 지형의 이음새, 재질 비율, 물체의 자세·크기·바닥과의 접촉 문제(공중에 뜨거나 파묻힌 경우 등)를 반복적으로 고쳐 나간다.
  5. 열대 해적 소굴, 협곡 부족 마을, 사막 전장, 눈 덮인 계곡 기지 등 서로 다른 프롬프트로 넓은 지역에 걸쳐 일관된 구조와 편집 가능한 개별 3D 자산을 갖춘 장면을 만들어냈다.
Figure 1: Overview of WorldClaw. Given an open-ended text prompt, WorldClaw constructs an explicit, explorable, and editable 3D world through a three-stage global-to-regional pipeline. (1) Intent Analysis and Planning translates the prompt into a structured scene specification. (2) Global Terrain Generation establishes a region-aware terrain foundation with coherent geometry, appearance, and spatial semantics. (3) Regional Object Generation and Placement populates planned regions with editable 3D assets and refines their arrangements and interactions with terrain.
Figure 1: Overview of WorldClaw. Given an open-ended text prompt, WorldClaw constructs an explicit, explorable, and editable 3D world through a three-stage global-to-regional pipeline. (1) Intent Analysis and Planning translates the prompt into a structured scene specification. (2) Global Terrain Generation establishes a region-aware terrain foundation with coherent geometry, appearance, and spatial semantics. (3) Regional Object Generation and Placement populates planned regions with editable 3D assets and refines their arrangements and interactions with terrain.
Figure 2: Overview of global terrain generation and refinement.(a) Initial Height-Field Generation constructs composite terrain geometry from the semantic layout map and region-specific terrain parameters and assigns materials to the corresponding regions.(b) Global Terrain Asset Scattering instantiates terrain-associated assets according to regional semantics and local surface conditions.(c) Terrain Refinement renders, inspects, and locally edits the terrain to correct geometric transitions, material scales, asset distributions, and rendering artifacts.
Figure 2: Overview of global terrain generation and refinement.(a) Initial Height-Field Generation constructs composite terrain geometry from the semantic layout map and region-specific terrain parameters and assigns materials to the corresponding regions.(b) Global Terrain Asset Scattering instantiates terrain-associated assets according to regional semantics and local surface conditions.(c) Terrain Refinement renders, inspects, and locally edits the terrain to correct geometric transitions, material scales, asset distributions, and rendering artifacts.

실제로 확인된 결과

  • 열대 해적 소굴, 협곡 부족 마을, 사막 전장, 눈 덮인 계곡 기지, 중세 마을, 설원 강변 마을, 사막 캠프, 일본풍 섬마을 등 다양한 프롬프트에서 전역적으로 일관된 공간 구성과 지역별 세부 콘텐츠를 갖춘 대규모 장면을 실제로 생성했다.
  • 기존 대표적인 텍스트 기반 3D 장면 생성 방법들과 같은 중세 마을 주제로 정성적 비교를 수행했다(수치 지표는 보고되지 않음).
  • 현재 오픈소스 언어모델은 실행 가능하면서 사용자 요구에 맞는 절차적 지형·재질 코드 생성에 자주 실패했고, 오픈소스 이미지 생성 모델도 쓸 만한 레이아웃 맵이나 물체 외형·자세 보존에 자주 실패했다고 저자들이 관찰했다.
  • 현재 파이프라인의 검증은 Claude Opus 4.8, GPT-Image-2, Hunyuan3D 같은 고성능 모델을 사용해야만 충분히 이루어진다고 저자들이 밝혔다.
Figure 3: Scene refinement. (a) Object Refinement processes objects from a report queue, evaluates pose, mesh quality, and scale against semantic and regional context, applies targeted edits, and verifies the result through re-rendering. (b) Terrain Refinement examines support-surface quality and object-terrain collisions, applies local co-deformation to defects such as floating, and updates the report after re-rendering.
Figure 3: Scene refinement. (a) Object Refinement processes objects from a report queue, evaluates pose, mesh quality, and scale against semantic and regional context, applies targeted edits, and verifies the result through re-rendering. (b) Terrain Refinement examines support-surface quality and object-terrain collisions, applies local co-deformation to defects such as floating, and updates the report after re-rendering.
Figure 4: Tropical pirate stronghold. An island terrain organizes dense vegetation, settlements, docks, and ships into distinct coastal regions. The figure presents the global composition, regional close-ups, local walk views, and their corresponding instance, depth, and normal renderings.
Figure 4: Tropical pirate stronghold. An island terrain organizes dense vegetation, settlements, docks, and ships into distinct coastal regions. The figure presents the global composition, regional close-ups, local walk views, and their corresponding instance, depth, and normal renderings.

어디에 쓸 수 있나

  • 게임 월드나 VR 환경의 초기 대규모 지형·건물 배치 초안을 텍스트만으로 빠르게 만드는 작업
  • 영화·애니메이션 프리비주얼용 대규모 야외 배경을 편집 가능한 3D 자산 형태로 준비하는 작업
  • 로봇 시뮬레이션이나 임베디드 AI 훈련을 위해 다양한 지형·물체 배치를 가진 3D 환경을 다수 생성하는 작업
  • 게임 엔진(예: Unreal Engine) 워크플로로 옮겨 추가 수정·애니메이션 작업을 이어갈 수 있는 텍스처드 메시 소스 제작
Figure 5: Canyon with tribal settlements. A continuous river connects the canyon, valley floor, vegetation, and settlement regions across substantial elevation changes. Global and regional views are complemented by local walk views and the corresponding instance, depth, and normal renderings.
Figure 5: Canyon with tribal settlements. A continuous river connects the canyon, valley floor, vegetation, and settlement regions across substantial elevation changes. Global and regional views are complemented by local walk views and the corresponding instance, depth, and normal renderings.
Figure 6: Thrilling battlefield in the desert. Layered rocky landforms surround open combat areas and populated compounds containing buildings, defensive structures, and vehicles. The figure shows the global and regional organization together with local walk views and their instance, depth, and normal renderings.
Figure 6: Thrilling battlefield in the desert. Layered rocky landforms surround open combat areas and populated compounds containing buildings, defensive structures, and vehicles. The figure shows the global and regional organization together with local walk views and their instance, depth, and normal renderings.

한계와 남은 검증

  • 파이프라인 각 단계가 대형 언어모델, 이미지 생성 모델, 3D 생성 모델 등 여러 외부 모델의 성능에 크게 의존하며, 저품질 백본을 쓰면 결과물의 시각적 품질이 그대로 떨어진다.
  • LLM이 생성한 Blender 코드가 척도 추정, 수치 파라미터, 노드 연결에서 오류를 자주 일으켜 지형 왜곡이나 부정확한 재질, 의도와 다른 배치로 이어지고, 이를 고치려면 여러 차례의 렌더링-검토-수정 반복이 필요하다.
  • 물체를 개별 생성·재구성하고 여러 차례 정제 루프를 도는 방식이라 추론 시간과 계산 비용이 크며, 이는 물체 수와 반복 횟수, 장면 규모가 커질수록 더 심해진다.
  • 단순한 장면에는 이 긴 파이프라인이 통짜 생성 방식보다 오히려 비효율적일 수 있다.
  • 현재는 생성형 3D 모델로 개별 오브젝트를 복원하기 때문에 명확한 부품 구조, 매개변수화된 구조, 관절(애니메이션) 정의, 상호작용 로직까지는 안정적으로 복원하지 못하며, 이는 향후 과제로 남아 있다.
Figure 7: Snow-covered mountain valley with style of Command & Conquer: Red Alert. The enclosing mountain terrain contains multiple regions populated with futuristic facilities, communication structures, and vehicles. Global, regional, and local walk views are shown with the corresponding instance, depth, and normal renderings.
Figure 7: Snow-covered mountain valley with style of Command & Conquer: Red Alert. The enclosing mountain terrain contains multiple regions populated with futuristic facilities, communication structures, and vehicles. Global, regional, and local walk views are shown with the corresponding instance, depth, and normal renderings.
Figure 8: Qualitative comparison with the representative text-driven 3D scene generation methods. All methods are conditioned on prompts that share the same medieval-village theme and comparable scene requirements, with the wording adapted when necessary to each method’s input format.
Figure 8: Qualitative comparison with the representative text-driven 3D scene generation methods. All methods are conditioned on prompts that share the same medieval-village theme and comparable scene requirements, with the wording adapted when necessary to each method’s input format.

왜 중요한가

게임, 영화, 가상현실, 로봇 시뮬레이션 모두 걸어 다닐 수 있고 부품 단위로 수정 가능한 3D 세계가 필요한데, 지금까지는 전체를 일관되게 유지하면서도 세부를 풍부하게 채우는 두 가지를 동시에 만족하기 어려웠다. WorldClaw는 이를 '먼저 전체 틀을 잡고, 필요한 곳만 세부를 채우는' 방식으로 풀어보려는 시도라서, 텍스트만으로 실사용 가능한 3D 콘텐츠를 만드는 작업 흐름에 참고가 될 수 있다.

Figure 9: Medieval-style village across diverse terrain. The scene combines snow-capped mountains, forested plains, waterways, and desert areas, with village buildings, windmills, vegetation, and animals distributed across the different regions. The figure shows the global layout, regional close-ups, local walk views, and their corresponding instance, depth, and normal renderings.
Figure 9: Medieval-style village across diverse terrain. The scene combines snow-capped mountains, forested plains, waterways, and desert areas, with village buildings, windmills, vegetation, and animals distributed across the different regions. The figure shows the global layout, regional close-ups, local walk views, and their corresponding instance, depth, and normal renderings.
Figure 10: Snow-covered riverside village. The village extends along both sides of a frozen river within a mountainous landscape. The figure shows the global layout, regional close-ups, and local walk views together with instance, depth, and normal renderings.
Figure 10: Snow-covered riverside village. The village extends along both sides of a frozen river within a mountainous landscape. The figure shows the global layout, regional close-ups, and local walk views together with instance, depth, and normal renderings.

이 논문의 용어

  • 에이전트(Agentic) · 스스로 계획을 세우고 도구를 실행하며 결과를 검토해 수정까지 하는 AI 프로그램
  • 높이 필드(Height Field) · 지형의 높낮이를 2D 격자 위 숫자로 표현한 것
  • 시맨틱 레이아웃 맵 · 색깔로 각 구역이 어떤 지형·용도인지 표시한 2D 지도
  • 이미지-투-3D · 2D 이미지를 입력받아 그 물체의 3D 모델(형태·질감)을 만들어내는 AI 기술
  • 텍스처드 메시(Textured Mesh) · 표면에 색과 질감 이미지가 입혀진, 편집 가능한 3D 형태 데이터
Figure 11: Desert adventure camp surrounded by dragons. The scene combines layered desert terrain, settlements, watchtowers, and dragons distributed around the camp. Global, regional, and local walk views are shown with the corresponding instance, depth, and normal renderings.
Figure 11: Desert adventure camp surrounded by dragons. The scene combines layered desert terrain, settlements, watchtowers, and dragons distributed around the camp. Global, regional, and local walk views are shown with the corresponding instance, depth, and normal renderings.
Figure 12: Island with Japanese-style towns. Multiple settlements are distributed across an island containing coastlines, vegetation, hills, and water. The figure presents the global and regional organization as well as local walk views and their instance, depth, and normal renderings.
Figure 12: Island with Japanese-style towns. Multiple settlements are distributed across an island containing coastlines, vegetation, hills, and water. The figure presents the global and regional organization as well as local walk views and their instance, depth, and normal renderings.

저자 · Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Chunchao Guo et al., arXiv:2608.05248, arxiv-nonexclusive