Projector Is All You Train
3D 인식 능력을 갖춘 AI를 만들 때, 거대 언어모델 본체는 건드릴 필요가 없다
멀티모달 AI(이미지·소리·3D 같은 다른 형태의 정보를 이해하는 AI)를 새로운 감각, 여기서는 3D 정보에 적응시킬 때 보통은 언어모델 본체와 그 사이를 잇는 작은 변환기(프로젝터)를 함께 학습시킨다. 이 연구는 프로젝터만 학습시켜도 본체까지 같이 학습시킨 것과 비슷하거나 더 나은 3D 이해 성능을 얻을 수 있음을 보였다. 게다가 프로젝터만 학습하면 속도가 약 2배 빠르고, 언어모델이 원래 갖고 있던 능력이 망가지는 부작용도 아예 생기지 않는다.
무엇을 했나
- 3D 점군(물체 표면을 점들의 집합으로 표현한 데이터) 데이터를 다루는 AI 모델(PointLLM 계열을 기준으로 삼음)을 대상으로, 언어모델 본체는 얼려두고 프로젝터라는 작은 연결 장치만 학습시키는 방식과, 본체 일부(LoRA 어댑터)와 프로젝터를 같이 학습시키는 방식을 비교했다
- Qwen3.5-4B, Qwen3.5-9B, Llama-3.1-8B-Instruct 세 종류의 언어모델 본체에 대해 동일한 A100 GPU 1대로 16시간씩 학습시켜 공정하게 비교했다
- ModelNet40, Objaverse, OmniObject3D 세 데이터셋의 3D 물체 분류 및 설명문 생성 과제에서, 프로젝터만 학습시킨 모델이 기존 PointLLM 모델보다 대체로 높은 점수를 받았고, 같은 조건에서 본체까지 학습시킨 모델과 비교해도 경쟁력 있는 성능을 보였다
- 프로젝터만 학습시키면 학습 속도(같은 시간에 처리하는 데이터 양)가 본체까지 학습시킬 때보다 약 2배 빨랐다
- 본체까지 같이 학습시키면 언어·시각·공간추론 관련 기존 벤치마크 성능이 떨어지는 부작용(원래 능력이 흐트러지는 현상)이 나타났지만, 프로젝터만 학습시킨 경우엔 본체를 아예 건드리지 않으므로 이런 부작용이 원천적으로 없었다

| ID | LM Backbone | Projector Trained | LoRA Fine-tuned |
|---|---|---|---|
| P-Llama8B | Llama-3.1-8B-Instruct | Yes | No |
| J-Llama8B | Llama-3.1-8B-Instruct | Yes | Yes |
| P-Qwen4B | Qwen3.5-4B | Yes | No |
| J-Qwen4B | Qwen3.5-4B | Yes | Yes |
| P-Qwen9B | Qwen3.5-9B | Yes | No |
| J-Qwen9B | Qwen3.5-9B | Yes | Yes |
![Figure 2: Evaluation results of projector-only-trained MLLMs against baseline PointLLM models. Chart (a) shows generative 3D object classification results on ModelNet40 (M40.), Objaverse (Obj.), and OmniObject3D (Omni.) objects under a zero-shot setting. There are two prompt types per benchmark: an instruction-style prompt (I, “What is this?”) and a completion-style prompt (C, “This is an object of”). Each entry reports accuracy judged by GPT-5.6 Luna [28]. Chart (b) shows the aggregate precision score on 3D object captioning tasks under the three judge LLMs. More details are found in Appendix B.](https://media.metallab.ai/papers/2608.19726/f1.png)
| Model | M40. (I) | M40. (C) | Obj. (I) | Obj. (C) | Omni. (I) | Omni. (C) |
|---|---|---|---|---|---|---|
| PointLLM-7B [43] | 54.94 | 55.55 | 59.83 | 60.43 | 34.78 | 35.70 |
| PointLLM-13B [43] | 57.66 | 56.81 | 61.73 | 60.90 | 34.63 | 34.01 |
| P-Llama8B | 61.95 | 63.57 | 65.60 | 62.33 | 49.19 | 40.69 |
| P-Qwen4B | 59.56 | 59.56 | 60.90 | 60.83 | 41.07 | 40.25 |
| P-Qwen9B | 59.28 | 62.80 | 74.07 | 67.77 | 48.33 | 38.76 |
| Mean | 58.68 | 59.66 | 64.43 | 62.45 | 41.60 | 37.88 |

| GPT-5.6 Luna | Claude Haiku 4.5 | Gemini 3.5 Flash-Lite | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | C | H ↓ | P | C | H ↓ | P | C | H ↓ | P | |
| PointLLM-7B [43] | 4.26 | 2.12 | 66.75 | 3.37 | 1.95 | 63.30 | 2.75 | 1.99 | 58.08 | |
| PointLLM-13B [43] | 4.35 | 2.10 | 67.43 | 3.47 | 1.91 | 64.46 | 2.82 | 1.92 | 59.49 | |
| P-Llama8B | 5.61 | 2.17 | 72.12 | 3.76 | 1.75 | 68.19 | 3.29 | 1.87 | 63.75 | |
| P-Qwen4B | 6.41 | 1.83 | 77.81 | 4.28 | 1.57 | 73.11 | 3.72 | 1.65 | 69.22 | |
| P-Qwen9B | 6.58 | 1.95 | 77.19 | 4.30 | 1.68 | 71.90 | 3.76 | 1.71 | 68.72 | |
| Mean | – | – | 72.26 | – | – | 68.19 | – | – | 63.85 |

| Llama-3.1-8B-Instruct | Qwen3.5-4B | Qwen3.5-9B | ||||
|---|---|---|---|---|---|---|
| Benchmark | P | J | P | J | P | J |
| Language | ||||||
| MMLU-Pro [39] | 37.32 | 11.21↓ | 45.04 | 45.48 | 51.38 | 51.30 |
| MMLU-Redux [15] | 60.37 | 22.83↓ | 69.13 | 70.53 | 74.57 | 74.27 |
| GPQA Diamond [35] | 33.33 | 24.24↓ | 33.33 | 41.41↑ | 45.96 | 49.49↑ |
| IFEval [51] | 72.46 | 10.35↓ | 80.41 | 38.82↓ | 83.55 | 67.84↓ |
| IFBench [29] | 26.33 | 16.33↓ | 29.33 | 17.67↓ | 32.67 | 28.33↓ |
| GSM8K [8] | 86.96 | 0.61↓ | 90.90 | 81.96↓ | 92.34 | 92.34 |
| WinoGrande [36] | 61.40 | 49.57↓ | 65.59 | 68.35↑ | 74.19 | 73.24 |
| OpenBookQA [26] | 81.60 | 27.60↓ | 86.20 | 86.90 | 90.30 | 92.20↑ |
| HumanEval [7] | 64.02 | 0.00↓ | 82.93 | 71.95↓ | 84.76 | 80.49↓ |
| Vision | ||||||
| MMMU [47] | – | – | 49.44 | 54.45↑ | 54.23 | 59.02↑ |
| MMMU-Pro [48] | – | – | 32.60 | 37.23↑ | 42.72 | 43.41 |
| MMMU-Pro Vision [48] | – | – | 31.45 | 34.16↑ | 40.46 | 40.64 |
| MMStar [6] | – | – | 53.27 | 61.67↑ | 65.87 | 65.67 |
| BabyVision [5] | – | – | 19.07 | 14.43↓ | 16.49 | 15.98 |
| RealWorldQA [42] | – | – | 74.12 | 69.54↓ | 75.42 | 76.08 |
| Spatial | ||||||
| ERQA [16] | – | – | 45.25 | 42.25↓ | 45.75 | 44.00↓ |
| EmbSpatialBench [14] | – | – | 75.14 | 75.16 | 76.59 | 76.95 |
| RefSpatialBench [50] | – | – | 20.94 | 1.81↓ | 38.27 | 27.80↓ |
| LingoQA [25] | – | – | 70.40 | 58.80↓ | 75.00 | 68.40↓ |
| Backbone | Encoder dim c | Tokens m | Backbone dim c′ | Proj. params | LoRA params |
|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct | 384 | 513 | 4096 | 10.89M | 41.94M |
| Qwen3.5-4B | 384 | 513 | 2560 | 7.74M | 32.46M |
| Qwen3.5-9B | 384 | 513 | 4096 | 10.89M | 43.28M |
| System message |
|---|
| Answer the question or follow the instruction regarding the given 3D object. |
| Answer the following question or carry out the instruction about the provided 3D object. |
| Given a 3D object, respond to the question or instruction about it. |
| Consider the 3D object and answer the question or follow the instruction that follows. |
| Respond to the question or instruction concerning the presented 3D object. |
| Using the given 3D object, answer the question or complete the instruction. |
| You are given a 3D object; answer the question or follow the instruction about it. |
| Examine the 3D object and answer the accompanying question or instruction. |
| Provide an answer to the question or complete the instruction about the given 3D object. |
| Based on the 3D object shown, answer the question or follow the instruction. |
| Address the question or instruction about the provided 3D object. |
| Hyperparameter | Value |
|---|---|
| Optimizer | AdamW, β=(0.9,0.999), ϵ=10−8 |
| Learning-rate schedule | constant, no warmup |
| Projector learning rate | 2×10−3 (1×10−3 for Llama8B) |
| Weight decay | 0.0 |
| Gradient clipping | global norm 1.0 |
| Micro-batch / accumulation / effective | 12 / 2 / 24 |
| Max response length | 512 tokens |
| Mixed precision | bfloat16 autocast |
| Random seed | 0 |
| Hardware | 1× NVIDIA A100-80GB |
| Gradient checkpointing | enabled |
| LoRA rank / α / dropout | 16 / 32 / 0.05 |
| LoRA target modules | all-linear |
| LoRA learning rate | 2×10−5 (2×10−4 for Llama8B) |
| ID | Total samples | Avg. throughput (h-1) | Avg. step rate (h-1) |
|---|---|---|---|
| P-Llama8B | 432,936 | 27,056 | 1,127 |
| J-Llama8B | 200,640 | 12,537 | 522 |
| P-Qwen4B | 480,768 | 30,045 | 1,252 |
| J-Qwen4B | 245,592 | 15,346 | 639 |
| P-Qwen9B | 325,176 | 20,321 | 847 |
| J-Qwen9B | 174,720 | 10,918 | 455 |
| Model | M40. (I) | M40. (C) | Obj. (I) | Obj. (C) | Omni. (I) | Omni. (C) |
|---|---|---|---|---|---|---|
| ShapeLLM-7B [32] | 18.76 | 17.95 | 30.23 | 31.37 | 15.28 | 18.75 |
| ShapeLLM-13B [32] | 22.45 | 21.60 | 40.67 | 38.90 | 28.66 | 33.33 |
| PointLLM-7B [43] | 54.94 | 55.55 | 59.83 | 60.43 | 34.78 | 35.70 |
| PointLLM-13B [43] | 57.66 | 56.81 | 61.73 | 60.90 | 34.63 | 34.01 |
| PointLLM-R [3] | 62.28 | 62.64 | 65.23 | 65.40 | 38.81 | 38.69 |
| MiniGPT-3D [37] | 63.41 | 62.84 | 63.97 | 63.00 | 41.10 | 39.27 |
| P-Llama8B | 61.95 | 63.57 | 65.60 | 62.33 | 49.19 | 40.69 |
| J-Llama8B | 49.72 | 57.37 | 66.80 | 59.53 | 42.34 | 37.17 |
| J-NoLoRA-Llama8B | 36.02 | 40.24 | 58.90 | 51.07 | 39.62 | 40.69 |
| P-Qwen4B | 59.56 | 59.56 | 60.90 | 60.83 | 41.07 | 40.25 |
| J-Qwen4B | 58.27 | 59.60 | 71.00 | 68.80 | 48.17 | 47.32 |
| J-NoLoRA-Qwen4B | 54.94 | 53.24 | 58.73 | 57.97 | 27.08 | 27.49 |
| P-Qwen9B | 59.28 | 62.80 | 74.07 | 67.77 | 48.33 | 38.76 |
| J-Qwen9B | 52.59 | 58.75 | 75.23 | 66.23 | 51.32 | 44.86 |
| J-NoLoRA-Qwen9B | 51.62 | 52.63 | 49.13 | 56.93 | 24.51 | 28.74 |
| Curriculum-Qwen4B | 53.48 | 54.34 | 74.80 | 75.63 | 46.08 | 46.73 |
| GPT-5.6 Luna | Claude Haiku 4.5 | Gemini 3.5 Flash-Lite | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | C | H ↓ | P | C | H ↓ | P | C | H ↓ | P |
| ShapeLLM-7B [32] | 2.23 | 1.52 | 59.47 | 1.65 | 1.46 | 53.15 | 1.03 | 1.69 | 37.85 |
| ShapeLLM-13B [32] | 2.68 | 1.84 | 59.32 | 2.04 | 1.68 | 54.93 | 1.39 | 2.02 | 40.81 |
| PointLLM-7B [43] | 4.26 | 2.12 | 66.75 | 3.37 | 1.95 | 63.30 | 2.75 | 1.99 | 58.08 |
| PointLLM-13B [43] | 4.35 | 2.10 | 67.43 | 3.47 | 1.91 | 64.46 | 2.82 | 1.92 | 59.49 |
| PointLLM-R [3] | 3.38 | 1.18 | 74.10 | 2.50 | 1.10 | 69.42 | 2.48 | 1.17 | 67.94 |
| MiniGPT-3D [37] | 4.51 | 2.01 | 69.16 | 3.19 | 1.93 | 62.32 | 2.70 | 1.95 | 58.07 |
| P-Llama8B | 5.61 | 2.17 | 72.12 | 3.76 | 1.75 | 68.19 | 3.29 | 1.87 | 63.75 |
| J-Llama8B | 6.09 | 2.35 | 72.15 | 4.01 | 1.95 | 67.28 | 3.41 | 2.11 | 61.79 |
| J-NoLoRA-Llama8B | 3.14 | 2.96 | 51.48 | 2.35 | 2.08 | 52.96 | 1.69 | 2.31 | 42.15 |
| P-Qwen4B | 6.41 | 1.83 | 77.81 | 4.28 | 1.57 | 73.11 | 3.72 | 1.65 | 69.22 |
| J-Qwen4B | 6.73 | 1.97 | 77.33 | 4.43 | 1.67 | 72.67 | 3.91 | 1.71 | 69.52 |
| J-NoLoRA-Qwen4B | 5.44 | 4.23 | 56.26 | 4.30 | 2.53 | 62.97 | 2.92 | 3.19 | 47.80 |
| P-Qwen9B | 6.58 | 1.95 | 77.19 | 4.30 | 1.68 | 71.90 | 3.76 | 1.71 | 68.72 |
| J-Qwen9B | 6.83 | 2.02 | 77.19 | 4.46 | 1.73 | 72.00 | 3.92 | 1.80 | 68.52 |
| J-NoLoRA-Qwen9B | 5.14 | 4.74 | 52.02 | 3.84 | 3.02 | 56.00 | 2.58 | 3.65 | 41.46 |
| Curriculum-Qwen4B | 6.52 | 1.95 | 76.96 | 4.36 | 1.67 | 72.33 | 3.76 | 1.74 | 68.38 |
| M40. (Avg.) | Obj. (Avg.) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | 2 h | 4 h | 8 h | 12 h | 16 h | 2 h | 4 h | 8 h | 12 h | 16 h |
| P-Qwen4B | 54.23 | 54.76 | 58.33 | 63.53 | 59.56 | 58.10 | 62.52 | 67.25 | 64.90 | 60.87 |
| J-Qwen4B | 46.84 | 51.64 | 51.03 | 52.09 | 58.93 | 58.20 | 66.50 | 73.28 | 75.05 | 69.90 |
| P-Qwen9B | 51.38 | 53.99 | 57.03 | 61.06 | 61.04 | 57.47 | 63.07 | 67.80 | 67.35 | 70.92 |
| J-Qwen9B | 39.49 | 46.84 | 54.52 | 58.53 | 55.67 | 51.20 | 60.48 | 71.08 | 67.90 | 70.73 |
| P-Llama8B | 53.85 | 57.22 | 57.80 | 56.14 | 62.76 | 51.48 | 57.58 | 66.10 | 70.07 | 63.97 |
| J-Llama8B | 45.50 | 48.28 | 47.49 | 56.62 | 53.55 | 47.80 | 54.02 | 62.28 | 65.23 | 63.17 |
| Object | ||
|---|---|---|
| PointLLM-7B | The 3D model represents a distinctive, green toy car that stands out due to its unconventional feature - a cornered driving unit. Designed for play, this model is aesthetically appealing with an eye-catching bright green color that would attract young children. The unique driving unit located at the corner provides an interesting twist to the traditional car design, promoting creative play and imaginative scenarios. However, despite its playful design, it is still a toy and should be handled with care. | The 3D object is a lively representation of a blue whale, which is the largest mammal on Earth. It is depicted in a playful cartoon style with an eye-catching cobalt-blue hue. The whale features proportional body parts, including a small dorsal fin, enormous mouth, and flat, wide head. It seems to be designed for entertainment or educational purposes, possibly for children, given its cartoon-like attributes and the significance of the whale in marine life. |
| PointLLM-13B | The 3D object model represents a small toy car, with the primary colour being a vibrant green. It features four black wheels that contrast well with the green body. Two red lights are situated at the front, possibly to mimic the headlights of an actual car. This model, being a toy, is typically used for children’s play, aiding in their motor skill development and imagination. The car’s simplistic design might be suitable for younger age groups. | This is a 3D model of a blue whale, the largest mammal on Earth, depicted in a cartoon style. The whale is distinguished by its small dorsal fin, enormous mouth, and flat, wide head. Despite the cartoon aesthetic, the model accurately represents the unique features of a blue whale. It could be used for educational purposes or in animated movies or documentaries. |
| P-Llama8B | The object is a green and white van with a boxy shape, featuring a prominent front grille, round headlights, and a rear window. It has a flat roof and a side door with a handle. The van has a distinctive green and white color scheme with black accents. The wheels are circular and black, and the vehicle has a distinctive front bumper. | The object is a sleek, streamlined shark with a pointed snout and a series of sharp, triangular fins along its back. The body is predominantly blue, with a lighter underside and a distinctive pattern of darker spots. The fins are a lighter shade, and the tail is long and narrow, tapering to a sharp point. The shark has a prominent dorsal fin and a smaller pectoral fin, with a series of gill slits visible along the sides. |
| J-Llama8B | The object is a green vehicle with a boxy design, featuring a large windshield, a rear window, and a front grille. It has a roof rack and a spare tire mounted on the rear. The vehicle appears to be a van, with a prominent front bumper and a rear section that includes a license plate. The wheels are circular, and the vehicle has a simple, utilitarian appearance. | The object is a stylized representation of a fish, characterized by a streamlined body with a pointed snout and a forked tail. The body is primarily blue with a white underbelly, and the fins are white. The fish has a smooth, abstract appearance with a minimalistic design. |
왜 중요한가
AI 개발자 입장에서는 새로운 데이터 형태(3D, 이미지 등)를 다룰 수 있게 만들 때 거대한 언어모델 전체를 다시 학습시킬 필요 없이 작은 연결 장치만 학습시키면 되므로, 계산 비용과 시간을 크게 아끼면서도 기존 언어 능력을 보존할 수 있다는 뜻이다. 이는 하나의 언어모델에 여러 감각(모달리티)용 프로젝터를 독립적으로 붙여 재사용하는 모듈식 AI 구축 방식의 가능성을 열어준다.
이 논문의 용어
- MLLM(멀티모달 대형 언어모델) · 텍스트뿐 아니라 이미지, 3D, 오디오 등 다양한 형태의 정보를 함께 이해하도록 만든 AI 모델
- 프로젝터 · 3D나 이미지 같은 다른 형태의 데이터를 언어모델이 이해할 수 있는 형태로 변환해주는 작은 신경망 연결부
- LoRA 어댑터 · 거대 모델 전체를 다시 학습시키지 않고 일부 작은 파라미터만 추가로 학습시켜 효율적으로 미세조정하는 기법
- 점군(point cloud) · 물체의 표면을 수많은 점들의 좌표(및 색상) 집합으로 표현한 3D 데이터 형식
- LLM-as-a-Judge · AI 모델의 답변이 맞았는지를 사람 대신 또 다른 대형 언어모델이 채점하게 하는 평가 방식
논문 원문 초록 (영문)
The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuning the backbone of an MLLM is necessary to adapt it to a new modality. Through experiments on 3D MLLMs, we find that training only the projector is sufficient to achieve strong multimodal performance relative to existing baseline models and our jointly trained MLLMs with the same encoder and backbone. We also show that joint training leads to undesirable drift in existing capabilities of the language model, which projector-only training avoids by definition. Furthermore, projector-only training has approximately twice the training sample throughput of joint training. We validate our findings across different language model backbones via 3D classification and captioning benchmarks as well as standard benchmarks evaluating language, vision, and spatial reasoning capabilities.
arXiv에서 원문 보기최신 논문
- Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms데이터 플랫폼 변경도 코드처럼 '설계도 조각'을 붙여서 검토하면 어떨까: 실험 설계 논문
- Are LLMs becoming similarly creative? Evidence from three years of models최신 AI 챗봇일수록 서로 비슷한 답을 내놓는다는 3년치 조사 결과
- Auditing Cross-Lingual Fairness in Language Model WatermarkingAI 생성 텍스트를 잡아내는 워터마크 기술이 영어 아닌 언어에서는 훨씬 부실하게 작동하고, 그 격차는 개별 언어가 아니라 언어 계열 단위로 나타난다
- TESTNAV: Pareto-Guided Search for Compositional Robustness TestingAI 모델을 여러 손상이 겹친 입력으로 시험할 때, 굳이 다 테스트하지 않고도 '진짜 위험한 실패'만 골라내는 탐색법
- Optimal Skill Selection for LLM Agents with Provable Bicriteria GuaranteesAI 에이전트에게 어떤 '스킬 문서'를 몇 개나 줘야 잘 작동하는지, 수학적으로 최적해를 보장하며 골라주는 방법
- Reliable Financial Named Entity Recognition under Domain Shift금융 AI가 서류체 문장에서 배운 자신감은 트위터로 가면 거짓말이 된다
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction텍스트가 빠지거나 망가져도, AI가 그 자리를 대신할 '가짜 텍스트'를 한 번에 만들지 않고 여러 번 고쳐가며 감정을 더 정확히 읽어낸다
METAL LAB 최신 기사
그림 출처: Nyx Iskandar et al., arXiv:2608.19726, CC BY 4.0