FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
AI 에이전트에게 '터미널 문제집'을 정합성 있게 만들어주는 파이프라인, FACET
터미널(명령줄) 작업을 수행하는 AI 에이전트를 훈련시키려면 지시문, 초기 환경, 정답 풀이, 채점기(verifier)가 서로 딱 맞아떨어지는 '문제 세트'가 대량으로 필요하다. FACET은 실제 공개된 작업 스킬들을 재조합해 그럴듯한 시나리오를 복원하고, 실행 가능한 도커 환경을 먼저 만든 뒤 그 환경을 기준으로 지시문·정답·채점기를 순서대로 생성해 서로 어긋나지 않게 만든다. 이렇게 만든 6,078개 작업으로 얻은 성공 사례 1,200개만 학습에 써도 Qwen3.5 계열 모델(4B~27B)의 Terminal-Bench 2.1 점수가 모두 눈에 띄게 올랐다.
무엇을 했나
- 71,000개 이상의 공개 에이전트 스킬(작업 절차)을 수집·정제한 뒤, 비슷한 목표를 가진 스킬들을 묶어 하나의 그럴듯한 사용자 시나리오로 복원한다.
- 시나리오를 목표·맥락·능력·상태·입출력 다섯 관점으로 정리해 정답 풀이 참조안과 지시문 참조안을 만들고, 둘이 같은 초기·최종 상태를 가리키는지 모델로 대조 검증한다.
- 실행 환경(도커 이미지, 파일, 서비스)을 먼저 만들고 빌드 오류가 나면 최대 3번까지 자동 수리한 뒤, 그 완성된 환경을 보면서 지시문→정답→채점기 순서로 생성해 셋이 실제 파일·경로·스키마를 공유하게 만든다.
- 최종적으로 환경 빌드, 채점기의 초기 상태 거부, 정답 실행, 채점기의 최종 상태 통과 네 조건을 모두 만족해야 문제로 채택되며, 실패 시 문제 전체가 아니라 원인이 된 부분(환경/정답/채점기)만 최대 5번 수리한다.
- 생성 순서 비교 실험에서 지시문→정답→채점기 순서(Forward)가 채점기를 먼저 만드는 방식보다 초기 유효율이 훨씬 높았고(46.5% vs 24.2%), 최종적으로 6,078개 검증된 작업과 1,200개 성공 궤적으로 미세조정한 Qwen3.5 모델들은 4B에서 17.60→24.72, 9B에서 27.34→35.58, 27B에서 40.82→47.57로 성능이 올라 15배 큰 모델(397B, 49.06점)에 근접했다.

| Dataset | Trajectory | Task | ||||
|---|---|---|---|---|---|---|
| #Traj. | Turns | #Tasks | Tests | P@1 | P@3 | |
| Nemotron-Terminal (29) | 5K | 6.12 | 15K | 6.18 | 40.67 | 48.00 |
| Endless-Terminals (13) | 200 | 4.53 | 2,492 | 5.51 | 83.00 | 87.00 |
| Terminal-Lego (33) | 32K | 5.77 | 15K | 16.60 | 47.00 | 49.00 |
| TerminalWorld (8) | 200 | 11.94 | 1,530 | 3.98 | 57.00 | 82.00 |
| Tmax (16) | 500 | 11.14 | 15K | 3.29 | 80.00 | 86.00 |
| FACET (ours) | 1.2K | 11.86 | 6K | 22.77 | 27.00 | 35.00 |
| Model | Size | Agent | Terminal-Bench 2.1 |
|---|---|---|---|
| Reported reference models | |||
| GPT-5.5 (xhigh) (31) | — | Terminus-2 | 78.00 |
| Claude Opus 4.7 (max) (31) | — | Terminus-2 | 66.10 |
| Gemini 3 Pro (high) (31) | — | Gemini CLI | 65.80 |
| Intern-S2-Preview-397B (7) | 397B | Terminus-2 | 67.42 |
| MiniMax M3 (19) | 428B | Terminus-2 | 66.00 |
| GLM-5.1 (max) (31) | 744B | Claude Code | 58.70 |
| Models evaluated under our setting | |||
| Qwen3.6-27B | 27B | Terminus-2 | 53.93 |
| Qwen3.5-397B-A17B | 397B | Terminus-2 | 49.06 |
| Kimi-K2.6 | 1T | Terminus-2 | 59.93 |
| DeepSeek-V4-Pro-Preview (high) | 1.6T | Terminus-2 | 73.03 |
| Qwen3.5 base models | |||
| Qwen3.5-4B | 4B | Terminus-2 | 17.60 |
| Qwen3.5-9B | 9B | Terminus-2 | 27.34 |
| Qwen3.5-27B | 27B | Terminus-2 | 40.82 |
| Fine-tuned models | |||
| FACET-Terminal-Qwen3.5-4B | 4B | Terminus-2 | 24.72 (+7.12) |
| FACET-Terminal-Qwen3.5-9B | 9B | Terminus-2 | 35.58 (+8.24) |
| FACET-Terminal-Qwen3.5-27B | 27B | Terminus-2 | 47.57 (+6.75) |

| Category | Skills | Share |
|---|---|---|
| AI, agents, and tools | 15,182 | 21.28% |
| Software, systems, and security | 15,059 | 21.11% |
| Data, analysis, and research | 12,409 | 17.39% |
| Documents, productivity, and workflows | 11,267 | 15.79% |
| Multimedia, creation, and publishing | 17,424 | 24.42% |
| Total | 71,341 | 100.00% |
| Parent | Fine-grained category | Count | Global | Within parent |
|---|---|---|---|---|
| AI, agents, and tools | Agent orchestration and automation | 2,782 | 3.90% | 18.32% |
| Prompt and model calls | 3,310 | 4.64% | 21.80% | |
| Skills, plugins, and extensions | 1,637 | 2.29% | 10.78% | |
| MCP and external tools | 2,579 | 3.62% | 16.99% | |
| Memory, RAG, and knowledge bases | 3,439 | 4.82% | 22.65% | |
| Multi-agent collaboration | 1,435 | 2.01% | 9.45% | |
| Software, systems, and security | Code generation and development | 3,065 | 4.30% | 20.35% |
| Testing, debugging, and code quality | 2,167 | 3.04% | 14.39% | |
| Git, build, and dependency management | 1,904 | 2.67% | 12.64% | |
| Deployment, containers, and DevOps | 2,069 | 2.90% | 13.74% | |
| System administration and CLI | 3,576 | 5.01% | 23.75% | |
| Security, privacy, and compliance | 2,278 | 3.19% | 15.13% | |
| Data, analysis, and research | JSON, YAML, and XML | 2,752 | 3.86% | 22.18% |
| CSV, Excel, and spreadsheets | 1,544 | 2.16% | 12.44% | |
| Databases and SQL | 1,627 | 2.28% | 13.11% | |
| Data cleaning, conversion, and validation | 1,595 | 2.24% | 12.85% | |
| Statistical analysis and metrics | 1,892 | 2.65% | 15.25% | |
| Visualization and dashboards | 1,419 | 1.99% | 11.44% | |
| Search, research, and extraction | 1,580 | 2.21% | 12.73% | |
| Documents, productivity, and workflows | Documents and Markdown | 1,496 | 2.10% | 13.28% |
| Reports, summaries, and briefs | 1,865 | 2.61% | 16.55% | |
| PDF, Office, and presentations | 1,429 | 2.00% | 12.68% | |
| Office and personal productivity | 1,592 | 2.23% | 14.13% | |
| Project, task, and schedule management | 1,682 | 2.36% | 14.93% | |
| System integration and automation | 1,597 | 2.24% | 14.17% | |
| Audit, checklists, and operation records | 1,606 | 2.25% | 14.25% | |
| Multimedia, creation, and publishing | Image generation and editing | 2,699 | 3.78% | 15.49% |
| Design, drawing, and visual assets | 1,858 | 2.60% | 10.66% | |
| Audio, speech, and music | 1,917 | 2.69% | 11.00% | |
| Video, animation, and captions | 2,491 | 3.49% | 14.30% |
| Stage | Count | Stage retention |
|---|---|---|
| Scenario–skill seeds | 7,852 | — |
| Seeds with first-build logs | 7,841 | 99.86% |
| Initial environment success | 6,630 | 84.56% |
| Environment repair recovery | 874 | — |
| Successful environments | 7,504 | 95.70% |
| Entering task validation | 7,446 | 99.23% |
| First-pass valid tasks | 2,856 | 38.35% |
| Task repair recovery | 3,222 | — |
| Final validated tasks | 6,078 | 81.63% |
| View | Statistic | Value |
|---|---|---|
| Command usage | cat occurrences | 18,168 (46.4%) |
| Top three commands | 27,207 (69.5%) | |
| Top ten commands | 34,598 (88.4%) | |
| Observation commands | 33,140 (84.7%) | |
| Turn dynamics | Observation-only first turn | 1,214 (95.6%) |
| Observation-only → observation-only | 4,324 (55.6%) | |
| Action-only → observation-only | 1,399 (53.1%) | |
| Action-only → action-only | 758 (28.7%) |
| Scheme | Reached validation | Initially valid | Final yield |
|---|---|---|---|
| Forward (Ours) | 99 | 46 (46.5%) | 83/100 |
| Reverse | 91 | 22 (24.2%) | 63/100 |
| Joint | 96 | 36 (37.5%) | 65/100 |
| Step | Stage | Operation |
|---|---|---|
| 1 | Materialize environment | Generate the Dockerfile, fixtures, dependencies, and initialization scripts from the reconstructed task specification. |
| 2 | Build and repair | Build the image with docker build. Build or initialization failures are returned to the environment-repair agent for at most three iterations. |
| 3 | Capture shared state | Start a temporary container, inspect the task workspace, and record the realized initial state e0. |
| 4 | Generate artifacts | Generate the final instruction, reference solution, and verifier using the same reconstructed specification and shared state e0. |
| 5 | Baseline validation | Start a clean container, copy and execute the verifier without running the solution, and require reward 0. |
| 6 | Oracle validation | Start another clean container, copy and execute solution/solve.sh, run tests/test.sh, and require reward 1. |
| 7 | Repair and revalidate | Classify a failure as an instruction, environment, solution, or verifier defect, apply the corresponding repair, and repeat the full validation procedure for at most five rounds. |
| 8 | Accept and clean up | Retain the task only after all validation conditions pass, then stop and remove temporary containers and images. |
| Setting | Value |
|---|---|
| Supervised fine-tuning | |
| Base models | Qwen3.5-4B, Qwen3.5-9B, Qwen3.5-27B |
| Training data | 1.2K complete successful trajectories |
| Training strategy | Full-parameter SFT with BF16 and ZeRO-3 |
| Epochs / effective batch size | 3 / 64 |
| Learning rate / schedule | 1×10−5 / cosine with 0.1 warmup ratio |
| Maximum sequence length | 32,768 tokens |
| Evaluation | |
| Benchmark / agent | Terminal-Bench 2.1 / Terminus-2 |
| Attempts / timeout | 3 per task / 2 hours per attempt |
| Temperature | 1.0 |
| Context length | FACET models: 32,768 tokens; other models: officially supported maximum |
| Pipeline | Packages | Validated | Yield | P@1 | P@3 | Avg. Cmds. |
|---|---|---|---|---|---|---|
| Baseline | 437 | 78 | 15.6% | 80.8% | 85.9% | 12.8 |
| TW | 449 | 139 | 27.8% | 58.8% | 64.0% | 17.0 |
| FACET (Ours) | 395 | 350 | 70.0% | 25.1% | 33.1% | 21.5 |
왜 중요한가
터미널 에이전트를 학습시키려면 사람이 일일이 만들기 힘든 '실행 가능하고 검증 가능한' 문제가 대량으로 필요한데, 지시문·환경·정답·채점기가 서로 어긋나면 학습 데이터 자체가 오염된다. FACET은 이런 정합성 문제를 구조적으로 해결하는 방법을 보여줘, 적은 데이터로도 효율적인 에이전트 훈련이 가능함을 시사한다.
이 논문의 용어
- 터미널 에이전트 · 명령줄(터미널)에서 파일 조작, 프로그램 설치, 명령어 실행 등을 수행하는 AI
- 검증기(verifier) · 작업이 올바르게 완료됐는지 자동으로 확인하는 실행 가능한 채점 코드
- Harbor 포맷 · 환경/정답/테스트/지시문을 표준화된 폴더 구조로 담는 작업 패키지 형식
- P@1 / P@3 · 한 번 또는 세 번 시도했을 때 문제를 통과하는 비율
- 미세조정(fine-tuning) · 이미 학습된 모델을 특정 데이터로 추가 학습시켜 성능을 개선하는 과정
논문 원문 초록 (영문)
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies, state tran
arXiv에서 원문 보기최신 논문
- FM-Bench: A Benchmark for Long-Horizon Management with Competing AgentsAI에게 축구 구단을 20년간 통째로 맡겨보니, 승패는 계산력이 아니라 '경영 감각'에서 갈렸다
- FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI SystemsAI가 '정답'을 맞혀도, 정작 증거는 못 찾는 경우가 대부분이었다 - 재무 이상거래 진단 AI의 숨은 약점
- Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models음성·소리를 알아듣는 AI, 지시문 학습 없이 '연결 다리'만 훈련해도 충분하다
- Adversarial Review: Structured Disagreement for Grounded Agentic Code ReviewAI 코딩 에이전트, 에이전트 늘리기보다 '검토자 vs 비판자' 한 쌍이 더 똑똑하게 일한다
- Looped Language Models Improve Compositional Tool Calling생각을 여러 번 되짚는 AI가 여러 개의 도구를 순서대로 엮어 쓰는 일도 더 잘한다
- Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement쇼핑몰 검색에서 이탈한 고객을 AI 에이전트가 다시 카톡(왓츠앱)으로 불러온 이야기
- GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks바이러스 유전자 서열을 코돈끼리 서로 옆에 등장하는 관계망(그래프)으로 바꿔 변이를 구분하는 법
- Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval이미지-긴문장 검색 AI가 '이미 다 맞혔다'고 착각해서 정작 헷갈리는 문제를 못 배우던 버릇을 고쳤다
METAL LAB 최신 기사
그림 출처: Kou Shi et al., arXiv:2608.18580, CC BY 4.0