FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
FACET:让终端命令行任务的“说明书、环境、答案、判分器”自动保持一致
要训练能在终端(命令行)里干活的AI智能体,需要大量任务,而每个任务的指令、初始环境、参考解法和自动判分器必须彼此吻合,否则训练数据本身就是错的。FACET把现实中公开的智能体技能重新组合还原成连贯的使用场景,先搭建并修复好可运行的Docker环境,再依次生成指令、解法、判分器,让三者都基于同一个真实环境状态。仅用6,078个通过验证的任务中挑出的1,200条成功轨迹做微调,Qwen3.5系列模型(4B到27B)在Terminal-Bench 2.1上的成绩全都明显提升。
他们做了什么
- 收集并清洗了超过71,000个公开的智能体技能,把目标相近的技能归并、还原成一个信息更丰富、连贯的用户场景。
- 从目标、场景、能力、状态、输入输出与工具五个维度描述场景,据此生成解法参考和指令参考,再用模型核对两者的初始状态和最终结果是否一致。
- 先构建执行环境(Docker镜像、文件、服务),构建失败时最多自动修复3次;环境跑通后,再依次生成指令、解法、判分器,三者都读取同一个真实环境状态,确保引用的文件、路径、格式一致。
- 只有当环境成功构建、判分器在初始状态下判定为未完成、参考解法能在干净环境中跑通、判分器在最终状态下判定通过这四条全部满足时,任务才会被采纳;验证失败时只针对性修复出问题的部分(环境/解法/判分器),最多修复5轮,而不是整体重做。
- 对比生成顺序发现,按“指令→解法→判分器”顺序生成(Forward)的初始有效率(46.5%)远高于先生成判分器的方式(24.2%);用最终1,200条成功轨迹微调后,Qwen3.5-4B从17.60升到24.72,9B从27.34升到35.58,27B从40.82升到47.57,已经非常接近参数量大15倍的397B模型(49.06分)。

| Dataset | Trajectory | Task | ||||
|---|---|---|---|---|---|---|
| #Traj. | Turns | #Tasks | Tests | P@1 | P@3 | |
| Nemotron-Terminal (29) | 5K | 6.12 | 15K | 6.18 | 40.67 | 48.00 |
| Endless-Terminals (13) | 200 | 4.53 | 2,492 | 5.51 | 83.00 | 87.00 |
| Terminal-Lego (33) | 32K | 5.77 | 15K | 16.60 | 47.00 | 49.00 |
| TerminalWorld (8) | 200 | 11.94 | 1,530 | 3.98 | 57.00 | 82.00 |
| Tmax (16) | 500 | 11.14 | 15K | 3.29 | 80.00 | 86.00 |
| FACET (ours) | 1.2K | 11.86 | 6K | 22.77 | 27.00 | 35.00 |
| Model | Size | Agent | Terminal-Bench 2.1 |
|---|---|---|---|
| Reported reference models | |||
| GPT-5.5 (xhigh) (31) | — | Terminus-2 | 78.00 |
| Claude Opus 4.7 (max) (31) | — | Terminus-2 | 66.10 |
| Gemini 3 Pro (high) (31) | — | Gemini CLI | 65.80 |
| Intern-S2-Preview-397B (7) | 397B | Terminus-2 | 67.42 |
| MiniMax M3 (19) | 428B | Terminus-2 | 66.00 |
| GLM-5.1 (max) (31) | 744B | Claude Code | 58.70 |
| Models evaluated under our setting | |||
| Qwen3.6-27B | 27B | Terminus-2 | 53.93 |
| Qwen3.5-397B-A17B | 397B | Terminus-2 | 49.06 |
| Kimi-K2.6 | 1T | Terminus-2 | 59.93 |
| DeepSeek-V4-Pro-Preview (high) | 1.6T | Terminus-2 | 73.03 |
| Qwen3.5 base models | |||
| Qwen3.5-4B | 4B | Terminus-2 | 17.60 |
| Qwen3.5-9B | 9B | Terminus-2 | 27.34 |
| Qwen3.5-27B | 27B | Terminus-2 | 40.82 |
| Fine-tuned models | |||
| FACET-Terminal-Qwen3.5-4B | 4B | Terminus-2 | 24.72 (+7.12) |
| FACET-Terminal-Qwen3.5-9B | 9B | Terminus-2 | 35.58 (+8.24) |
| FACET-Terminal-Qwen3.5-27B | 27B | Terminus-2 | 47.57 (+6.75) |

| Category | Skills | Share |
|---|---|---|
| AI, agents, and tools | 15,182 | 21.28% |
| Software, systems, and security | 15,059 | 21.11% |
| Data, analysis, and research | 12,409 | 17.39% |
| Documents, productivity, and workflows | 11,267 | 15.79% |
| Multimedia, creation, and publishing | 17,424 | 24.42% |
| Total | 71,341 | 100.00% |
| Parent | Fine-grained category | Count | Global | Within parent |
|---|---|---|---|---|
| AI, agents, and tools | Agent orchestration and automation | 2,782 | 3.90% | 18.32% |
| Prompt and model calls | 3,310 | 4.64% | 21.80% | |
| Skills, plugins, and extensions | 1,637 | 2.29% | 10.78% | |
| MCP and external tools | 2,579 | 3.62% | 16.99% | |
| Memory, RAG, and knowledge bases | 3,439 | 4.82% | 22.65% | |
| Multi-agent collaboration | 1,435 | 2.01% | 9.45% | |
| Software, systems, and security | Code generation and development | 3,065 | 4.30% | 20.35% |
| Testing, debugging, and code quality | 2,167 | 3.04% | 14.39% | |
| Git, build, and dependency management | 1,904 | 2.67% | 12.64% | |
| Deployment, containers, and DevOps | 2,069 | 2.90% | 13.74% | |
| System administration and CLI | 3,576 | 5.01% | 23.75% | |
| Security, privacy, and compliance | 2,278 | 3.19% | 15.13% | |
| Data, analysis, and research | JSON, YAML, and XML | 2,752 | 3.86% | 22.18% |
| CSV, Excel, and spreadsheets | 1,544 | 2.16% | 12.44% | |
| Databases and SQL | 1,627 | 2.28% | 13.11% | |
| Data cleaning, conversion, and validation | 1,595 | 2.24% | 12.85% | |
| Statistical analysis and metrics | 1,892 | 2.65% | 15.25% | |
| Visualization and dashboards | 1,419 | 1.99% | 11.44% | |
| Search, research, and extraction | 1,580 | 2.21% | 12.73% | |
| Documents, productivity, and workflows | Documents and Markdown | 1,496 | 2.10% | 13.28% |
| Reports, summaries, and briefs | 1,865 | 2.61% | 16.55% | |
| PDF, Office, and presentations | 1,429 | 2.00% | 12.68% | |
| Office and personal productivity | 1,592 | 2.23% | 14.13% | |
| Project, task, and schedule management | 1,682 | 2.36% | 14.93% | |
| System integration and automation | 1,597 | 2.24% | 14.17% | |
| Audit, checklists, and operation records | 1,606 | 2.25% | 14.25% | |
| Multimedia, creation, and publishing | Image generation and editing | 2,699 | 3.78% | 15.49% |
| Design, drawing, and visual assets | 1,858 | 2.60% | 10.66% | |
| Audio, speech, and music | 1,917 | 2.69% | 11.00% | |
| Video, animation, and captions | 2,491 | 3.49% | 14.30% |
| Stage | Count | Stage retention |
|---|---|---|
| Scenario–skill seeds | 7,852 | — |
| Seeds with first-build logs | 7,841 | 99.86% |
| Initial environment success | 6,630 | 84.56% |
| Environment repair recovery | 874 | — |
| Successful environments | 7,504 | 95.70% |
| Entering task validation | 7,446 | 99.23% |
| First-pass valid tasks | 2,856 | 38.35% |
| Task repair recovery | 3,222 | — |
| Final validated tasks | 6,078 | 81.63% |
| View | Statistic | Value |
|---|---|---|
| Command usage | cat occurrences | 18,168 (46.4%) |
| Top three commands | 27,207 (69.5%) | |
| Top ten commands | 34,598 (88.4%) | |
| Observation commands | 33,140 (84.7%) | |
| Turn dynamics | Observation-only first turn | 1,214 (95.6%) |
| Observation-only → observation-only | 4,324 (55.6%) | |
| Action-only → observation-only | 1,399 (53.1%) | |
| Action-only → action-only | 758 (28.7%) |
| Scheme | Reached validation | Initially valid | Final yield |
|---|---|---|---|
| Forward (Ours) | 99 | 46 (46.5%) | 83/100 |
| Reverse | 91 | 22 (24.2%) | 63/100 |
| Joint | 96 | 36 (37.5%) | 65/100 |
| Step | Stage | Operation |
|---|---|---|
| 1 | Materialize environment | Generate the Dockerfile, fixtures, dependencies, and initialization scripts from the reconstructed task specification. |
| 2 | Build and repair | Build the image with docker build. Build or initialization failures are returned to the environment-repair agent for at most three iterations. |
| 3 | Capture shared state | Start a temporary container, inspect the task workspace, and record the realized initial state e0. |
| 4 | Generate artifacts | Generate the final instruction, reference solution, and verifier using the same reconstructed specification and shared state e0. |
| 5 | Baseline validation | Start a clean container, copy and execute the verifier without running the solution, and require reward 0. |
| 6 | Oracle validation | Start another clean container, copy and execute solution/solve.sh, run tests/test.sh, and require reward 1. |
| 7 | Repair and revalidate | Classify a failure as an instruction, environment, solution, or verifier defect, apply the corresponding repair, and repeat the full validation procedure for at most five rounds. |
| 8 | Accept and clean up | Retain the task only after all validation conditions pass, then stop and remove temporary containers and images. |
| Setting | Value |
|---|---|
| Supervised fine-tuning | |
| Base models | Qwen3.5-4B, Qwen3.5-9B, Qwen3.5-27B |
| Training data | 1.2K complete successful trajectories |
| Training strategy | Full-parameter SFT with BF16 and ZeRO-3 |
| Epochs / effective batch size | 3 / 64 |
| Learning rate / schedule | 1×10−5 / cosine with 0.1 warmup ratio |
| Maximum sequence length | 32,768 tokens |
| Evaluation | |
| Benchmark / agent | Terminal-Bench 2.1 / Terminus-2 |
| Attempts / timeout | 3 per task / 2 hours per attempt |
| Temperature | 1.0 |
| Context length | FACET models: 32,768 tokens; other models: officially supported maximum |
| Pipeline | Packages | Validated | Yield | P@1 | P@3 | Avg. Cmds. |
|---|---|---|---|---|---|---|
| Baseline | 437 | 78 | 15.6% | 80.8% | 85.9% | 12.8 |
| TW | 449 | 139 | 27.8% | 58.8% | 64.0% | 17.0 |
| FACET (Ours) | 395 | 350 | 70.0% | 25.1% | 33.1% | 21.5 |
为什么重要
训练终端智能体需要大量可执行、可验证的任务,人工编写成本太高,但只要指令、环境、解法、判分器之间稍有错位,训练数据就会失真甚至误导模型。FACET展示了一种结构化的方法来保证这些环节相互一致,说明用少量高质量数据也能有效提升智能体能力。
本文术语
- 终端智能体(terminal agent) · 能在命令行环境中操作文件、安装依赖、执行程序的AI系统
- 判分器(verifier) · 用来自动检查任务是否完成正确的可执行代码
- Harbor格式 · 将环境、解法、测试、指令统一打包成标准文件夹结构的任务格式
- P@1 / P@3 · 智能体尝试1次或3次时任务的通过率
- 微调(fine-tuning) · 在已训练好的模型基础上,用特定数据继续训练以提升其在某任务上的表现
论文原文摘要(英文)
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies, state tran
在 arXiv 阅读最新论文
- FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents让AI连续管理一家足球俱乐部20年后发现,胜负关键不在模型大小,而在经营习惯
- FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI SystemsAI经常能说对财务对账出错的原因,却拿不出真正的证据
- Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models语言模型全程冻结,只训练一个小连接器,也能做出好用的听觉理解AI
- Adversarial Review: Structured Disagreement for Grounded Agentic Code ReviewAI代码审查:与其堆更多智能体,不如让一个审查者和一个批评者互相较真
- Looped Language Models Improve Compositional Tool Calling会反复回想自己答案的AI模型,更擅长按顺序组合调用多个工具
- Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-EngagementAI智能体追着离场用户发WhatsApp,把逛而不买的顾客拉回来
- GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks把病毒基因序列变成密码子关系网络图,用来区分新冠变异株
- Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval图文检索AI误以为自己已经全学会了,结果学不会区分那些几乎一样的描述句子
METAL LAB 最新报道
图片来源: Kou Shi et al., arXiv:2608.18580, CC BY 4.0