Repo0: Design-Driven Zero-to-All Code Generation
要让AI仅凭自然语言需求就从零搭建整个代码仓库,架构设计不能一次画完就定型,而要在写代码的过程中持续修改
现有的代码生成智能体擅长在文件结构和模块边界已经定好的情况下填代码,但一旦只给自然语言需求、要求它自己设计整个仓库架构,就容易出问题。Repo0维护两张互相关联的图,一张记录需求之间的关系,一张记录实现组件之间的依赖,在编码过程中不断对组件结构进行拆分、合并、修订,而不是一开始就固定蓝图。在六个改编自真实开源项目的基准仓库上,相比最强的现有方法,功能覆盖率最高提升20.08个百分点,测试通过率最高提升29.74个百分点。
他们做了什么
- 以往方法如RPG先分析一次需求、生成一份仓库设计蓝图,然后严格按蓝图生成代码,但现实开发中很多模块划分的问题要写代码时才会暴露出来。
- Repo0维护一张需求级图(requirement-level DAG,记录需求单元之间的关联)和一张组件级图(component-level DAG,记录实现模块及其依赖),并用对齐关系把两者连接起来;需求图相对固定,组件图则持续演化。
- 依据内聚度(一个组件内部聚合的职责是否真正相关)和耦合度(两个组件的职责是否重叠)这两个指标,Repo0反复执行拆分、合并、修订、保留四种结构动作,直到不再需要改动为止,即达到结构收敛。
- 结构收敛后,采用测试驱动开发的方式生成代码:先根据对齐的需求写测试,再填充实现去满足测试,失败时只对相关部分做局部修复。
- 在改编自scikit-learn、pandas、sympy、statsmodels、requests、django的六个基准仓库上,分别用GPT-5 mini和DeepSeek V3.2测试,Repo0在所有设置下都取得了最高的功能覆盖率和测试通过率。

| Real Repo | Para. Name | #Files | LOC | Task Counts |
|---|---|---|---|---|
| scikit-learn | MLKit-Py | 185 | 65,972 | 236 |
| pandas | TableKit | 217 | 106,447 | 175 |
| sympy | SymbolicMath | 699 | 218,924 | 192 |
| statsmodels | StatModeler | 271 | 83,325 | 234 |
| requests | HttpEasy | 17 | 2,793 | 50 |
| django | PyWebEngine | 681 | 109,457 | 165 |

| Model | Method | requests | statsmodels | django | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Cov. (%) | Nov. (%) | Pass./Vot. (%) | Cov. (%) | Nov. (%) | Pass./Vot. (%) | Cov. (%) | Nov. (%) | Pass./Vot. (%) | ||
| GPT-5 mini | mini-SWE-agent | 68.18 | 2.63 | 4.11 / 27.40 | 18.18 | 9.09 | 0.00 / 31.86 | 47.92 | 6.96 | 37.04 / 44.44 |
| Paper2Code | 95.50 | 7.20 | 24.66 / 24.66 | 44.32 | 24.13 | 4.42 / 30.09 | 66.67 | 15.30 | 30.04 / 78.60 | |
| RPG | 90.91 | 13.70 | 31.51 / 95.89 | 70.40 | 13.80 | 77.90 / 92.00 | 60.42 | 11.58 | 47.33 / 74.07 | |
| Repo0 | 100.00 | 18.20 | 50.98 / 100.00 | 80.68 | 11.48 | 85.51 / 98.65 | 80.50 | 13.59 | 74.36 / 97.12 | |
| DeepSeek V3.2 | mini-SWE-agent | 86.36 | 15.34 | 21.92 / 47.95 | 59.09 | 23.39 | 2.65 / 53.98 | 33.33 | 9.38 | 10.70 / 47.33 |
| Paper2Code | 90.91 | 11.80 | 4.11 / 56.16 | 14.77 | 5.00 | 49.56 / 61.95 | 62.50 | 43.46 | 7.82 / 53.50 | |
| RPG | 95.45 | 9.23 | 61.64 / 90.41 | 64.70 | 13.70 | 39.29 / 73.57 | 68.75 | 26.50 | 46.50 / 69.55 | |
| Repo0 | 100.00 | 24.77 | 78.08 / 100.00 | 78.41 | 14.10 | 69.03 / 86.46 | 79.17 | 14.29 | 74.07 / 93.83 | |
| Human Developer | Gold Project | 100.00 | – | 94.12 / 100.00 | 100 | – | 94.15 / 100.00 | 100.00 | – | 96.34 / 100.00 |

| Repository | Setting | Cov. (%) | Nov. (%) | Pass./Vot. (%) |
|---|---|---|---|---|
| requests | Repo0 | 100.00 | 18.20 | 50.98 / 100.00 |
| w/o Requirement Context | 95.45 (-4.55) | 9.05 (-9.15) | 45.39 (-5.59) / 87.86 (-12.14) | |
| w/o Component-Graph Ordering | 100.00 (0.00) | 12.85 (-5.35) | 50.98 (0.00) / 100.00 (0.00) | |
| w/o Dual-DAG | 95.45 (-4.55) | 9.14 (-9.06) | 48.72 (-2.26) / 87.86 (-12.14) | |
| w/o Structural Evolution | 94.32 (-5.68) | 8.94 (-9.26) | 42.51 (-8.47) / 82.14 (-17.86) | |
| statsmodels | Repo0 | 80.68 | 11.48 | 85.51 / 98.65 |
| w/o Requirement Context | 67.92 (-12.76) | 11.69 (+0.21) | 85.51 (0.00) / 98.65 (0.00) | |
| w/o Component-Graph Ordering | 80.68 (0.00) | 15.57 (+4.09) | 55.51 (-30.00) / 88.65 (-10.00) | |
| w/o Dual-DAG | 78.55 (-2.13) | 14.66 (+3.18) | 85.51 (0.00) / 95.32 (-3.33) | |
| w/o Structural Evolution | 75.35 (-5.33) | 13.50 (+2.02) | 73.51 (-12.00) / 93.65 (-5.00) | |
| django | Repo0 | 87.50 | 13.59 | 74.36 / 100.00 |
| w/o Requirement Context | 78.29 (-9.21) | 12.10 (-1.49) | 64.36 (-10.00) / 100.00 (0.00) | |
| w/o Component-Graph Ordering | 83.56 (-3.94) | 11.58 (-2.01) | 67.70 (-6.66) / 100.00 (0.00) | |
| w/o Dual-DAG | 82.24 (-5.26) | 12.40 (-1.19) | 64.36 (-10.00) / 93.33 (-6.67) | |
| w/o Structural Evolution | 81.58 (-5.92) | 10.94 (-2.65) | 61.03 (-13.33) / 91.66 (-8.34) |
为什么重要
一个能从零搭建整个项目的代码智能体要真正好用,离不开如何拆分文件和模块的软件设计判断力,这项研究用实验说明这种设计判断必须在编码过程中持续修正,而不能一次性定死。这对想用智能体从需求文档直接搭建新项目脚手架的开发者和工具建设者有直接参考价值。
本文术语
- Dual-DAG · 两张相互关联的图,一张记录需求之间的关系,一张记录实现组件之间的依赖
- DAG(有向无环图) · 由单向箭头连接、且不会绕回起点形成回路的图结构
- 内聚度(Cohesion) · 衡量一个组件内部所聚合的需求彼此关联程度的指标,数值低说明这个组件职责杂乱
- 耦合度(Coupling) · 衡量两个组件所负责的需求重叠程度的指标,数值高意味着可能需要合并
- 测试驱动开发(TDD) · 先写测试代码,再编写实现代码去满足这些测试的开发方式
- 功能覆盖率/通过率 · 衡量生成的仓库覆盖了原需求中多少功能,以及能通过多少参考测试的指标

论文原文摘要(英文)
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.
在 arXiv 阅读最新论文
- LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment在正式微调前先偷看几步训练的梯度,让LoRA的初始化更聪明
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction文本信息缺失或损坏时,这个AI不靠一次性猜测,而是反复修正猜测结果,从而更准确地判断情绪
- Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems要测试访谈式对话系统需要大量不同性格的虚拟用户,这项研究用大语言模型自动生成这些虚拟用户人设
- Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured Partitioning别再机械切分时间序列,按语义把它切成有意义的块
- Reliable Financial Named Entity Recognition under Domain ShiftAI在正式文件里学到的自信,一到推特上就变得不可信
- Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis让AI分析脑影像数据时,把“为什么这个结论可信”也一并记录下来
- GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-Hailing滴滴把打车派单从预测-计算-匹配三段式流程改成一次生成完成,线上效果提升明显
METAL LAB 最新报道
图片来源: Silin Chen et al., arXiv:2608.19854, CC BY 4.0