Repo0: Design-Driven Zero-to-All Code Generation
To have an AI build an entire software repository from scratch out of plain-language requirements, the design has to keep changing while the code is being written, not get locked in upfront
Existing code-generation agents are good at filling in code once a repository's file structure and module boundaries are already fixed, but they struggle when given only natural-language requirements and asked to invent the whole repository architecture themselves. Repo0 keeps two linked graphs, one for requirements and one for implementation components, and keeps splitting, merging, and revising the component structure as coding proceeds instead of freezing a blueprint upfront. Across six real-world repositories, it beat the strongest prior baseline by up to 20.08 percentage points in functionality coverage and up to 29.74 points in test pass rate.
What they did
- Prior approaches like RPG analyze the requirements once, produce a repository blueprint, and then generate code strictly following that fixed plan, but real development isn't like that: flawed module boundaries often only become apparent once code is actually being written.
- Repo0 maintains a requirement-level DAG (a graph of how requirement units relate to each other) and a separate component-level DAG (a graph of implementation modules and their dependencies), connected by an alignment relation; the requirement graph is kept relatively fixed while the component graph keeps evolving.
- Using cohesion (whether a component's grouped responsibilities are actually related) and coupling (whether two components overlap in responsibility) as metrics, Repo0 repeatedly applies four structural actions, split, merge, revise, and save, until no more changes are needed and the structure converges.
- After convergence, code is generated with test-driven development: tests are written from the aligned requirements first, then implementation is filled in to pass them, with failures triggering localized repairs.
- Tested on six benchmark repositories adapted from scikit-learn, pandas, sympy, statsmodels, requests, and django using GPT-5 mini and DeepSeek V3.2, Repo0 achieved the highest functionality coverage and pass rate in every setting.

| Real Repo | Para. Name | #Files | LOC | Task Counts |
|---|---|---|---|---|
| scikit-learn | MLKit-Py | 185 | 65,972 | 236 |
| pandas | TableKit | 217 | 106,447 | 175 |
| sympy | SymbolicMath | 699 | 218,924 | 192 |
| statsmodels | StatModeler | 271 | 83,325 | 234 |
| requests | HttpEasy | 17 | 2,793 | 50 |
| django | PyWebEngine | 681 | 109,457 | 165 |

| Model | Method | requests | statsmodels | django | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Cov. (%) | Nov. (%) | Pass./Vot. (%) | Cov. (%) | Nov. (%) | Pass./Vot. (%) | Cov. (%) | Nov. (%) | Pass./Vot. (%) | ||
| GPT-5 mini | mini-SWE-agent | 68.18 | 2.63 | 4.11 / 27.40 | 18.18 | 9.09 | 0.00 / 31.86 | 47.92 | 6.96 | 37.04 / 44.44 |
| Paper2Code | 95.50 | 7.20 | 24.66 / 24.66 | 44.32 | 24.13 | 4.42 / 30.09 | 66.67 | 15.30 | 30.04 / 78.60 | |
| RPG | 90.91 | 13.70 | 31.51 / 95.89 | 70.40 | 13.80 | 77.90 / 92.00 | 60.42 | 11.58 | 47.33 / 74.07 | |
| Repo0 | 100.00 | 18.20 | 50.98 / 100.00 | 80.68 | 11.48 | 85.51 / 98.65 | 80.50 | 13.59 | 74.36 / 97.12 | |
| DeepSeek V3.2 | mini-SWE-agent | 86.36 | 15.34 | 21.92 / 47.95 | 59.09 | 23.39 | 2.65 / 53.98 | 33.33 | 9.38 | 10.70 / 47.33 |
| Paper2Code | 90.91 | 11.80 | 4.11 / 56.16 | 14.77 | 5.00 | 49.56 / 61.95 | 62.50 | 43.46 | 7.82 / 53.50 | |
| RPG | 95.45 | 9.23 | 61.64 / 90.41 | 64.70 | 13.70 | 39.29 / 73.57 | 68.75 | 26.50 | 46.50 / 69.55 | |
| Repo0 | 100.00 | 24.77 | 78.08 / 100.00 | 78.41 | 14.10 | 69.03 / 86.46 | 79.17 | 14.29 | 74.07 / 93.83 | |
| Human Developer | Gold Project | 100.00 | – | 94.12 / 100.00 | 100 | – | 94.15 / 100.00 | 100.00 | – | 96.34 / 100.00 |

| Repository | Setting | Cov. (%) | Nov. (%) | Pass./Vot. (%) |
|---|---|---|---|---|
| requests | Repo0 | 100.00 | 18.20 | 50.98 / 100.00 |
| w/o Requirement Context | 95.45 (-4.55) | 9.05 (-9.15) | 45.39 (-5.59) / 87.86 (-12.14) | |
| w/o Component-Graph Ordering | 100.00 (0.00) | 12.85 (-5.35) | 50.98 (0.00) / 100.00 (0.00) | |
| w/o Dual-DAG | 95.45 (-4.55) | 9.14 (-9.06) | 48.72 (-2.26) / 87.86 (-12.14) | |
| w/o Structural Evolution | 94.32 (-5.68) | 8.94 (-9.26) | 42.51 (-8.47) / 82.14 (-17.86) | |
| statsmodels | Repo0 | 80.68 | 11.48 | 85.51 / 98.65 |
| w/o Requirement Context | 67.92 (-12.76) | 11.69 (+0.21) | 85.51 (0.00) / 98.65 (0.00) | |
| w/o Component-Graph Ordering | 80.68 (0.00) | 15.57 (+4.09) | 55.51 (-30.00) / 88.65 (-10.00) | |
| w/o Dual-DAG | 78.55 (-2.13) | 14.66 (+3.18) | 85.51 (0.00) / 95.32 (-3.33) | |
| w/o Structural Evolution | 75.35 (-5.33) | 13.50 (+2.02) | 73.51 (-12.00) / 93.65 (-5.00) | |
| django | Repo0 | 87.50 | 13.59 | 74.36 / 100.00 |
| w/o Requirement Context | 78.29 (-9.21) | 12.10 (-1.49) | 64.36 (-10.00) / 100.00 (0.00) | |
| w/o Component-Graph Ordering | 83.56 (-3.94) | 11.58 (-2.01) | 67.70 (-6.66) / 100.00 (0.00) | |
| w/o Dual-DAG | 82.24 (-5.26) | 12.40 (-1.19) | 64.36 (-10.00) / 93.33 (-6.67) | |
| w/o Structural Evolution | 81.58 (-5.92) | 10.94 (-2.65) | 61.03 (-13.33) / 91.66 (-8.34) |
Why it matters
For an AI agent that builds a whole project from scratch to be genuinely useful, it needs real software-design judgment about how to split files and modules, and this work provides evidence that such design must keep being revised during coding rather than fixed once at the start. It's directly relevant to anyone building developer tools or agents intended to scaffold new projects from requirements.
Terms in this paper
- Dual-DAG · a pair of linked graphs, one tracking relationships among requirements and one tracking dependencies among implementation components
- DAG (Directed Acyclic Graph) · a graph made of one-way connections that never loop back to where they started
- Cohesion · a measure of how related the requirements grouped inside one component actually are; low cohesion means the component is doing unrelated things
- Coupling · a measure of how much two components' responsibilities overlap; high coupling flags candidates for merging
- Test-driven development (TDD) · writing tests before the implementation, then writing code to make those tests pass
- Functionality Coverage / Pass Rate · metrics for how much of the required functionality the generated repository actually has, and how many reference tests it passes

Original abstract (English)
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.
Read on arXivLatest papers
- Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platformsTreating data-platform changes like reviewable spec snippets instead of code diffs: an experiment design paper
- Are LLMs becoming similarly creative? Evidence from three years of modelsNewer AI chatbots are giving increasingly similar answers to each other, three years of data show
- Auditing Cross-Lingual Fairness in Language Model WatermarkingAI text watermarks that are supposed to catch machine-written content work far less reliably in many non-English languages, and the gap tracks language families, not individual languages
- TESTNAV: Pareto-Guided Search for Compositional Robustness TestingA smarter way to test AI models against combined real-world glitches, without checking every possible combination
- Optimal Skill Selection for LLM Agents with Provable Bicriteria GuaranteesA method that picks which 'skill documents' to feed an AI coding agent, with mathematically guaranteed near-optimal results
- Reliable Financial Named Entity Recognition under Domain ShiftAn AI's confidence trained on formal filings turns unreliable once it reads tweets
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy CorrectionWhen text input is missing or broken, this AI doesn't guess once and move on—it revises its guess step by step to read emotions more reliably
Latest from METAL LAB
- Google Discover adds chatbot that adjusts your feed based on spoken preferences
- OpenAI Closes In on Anthropic Again in Enterprise Spending Share
- Meta Unveils First 10 Tasks in WildArtifactBench, a Benchmark for AI Agents
- Musk: "Optimus + Grok will one day handle healthcare for all humanity"
- 35% of Web Pages Published Since ChatGPT Show Signs of AI Authorship
Figures: Silin Chen et al., arXiv:2608.19854, CC BY 4.0