Repo0: Design-Driven Zero-to-All Code Generation
To have an AI build an entire software repository from scratch out of plain-language requirements, the design has to keep changing while the code is being written, not get locked in upfront
Existing code-generation agents are good at filling in code once a repository's file structure and module boundaries are already fixed, but they struggle when given only natural-language requirements and asked to invent the whole repository architecture themselves. Repo0 keeps two linked graphs, one for requirements and one for implementation components, and keeps splitting, merging, and revising the component structure as coding proceeds instead of freezing a blueprint upfront. Across six real-world repositories, it beat the strongest prior baseline by up to 20.08 percentage points in functionality coverage and up to 29.74 points in test pass rate.
What they did
- Prior approaches like RPG analyze the requirements once, produce a repository blueprint, and then generate code strictly following that fixed plan, but real development isn't like that: flawed module boundaries often only become apparent once code is actually being written.
- Repo0 maintains a requirement-level DAG (a graph of how requirement units relate to each other) and a separate component-level DAG (a graph of implementation modules and their dependencies), connected by an alignment relation; the requirement graph is kept relatively fixed while the component graph keeps evolving.
- Using cohesion (whether a component's grouped responsibilities are actually related) and coupling (whether two components overlap in responsibility) as metrics, Repo0 repeatedly applies four structural actions, split, merge, revise, and save, until no more changes are needed and the structure converges.
- After convergence, code is generated with test-driven development: tests are written from the aligned requirements first, then implementation is filled in to pass them, with failures triggering localized repairs.
- Tested on six benchmark repositories adapted from scikit-learn, pandas, sympy, statsmodels, requests, and django using GPT-5 mini and DeepSeek V3.2, Repo0 achieved the highest functionality coverage and pass rate in every setting.

| Real Repo | Para. Name | #Files | LOC | Task Counts |
|---|---|---|---|---|
| scikit-learn | MLKit-Py | 185 | 65,972 | 236 |
| pandas | TableKit | 217 | 106,447 | 175 |
| sympy | SymbolicMath | 699 | 218,924 | 192 |
| statsmodels | StatModeler | 271 | 83,325 | 234 |
| requests | HttpEasy | 17 | 2,793 | 50 |
| django | PyWebEngine | 681 | 109,457 | 165 |

| Model | Method | requests | statsmodels | django | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Cov. (%) | Nov. (%) | Pass./Vot. (%) | Cov. (%) | Nov. (%) | Pass./Vot. (%) | Cov. (%) | Nov. (%) | Pass./Vot. (%) | ||
| GPT-5 mini | mini-SWE-agent | 68.18 | 2.63 | 4.11 / 27.40 | 18.18 | 9.09 | 0.00 / 31.86 | 47.92 | 6.96 | 37.04 / 44.44 |
| Paper2Code | 95.50 | 7.20 | 24.66 / 24.66 | 44.32 | 24.13 | 4.42 / 30.09 | 66.67 | 15.30 | 30.04 / 78.60 | |
| RPG | 90.91 | 13.70 | 31.51 / 95.89 | 70.40 | 13.80 | 77.90 / 92.00 | 60.42 | 11.58 | 47.33 / 74.07 | |
| Repo0 | 100.00 | 18.20 | 50.98 / 100.00 | 80.68 | 11.48 | 85.51 / 98.65 | 80.50 | 13.59 | 74.36 / 97.12 | |
| DeepSeek V3.2 | mini-SWE-agent | 86.36 | 15.34 | 21.92 / 47.95 | 59.09 | 23.39 | 2.65 / 53.98 | 33.33 | 9.38 | 10.70 / 47.33 |
| Paper2Code | 90.91 | 11.80 | 4.11 / 56.16 | 14.77 | 5.00 | 49.56 / 61.95 | 62.50 | 43.46 | 7.82 / 53.50 | |
| RPG | 95.45 | 9.23 | 61.64 / 90.41 | 64.70 | 13.70 | 39.29 / 73.57 | 68.75 | 26.50 | 46.50 / 69.55 | |
| Repo0 | 100.00 | 24.77 | 78.08 / 100.00 | 78.41 | 14.10 | 69.03 / 86.46 | 79.17 | 14.29 | 74.07 / 93.83 | |
| Human Developer | Gold Project | 100.00 | – | 94.12 / 100.00 | 100 | – | 94.15 / 100.00 | 100.00 | – | 96.34 / 100.00 |

| Repository | Setting | Cov. (%) | Nov. (%) | Pass./Vot. (%) |
|---|---|---|---|---|
| requests | Repo0 | 100.00 | 18.20 | 50.98 / 100.00 |
| w/o Requirement Context | 95.45 (-4.55) | 9.05 (-9.15) | 45.39 (-5.59) / 87.86 (-12.14) | |
| w/o Component-Graph Ordering | 100.00 (0.00) | 12.85 (-5.35) | 50.98 (0.00) / 100.00 (0.00) | |
| w/o Dual-DAG | 95.45 (-4.55) | 9.14 (-9.06) | 48.72 (-2.26) / 87.86 (-12.14) | |
| w/o Structural Evolution | 94.32 (-5.68) | 8.94 (-9.26) | 42.51 (-8.47) / 82.14 (-17.86) | |
| statsmodels | Repo0 | 80.68 | 11.48 | 85.51 / 98.65 |
| w/o Requirement Context | 67.92 (-12.76) | 11.69 (+0.21) | 85.51 (0.00) / 98.65 (0.00) | |
| w/o Component-Graph Ordering | 80.68 (0.00) | 15.57 (+4.09) | 55.51 (-30.00) / 88.65 (-10.00) | |
| w/o Dual-DAG | 78.55 (-2.13) | 14.66 (+3.18) | 85.51 (0.00) / 95.32 (-3.33) | |
| w/o Structural Evolution | 75.35 (-5.33) | 13.50 (+2.02) | 73.51 (-12.00) / 93.65 (-5.00) | |
| django | Repo0 | 87.50 | 13.59 | 74.36 / 100.00 |
| w/o Requirement Context | 78.29 (-9.21) | 12.10 (-1.49) | 64.36 (-10.00) / 100.00 (0.00) | |
| w/o Component-Graph Ordering | 83.56 (-3.94) | 11.58 (-2.01) | 67.70 (-6.66) / 100.00 (0.00) | |
| w/o Dual-DAG | 82.24 (-5.26) | 12.40 (-1.19) | 64.36 (-10.00) / 93.33 (-6.67) | |
| w/o Structural Evolution | 81.58 (-5.92) | 10.94 (-2.65) | 61.03 (-13.33) / 91.66 (-8.34) |
Why it matters
For an AI agent that builds a whole project from scratch to be genuinely useful, it needs real software-design judgment about how to split files and modules, and this work provides evidence that such design must keep being revised during coding rather than fixed once at the start. It's directly relevant to anyone building developer tools or agents intended to scaffold new projects from requirements.
Terms in this paper
- Dual-DAG · a pair of linked graphs, one tracking relationships among requirements and one tracking dependencies among implementation components
- DAG (Directed Acyclic Graph) · a graph made of one-way connections that never loop back to where they started
- Cohesion · a measure of how related the requirements grouped inside one component actually are; low cohesion means the component is doing unrelated things
- Coupling · a measure of how much two components' responsibilities overlap; high coupling flags candidates for merging
- Test-driven development (TDD) · writing tests before the implementation, then writing code to make those tests pass
- Functionality Coverage / Pass Rate · metrics for how much of the required functionality the generated repository actually has, and how many reference tests it passes

Original abstract (English)
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.
Read on arXivLatest papers
- LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive AlignmentPeeking at a few early training gradients before fine-tuning starts to set up LoRA smarter
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI assistants would rather double-check facts than ask you a question, even when asking is the right call
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy CorrectionWhen text input is missing or broken, this AI doesn't guess once and move on—it revises its guess step by step to read emotions more reliably
- Generating Diverse Personas for User Simulators to Test Interview Dialogue SystemsTo test interview-style chatbots you need many different fake users, so this work has an LLM automatically generate those fake user personalities
- Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured PartitioningA new way to slice time series into meaningful chunks instead of arbitrary equal-length pieces
- Reliable Financial Named Entity Recognition under Domain ShiftAn AI's confidence trained on formal filings turns unreliable once it reads tweets
- Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysisA system that makes AI show its work when analyzing brain-imaging data, not just deliver an answer
- GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-HailingDiDi replaced its multi-step ride-hailing dispatch pipeline with one generative model and saw real-world gains
Latest from METAL LAB
- Google Discover adds chatbot that adjusts your feed based on spoken preferences
- OpenAI Closes In on Anthropic Again in Enterprise Spending Share
- Meta Unveils First 10 Tasks in WildArtifactBench, a Benchmark for AI Agents
- Musk: "Optimus + Grok will one day handle healthcare for all humanity"
- 35% of Web Pages Published Since ChatGPT Show Signs of AI Authorship
Figures: Silin Chen et al., arXiv:2608.19854, CC BY 4.0