월~금 오전 7시, 일요일 오전 8시 — AI 뉴스와 용어를 보내드립니다메일로 받아보기

METAL LAB

모델은 그대로 두고 '틀(하네스)'만 자연선택으로 진화시켜 AI 에이전트 성능을 끌어올렸다

arXiv:2608.075452026-07-30

DarwinX: Evolving Agent Harnesses Through Natural Selection

모델은 그대로 두고 '틀(하네스)'만 자연선택으로 진화시켜 AI 에이전트 성능을 끌어올렸다

DarwinX는 LLM 모델 가중치는 전혀 바꾸지 않고, 프롬프트·도구·기술 문서·제어 흐름 같은 '하네스'만 여러 변형(variant)의 집단으로 놓고 자연선택처럼 골라낸다. 기존 성과를 해치지 않으면서 새로운 과제를 풀 수 있는 변형만 살리고, 서로 다른 장점을 가진 계열들을 교배(재조합)시켜 합친다. 터미널 작업, 웹 자동화, 코드 수정 등 네 가지 벤치마크에서 평균 17점 정도의 향상을 보였다고 보고한다.

METAL LAB 해설 도표

DarwinX 선택 루프 구조

증거 상태측정 결과가 보고됨

  1. 고정된 모델GPT-5.5, GPT-5.6, Opus 4.8 등 모델 가중치는 전혀 바꾸지 않고 그대로 둔다
  2. 하네스 변형 생성실패 분석, 교사 시연, 자기 대조 신호로 프롬프트·도구·제어 흐름을 조금씩 편집한 변형을 만든다
  3. 보존-확장 심사기존에 풀던 과제를 크게 해치지 않고 새 과제를 풀 때만 변형을 승격시키는 규칙으로 걸러낸다
  4. 보관소와 재조합승격되지 않은 변형도 보관소에 남겨 서로 다른 강점을 가진 계보를 합쳐 더 나은 자식 변형을 만든다
  5. 네 가지 벤치마크 검증Terminal-Bench 2.1, TerminalWorld, WebArena-Infinity, SWE-bench Verified로 옮겨졌을 때도 성능이 유지되는지 측정한다
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 기존 자기개선 에이전트들은 한 줄기(단일 계보)로만 수정을 이어가다 초반 선택에 갇히거나, 한 과제를 고치면 다른 과제 성능이 몰래 떨어지는 문제가 있었다.
  2. DarwinX는 하네스 변형들을 하나의 집단(population)으로 관리하는 '보관소(archive)'를 두고, 새 변형은 기존에 풀던 문제를 깨지 않으면서 새 문제를 풀 때만 승격시키는 '보존-확장 규약'을 적용한다.
  3. 실패 사례 분석, 정답 시연(교사 신호), 자기 성공/실패 대조라는 세 가지 학습 신호를 하나의 편집 인터페이스로 통합해 하네스를 고친다.
  4. 정답지나 사람이 고른 승자 없이, 각 벤치마크 자체의 채점기(verifier)로 측정한 성공률만으로 우열을 가린다.
  5. Terminal-Bench 2.1에서 기본 83.2%로 +7.7점, 더 강한 모델에서는 검증된 최고 기록 84.7%까지 올랐고, TerminalWorld의 미학습 과제에서도 68.3%로 다른 모든 상용 에이전트를 앞섰다.
Figure 1: With the base model frozen, evolving the harness alone matches or beats the strongest prior agent on four benchmarks. Left: variants survive on measured fitness (avg@k, no gold solutions) and complementary survivors are merged. Right: bars within a panel share one frozen model; the hatched bar is zero-shot transfer of the Terminal-Bench 2.1 harness. y-ranges are truncated and differ per panel.
Figure 1: With the base model frozen, evolving the harness alone matches or beats the strongest prior agent on four benchmarks. Left: variants survive on measured fitness (avg@k, no gold solutions) and complementary survivors are merged. Right: bars within a panel share one frozen model; the hatched bar is zero-shot transfer of the Terminal-Bench 2.1 harness. y-ranges are truncated and differ per panel.
Table 1: Positioning among self-improving-agent methods (✓ present, ∼ restricted, ✗ absent). All share the same inner loop and differ in the selection wrapped around it: Search and Selection group the two failure modes we target, path dependence and cross-task interference. Appendix Table 7 gives the mechanism behind each mark.
EditsSearchSelectionSignal
Methodtools & control flowtools &control flowpopulation archivepopulationarchivecross-lineage mergecross-lineagemergebounded regressionboundedregressionnoise-aware avg@knoise-awareavg@kteacher & self signalsteacher &self signals
tools &
control flow
population
archive
cross-lineage
merge
bounded
regression
noise-aware
avg@k
teacher &
self signals
Optimizers over one designated artifact
OPRO/PromptBreeder/TextGrad∼†
ADAS/AFlow/GPTSwarm
SkillOpt
Agents that edit their own scaffold
SICA
DGM
HarnessX
DarwinX (ours)
Figure 2: DarwinX’s selection loop with the model frozen. Left: the preserve-and-extend contract. Middle: the archive of alternative lineages. Right: shared memory carried across generations.
Figure 2: DarwinX’s selection loop with the model frozen. Left: the preserve-and-extend contract. Middle: the archive of alternative lineages. Right: shared memory carried across generations.
Table 4: WebArena-Infinity per-application audit-clean pass@1 on the official 10-application, 1,260-task real suite; Δ is Monet (DarwinX)’s gain over base Monet. Baseline provenance and raw pre-audit scores are in Appendix D.3.
ApplicationKimiQwenGemini+BUGPT-5.5+BUMonet (base)Monet (DarwinX)Δ
Elation clinical records50.054.281.792.595.896.7+0.9
Elation prescriptions23.341.780.890.820.095.0+75.0
GitLab plan and track39.337.163.677.963.697.9+34.3
Gmail70.056.775.085.025.098.3+73.3
Gmail accounts and contacts40.033.361.787.521.791.7+70.0
Handshake career exploration50.050.550.583.536.584.0+47.5
Linear account settings54.265.873.381.743.394.2+50.9
PayPal wallet70.771.488.690.049.395.7+46.4
Superhuman general15.025.850.080.831.787.5+55.8
Xero invoicing52.555.880.893.339.296.7+57.5
Overall43.348.369.386.143.593.0+49.5
Figure 3: DarwinX’s per-generation operators. Left: the mutation loop and the three learning signals that drive it. Middle: variants classified by how their solved set changes, where those preserving inherited solves stay eligible for recombination while the rest contribute only distilled lessons. Right: the merge operator and its acceptance criterion.
Figure 3: DarwinX’s per-generation operators. Left: the mutation loop and the three learning signals that drive it. Middle: variants classified by how their solved set changes, where those preserving inherited solves stay eligible for recombination while the rest contribute only distilled lessons. Right: the merge operator and its acceptance criterion.
Table 6: TB2.1 skill-bundle diff: the seven skills the evolved lineage adds over base Monet, all in one verification / artifact-contract family. Skills are co-selected, so this attributes composition, not per-skill effect.
Evolved skillsRole
verifier-contract contract-candidateDerive the task’s acceptance contract and check the solution against it before finalizing.
graded-artifact-final-check artifact-verification-loopVerify the graded artifact (output file, format, and values) and iterate a fix-and-recheck loop.
real-tool-artifact tool-grounded-artifactGround outputs in real tool execution rather than asserted or simulated results.
security-contract-repairRepair the solution against security and contract checks.
Figure 7: WebArena-Infinity evolution as optimization. (b) shows the same run as a lineage tree: accepted (blue) and reverted (grey) variants, the primary lineage (gold, base → evolved), and recombination edges (dashed).
Figure 7: WebArena-Infinity evolution as optimization. (b) shows the same run as a lineage tree: accepted (blue) and reverted (grey) variants, the primary lineage (gold, base → evolved), and recombination edges (dashed).
Table 7: Mechanism-level comparison with the closest self-improving-agent systems, expanding Table 1. “Promotion rule” is the evidence a candidate must produce before it is kept; “cross-task interference” is how each method prevents an edit that helps one task family from silently regressing another.
MethodEditable surfaceSearch structurePromotion ruleCross-task interference
Optimizers over one designated artifact
OPRO, PromptBreeder, TextGradInstruction text; tools and control flow stay fixed.Iterative keep-best, or a genetic population with prompt crossover.Scalar score on a fixed development set, or a textual gradient from failures.Not addressed; the search targets a single task or few-shot pool.
ADAS, AFlow, GPTSwarmThe composition graph over otherwise-fixed components.Archive of past workflows, or MCTS over graph edits.Mean accuracy on the target benchmark.One benchmark per search; no per-task preservation check.
SkillOptAn external skill document.Single-lineage keep-best, framed as a gradient-descent analogy.A held-out validation point estimate gates each edit.One domain at a time.
Agents that edit their own scaffold
SICAThe agent’s own source code.A single lineage of self-modification.Benchmark reward on the current coding suite.One coding domain; the authors report an early-edit plateau.
DGMAgent source code.Open-ended archive, stochastic single-parent mutation, no merge operator.Score against the parent on a task subset that grows with confidence.Staged subsets, but no explicit preservation contract.
HarnessX‡A typed harness: prompts, tools, and control flow.Staged single-lineage pipeline; variants are kept isolated from each other.Per-edit gate on the average score plus a seesaw test.Isolation keeps task families apart, so sub-threshold regressions still accumulate.
DarwinX (ours)The full harness: a skill layer (prompts, memory, distilled knowledge) and a code layer (tools, control flow, agent loop).Population archive of typed nodes; parents sampled by cumulative lineage gain; complementary specialists merged into inherited children.Fitness enabler (g>0, R≤δ) adjudicated by a verifier, then avg@k confirmation and a preservation probe before a node may steer search.Specialists are retained and recombined rather than isolated, and the preservation probe bounds what any promotion may cost.
(b) Archive lineage tree (node size ∝ screening score).
(b) Archive lineage tree (node size ∝ screening score).
Table 8: Benchmark-specific base models, evolution and reporting protocols. The base model is frozen throughout every matched comparison, so the harness is the only thing DarwinX changes.
BenchmarkFrozen baseEvolution dataReport dataSelection signalReport metric
TB2.1GPT-5.5∘89 verifier taskssame 89 tasksavg@3 screen, avg@5 confirmavg@5
TerminalWorldOpus 4.8*94 train tasks41 held-out tasksadaptive avg@k subsetspass@1
WAIGPT-5.5300 synthetic intents1,260 real tasksLLM judge, avg@3/avg@5deterministic pass@1
SWE-V (transfer)Opus 4.8none (transfer target)500 issuesn/a (frozen)official pass@1
Figure 8: Invalid trajectories before vs. after evolution, by application (left) and mechanism (right). Evolution cuts invalid trajectories from 293 to 17, leaving only raw-state mutations.
Figure 8: Invalid trajectories before vs. after evolution, by application (left) and mechanism (right). Evolution cuts invalid trajectories from 293 to 17, leaving only raw-state mutations.
Table 10: Construction of the 300-intent WAI evolution split. The benchmark’s own task suites are never read at any stage, so deduplication is intra-pool only.
StageIntentsOperation
Seeds1,080Two models propose intents from each application’s description document
Raw synthesis11,279Dual-model expansion of the seeds under per-application quotas (18 applications)
Filtered8,013Quality heuristics, Jaccard near-duplicate removal, TF–IDF uniqueness pruning
Served apps5,332Restricted to the 12 applications with a served instance
Evolution split300Balanced draw of 25 intents per served application
Figure 9: Per-application success on the real WAI suite, raw pre-audit (left) vs. audit-clean (right). The audit collapses the base’s inflated successes (→43.5%) while leaving Monet (DarwinX) nearly unchanged (→93.0%).
Figure 9: Per-application success on the real WAI suite, raw pre-audit (left) vs. audit-clean (right). The audit collapses the base’s inflated successes (→43.5%) while leaving Monet (DarwinX) nearly unchanged (→93.0%).
Table 11: Regex patterns used to detect invalid state access or mutation.
#Regex (paraphrased)Catches
1/apistate/Any read or write of the /api/state scoring endpoint
2_pushStateToServer | getSerializableState | resetToSeedData | __APP_STATE__ | window.state=State-sync internals
3localStorage/sessionStorage setItem/removeItem/clearDirect storage writes
Table 12: Examples of valid and invalid agent operations.
OperationClassificationReason
Admissible
Click through the UI to add a patient tagVALID_UINormal product operation
Inspect /js/state.js after the app loads or references itVALIDClient-exposed app knowledge
Read AppState.patients in the browser runtimeVALIDRuntime application observation
POST /patients/123/tags with the current user’s sessionVALID_PRODUCT_APIDomain-level app API
AppState.addPatientTag(123, "Flu-Season")VALID_INTERNAL_APP_COMMANDApp-defined semantic mutator
AppState.currentSection = "patients" to navigateVALIDInternal navigation; no scored-state fabrication
Inadmissible
cat <apps>/.../js/state.js via the filesystemINVALID_PRIVILEGED_KNOWLEDGELearns from the co-located source tree
Read .../server.py to discover hidden endpointsINVALID_PRIVILEGED_KNOWLEDGEHost-only server details
Fetch server source through path traversal or unintended static exposureINVALID_EXPLOIT_OR_PRIVILEGE_ESCALATIONBrowser use does not legitimize an unintended disclosure
patient.tags.push("Flu-Season"); AppState.notify()INVALID_RAW_STATE_MUTATIONDirectly edits the scored representation
Replace local storage with hand-edited serialized stateINVALID_RAW_STATE_MUTATIONBypasses application semantics
Send a fabricated full state to /api/stateINVALID_EVALUATION_PLANE_ACCESSUses benchmark synchronization as a control channel
sqlite3 app.db ’UPDATE patients ...’INVALID_DIRECT_DATABASE_MANIPULATIONHost access plus a direct database edit
Inject SQL through a request to bypass authorization or alter rowsINVALID_EXPLOIT_OR_PRIVILEGE_ESCALATIONExploit rather than authorized semantics
Forge an admin token and call an otherwise valid APIINVALID_EXPLOIT_OR_PRIVILEGE_ESCALATIONExceeds the current user’s authority
Table 13: Contract-oriented browser skills added by evolution.
SkillWhen to useCore procedure
web_task_contractGeneral browser/web-UI tasks; durable state change or exact filtered/sorted/read answerDerive an explicit acceptance contract (target, operation, exact final observable state, forbidden substitutions, persistence) → find a valid path, UI first → execute the smallest change → verify both rendered UI and backing state, and reload to confirm persistence → if a target seems missing, prove “not found” from ≥2 independent app surfaces before declaring a no-op.
filtered_list_report_contractCount / latest / oldest / value questions over lists and tablesPreserve the active collection scope (tab, status, search, project, date range) while applying the requested filter; count across the whole scoped set (not just the rendered page); answer with only the requested value.
browser_spa_state_contractDurable state changes where visible controls are missing/ambiguousDerive the exact field-level contract; inspect app-owned stores/reducers/action helpers; mutate through the app’s own action/persistence path; then read back both state and UI.
browser_config_contractDurable configuration records (filters, rules, reminders, routing)Prove every field (condition, action, enabled, timing, channel, persistence), not a partial visible match.
Table 14: The evolved browser prompt replaces an absolute UI-only rule with a bounded semantic fallback and persistence verification.
AspectBase prompt (before)Evolved prompt (after)
Interaction policy“Interact ONLY through the UI…Do NOT write application state directly or touch /api/state.”“Prefer real UI controls first…If a bounded audit proves no visible UI path can satisfy a durable state-changing task, you may inspect app-owned stores, reducers, loaded modules, public helper methods, and readback paths, then use the app’s own exposed action/update helper for the smallest targeted mutation. Do NOT touch /api/state, write local/session storage, use seed/reset helpers, or call state-sync internals.”
Finishing (verification)“For a state-changing task, make the change in the UI, screenshot to confirm, then stop.”“For a state-changing task, verify both app-owned state/readback and the rendered UI; reload or navigate away/back to confirm persistence, then stop.”

실제로 확인된 결과

  • Terminal-Bench 2.1(89개 과제)에서 GPT-5.5 기반 base Monet 75.5%에서 DarwinX 진화 후 83.2%로 +7.7점 상승했고, 더 강한 GPT-5.6 기반에서는 84.7%로 공개 검증 리더보드 최상위권과 동등하거나 앞섰다.
  • TerminalWorld의 94개 학습 과제 이외의 41개 미학습(held-out) 과제에서 Opus 4.8 기반 Monet(DarwinX)이 28/41(68.3%)을 풀어, 평가에 포함된 모든 기성 에이전트보다 높았고 진화 전(25/41) 대비 +7.3점이었다.
  • WebArena-Infinity에서는 합성 의도 300개로만 진화시켰음에도 실제 1,260개 실과제 pass@1이 검증 이후 기준 43.5%에서 93.0%로 상승했다.
  • Terminal-Bench 2.1에서 진화시킨 하네스를 그대로 SWE-bench Verified(500개 이슈)에 적용해도 성능이 이전되었다.
  • 제출물 감사 결과 하네스 자체가 채점기를 속이는 사례는 발견되지 않았고, 무효 궤적 수가 진화 전후 293건에서 17건으로 줄었다.

어디에 쓸 수 있나

  • 코드/터미널 작업 에이전트나 웹 자동화 에이전트의 프롬프트·도구·제어 흐름을 개선하는 파이프라인 설계에 참고할 수 있다.
  • 모델 가중치를 재학습하지 않고 평가용 컴퓨팅 자원을 활용해 에이전트 성능을 끌어올리려는 조직에 적용 가능성이 있다.
  • 자체 검증기(verifier)가 있는 과제 도메인에서 정답 라벨 없이 에이전트 개선 루프를 구축하는 데 참고할 수 있다.

한계와 남은 검증

  • 평가에는 자체 채점기가 있는 벤치마크가 필요하며, 실제 배포 환경에는 이런 검증기가 흔히 없다는 점을 저자들도 인정한다.
  • avg@k 방식은 후보마다 여러 번 반복 실행이 필요해 오프라인 정기 작업으로는 가능하지만 요청 단위 실시간 진화에는 비용이 크다.
  • 모델과 하네스를 동시에 바꾸는 co-evolution, 규정 준수(compliance) 제약을 보존 대상으로 쓰는 일반화된 적용 등은 아직 검증되지 않았고 향후 실험으로 남겨두었다.
  • 서로 다른 모델 세대 간 하네스가 얼마나 유지되는지, 웜 아카이브에서 재선택에 몇 세대가 필요한지는 아직 측정되지 않았다.
  • Terminal-Bench 2.1의 스킬 묶음 기여도 분석은 스킬들이 함께 선택되었기 때문에 개별 스킬의 인과 효과가 아닌 조합 전체의 기여로만 해석해야 한다.

왜 중요한가

모델 가중치를 새로 학습시키지 않고도 프롬프트와 도구 구성만 자연선택식으로 개선해 실질적인 성능 향상을 얻을 수 있다는 것을 보여준다. 모델 교체 주기와 무관하게 재사용 가능한 '하네스'라는 자산을 축적할 수 있다는 뜻이어서, 모델 업그레이드 비용을 줄이는 실무적 함의가 있다.

이 논문의 용어

  • 하네스(harness) · LLM을 감싸는 프롬프트, 도구, 기술 문서, 제어 흐름 등 에이전트의 운영 틀
  • avg@k · 같은 과제를 k번 반복 시도해 평균 성공률을 매기는 측정 방식
  • 보존-확장 규약(preserve-and-extend contract) · 기존에 풀던 과제 성능을 크게 해치지 않으면서 새로운 과제를 풀 때만 변형을 승격시키는 규칙
  • 재조합(recombination) · 서로 다른 장점을 가진 하네스 계보를 합쳐 두 계보의 강점을 모두 가진 자식 변형을 만드는 연산
  • pass@1 · 한 번의 시도로 과제를 성공시켰는지를 보는 지표

저자 · Yifan Zhang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Yifan Zhang et al., arXiv:2608.07545, CC BY 4.0