SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다
연구자들이 화학, 생물학, 물리학 등 20개 과학 분야의 실제 깃허브 저장소 98곳에서 뽑은 119개 수리 과제로 'SWE-bench Science'라는 새 시험지를 만들었다. 가장 성적이 좋았던 AI 코딩 에이전트(Claude Code + Opus-5)조차 정확 성공률(pass@1)이 50%를 넘지 못했다. 연구팀은 왜 실패하는지 네 가지 유형으로 분석하고, 과학적 배경지식을 알려주는 것이 항상 도움이 되는 건 아니라는 점도 확인했다.
무엇을 했나
- 기존 코딩 벤치마크는 단순 함수 작성이나 일반 소프트웨어 버그 수정 위주였는데, 이번 연구는 실제 과학 연구용 소프트웨어(시뮬레이션, 실험 데이터 처리 코드 등)를 고치는 능력을 테스트했다
- 과제를 세 유형으로 나눴다: 알려진 버그를 고치는 '이슈 기반', 원인을 스스로 추리해야 하는 '전문가 탐색형', 여러 모듈을 연결해야 하는 '엔지니어링 통합형'
- GPT-5.6-sol, Claude-Opus-5, DeepSeek-V4-Pro 등 8개 AI 에이전트 조합을 테스트한 결과, 눈에 보이는 공개 테스트는 대부분 90%대로 잘 통과했지만 숨겨진 정밀 검증(비공개 테스트)에서는 성공률이 크게 떨어졌다
- 실패 원인을 지식 부족, 겉핥기식 수리, 통합 누락, 일반화 실패의 네 가지로 분류해 분석했다
- 과학적 배경 설명을 추가로 주는 실험에서, 어떤 모델(DeepSeek-V4-flash)은 성적이 올랐지만 다른 모델(GPT-5.6-sol)은 오히려 정확 성공률이 떨어졌다. 즉 정보를 더 준다고 무조건 더 잘 고치는 게 아니었다

| Scientific domain | Tasks |
|---|---|
| Chemistry | 24 |
| Materials Science and Engineering | 16 |
| Biology | 13 |
| Biomedical Engineering | 12 |
| Physics | 11 |
| Mathematics | 7 |
| Astronomy | 7 |
| Atmospheric Science | 5 |
| Civil Engineering | 5 |
| Surveying and Mapping Science and Technology | 3 |
| Geophysics | 3 |
| Mechanics | 3 |
| Electrical Engineering | 2 |
| Marine Science | 2 |
| Geography | 1 |
| Aeronautical and Astronautical Science and Technology | 1 |
| Nuclear Science and Technology | 1 |
| Computer Science and Technology | 1 |
| Statistics | 1 |
| Information and Communication Engineering | 1 |
| Total | 119 |

왜 중요한가
AI가 과학 논문을 뒷받침하는 코드까지 대신 고쳐주는 시대가 다가오는데, 겉보기 테스트만 통과시키고 실제 과학적 타당성은 놓치는 '눈속임 수리'가 흔하다는 걸 이 연구가 구체적으로 보여준다. 연구 소프트웨어를 다루는 개발자나 AI 도구 개발자 모두 이런 실패 유형을 알아야 더 신뢰할 수 있는 도구를 만들 수 있다.

이 논문의 용어
- pass@1 · AI가 한 번 시도해서 숨겨진 모든 정밀 테스트를 완벽히 통과했는지 보는 엄격한 성공률 지표
- 이슈 기반(Issue-driven) 과제 · 실제 보고된 버그를 그대로 고치는 유형의 과제
- 전문가 탐색형(Expert-exploratory) 과제 · 원인이 알려지지 않은 상태에서 AI가 스스로 원인을 찾아내야 하는 유형
- 엔지니어링 통합형(Engineering-integration) 과제 · 여러 파일과 모듈을 연결해 전체 기능을 완성해야 하는 유형
- 공개/비공개 테스트 · AI가 미리 볼 수 있는 검증용 테스트와, 제출 후에야 적용되는 숨겨진 정밀 검증 테스트

본문에 싣지 못한 그림
- Figure 1
논문 원문 초록 (영문)
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software. We introduce \textbf{SWE-bench Science}, a repository-level benchmark for scientific software engineering comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. Each task is organized into one of three paradigms: Issue-driven, Expert-exploratory, and Engineering-integration. Even the best-performing agent, \textbf{Claude Code with Opus-5 (max), achieves a pass@1 below 50\%}, highlighting the substantial challenges posed by scientific software engineering. We identify four recurring failure mechanisms: deficits in scientific knowledge or abstraction, misguided exploration or surface-level repair, incomplete repair coverage or system integration, and failures to generalize scientific knowledge beyond observed cases in our analysis. We further conduct a paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context. The results show that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success. Together, SWE-bench Science provides a broad testbed for studying both the capabilities and failure mechanisms of coding agents in scientific software engineering.
arXiv에서 원문 보기최신 논문
- Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages이미지 보고 답하는 AI, 시스템 지시(규칙)까지 지키게 하면 성능이 뚝 떨어지고 사용자가 규칙을 어기라고 우기면 더 쉽게 무너진다
- EXIMO: VLM Guided Exploration of VLA Policies로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction텍스트가 빠지거나 망가져도, AI가 그 자리를 대신할 '가짜 텍스트'를 한 번에 만들지 않고 여러 번 고쳐가며 감정을 더 정확히 읽어낸다
- Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder러시아어로 1C 회계 소프트웨어 코드를 찾아주는 첫 검색 벤치마크와 전용 AI 모델이 나왔다
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI 비서가 뭘 기억할지 결정할 때, '물어봐야 할 순간'에 되레 세상에 확인하고 넘어간다
- Reliable Financial Named Entity Recognition under Domain Shift금융 AI가 서류체 문장에서 배운 자신감은 트위터로 가면 거짓말이 된다
- When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models이미지와 상관없는 텍스트 한 줄만 끼워 넣어도 멀티모달 AI 답이 규칙적으로 흔들린다
- Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured Partitioning시계열을 일정 구간으로 자르지 말고, 의미 단위로 잘라서 예측하자는 새 방법
METAL LAB 최신 기사
그림 출처: Zhipeng Xu et al., arXiv:2608.19799, CC BY 4.0