METAL for iPhone

AI 뉴스, 이제 앱에서 읽으세요.

METAL 앱을 다운로드하고 매일 새로운 AI 기사를 만나보세요.

App Store에서 다운로드

iPhone용 앱 · 무료 다운로드

iPhone의 App Store에서도 ‘메탈 AI 매거진’을 검색할 수 있습니다.

METAL

클로드, 신약 결합 단백질 설계에서 업계 평균 웃도는 성공률 기록

앤스로픽이 15개 표적 중 14개에 대해 클로드가 자율 설계한 단백질 바인더를 외부 실험실에서 검증했다

클로드, 신약 결합 단백질 설계에서 업계 평균 웃도는 성공률 기록 · Image: METAL LAB

요약

  • 클로드가 인간 전문가의 프롬프트 하나로 15개 표적 단백질 중 14개에 결합 분자를 자율 설계했다
  • Adaptyv Bio와 Twist Bioscience가 독립 제작·검증한 결과 22~35%가 실제로 결합에 성공해 업계 평균 10~15%를 넘어섰다
  • 앤스로픽은 이를 발판 삼아 항체부터 저분자화합물까지 신약 개발 전 과정을 자동화하는 작업을 이어가고 있다
원문 영상
설계 성공 표적 수
15개 표적 중 14개에 결합 단백질 설계
검증 파트너
Adaptyv Bio, Twist Bioscience (독립 제작·검증)
업계 평균 성공률
10~15%
클로드 설계 성공률
구성에 따라 22~35%
사용 모델
Opus 4.8, Mythos Preview
대표 결합력 사례
EGFR 1.7pM, VEGF-A 1.6pM, TREM2 1.1pM (Mythos Preview 설계)
생명과학 연구용 최상위 모델
Opus 5

클로드가 설계한 단백질, 실험실에서 검증받다

앤스로픽(Anthropic)이 자사 모델 클로드(Claude)에게 신약 개발의 첫 관문인 '단백질 바인더 설계'를 맡기고 그 결과를 X 게시물을 통해 공개했다. 사람이 작성한 단백질 설계 프롬프트 하나로, 클로드는 15개 표적 단백질 중 14개에 대해 결합 가능한 새 단백질을 처음부터(de novo) 설계했다. 이렇게 나온 설계는 앤스로픽이 직접 검증하지 않았다. 제3의 실험실인 Adaptyv Bio와 Twist Bioscience가 독립적으로 단백질을 합성하고 실제 결합 여부를 테스트했다.

단백질 바인더가 신약 개발의 첫 관문인 이유

약물 대부분은 몸속 특정 표적에 달라붙어 그 기능을 막거나 바꾸는 방식으로 작동한다. 이 표적에 딱 맞게 달라붙는 분자를 설계하는 일이 신약 개발의 출발점인데, 지금까지는 표적 하나당 전문가가 몇 주에서 몇 달을 들여 수많은 후보를 걸러내야 했다. 단백질 바인더 설계는 실제 약물 설계보다는 쉬운 과제지만, AI가 이 단계를 얼마나 잘 해내는지를 가늠하는 유용한 시험대로 쓰인다. 지금 이 분야에서 사람이 설계한 바인더가 실제로 결합에 성공하는 비율은 대개 10~15% 수준으로 알려져 있다.

여러 단백질 구조 모형이 3행 3열 배열로 배치된 모습
이미지: @AnthropicAI (X)

클로드에게 실제로 준 지시 — 48시간 자율 캠페인

이번 공개에서 성공률만큼 눈여겨볼 대목은 클로드가 이 설계를 '어떻게' 해냈느냐다. 사람이 표적마다 후보를 돌린 게 아니라, 클로드가 최상위 오케스트레이터(orchestrator) — 여러 하위 에이전트를 부리는 지휘자 — 로 서서 캠페인 전체를 스스로 굴렸다. 콜센터에 비유하면, 사람이 상담원을 한 명씩 지시하는 대신 팀장 AI 하나에게 "48시간 안에 끝내라"고 던지고 방을 나간 셈이다.

앤스로픽이 클로드에게 준 시작 지시(kickoff)는 짧고 단호하다. 14개 표적을 한 번에 다루는 다중 표적 캠페인의 킥오프 프롬프트 전문은 이렇다.

Execute the 48-hour, $50,000 de novo miniprotein binder design campaign exactly as specified in the campaign prompt. The prompt is in your system context and is also attached to this message as a markdown file (the two are the same document; the system-context copy is authoritative and is what every sub-agent at every depth carries). The two figures the prompt references (Figure 1 and Figure 2, from corpus folder "06 Prompt Figures") are also attached. You are the top-level orchestrator. Do not ask me any questions or wait for approval.

This is a FRESH campaign in a FRESH project, in a workspace newly created for this run. It is not a resume of any prior run; at kickoff there is no prior campaign state of yours to reconcile against. The Slack channel and the Drive deliverables folder are shared with other independent campaigns and the prompt's Logistics and Isolation sections govern how to treat them. Any Modal volumes that already exist on the account belong to other campaigns and the same rules apply. compute_details for byoc:modal starts empty in this workspace; keep it that way (do not write to compute_details at any point; record your own environment findings as artifacts in this project instead).

A setup session was run in this workspace immediately before this message to pre-approve network domains and verify connectors read-only. It left NO persistent campaign state (no campaign volumes, no governor file, no Slack posts, no Drive activity). It is not part of your campaign and will not appear in your project's frames table or artifact store.

Clock: T0 is the timestamp of this message. End time is T0 + 48 hours.

Your first actions, in this order: (a) determine your model identifier from host.current_model() and your orchestrator root frame_id; (b) dispatch the SETUP sub-agent via host.delegate and BLOCK on its return: SETUP (and only SETUP) creates the four campaign Modal volumes named campaign-<model>-<YYYY-MM-DD>-<frameid8>-{state,ledger,out,novelty}, seeds /state/governor.json with {ceiling:325, basis:"BOOTSTRAP", set_at:T0_utc}, writes /state/lib/submit_gate.py, and returns {status:OK, gate_sha256, governor_sha256}; (c) immediately after SETUP returns OK, call host.compute.set_concurrency_limit(325) once; (d) dispatch the CLOCK long-lived singleton; (e) post your kickoff message as a NEW top-level message in the Slack channel which starts your campaign thread (every subsequent post is a reply in that thread); (f) create exactly ONE new sub-folder inside the shared Drive deliverables folder, named "Campaign <model> <YYYY-MM-DD> <frameid8>", and put all your Drive deliverables there. State your model identifier, T0 and the campaign end time (ISO-8601 UTC), your thread permalink, your sub-folder name, and your four volume names in the kickoff post.

Begin now. Good luck!

표적 하나에 24시간을 쓰는 단일 표적 캠페인은 예산과 시계만 다르고 구조가 같다. 킥오프 프롬프트 전문은 이렇다.

Execute the 24-hour, $10,000 de novo miniprotein binder design campaign exactly as specified in the campaign prompt. The prompt is in your system context and is also attached to this message as a markdown file (the two are the same document; the system-context copy is authoritative and is what every sub-agent at every depth carries). The two figures the prompt references (Figure 1 and Figure 2, from corpus folder "06 Prompt Figures") are also attached. You are the top-level orchestrator. Do not ask me any questions or wait for approval.

This is a FRESH campaign in a FRESH project, in a workspace newly created for this run. It is not a resume of any prior run; at kickoff there is no prior campaign state of yours to reconcile against. The Slack channel, the Drive deliverables folder, and the Modal account are shared with other independent campaigns — including several concurrent single-target campaigns against other targets that started at or near your T0 — and the prompt's Logistics and Isolation sections govern how to treat them. You will observe those campaigns' Modal volumes, apps, and running sandboxes, their Slack threads, and their Drive sub-folders: do not read, write, delete, terminate, or post into any of them, and do not treat their existence as an anomaly to report or reconcile. Your governor and submit_gate() count live GPU sandboxes filtered by YOUR project_tag only (per the prompt); account-wide GPU load outside that tag is expected and is never a reason to throttle, halt, or raise WATCHDOG. Any Modal volumes that already exist on the account belong to other campaigns and the same rules apply. compute_details for byoc:modal starts empty in this workspace; keep it that way (do not write to compute_details at any point; record your own environment findings as artifacts in this project instead).

A setup session was run in this workspace immediately before this message to pre-approve network domains and verify connectors read-only. It left NO persistent campaign state (no campaign volumes, no governor file, no Slack posts, no Drive activity). It is not part of your campaign and will not appear in your project's frames table or artifact store.

Clock: T0 is the timestamp of this message. End time is T0 + 24 hours.

Your first actions, in this order: (a) determine your model identifier from host.current_model(), your orchestrator root frame_id, and your <target> string as the filename stem of the attached campaign-prompt markdown file (e.g. "TREM2", "GDF-8", "Cas9" — use it verbatim in Slack headers and the Drive folder name; lowercase it for the Modal volume-name slug); (b) dispatch the SETUP sub-agent via host.delegate and BLOCK on its return: SETUP (and only SETUP) creates the four campaign Modal volumes named campaign-<target>-<model>-<YYYY-MM-DD>-<frameid8>-{state,ledger,out,novelty}, seeds /state/governor.json with {ceiling:150, basis:"BOOTSTRAP", set_at:T0_utc}, and writes /state/lib/submit_gate.py, returning {status:OK, gate_sha256, governor_sha256}; (c) dispatch the CLOCK long-lived singleton; (d) post your kickoff message as a NEW top-level message in the Slack channel with header "Campaign Kickoff (<target>): <model> <YYYY-MM-DD> <frameid8>", which starts your campaign thread (every subsequent post is a reply in that thread); (e) create exactly ONE new sub-folder inside the shared Drive deliverables folder, named "Campaign <target> <model> <YYYY-MM-DD> <frameid8>", and put all your Drive deliverables there. State your model identifier, your target, T0 and the campaign end time (ISO-8601 UTC), your thread permalink, your sub-folder name, and your four volume names in the kickoff post.

Begin now. Good luck!

지시문을 뜯어보면 클로드가 받은 자율성의 폭이 드러난다. 시작과 동시에 스스로 자기 모델 식별자를 확인하고, SETUP 하위 에이전트를 띄워 작업용 저장 공간과 제출 게이트(제 손으로 만든 필터)를 세팅한 뒤, 시계 역할의 CLOCK 에이전트를 상시로 돌리고, 슬랙 채널에 캠페인 스레드를 열어 진행 상황을 스스로 보고했다. "나에게 아무것도 묻지 말고 승인도 기다리지 말라"는 문장이 이 실험의 성격을 요약한다. 동시에 도는 다른 캠페인의 자원은 건드리지 말라는 격리 규칙, 한 번에 돌릴 수 있는 작업 수를 스스로 제한하는 거버너(governor) 상한선까지 지시에 박혀 있다.

성공률로 보는 클로드의 설계 능력

앤스로픽이 공개한 수치에 따르면 클로드의 설계는 구성에 따라 22%에서 35% 사이의 성공률을 보였다. 15개 표적 전체를 합친 결과에서 Opus 4.8은 다중 표적 방식으로 390건 중 88건(약 22.6%), 프리뷰 모델인 Mythos Preview는 같은 방식으로 390건 중 104건(약 26.7%)이 결합에 성공했다. 표적 하나에 집중한 단일 표적 방식에서는 Mythos Preview가 450건 중 158건, 약 35.1%까지 성공률을 끌어올렸다.

구성성공률막대
업계 평균(기존 방식)10~15%12
Opus 4.8 (다중 표적)22.6%23
Mythos Preview (다중 표적)26.7%27
Mythos Preview (단일 표적)35.1%35

표적별로 보면 편차가 컸다. TREM2는 세 가지 구성 모두 76~83%에 달하는 높은 성공률을 보인 반면 15-PGDH나 MBP 같은 표적은 0~1% 수준에 머물렀다. 결합력을 나타내는 지표인 Kd값(낮을수록 강하게 결합)에서도 Mythos Preview가 설계한 EGFR 결합체는 1.7pM, VEGF-A는 1.6pM, TREM2는 1.1pM을 기록해 기존에 발표된 최고 성능의 de novo 바인더보다 여러 배 강하게 결합한 사례도 나왔다.

아직 넘어야 할 산

앤스로픽 스스로도 이 결과에 선을 그었다. 단백질 바인더는 약이 아니다. 표적에 강하게 달라붙는 분자를 만드는 건 약물 후보 물질을 개발하는 과정의 첫 단계일 뿐이고, 그 약물 후보가 사람에게 안전하고 효과적이라는 걸 입증하기까지는 훨씬 더 많은 단계가 남아 있다. 앤스로픽은 이번 결과를 발판 삼아 항체부터 저분자화합물까지 신약 개발 전 과정을 클로드가 처음부터 끝까지 수행하도록 훈련시키고 있다고 밝혔다. 과학자들이 자사의 가장 뛰어난 모델을 쓸 수 있는 접근 프로그램도 조만간 공개하겠다고 예고했으며, 생명과학 연구용으로는 Opus 5가 현재 가장 뛰어난 모델이라고 설명했다. 앤스로픽은 실험 프롬프트와 데이터도 함께 오픈소스로 공개했다.

여러 단백질 타깃별 히트율을 비교한 막대그래프 두 개가 세로로 배치됨
이미지: @AnthropicAI (X)

허깅페이스에 통째로 올라온 원본

앤스로픽은 트윗과 성공률 표만 낸 게 아니라, 설계·측정·구조 데이터 전량을 허깅페이스에 데이터셋으로 올렸다(라이선스 CC BY 4.0). 그동안 허깅페이스에서 모델 파일만 좇느라 놓치기 쉬웠던, 기업이 통째로 공개한 1차 자료다.

무엇이 들어 있는지만 봐도 규모가 가늠된다. 설계 결과를 정리한 표(parquet) 20개, 예측기 10종에 시드 5개를 곱해 만든 단백질 구조 모델 113,550개, 그리고 앞서 본 킥오프와 함께 표적별 단일 프롬프트 16개, 다중 표적 캠페인 프롬프트 전문이 함께 실렸다. 표들은 전부 고유 식별자(uuid)로 연결돼 어떤 모델이 어떤 표적의 몇 번째 후보를 냈고 실험에서 어떻게 나왔는지까지 따라갈 수 있다.

표적도 뭉뚱그리지 않았다. 다중 표적 캠페인이 다룬 14개만 봐도 EGFR·PD-L1·IL-7Rα 같은 항암·면역 표적, TNF-α(자가면역), VEGF-A(혈관), TREM2·TrkA(신경), 유전자 가위 SpCas9, 니파 바이러스 외피 단백질까지 걸쳐 있다. 근육 성장을 억제하는 GDF-8(마이오스타틴)에는 사촌 격인 GDF-11에는 붙지 말아야 한다는 선택성 조건까지 프롬프트에 못 박아 뒀다 — 약이 엉뚱한 표적을 건드리면 안 되는 실제 신약 개발의 까다로움이 그대로 과제에 들어간 것이다.

설계에 쓴 도구도 열려 있다. 클로드는 RFdiffusion·BindCraft·AlphaProteo·BoltzGen처럼 이미 공개된 단백질 설계·구조 예측 모델을 상업 라이선스 범위 안에서 골라 조합했다. AI가 밑바닥부터 새 방법을 발명한 게 아니라, 흩어진 오픈소스 도구를 스스로 엮어 파이프라인을 짠 셈이다.

에디터의 시선

이번 발표에서 가장 눈에 띄는 대목은 성공률 수치 자체보다 '독립 검증'이라는 절차다. 앤스로픽이 자체적으로 결합률을 측정했다면 신뢰하기 어려웠을 결과지만, Adaptyv Bio와 Twist Bioscience라는 제3의 실험실이 직접 단백질을 합성하고 테스트했다는 점에서 이 수치는 실험실 바깥 검증을 거친 셈이다. AI 모델 회사가 자체 벤치마크로 우수성을 주장하는 일이 흔한 요즘, 생물학처럼 물리적 실험이 반드시 필요한 영역에서는 이런 외부 검증 절차가 신뢰의 기준이 될 가능성이 크다.

세대 비교로 보면 흥미로운 지점이 하나 더 있다. 표에 등장한 'Mythos Preview'라는 이름의 모델이 여러 표적에서 기존 Opus 4.8보다 높은 성공률과 더 강한 결합력을 보였다는 점이다. 아직 정식으로 소개되지 않은 프리뷰 모델로 보이는데, 다음 세대 클로드가 생명과학 영역에서 어떤 성능을 보일지 가늠할 수 있는 단서로 읽힌다. AI 모델이 코딩이나 글쓰기를 넘어 실험실 워크플로에 붙는 사례를 계속 지켜보면, 매번 '사람이 몇 주 걸리던 일을 몇 시간으로 줄였다'는 식의 서사가 반복된다. 이번에도 그 흐름과 크게 다르지 않다.

다만 실무적으로 냉정해질 필요가 있다. 국내 바이오·제약 업계가 당장 이 결과를 보고 AI로 신약 후보를 뽑아내겠다고 나서는 건 이르다. 단백질 바인더 설계는 신약 개발의 수십 단계 중 첫 단계에 불과하고, 안전성·효능 검증까지 가려면 이번 실험과는 비교할 수 없는 시간과 비용이 든다. 지금 시점에서 실무진이 할 수 있는 건 이런 AI 설계 파이프라인을 초기 후보 물질 스크리닝 단계에 보조 도구로 검토하는 정도이지, 전체 프로세스를 대체할 단계는 아니다.

앞으로 몇 달 안에는 앤스로픽이 예고한 과학자용 접근 프로그램이 구체화될 가능성이 크고, 항체·저분자화합물 등 다른 약물 유형으로 실험 범위를 넓힌 후속 결과도 나올 것으로 보인다. 생명과학 분야에서 AI 모델 간 경쟁이 코딩·이미지 생성 못지않게 치열해지는 흐름을 이번 발표가 앞당길 것이다.

김현국

METAL 발행인 · 아카이브저널 발행인 · 이라선 운영

매거진·서점·교육을 오래 해 온 편집자예요. 라이프스타일 매거진 <아카이브저널>을 발행하고, 아트 서점 <이라선>을 운영했으며, 라이카 아카데미에서 디렉터로 사진 교육 프로그램을 만들고 이끌었어요. 그 시선으로 전 세계 AI 소식을 독자에게 전하는 메탈(METAL) 매거진 편집장입니다.

이 에디터의 기사 더 보기 →

공유

댓글