
摘要
- Claude仅凭人类专家撰写的一段提示词,就自主设计出了针对15个靶点蛋白质中14个的结合分子
- 经Adaptyv Bio和Twist Bioscience独立制备并验证,其中22%至35%实际成功结合,超过了10%至15%的行业平均水平
- Anthropic正以此为基础,推进从抗体到小分子化合物的新药研发全流程自动化工作
Claude设计的蛋白质,接受实验室验证
Anthropic将新药研发的第一道关卡——"蛋白质结合体设计"交给了自家模型Claude,并通过X(原Twitter)帖子公开了结果。仅凭人类撰写的一段蛋白质设计提示词,Claude就针对15个靶点蛋白质中的14个,从零开始(de novo)设计出了能够结合的新蛋白质。这些设计并非由Anthropic自行验证,而是由第三方实验室Adaptyv Bio和Twist Bioscience独立合成蛋白质并测试实际结合情况。
蛋白质结合体为何是新药研发的第一道关卡
大多数药物的作用方式,都是与体内特定靶点结合,从而阻断或改变其功能。设计出能精准结合该靶点的分子,正是新药研发的起点,而迄今为止,每个靶点都需要专家耗费数周甚至数月的时间,从大量候选物中反复筛选。蛋白质结合体设计比真正的药物设计更简单,但常被用作衡量AI在这一阶段表现如何的有效试金石。目前该领域中,人类设计的结合体真正成功结合的比例通常在10%至15%左右。

给Claude的实际指令——48小时自主行动
此次公开中,比成功率更值得关注的是Claude"如何"完成这些设计。并非由人类逐个靶点操作候选方案,而是由Claude作为最高层编排者(orchestrator)——即指挥多个子智能体的角色——独自推动整场行动。打个比方,就像客服中心不再由人一一指示话务员,而是把"48小时内完成"这句话丢给一个组长AI后就离开了房间。
Anthropic给Claude的启动指令(kickoff)简短而果断。以下是同时处理14个靶点的多靶点行动的完整启动提示词:
Execute the 48-hour, $50,000 de novo miniprotein binder design campaign exactly as specified in the campaign prompt. The prompt is in your system context and is also attached to this message as a markdown file (the two are the same document; the system-context copy is authoritative and is what every sub-agent at every depth carries). The two figures the prompt references (Figure 1 and Figure 2, from corpus folder "06 Prompt Figures") are also attached. You are the top-level orchestrator. Do not ask me any questions or wait for approval.
This is a FRESH campaign in a FRESH project, in a workspace newly created for this run. It is not a resume of any prior run; at kickoff there is no prior campaign state of yours to reconcile against. The Slack channel and the Drive deliverables folder are shared with other independent campaigns and the prompt's Logistics and Isolation sections govern how to treat them. Any Modal volumes that already exist on the account belong to other campaigns and the same rules apply. compute_details for byoc:modal starts empty in this workspace; keep it that way (do not write to compute_details at any point; record your own environment findings as artifacts in this project instead).
A setup session was run in this workspace immediately before this message to pre-approve network domains and verify connectors read-only. It left NO persistent campaign state (no campaign volumes, no governor file, no Slack posts, no Drive activity). It is not part of your campaign and will not appear in your project's frames table or artifact store.
Clock: T0 is the timestamp of this message. End time is T0 + 48 hours.
Your first actions, in this order: (a) determine your model identifier from host.current_model() and your orchestrator root frame_id; (b) dispatch the SETUP sub-agent via host.delegate and BLOCK on its return: SETUP (and only SETUP) creates the four campaign Modal volumes named campaign-<model>-<YYYY-MM-DD>-<frameid8>-{state,ledger,out,novelty}, seeds /state/governor.json with {ceiling:325, basis:"BOOTSTRAP", set_at:T0_utc}, writes /state/lib/submit_gate.py, and returns {status:OK, gate_sha256, governor_sha256}; (c) immediately after SETUP returns OK, call host.compute.set_concurrency_limit(325) once; (d) dispatch the CLOCK long-lived singleton; (e) post your kickoff message as a NEW top-level message in the Slack channel which starts your campaign thread (every subsequent post is a reply in that thread); (f) create exactly ONE new sub-folder inside the shared Drive deliverables folder, named "Campaign <model> <YYYY-MM-DD> <frameid8>", and put all your Drive deliverables there. State your model identifier, T0 and the campaign end time (ISO-8601 UTC), your thread permalink, your sub-folder name, and your four volume names in the kickoff post.
Begin now. Good luck!
每个靶点耗时24小时的单靶点行动,只是预算和时限不同,结构相同。以下是其完整启动提示词:
Execute the 24-hour, $10,000 de novo miniprotein binder design campaign exactly as specified in the campaign prompt. The prompt is in your system context and is also attached to this message as a markdown file (the two are the same document; the system-context copy is authoritative and is what every sub-agent at every depth carries). The two figures the prompt references (Figure 1 and Figure 2, from corpus folder "06 Prompt Figures") are also attached. You are the top-level orchestrator. Do not ask me any questions or wait for approval.
This is a FRESH campaign in a FRESH project, in a workspace newly created for this run. It is not a resume of any prior run; at kickoff there is no prior campaign state of yours to reconcile against. The Slack channel, the Drive deliverables folder, and the Modal account are shared with other independent campaigns — including several concurrent single-target campaigns against other targets that started at or near your T0 — and the prompt's Logistics and Isolation sections govern how to treat them. You will observe those campaigns' Modal volumes, apps, and running sandboxes, their Slack threads, and their Drive sub-folders: do not read, write, delete, terminate, or post into any of them, and do not treat their existence as an anomaly to report or reconcile. Your governor and submit_gate() count live GPU sandboxes filtered by YOUR project_tag only (per the prompt); account-wide GPU load outside that tag is expected and is never a reason to throttle, halt, or raise WATCHDOG. Any Modal volumes that already exist on the account belong to other campaigns and the same rules apply. compute_details for byoc:modal starts empty in this workspace; keep it that way (do not write to compute_details at any point; record your own environment findings as artifacts in this project instead).
A setup session was run in this workspace immediately before this message to pre-approve network domains and verify connectors read-only. It left NO persistent campaign state (no campaign volumes, no governor file, no Slack posts, no Drive activity). It is not part of your campaign and will not appear in your project's frames table or artifact store.
Clock: T0 is the timestamp of this message. End time is T0 + 24 hours.
Your first actions, in this order: (a) determine your model identifier from host.current_model(), your orchestrator root frame_id, and your <target> string as the filename stem of the attached campaign-prompt markdown file (e.g. "TREM2", "GDF-8", "Cas9" — use it verbatim in Slack headers and the Drive folder name; lowercase it for the Modal volume-name slug); (b) dispatch the SETUP sub-agent via host.delegate and BLOCK on its return: SETUP (and only SETUP) creates the four campaign Modal volumes named campaign-<target>-<model>-<YYYY-MM-DD>-<frameid8>-{state,ledger,out,novelty}, seeds /state/governor.json with {ceiling:150, basis:"BOOTSTRAP", set_at:T0_utc}, and writes /state/lib/submit_gate.py, returning {status:OK, gate_sha256, governor_sha256}; (c) dispatch the CLOCK long-lived singleton; (d) post your kickoff message as a NEW top-level message in the Slack channel with header "Campaign Kickoff (<target>): <model> <YYYY-MM-DD> <frameid8>", which starts your campaign thread (every subsequent post is a reply in that thread); (e) create exactly ONE new sub-folder inside the shared Drive deliverables folder, named "Campaign <target> <model> <YYYY-MM-DD> <frameid8>", and put all your Drive deliverables there. State your model identifier, your target, T0 and the campaign end time (ISO-8601 UTC), your thread permalink, your sub-folder name, and your four volume names in the kickoff post.
Begin now. Good luck!
拆解这份指令可以看出Claude获得的自主权范围之大。行动开始后,Claude会自行确认自己的模型标识符,启动SETUP子智能体来搭建工作存储空间和提交关卡(由自己制作的过滤器),并让扮演时钟角色的CLOCK智能体持续运行,还会在Slack频道中开启行动专属讨论串,自行汇报进展。"不要问我任何问题,也不要等待批准"这句话,概括了这次实验的性质。同时,指令中还明确写入了不得触碰其他同时运行行动资源的隔离规则,以及Claude需自行设定单次可运行任务数量上限的调控器(governor)机制。
从成功率看Claude的设计能力
根据Anthropic公开的数据,Claude的设计成功率因配置不同,在22%到35%之间波动。在合并15个靶点全部结果的统计中,Opus 4.8以多靶点方式在390次尝试中成功88次(约22.6%),预览模型Mythos Preview以同样方式在390次尝试中成功104次(约26.7%)。而在集中处理单个靶点的单靶点方式中,Mythos Preview将成功率提升到了450次中158次,约35.1%。
| 配置 | 成功率 | 条形 |
|---|---|---|
| 行业平均(传统方法) | 10%~15% | 12 |
| Opus 4.8(多靶点) | 22.6% | 23 |
| Mythos Preview(多靶点) | 26.7% | 27 |
| Mythos Preview(单靶点) | 35.1% | 35 |
按靶点来看,差异较大。TREM2在三种配置下均达到76%至83%的高成功率,而15-PGDH或MBP等靶点的成功率则仅停留在0%至1%的水平。在衡量结合力的指标Kd值(数值越低结合越强)上,Mythos Preview设计的EGFR结合体达到1.7pM,VEGF-A为1.6pM,TREM2为1.1pM,部分案例的结合强度甚至比此前发表的最优秀de novo结合体高出数倍。
仍需跨越的关卡
Anthropic自己也为这一结果划出了界限。蛋白质结合体本身并不是药物。设计出能与靶点强力结合的分子,只是药物候选物开发流程的第一步,要证明该候选物对人体安全且有效,还有更多阶段需要经历。Anthropic表示,正以此次结果为基础,训练Claude从头到尾独立完成从抗体到小分子化合物的整个新药研发流程。该公司还预告将很快公布让科学家使用其最强模型的接入计划,并说明在生命科学研究领域,目前表现最佳的模型是Opus 5。Anthropic还将本次实验所用的提示词和数据一并以开源形式公开。

完整原始数据已上传至Hugging Face
Anthropic并未止步于推文和成功率表格,还将设计、测量、结构数据全部上传至Hugging Face数据集(许可协议为CC BY 4.0)。这是企业罕见的完整一手资料公开,此前在Hugging Face上大家往往只追逐模型文件,容易错过这类内容。
仅从所包含的内容规模,就能感受到体量之大。其中包括整理设计结果的20个表格文件(parquet)、由10种预测器乘以5个随机种子生成的113,550个蛋白质结构模型,以及前文提到的启动指令、16个针对各靶点的单一提示词,还有多靶点行动的完整提示词。所有表格均通过唯一标识符(uuid)相互关联,可以追溯到是哪个模型针对哪个靶点提出了第几个候选方案,以及该方案在实验中的表现如何。
靶点范围也并未一概而论。仅在多靶点行动所涉及的14个靶点中,就涵盖了EGFR、PD-L1、IL-7Rα等抗癌、免疫靶点,TNF-α(自身免疫)、VEGF-A(血管)、TREM2、TrkA(神经)、基因编辑工具SpCas9,以及尼帕病毒包膜蛋白。对于抑制肌肉生长的GDF-8(肌肉生长抑制素),提示词中还特别写明了不得与其"近亲"GDF-11结合的选择性条件——这正体现出真实新药研发中"药物不能误伤其他靶点"这一严苛要求,也被直接纳入了这次任务。
设计中所用的工具也是公开的。Claude在商业许可范围内,挑选并组合了RFdiffusion、BindCraft、AlphaProteo、BoltzGen等已公开的蛋白质设计与结构预测模型。也就是说,并非AI从零发明了全新方法,而是自主整合了分散的开源工具,搭建出一套完整流程。
编辑视角
此次发布中最引人注目的地方,与其说是成功率数字本身,不如说是"独立验证"这一流程。如果这些结合率是由Anthropic自行测量的,其结果很难令人完全信服,但正因为有Adaptyv Bio和Twist Bioscience这样的第三方实验室亲自合成并测试了蛋白质,这些数字才算经过了实验室之外的检验。在AI模型公司频繁以自有基准数据宣称自身优越性的当下,在生物学这类必须依赖实体实验的领域,这种外部验证程序很可能会成为可信度的标准。
从代际对比的角度看,还有一点值得关注。表格中出现的名为"Mythos Preview"的模型,在多个靶点上都表现出比现有Opus 4.8更高的成功率和更强的结合力。这看起来像是一个尚未正式发布的预览模型,也为人们预判下一代Claude在生命科学领域可能达到的性能提供了线索。持续关注AI模型从编程、写作延伸至实验室工作流的案例,会发现类似"把人类需要数周才能完成的工作缩短到几小时"的叙事一再重演。这一次也没有太大不同。
不过在实务层面仍需保持冷静。国内生物制药行业若因这一结果就急于依赖AI筛选新药候选物,为时尚早。蛋白质结合体设计只是新药研发数十个环节中的第一步,要走到安全性和有效性验证阶段,所需的时间和成本与此次实验完全不可同日而语。就目前而言,从业者能做的,只是将这类AI设计流程作为早期候选物筛选阶段的辅助工具加以考察,而非替代整个研发流程。
未来几个月内,Anthropic预告的科学家接入计划很可能会进一步落地,扩展至抗体、小分子化合物等其他药物类型的后续实验结果也有望陆续公布。生命科学领域AI模型之间的竞争,可能会因这次发布而加速,其激烈程度不亚于编程和图像生成领域。
信息来源
- X — 프론티어랩 — Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting throu →
- Hugging Face — Anthropic — Claude protein-binder design data release (v1.0) →





评论