
Summary
- Given a single prompt written by a human expert, Claude autonomously designed binding molecules for 14 of 15 target proteins
- Adaptyv Bio and Twist Bioscience independently synthesized and tested the designs, finding that 22-35% actually bound successfully, above the industry average of 10-15%
- Building on this result, Anthropic is continuing work to automate the entire drug discovery process, from antibodies to small molecules
Claude's protein designs put to the test in the lab
Anthropic gave its model Claude the task of "protein binder design" — the first gate in drug discovery — and shared the results in an X post. Using a single protein-design prompt written by a human, Claude designed, from scratch (de novo), new proteins capable of binding to 14 of 15 target proteins. Anthropic did not verify these designs itself. Third-party labs Adaptyv Bio and Twist Bioscience independently synthesized the proteins and tested whether they actually bound.
Why protein binder design is the first gate in drug discovery
Most drugs work by attaching to a specific target in the body and blocking or altering its function. Designing a molecule that fits precisely onto that target is the starting point of drug discovery, and until now this has required experts to spend weeks to months filtering through countless candidates for each target. Protein binder design is an easier task than actual drug design, but it serves as a useful benchmark for gauging how well AI can handle this stage. In this field, the success rate for human-designed binders that actually bind is generally known to be around 10-15%.

What Claude was actually told to do — a 48-hour autonomous campaign
Just as notable as the success rate is how Claude carried out this design work. Rather than a human running candidates for each target one at a time, Claude itself acted as the top-level orchestrator — a conductor directing multiple sub-agents — and ran the entire campaign autonomously. To use a call-center analogy, instead of a human directing agents one by one, a single team-lead AI was told "finish this within 48 hours" and left to work on its own.
The kickoff instruction Anthropic gave Claude is short and firm. Here is the full kickoff prompt for the multi-target campaign, which handled 14 targets at once.
Execute the 48-hour, $50,000 de novo miniprotein binder design campaign exactly as specified in the campaign prompt. The prompt is in your system context and is also attached to this message as a markdown file (the two are the same document; the system-context copy is authoritative and is what every sub-agent at every depth carries). The two figures the prompt references (Figure 1 and Figure 2, from corpus folder "06 Prompt Figures") are also attached. You are the top-level orchestrator. Do not ask me any questions or wait for approval.
This is a FRESH campaign in a FRESH project, in a workspace newly created for this run. It is not a resume of any prior run; at kickoff there is no prior campaign state of yours to reconcile against. The Slack channel and the Drive deliverables folder are shared with other independent campaigns and the prompt's Logistics and Isolation sections govern how to treat them. Any Modal volumes that already exist on the account belong to other campaigns and the same rules apply. compute_details for byoc:modal starts empty in this workspace; keep it that way (do not write to compute_details at any point; record your own environment findings as artifacts in this project instead).
A setup session was run in this workspace immediately before this message to pre-approve network domains and verify connectors read-only. It left NO persistent campaign state (no campaign volumes, no governor file, no Slack posts, no Drive activity). It is not part of your campaign and will not appear in your project's frames table or artifact store.
Clock: T0 is the timestamp of this message. End time is T0 + 48 hours.
Your first actions, in this order: (a) determine your model identifier from host.current_model() and your orchestrator root frame_id; (b) dispatch the SETUP sub-agent via host.delegate and BLOCK on its return: SETUP (and only SETUP) creates the four campaign Modal volumes named campaign-<model>-<YYYY-MM-DD>-<frameid8>-{state,ledger,out,novelty}, seeds /state/governor.json with {ceiling:325, basis:"BOOTSTRAP", set_at:T0_utc}, writes /state/lib/submit_gate.py, and returns {status:OK, gate_sha256, governor_sha256}; (c) immediately after SETUP returns OK, call host.compute.set_concurrency_limit(325) once; (d) dispatch the CLOCK long-lived singleton; (e) post your kickoff message as a NEW top-level message in the Slack channel which starts your campaign thread (every subsequent post is a reply in that thread); (f) create exactly ONE new sub-folder inside the shared Drive deliverables folder, named "Campaign <model> <YYYY-MM-DD> <frameid8>", and put all your Drive deliverables there. State your model identifier, T0 and the campaign end time (ISO-8601 UTC), your thread permalink, your sub-folder name, and your four volume names in the kickoff post.
Begin now. Good luck!
The single-target campaign, which spent 24 hours on one target, has the same structure but a different budget and clock. The full kickoff prompt is as follows.
Execute the 24-hour, $10,000 de novo miniprotein binder design campaign exactly as specified in the campaign prompt. The prompt is in your system context and is also attached to this message as a markdown file (the two are the same document; the system-context copy is authoritative and is what every sub-agent at every depth carries). The two figures the prompt references (Figure 1 and Figure 2, from corpus folder "06 Prompt Figures") are also attached. You are the top-level orchestrator. Do not ask me any questions or wait for approval.
This is a FRESH campaign in a FRESH project, in a workspace newly created for this run. It is not a resume of any prior run; at kickoff there is no prior campaign state of yours to reconcile against. The Slack channel, the Drive deliverables folder, and the Modal account are shared with other independent campaigns — including several concurrent single-target campaigns against other targets that started at or near your T0 — and the prompt's Logistics and Isolation sections govern how to treat them. You will observe those campaigns' Modal volumes, apps, and running sandboxes, their Slack threads, and their Drive sub-folders: do not read, write, delete, terminate, or post into any of them, and do not treat their existence as an anomaly to report or reconcile. Your governor and submit_gate() count live GPU sandboxes filtered by YOUR project_tag only (per the prompt); account-wide GPU load outside that tag is expected and is never a reason to throttle, halt, or raise WATCHDOG. Any Modal volumes that already exist on the account belong to other campaigns and the same rules apply. compute_details for byoc:modal starts empty in this workspace; keep it that way (do not write to compute_details at any point; record your own environment findings as artifacts in this project instead).
A setup session was run in this workspace immediately before this message to pre-approve network domains and verify connectors read-only. It left NO persistent campaign state (no campaign volumes, no governor file, no Slack posts, no Drive activity). It is not part of your campaign and will not appear in your project's frames table or artifact store.
Clock: T0 is the timestamp of this message. End time is T0 + 24 hours.
Your first actions, in this order: (a) determine your model identifier from host.current_model(), your orchestrator root frame_id, and your <target> string as the filename stem of the attached campaign-prompt markdown file (e.g. "TREM2", "GDF-8", "Cas9" — use it verbatim in Slack headers and the Drive folder name; lowercase it for the Modal volume-name slug); (b) dispatch the SETUP sub-agent via host.delegate and BLOCK on its return: SETUP (and only SETUP) creates the four campaign Modal volumes named campaign-<target>-<model>-<YYYY-MM-DD>-<frameid8>-{state,ledger,out,novelty}, seeds /state/governor.json with {ceiling:150, basis:"BOOTSTRAP", set_at:T0_utc}, and writes /state/lib/submit_gate.py, returning {status:OK, gate_sha256, governor_sha256}; (c) dispatch the CLOCK long-lived singleton; (d) post your kickoff message as a NEW top-level message in the Slack channel with header "Campaign Kickoff (<target>): <model> <YYYY-MM-DD> <frameid8>", which starts your campaign thread (every subsequent post is a reply in that thread); (e) create exactly ONE new sub-folder inside the shared Drive deliverables folder, named "Campaign <target> <model> <YYYY-MM-DD> <frameid8>", and put all your Drive deliverables there. State your model identifier, your target, T0 and the campaign end time (ISO-8601 UTC), your thread permalink, your sub-folder name, and your four volume names in the kickoff post.
Begin now. Good luck!
Breaking down the instructions reveals the scope of autonomy Claude was given. As soon as it started, it identified its own model identifier, spun up a SETUP sub-agent to set up storage space and a submission gate (a filter it built itself), kept a CLOCK agent running at all times, and opened a campaign thread on Slack to report on progress on its own. The line "Do not ask me any questions or wait for approval" sums up the nature of this experiment. At the same time, the instructions also embed isolation rules against touching other concurrently running campaigns' resources, and a governor cap that lets the system limit its own concurrent workload.
What the success rates say about Claude's design ability
According to the figures Anthropic disclosed, Claude's designs achieved success rates ranging from 22% to 35% depending on the configuration. Across all 15 targets combined, Opus 4.8 succeeded on 88 of 390 attempts (about 22.6%) using the multi-target approach, while the preview model Mythos Preview succeeded on 104 of 390 (about 26.7%) using the same approach. In the single-target approach, which focused on one target at a time, Mythos Preview pushed its success rate up to 158 of 450, or about 35.1%.
| Configuration | Success rate | Bar |
|---|---|---|
| Industry average (existing methods) | 10-15% | 12 |
| Opus 4.8 (multi-target) | 22.6% | 23 |
| Mythos Preview (multi-target) | 26.7% | 27 |
| Mythos Preview (single-target) | 35.1% | 35 |
The variance across individual targets was large. TREM2 showed high success rates of 76-83% across all three configurations, while targets like 15-PGDH and MBP stayed at just 0-1%. In terms of Kd values — the metric for binding strength, where lower is stronger — Mythos Preview's designs for EGFR bound at 1.7pM, VEGF-A at 1.6pM, and TREM2 at 1.1pM, in some cases binding several times more strongly than the best previously published de novo binders.
Hurdles that remain
Anthropic itself drew a clear line around these results. A protein binder is not a drug. Creating a molecule that binds strongly to a target is only the first step in developing a drug candidate, and there are far more steps remaining to prove that a drug candidate is safe and effective in humans. Anthropic said it is building on this result to train Claude to carry out the entire drug discovery process end to end, from antibodies to small molecules. The company also said it plans to soon announce an access program that would let scientists use its top-tier models, and noted that Opus 5 is currently its best model for life-science research. Anthropic also open-sourced the experimental prompts and data.

The raw data, uploaded in full to Hugging Face
Anthropic didn't stop at a tweet and a success-rate table — it uploaded the full set of design, measurement, and structural data as a dataset on Hugging Face (licensed under CC BY 4.0). This is a rare case of a company publishing raw primary data in full, at a time when people tend to go to Hugging Face only for model files.
The scale is clear just from what's included: 20 parquet tables summarizing the design results, 113,550 protein structure models generated by combining 10 predictor types with 5 seeds each, and, alongside the kickoff prompts shown above, the full text of 16 single-target prompts and the multi-target campaign prompt. All the tables are linked by unique identifiers (UUIDs), making it possible to trace which model produced which candidate for which target, and how it performed experimentally.
The targets weren't lumped together either. The 14 targets covered by the multi-target campaign alone span cancer and immune targets like EGFR, PD-L1, and IL-7Rα, the autoimmune target TNF-α, the vascular target VEGF-A, neurological targets TREM2 and TrkA, the gene-editing enzyme SpCas9, and the Nipah virus envelope protein. For GDF-8 (myostatin), which suppresses muscle growth, the prompt even specified a selectivity requirement that it must not bind to its close relative GDF-11 — building in exactly the kind of real-world drug-discovery constraint where a drug must not hit the wrong target.
The tools used for design were also open. Claude selected and combined already-public protein design and structure prediction models such as RFdiffusion, BindCraft, AlphaProteo, and BoltzGen, staying within their commercial license terms. Rather than inventing an entirely new method from scratch, the AI assembled a pipeline by stitching together scattered open-source tools on its own.
Editor's view
The most notable part of this announcement isn't the success-rate figures themselves but the "independent verification" process behind them. Had Anthropic measured the binding rates itself, the results would have been hard to trust — but because third-party labs Adaptyv Bio and Twist Bioscience actually synthesized and tested the proteins, these numbers have effectively been verified outside the lab that produced them. At a time when it's common for AI model companies to claim superiority based on their own benchmarks, this kind of external verification is likely to become the standard for credibility in fields like biology, where physical experiments are unavoidable.
There's another interesting point in the generational comparison. A model called "Mythos Preview" appearing in the table outperformed the existing Opus 4.8 on multiple targets, both in success rate and binding strength. It appears to be a preview model that hasn't been formally introduced yet, offering a hint of what the next generation of Claude might be capable of in the life sciences. As AI models increasingly attach themselves to lab workflows beyond coding or writing, the same narrative keeps recurring — "what used to take humans weeks now takes hours." This case is largely consistent with that pattern.
That said, a dose of practical caution is warranted. It would be premature for Korea's bio-pharma industry to look at this result and rush to use AI to generate drug candidates right away. Protein binder design is just the first of dozens of steps in drug discovery, and reaching safety and efficacy validation requires time and cost on a scale far beyond what this experiment covered. At this point, the most practitioners can realistically do is consider this kind of AI design pipeline as a supplementary tool for early candidate screening — not as a replacement for the overall process.
In the months ahead, Anthropic's promised access program for scientists is likely to take concrete shape, and follow-up results extending the experiment to other drug modalities such as antibodies and small molecules are likely to emerge. This announcement is likely to accelerate a trend in which competition among AI models in the life sciences becomes as intense as it already is in coding and image generation.
Sources
- X — 프론티어랩 — Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting throu →
- Hugging Face — Anthropic — Claude protein-binder design data release (v1.0) →





Comments