Back to Blog

Anthropic's AI Protein Design Run, Number by Number: 354 Binders, 27%, and What It Doesn't Prove

Prateek SinghAugust 20, 202612 min read
Anthropic's AI Protein Design Run, Number by Number: 354 Binders, 27%, and What It Doesn't Prove

Anthropic gave Claude an entire protein binder design campaign — target research, epitope choice, tool orchestration, ranking — and synthesised every design it returned. 354 of 1,320 bound. Here is every figure from the paper, including the three targets where it failed.

On August 18, 2026, Anthropic published a paper claiming something that is easy to misread in either direction. It did not cure anything. It did not design a drug. What it did was hand a language model the entire job of a protein binder design campaign — target research, epitope choice, tool installation, compute budgeting, ranking — and then synthesise every single design the model returned, without editing any of them, and measure what stuck.

The paper is Autonomous de novo protein binder design with Claude, authored by Amir Shanehsazzadeh at Anthropic. Two Claude models ran the campaigns: Claude Opus 4.8 and Mythos Preview, the invitation-only model that has since been succeeded by Claude Mythos 5. Two independent contract research organisations did the wet-lab work without knowing which model, campaign, or rank any sequence came from.

Here is every number that matters, including the ones that went badly.

What was actually automated

A binder design campaign normally eats weeks of a specialist's time per target. You pick which region of the target to model, decide which surface to bind, install and configure a fast-moving stack of diffusion and structure-prediction models, filter candidates, decide whether a redesign round is worth the GPU spend, and finally choose which handful to send for synthesis.

Anthropic wrote that working knowledge into a single protocol prompt of roughly 16,000 words and handed it to Claude as a system prompt. Critically, the prompt names no epitope, no scaffold, no target construct, and no sequence for any target. Humans did four things: chose the 16 targets and passed each one to Claude as nothing more than a name, a UniProt accession, an organism and an oligomeric state; supplied a cloud GPU account with a fixed budget and time limit; placed the synthesis orders; and interpreted the binding data at the end.

Everything in between — which region of the protein to attack, which of ten generators to run, how stringently to filter, which 30 designs to deliver and in what order — was Claude's. Campaigns ran 24 to 48 hours unattended. The paper notes that after infrastructure outages, "only short, non-technical instructions were needed to resume."

The headline: 354 of 1,320 designs bound

Sixteen targets went in. One of them — the mature GDF-8 dimer — aggregated under assay conditions and gave uninterpretable readings at both CROs, so its 120 designs are excluded from every count. That leaves 1,320 designs across 15 targets.

Claude produced confirmed binders against 14 of those 15 targets. In total 354 designs bound, a pooled hit rate of 26.8%. Anthropic puts the typical rate for a published campaign at 10–15%.

Hit rate: what bound, by campaign Share of delivered designs that showed concentration-dependent binding at one or both CROs Typical published binder campaign (Anthropic's stated baseline) 10–15% Claude Opus 4.8 — all 13 targets at once, 48 h (88 of 390) 22.6% Mythos Preview — all 13 targets at once, 48 h (104 of 390) 26.7% All campaigns pooled (354 of 1,320) 26.8% Mythos Preview — one target at a time, 24 h (158 of 450) 35.1% Source: Shanehsazzadeh, Autonomous de novo protein binder design with Claude, 18 Aug 2026, Table 1 and text.

Two campaign formats ran. In the multi-target format, each model designed against 13 targets simultaneously inside one 48-hour session. In the single-target format, Mythos Preview got a dedicated 24-hour session per target — with roughly 2.8× the compute per target, which the paper is careful to say means focus and budget cannot be separated as explanations for the higher rate.

Claude's own ranking predicted what would bind

This is the result I find hardest to dismiss, and it gets less attention than the headline. Each campaign delivered 30 designs in rank order — Claude's own judgement of which were most likely to work, made before anything was synthesised.

If that ranking were noise, the hit rate would be flat across the list. It is not. The single design Claude ranked first in each campaign bound 49% of the time. Across the top five it was 44%, across the top ten 39%, and across all 30 it was 28%.

Claude's own ranking carried real signal Hit rate by the rank Claude assigned each design before anything was synthesised Rank 1 only — the single design Claude put first per campaign 49% Top 5 ranks 44% Top 10 ranks 39% All 30 delivered designs 28% A ranking that meant nothing would sit flat at 28% across all four rows.

In other words, if you had only ordered the one design Claude put at the top of each list, roughly half of them would have worked. That is a model that knows something about its own output.

RBX1: the only clean comparison against humans

Most of the paper has no human control group, and Anthropic says so plainly. There is one exception, and it is the most quotable result in the paper.

RBX1, a subunit of an E3 ubiquitin ligase, was recently the subject of an open design competition. Across 245 de novo designs submitted by human entrants, 9 bound — a 3.7% hit rate. Claude's three campaigns against the same target produced 28 binders from 90 designs.

RBX1: the one head-to-head against humans Same target, same assay plate — an open design competition versus Claude's campaigns Open competition entrants (9 of 245 designs bound) 3.7% Claude, across three campaigns (28 of 90 designs bound) 31.1% Claude's tightest RBX1 binder: KD 3.9 nM. The competition's winner, re-run on the same plate: 45 nM.

Affinity told the same story. Claude's tightest RBX1 binder came in at a dissociation constant of 3.9 nM. The competition's winning entry, re-synthesised and measured on the same plate under the same conditions, was 45 nM — roughly ten times weaker.

One caveat worth stating: the competition entrants were working under their own constraints and design budgets, not a matched protocol. It is a real comparison, not a controlled one.

How tight were the binders?

"Bound" is a low bar on its own. The useful question is how tightly. Of the 354 binders, 142 came in under 100 nM, 78 under 10 nM — the rough threshold at which a binder starts being useful as a reagent — and 38 under 1 nM. Six TREM2 binders hit the assay's detection floor and are reported only as below 100 pM, meaning the instrument could not tell how tight they actually were.

How tight were the binders? Equilibrium dissociation constant (KD) for the 354 designs that bound — lower is tighter Bound at all (354 designs) 354 KD below 100 nM 142 KD below 10 nM — the usual bar for a strong binder 78 KD below 1 nM 38 Six TREM2 binders sat at the assay's detection floor and are reported only as below 100 pM.

Anthropic reports high-affinity binders (KD under 10 nM) against at least six targets, and designs matching or beating the best previously reported affinity against at least four. There is also a result nobody optimised for: cross-species reactivity was only a secondary objective in the prompt, yet 130 of the 233 binders tested against the mouse version of their target bound that too — which is precisely the property you need for a molecule to be testable in animals.

Where it failed, and why that matters more

Three targets went badly, and the failures are more informative than the successes.

Against MBP (maltose-binding protein), all 90 designs failed. Against 15-PGDH, 1 of 30 bound. Against TNFα — a compact homotrimer that is already the target of five approved biologics — 12 of 150 designs bound, an 8% rate, and every one of them came from Opus 4.8. Both Mythos Preview campaigns against TNFα returned 0 for 60.

Where the campaigns failed Designs that bound, out of designs delivered, on the targets that went badly MBP — maltose-binding protein 0 of 90 15-PGDH 1 of 30 TNFα — every binder came from Opus 4.8; Mythos Preview got 0 of 60 12 of 150 Pooled average, all 15 targets 354 of 1,320 Confidence scores missed these — MBP and BBF-14 designs scored much like the ones that worked.

The uncomfortable part is the confidence scoring. Claude's in-silico predictions for the MBP and BBF-14 designs scored about as well as its predictions for targets that worked. The system had no idea it was failing. In a field where the whole promise is "filter in silico, synthesise only the winners," a confidence score that fails silently on hard targets is the thing that will cost people money.

The stack underneath, and what it cost

Everything Claude used is open source, which is the part of this paper with the longest tail. Ten structure generators contributed designs that were ordered: RFdiffusion (118 designs) and RFdiffusion3 (267), PXDesign (358 — the single largest contributor), Genie 3 (185), FreeBindCraft (135), BoltzGen (134), Proteina-Complexa (100), FoldCraft (14), BoltzDesign1 (2) and Protein Hunter (2).

Sequences came overwhelmingly from SolubleMPNN, the soluble variant of ProteinMPNN — 1,133 of the tested designs — with 111 from SolubleCaliby and 21 from stock ProteinMPNN. Ranking used an ensemble of ESMFold2, ESMFold2-Fast and Protenix v2. Notably, AlphaFold 3's weights, Rosetta/PyRosetta and ESM3 were all excluded on licensing grounds, and the campaigns worked anyway.

Compute was not trivial but not exotic: about 12,500 NVIDIA H100-hours for a 48-hour multi-target campaign, and about 2,500 H100-hours per target in the 24-hour single-target format, run on rented cloud GPUs.

How much of this is the model, and how much is the prompt?

Here is the honest tension in the paper. The 16,000-word protocol prompt is itself a substantial piece of expert work — and only about a third of it is science.

Two thirds of the 16,000-word prompt is not science Anatomy of the protocol prompt every campaign received as its system prompt Science and tooling — target dossiers, epitope choice, design strategy, filters, ranking 34.2% Orchestration and verification — sub-agent delegation, clock discipline, ranking rules 34.7% Operations — compute budget, pacing governor, reporting 31.1% The expertise that made the campaigns work is partly in the model and partly written into this prompt by hand.

Science and tooling — target dossiers, epitope selection, design strategy, filters, the ranking score — is 34.2% of the prompt. Orchestration and verification is 34.7%, and operations (compute budget, a pacing governor, reporting) is another 31.1%. Two thirds of what makes these campaigns work is teaching the model to run a 48-hour job without falling over: delegate to sub-agents, watch the clock, verify its own outputs, don't burn the budget in hour six.

That cuts both ways. A skeptic can say the human expertise never left — it just moved into the prompt. That is fair, and the paper concedes it. But the same frozen prompt served all 16 targets without modification, which is the claim that actually matters: the expertise was written down once and then applied to targets it had never seen.

What this is not

The paper's own limitations section is unusually direct, and it deserves to be read before the headlines.

The evidence is binding, not structure and not function. Not a single design was structurally resolved. Every pose in the paper is a prediction. No design was tested for biological activity — a molecule that sticks to TNFα is not a molecule that does anything useful to TNFα.

The designs are not fully independent. Sequence variants of one backbone were each counted as a design. Counting only the best-ranked sequence from each of the 809 generated backbones, 200 bound — a 24.7% hit rate rather than 26.8%. The paper reports this itself and notes the main comparisons hold.

Each configuration ran exactly once. Model, campaign format and run-to-run chance are confounded. This is why Anthropic describes campaigns rather than ranking models — you cannot conclude from this data that Mythos Preview is better than Opus 4.8 at protein design, only that these particular runs went this way.

There was no parallel human campaign as a control, and most of the chosen targets are extensively characterised in the literature — meaning the model had plenty to work from.

Outside critics have pushed harder. Martin Shkreli, the convicted former pharmaceutical executive, called the work "not impressive" on X, arguing the affinities are unremarkable for this class of molecule and pointing out that none of the targets are intracellular — if you need something to bind a protein on the outside of a cell, a monoclonal antibody already does that job. It is a real objection: extracellular targets are the easier half of the problem, and the hard, undruggable-by-antibody targets sit inside the cell.

Anthropic's own framing is the one to keep: protein binders are not drugs. A high-affinity binder is the first step of a process that kills most candidates much later.

Why it still matters

Strip out the hype in both directions and one claim survives: a general-purpose language model, given a written protocol and a GPU budget, ran a multi-day computational biology campaign end to end and produced molecules that worked in someone else's lab, at a rate above the published norm, on targets it had never seen.

The tools it used are free. The prompts, all 1,440 design models, the per-design provenance and both CROs' binding data are released on Hugging Face under CC BY 4.0, with the scripts under MIT. Any lab with targets of interest and no computational protein design expertise can, in principle, rerun the protocol as it stands — or, more usefully, try to beat it.

That is the actual news. Not that AI designed a protein — that has happened before — but that the expensive, judgement-heavy orchestration around protein design turned out to be writable down, and that the write-up is public.

Credits and sources

The research is Anthropic's, and the paper is the work of Amir Shanehsazzadeh and the Claude Science team. The experimental validation — the part that makes this more than a simulation — was done independently by Adaptyv Bio (Lausanne) by surface plasmon resonance across five target concentrations, and by Twist Bioscience in a separate format; neither saw the other's data, the design models, or which model or rank produced a sequence. Adaptyv Bio returned usable results for 1,296 of the 1,320 designs.

The design stack belongs to the open-source structural biology community — the RFdiffusion, ProteinMPNN, BindCraft, Genie, Boltz, Protenix and ESMFold authors, among others. None of this campaign happens without a decade of their work being freely available.

All figures in this article were built from the numbers in the paper. Where the paper and the press release differ in rounding, the paper wins.

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article

Related Posts