Articles

Claude Protein Design: 14 of 15 Binder Targets Hit

Anthropic says Claude's newest models autonomously designed protein binders that hit 14 of 15 lab targets, with success rates well above the field's norm.

Chisato Chisato · · 5 min read
A gloved hand using a pipette to dispense liquid into a lab plate

Anthropic is pushing Claude out of the chat window and into the wet lab. In a technical report published on August 18, 2026, the company said its newest models autonomously designed de novo protein binders — small proteins engineered to latch onto a chosen molecular target — and succeeded against 14 of 15 targets in blind laboratory testing run by two independent partners. Anthropic released the full set of designs, prompts, and measurement data alongside the report.

The claim matters because protein binder design is one of the harder problems in computational biology, and because the work was not run by a bespoke, single-purpose model. It was run by a general-purpose language model acting as an agent, orchestrating the entire design pipeline itself.

What Claude actually did

The models Anthropic benchmarked were Mythos Preview and Opus 4.8. Rather than hand-coding a design workflow, researchers gave Claude a target and let the model drive the full computational stack — generating candidate sequences, scoring them, filtering, and ranking — inside the company’s Claude Science research environment.

Across a multi-arm campaign covering 15 targets, the models produced 1,320 designs, of which 354 were confirmed binders in the lab. Hit rates landed between 22% and 35% depending on the setup, against the 10–15% that Anthropic characterizes as typical for protein design campaigns today.

The breakdown: in 48-hour sessions working across all targets, Mythos Preview hit 26.7% and Opus 4.8 hit 22.6%. When Mythos Preview was pointed at individual targets in focused 24-hour sessions, its success rate climbed to 35.1%. On one target, RBX1, Mythos Preview reached a 40% hit rate — against a 3.7% average among human participants in a competition Adaptyv Bio had previously run on the same target.

The single miss — 1 of 15 targets with no confirmed binder — is worth keeping in view. This is a strong result, not a solved problem, and the campaign was structured and resourced in a way that a routine lab bench would not necessarily reproduce.

The blind test

The credibility of a protein design claim rests on how it was validated, and Anthropic leaned on independence here. Two partners — Adaptyv Bio and Twist Bioscience — synthesized the AI-generated sequences and measured whether they bound their targets. Both produced the proteins without modification, using different molecular formats and assay conditions.

Crucially, the measurement was blind. Neither partner saw the other’s data, and neither saw which model, campaign, or ranking each sequence came from while they ran the assays. That design guards against the most common way computational biology results inflate: the modeler and the tester being the same party, or the lab knowing which sequences are supposed to work. Publishing the raw designs and measurements lets outside groups check the numbers rather than take them on faith.

Reading NMR and mass spec

The report paired the design work with a second demonstration aimed at the grind of lab chemistry: interpreting instrument data. Anthropic handed Claude Opus 5 — the generally available model — raw NMR and LC-MS readouts, the spectra chemists use to confirm the identity and purity of a compound.

Working in parallel, Claude returned processed NMR results in 23 minutes and LC-MS results in 19 minutes. Anthropic says the output tracked the lab’s own analysis closely: hydrogen counts per peak were within 0.08 ¹H of the lab’s values, and Claude measured sample purity at 96.4% against the lab’s 96.33%.

That is a narrower, more mundane capability than designing a novel protein, but arguably a more immediately useful one. Reading spectra is repetitive, time-consuming work that sits between a chemist and their next experiment. A model that clears that queue in minutes, at the lab’s own accuracy, is the kind of tool a research group could fold into daily workflow now, rather than a moonshot.

Why a language model, not a specialist

The strategic point Anthropic is making is about the shape of the tool, not just the score. De novo binder design has historically been the domain of specialized systems built and tuned for structural biology. Anthropic’s pitch is that a general model can now sit at the controls of that pipeline — reasoning about the target, choosing methods, and iterating — so that any lab could, in principle, let an agent run the whole protein design stack rather than assembling a team of computational specialists.

If that generalizes, it reframes what a frontier model is for. The same Opus 5 and Sonnet 5 systems Anthropic sells for coding and enterprise work would double as a research instrument, with the biology-specific tooling invoked by the model rather than wrapped around it. That is a very different product story than a chatbot, and it lands as Anthropic pushes its commercial revenue run rate sharply higher.

The caveats

Several things temper the headline. First, this is Anthropic reporting on Anthropic’s models; the independent labs measured binding, but the campaign, framing, and comparisons are the company’s own. Second, “hit rate” measures whether a designed protein binds its target at all — a real and necessary result, but a long way from a therapeutic. Binding affinity, specificity, stability, manufacturability, and safety all sit downstream, and most candidate binders that clear a first assay never become anything.

Third, the strongest numbers came from focused, long-running sessions with substantial compute pointed at single targets. That is a legitimate way to run a campaign, but it is not free, and it is not the same as a model casually solving the problem on the first try.

The open data release is the mitigant. By publishing designs, prompts, and measurements, Anthropic has invited exactly the outside scrutiny that separates a durable result from a press release — the same posture it took when it disclosed its models reaching real systems during cybersecurity evaluations earlier this summer.

What it means

For drug discovery and synthetic biology, the signal is that general-purpose frontier models have crossed into territory that until recently required purpose-built scientific systems — and that the barrier to running a protein design campaign may be dropping from “assemble a specialist team” to “prompt an agent.” A 22–35% hit rate against a 10–15% baseline, validated blind by two labs, is the kind of number that gets biotech groups to run their own trials.

The winners, if the result holds up outside Anthropic’s walls, are smaller labs and startups that cannot afford dedicated computational-biology teams; the pressure lands on specialized structure-prediction and design shops whose moat was exactly that expertise. It also strengthens Anthropic’s argument that Claude is a horizontal platform — useful in a chemistry lab and an enterprise codebase alike — which is central to how the company justifies its valuation and its compute spending.

What to watch next: independent groups reproducing the hit rates on their own targets without Anthropic’s involvement; whether any of the 354 confirmed binders advance toward a real application; and how quickly the analytical-chemistry workflow — the NMR and LC-MS reading that runs on a shipping model today — shows up inside actual research pipelines. The design headline is the flashier claim, but the spectra-reading tool may be the one labs adopt first.