Claude autonomous protein design doubles typical binder hit rate

Deep blue and slate-gray curves converge around soft teal circles, suggesting an autonomous design workflow converging on successful protein binders.
Anthropic’s autonomous protein-design campaign combined specialist open-source tools with Claude-led orchestration before independent wet-lab validation.

Anthropic published wet-lab results Tuesday from a Claude autonomous protein design campaign that produced confirmed binders at roughly twice the industry’s typical hit rate across 15 clinically significant targets. Adaptyv Bio and Twist Bioscience physically synthesized and tested the designs without modification, marking the first independently verified benchmark of an AI-orchestrated campaign against the human-expert baseline.

Hit Rates and Headline Numbers

Across 15 clinically significant protein targets — including PD-L1 (a checkpoint protein central to cancer immunotherapy), TREM2 (implicated in Alzheimer’s disease), TNF-alpha (the target of blockbuster anti-inflammatory drugs including Humira), and EGFR (a well-established oncology target) — Claude’s autonomous design campaigns produced 354 confirmed binders from 1,320 total designs. That translates to an overall hit rate of 22.6% to 35.1%, depending on session configuration, against an industry baseline of 10% to 15% for de novo protein binder campaigns.

Fourteen of 15 targets yielded at least one successful design. Against TREM2, 72 of 90 Claude-designed proteins bound — an 80% hit rate on a target directly relevant to Alzheimer’s disease research. Against VEGF-A, a target in oncology, 54 of 90 bound. The campaign also produced 15 confirmed beta-sheet binders containing at least 20% beta-sheet structure across six targets — a harder design class than the alpha-helical bundles most computational methods favor.

How the Campaign Actually Ran

Claude is not a protein design model and does not compute protein structures itself. It acted as an autonomous orchestration layer above a collection of open-source specialist tools that the field already uses. Starting from a single human-written protocol of approximately 30,000 tokens, Claude chose which site on each target protein to attack, selected backbone scaffold generation methods (drawing from PXDesign, RFdiffusion3, Genie 3, FreeBindCraft, BoltzGen, RFdiffusion, and Proteina-Complexa), ran the resulting structures through sequence design via SolubleMPNN, applied multiple rounds of in silico optimization, and used ESMFold2 and Protenix v2 to screen candidates for predicted binding before submitting 30 final designs per target.

No additional scientific, technical, or operational guidance was given after the sessions began. The compute budget was approximately $50,000 per multi-target 48-hour campaign and $10,000 per single-target 24-hour session, run through the Modal cloud platform. The prompt itself encodes roughly one-third scientific guidance and reading list, with the remaining two-thirds covering scheduling, sub-agent delegation, verification, and budget discipline. Anthropic has released the prompt and all design data on Hugging Face.

Where Claude Beat Human Competition Winners

The most concrete competitive benchmark involves RBX1, a small protein that drives targeted destruction of specific regulatory proteins inside cells. RBX1 was the subject of a GEM Workshop design competition; when Mythos Preview ran in single-target mode, it produced binders at a 40% hit rate. The same target was the subject of an open competition run by Adaptyv Bio, where 245 entrants achieved a 3.7% hit rate. Anthropic had the competition’s winning design physically reproduced and tested on the same assay plate as Claude’s top candidate.

The winner bound at approximately 45 nM. Claude’s top design bound at approximately 3.9 nM — roughly ten times more tightly, as The Decoder’s RBX1 comparison confirmed independently. Against TNF-alpha, a challenging multimeric target where multiple human expert groups have reported zero hits, Opus 4.8 succeeded where Mythos Preview failed, producing 12 binders — several of which bound not just human TNF-alpha but also the cynomolgus monkey and mouse equivalents, a cross-reactive property that matters for animal testing. Anthropic said it does not fully understand why a less capable model outperformed the more capable one on this specific target.

What the Pipeline Cannot Yet Do

Two targets exposed clear failure modes. Maltose-binding protein (MBP) is a large, flexible bacterial protein with a smooth, water-loving surface that gives a designed binder very little structural grip. None of Claude’s 90 MBP designs was confirmed as a binder, though one showed a weak, reproducible binding signal. BBF-14 is a protein that does not exist in nature — it was itself computationally designed and is used as a benchmark specifically because it has no evolutionary history for design tools to draw on. Claude produced three weakly binding designs against BBF-14, one from each design method, each built on a different scaffold.

The folding prediction confidence scores for MBP and BBF-14 were not meaningfully lower than those for successful targets — meaning the pipeline’s internal quality filter did not warn of either failure. No parallel campaign by human protein design experts using the same tools and budget was run as a control, and each target-model-format combination ran only once, making it impossible to separate genuine model performance from run-to-run variation. Anthropic has not submitted the work to peer review and said it plans more extensive characterization to confirm hit rates and affinity measurements.

A Second Experiment: Analytical Chemistry in 23 Minutes

The protein design results appeared alongside a separate experiment testing a different kind of scientific work. Claude Opus 5, available to general users and not subject to the biology access restrictions that gate Mythos Preview, was given raw output files from a contract laboratory for a routine quality-control sample. The prompt for each was two sentences. No vendor software and no manual setup were required. Working in parallel, Opus 5 returned processed NMR and LC-MS results in 23 and 19 minutes, respectively, with its hydrogen count per NMR peak matching the laboratory’s own result within 0.0.

“This is not a computational prediction. The proteins were made and measured.”

For investors tracking the convergence of foundation models and biotech tooling, the verifiable wet-lab confirmation matters more than the headline hit-rate figures. Adaptyv Bio and Twist Bioscience physically produced and tested Claude’s designs without modification, separating this benchmark from prior AI protein design claims that relied on computational surrogates. Two clear misses on MBP and BBF-14 — where fold-confidence scores did not predict wet-lab failure — also temper the result and signal the limits of current self-driving pipelines. The full data, prompt, and campaign details are available for independent reanalysis. The Claude autonomous protein design campaign, documented at https://www.techtimes.com/articles/325081/20260820/claude-runs-autonomous-protein-design-campaign-wet-lab-confirms-twice, establishes a verifiable performance benchmark for Claude autonomous protein design.

Leave a Comment

Your email address will not be published. Required fields are marked *