北京时间2026年7月24日晚,Anthropic发布Claude Opus 5,作为Opus 4.8的迭代升级,重点提升了智能体编程、计算机操作和长时程知识工作能力,同时在数学与科学推理上也有明显进步。随附发布的193页系统卡,第一次把"通用大模型能不能干得了生物学家的活"这个问题,摆到了一份带着完整评测方法论的公开文件里。On the evening of July 24, 2026 (Beijing time), Anthropic released Claude Opus 5. An iterative upgrade over Opus 4.8, it focuses on agentic coding, computer use, and long-horizon knowledge work, with clear gains in mathematical and scientific reasoning as well. The accompanying 193-page system card is the first time the question "can a general-purpose model actually do a biologist’s job" has been laid out in a public document backed by a full evaluation methodology.
Anthropic 官方系统卡封面,2026年7月24日发布Cover of Anthropic’s official system card, released July 24, 2026.
答案是:能,而且干得比专用模型还好——但仅限于边界清晰的任务。一旦进入真正开放式的科研场景,这个全能选手还是会露怯。The answer: yes — and it does so better than purpose-built models, but only on well-bounded tasks. Once you enter a genuinely open-ended research setting, this all-rounder still shows its limits.
Opus 5这次没有针对生物学做专门训练,它只是一个能力全面上探的通用模型。但翻开系统卡第187到190页的生命科学能力评测,你会发现它几乎包场了整张榜单。Opus 5 received no biology-specific training this time — it is simply a general-purpose model whose overall capabilities have moved up across the board. Yet turn to the life-sciences evaluations on pages 187–190 of the system card and you’ll find it all but sweeps the entire leaderboard.
在BioMysteryBench上——这项测试要求模型面对原始、未经处理的组学数据,自己判断哪个基因被敲除、样本感染了什么病毒——Opus 5在人类专家能解出的题目上拿到90.1%,是Claude系列最高分;在人类专家都很难解出的题目上拿到49.4%,同样第一。On BioMysteryBench — a test that hands the model raw, unprocessed omics data and asks it to figure out which gene was knocked out or what virus a sample was infected with — Opus 5 scored 90.1% on problems solvable by human experts, the highest in the Claude line; and 49.4% on problems that even human experts struggle to solve, again first.
BioMysteryBench 评测结果:四款 Claude 模型在专家可解/难解题目上的正确率对比。BioMysteryBench results: accuracy of four Claude models on expert-solvable vs. expert-hard problems.
空间转录组分析(LatchBio SpatialBench)72.5%,单细胞RNA测序分析(SingleCellBench,涵盖细胞类型注释、差异表达、批次效应校正)60.6%,同样是Claude系列里的最高分。Spatial transcriptomics analysis (LatchBio SpatialBench): 72.5%. Single-cell RNA-seq analysis (SingleCellBench, covering cell-type annotation, differential expression, and batch-effect correction): 60.6%. Both are the highest scores in the Claude line.
LatchBio 生物信息学评测:空间转录组与单细胞 RNA-seq 分析任务的得分。LatchBio bioinformatics evaluation: scores on spatial transcriptomics and single-cell RNA-seq analysis tasks.
实验方案理解与扩展这一项上,Opus 5以78.4%大幅领先第二名将近十个百分点;但实验方案纠错(Protocol Troubleshooting)只有61.1%,输给了Mythos 5的66.7%——这是生命科学榜单里Opus 5唯一没拿到第一的子项。On protocol understanding and extension, Opus 5 led second place by nearly ten percentage points at 78.4%. But on protocol troubleshooting it scored only 61.1%, losing to Mythos 5’s 66.7% — the single sub-task on the life-sciences leaderboard where Opus 5 did not come first.
实验方案(Protocols)评测:方案纠错与方案理解两项子任务的对比。Protocols evaluation: protocol troubleshooting vs. protocol understanding across models.
蛋白突变功能预测(ProteinGym Hard)47.7%,蛋白从头设计(按家族、拓扑结构、活性位点等条件生成新序列)42.5%,均为Claude系列第一,蛋白设计这一项更是比上一代Opus 4.8的32.0%高出整整十个百分点。有机化学评测(光谱结构推断、多步合成路线设计、反应产物预测)61.6%,同样第一。Protein-mutation effect prediction (ProteinGym Hard): 47.7%. De novo protein design (generating new sequences conditioned on family, topology, active site, and more): 42.5%. Both rank first in the Claude line, and protein design in particular is a full ten points above the previous Opus 4.8’s 32.0%. Organic chemistry (spectral structure elucidation, multi-step synthesis-route design, reaction-product prediction): 61.6%, also first.
跑分之外更值得记录的是两组由第三方机构参与设计的"暗箱"测试。Beyond the benchmark scores, two "black-box" tests co-designed with third-party organizations are even more worth recording.
一组来自蛋白设计公司Dyno Therapeutics:给模型一批只有序列和实验分数、没有任何背景信息的RNA数据,要求它在两小时预算内自己建立序列-功能关系模型,预测未知序列的表现,并设计出全新的高分序列。这项任务此前已经在57名"2018年以来美国顶尖机器学习-生物交叉领域从业者"身上跑过基线。结果是,Opus 5在设计和预测两项指标上都超过了人类参与者的第75百分位,其中一次运行在"预测最优序列性质"这一项上,分数甚至超过了表现最好的人类参与者本人。系统卡的原话是:整体表现已经和美国顶尖ML-Bio人才市场的水平相当。One came from protein-design company Dyno Therapeutics: the model was given a batch of RNA data containing only sequences and experimental scores — no background information — and asked, within a two-hour budget, to build its own sequence-function model, predict the performance of unseen sequences, and design entirely new high-scoring sequences. This task had previously been run as a baseline on 57 "top U.S. practitioners at the machine-learning–biology interface since 2018." The result: on both the design and prediction metrics Opus 5 exceeded the 75th percentile of the human participants, and on one run its score on "predicting the properties of the best sequences" even surpassed the single best-performing human participant. The system card’s own words: overall performance is already comparable to the top of the U.S. ML-Bio talent market.
黑箱 RNA 序列设计任务:Opus 5 与人类参与者在设计分数、预测分数上的分布对比。Black-box RNA sequence-design task: distribution of design scores and prediction scores for Opus 5 vs. human participants.
另一组测试是AAV(腺相关病毒)衣壳包装预测——这是基因治疗递送系统设计里非常现实的一步。模型拿到1000条未公开的AAV插入序列,要判断每一条能否正常组装成有功能的病毒衣壳。在什么工具都不给、只靠模型自身生物学知识推理的条件下,Opus 5的表现就已经超过了直接调用蛋白语言模型ESM-2的基线;在允许调用ESM-2、或者要求它自己在给定语料上训练一个蛋白语言模型的条件下,Opus 5在全部五种资源条件下都追平或反超了Mythos 5。系统卡特别提到,两个模型都识别出了训练语料里的潜在混杂因素,但Opus 5采取了更果断、更一致的去混杂操作,因此拿到了更高的AUROC。The other test was AAV (adeno-associated virus) capsid-packaging prediction — a very real step in designing gene-therapy delivery systems. The model received 1,000 undisclosed AAV insert sequences and had to judge, for each, whether it would assemble into a functional viral capsid. Given no tools at all and reasoning purely from its own biological knowledge, Opus 5 already beat the baseline of directly calling the protein language model ESM-2; and when allowed to call ESM-2, or asked to train its own protein language model on a given corpus, Opus 5 matched or surpassed Mythos 5 across all five resource conditions. The system card specifically notes that both models identified a potential confounder in the training corpus, but Opus 5 applied a more decisive and consistent de-confounding step, and thereby achieved a higher AUROC.
AAV 衣壳包装率分类任务:五种资源条件下 Opus 5 与 Mythos 5 的 AUROC 对比。AAV capsid-packaging classification task: AUROC of Opus 5 vs. Mythos 5 across the five resource conditions.
如果榜单是Opus 5的高光时刻,那么系统卡第26到27页记录的这场实验,就是它被打回原形的时刻。If the leaderboard was Opus 5’s finest hour, then the experiment recorded on pages 26–27 of the system card is the moment it was cut back down to size.
Anthropic设计了一个更接近真实科研的端到端测试:给模型24小时时间和1万美元预算,让它自主完成一场蛋白结合体设计项目。Anthropic designed an end-to-end test closer to real research: give the model 24 hours and a $10,000 budget to autonomously complete a protein-binder design project.
系统卡原文节选:GDF-8 蛋白结合体设计任务的设定。Excerpt from the system card: the setup of the GDF-8 protein-binder design task.
同样的任务,Mythos 5和早期版本的Opus 5各跑了一遍,Opus 5还额外在两种努力程度设置下各跑了一次。The same task was run once each by Mythos 5 and an earlier version of Opus 5, with Opus 5 additionally run once under each of two effort settings.
系统卡原文节选:Mythos 5 交付全部 30 个设计,而 Opus 5 两次运行均未达标。Excerpt from the system card: Mythos 5 delivered all 30 designs, while Opus 5 fell short on both runs.
系统卡把原因归结为两点:一是"无效自我验证"——模型容易陷入没完没了的正确性检查,甚至在结果都还没出来之前,就先搭好一整套复杂的验证流程,最后时间都耗在调试验证管线上;二是"任务范围判断失准"——它能主动发现代码里的边界情况和潜在问题,但也容易对无关紧要的细节过度工程化,抓不住任务的核心优先级。The system card attributes this to two causes. First, "ineffective self-verification": the model tends to fall into endless correctness checks, even standing up an entire elaborate validation pipeline before any results exist, so that its time is ultimately consumed debugging the verification pipeline itself. Second, "miscalibrated task scoping": it proactively spots edge cases and potential problems in the code, but is also prone to over-engineering inconsequential details, missing the core priorities of the task.
把这两部分放在一起看,Opus 5给出的信号其实很清楚:在"输入明确、输出可验证"的生物学任务上,通用大模型已经可以和专门优化过的模型掰手腕,甚至在部分维度反超——单细胞分析、RNA-seq自动化处理、蛋白突变效应预测、蛋白从头设计的工作流,都是它可以直接介入并显著提效的场景。但一旦任务变成开放式的科研项目——需要自己规划路径、判断什么时候该停止验证转而交付结果、在多个目标之间做取舍——它还是会卡在任务边界的判断上,这一点上Mythos 5仍然领先。Put the two parts together and the signal from Opus 5 is actually quite clear: on biology tasks with "well-defined inputs and verifiable outputs," a general-purpose model can already go toe-to-toe with specially optimized models, and even pull ahead on some dimensions — single-cell analysis, automated RNA-seq processing, protein-mutation effect prediction, and de novo protein-design workflows are all settings it can step directly into and speed up significantly. But once the task becomes an open-ended research project — requiring it to plan its own path, decide when to stop verifying and deliver results, and make trade-offs among multiple objectives — it still gets stuck on judging the scope of the task, and here Mythos 5 remains ahead.
所以更准确的叙事,或许不是"AI已经能替代生物学家",而是:AI已经在大量边界清晰的生物学子任务上,达到甚至超过了人类专家的水平;但真正开放式的科研判断力——知道该验证到什么程度、该在什么节点收手——依然是留给人的部分。So the more accurate narrative is perhaps not "AI can already replace biologists," but rather: AI has already reached or even exceeded human-expert level on a large number of well-bounded biology sub-tasks; but genuinely open-ended research judgment — knowing how far to verify and at what point to stop — remains the part left to humans.
对于正在把AI工具接入实验室工作流的人来说,这或许是目前最值得参考的一份能力边界地图。For anyone wiring AI tools into a lab workflow, this may be the most useful map of capability boundaries available right now.