我们正在 Anthropic 组建一个新的生命科学研究团队和实验室。我们的重点是利用 Claude 开展基础生物学研究:探索 DNA 数据集以识别尚未被表征的蛋白质家族,大规模生成假设,并通过实验室实验对其进行验证。本文介绍这项工作背后的团队,并分享早期成果——在这些成果中,Claude 仅凭我们科学家提供的高层方向指引,就发现了一种具有类似 CRISPR 特性的新型酶系统。
许多彻底改变生物学和医学的发现,都始于科学家在自然界令人惊叹的分子机器多样性中注意到某种异常现象。限制性内切酶,即能在特定短序列处切割 DNA 的蛋白质,是在细菌免疫系统中被发现的,它们在那里摧毁入侵病毒的 DNA。研究人员意识到,他们可以利用这些酶在选定的位置切割 DNA,并将基因从一个生物体拼接至另一个生物体,由此开启了生物技术产业。Taq 聚合酶是一种能在高温下复制 DNA 的酶,它是在黄石公园一处温泉中的细菌里被鉴定出来的。它成为了 PCR 的基础,而 PCR 这种 DNA 复制方法被广泛应用于现代诊断的诸多领域。CRISPR 最初是作为某些细菌 DNA 中一段不寻常的重复序列被注意到的,如今已成为基因编辑药物的基石。
2026 年春季,我们组建了一个研究团队,旨在探究通用 AI 模型能否将此类发现系统化并加速这一进程。我们相信,这种加速将来自一种全新的生物学研究方式的建立——在这一方式中,智能体在流程的每一步都与人类协作。开发这种新的工作方式,要求我们建立自己的实验室,并组建一个统一的团队,负责从在生物学领域训练 Claude 到在实验室中运行实验的所有工作。
今天,我们分享我们首批研究项目之一的早期成果。在该项目中,Claude 自主发现了一个新型酶系统,它与一系列 DNA 重复序列相关联,其模式令人联想到 CRISPR。尽管我们尚不清楚它的功能,但 Claude 发现的这个系统具有一组特征,这些特征此前只在少数其他系统中被同时发现过,而所有这些系统都是可编程的,并能执行诸如剪切、复制和粘贴 DNA 之类的操作。
除了已经变革了科学与医学的 CRISPR 之外,目前还有若干此类系统正在作为前景广阔的工具进行开发。
Claude 发现的这个系统基于一种逆转录酶(RT),这类酶能将 RNA 复制为 DNA。虽然这种存在于巨型噬菌体中的基础 RT 在此前的研究中已被识别,但 Claude 似乎是第一个注意到该系统的决定性特征——一组相关联的非编码 DNA 序列,以及一个功能未知的额外附属蛋白。
在审阅了这篇预印本之后,CRISPR 基因组编辑的先驱之一、MIT 和 Broad Institute 教授 Feng Zhang 表示:
这是一个令人振奋的范例,展示了 AI 智能体如何为生物学发现做出贡献。识别出与逆转录酶相关联的 RNA 重复序列确实引人入胜,值得进一步研究。我希望这项工作能鼓励更多科学家探索 AI 如何支持他们的研究。
我们给 Claude 一个提示词,让它在庞大的 DNA 序列数据库中搜索有趣的新型 RT 实例。我们的参与仅限于最初的提示词和实验室工作,而 Claude 智能体则梳理数据库、研究不同的 RT 家族,并运用自己的判断来识别有趣的候选对象。
在大约 950 个智能体使用 2.1 亿 token 对该数据进行了 21 小时的搜索后,其中一个智能体发现了某种引人注目的东西:一段重复的 DNA 序列模式,出现在一个外形奇特的 RT 基因旁边。经过进一步分析和我们实验室的测试,我们认识到这一模式标记了一种此前未被表征的酶系统,存在于噬菌体(感染细菌的病毒)中,我们将其称为阵列相关逆转录酶(ART)。
我们理解 ART 主要功能的工作仍在进行中。然而,我们认为尽早分享此类发现很重要,这既能展示 Claude 的能力,也能让更广泛的社区了解我们正在研究的内容。我们已发布了一篇预印本(此处),对此进行了更详细的讨论。
关于我们的实验室
我们是一支科学家团队,整个职业生涯都在探索不寻常的蛋白质,并专注于运用计算方法系统性地读取 DNA、解读其演化过程,并挑选出值得进一步表征的生物系统。在加入 Anthropic 之前,我们的研究帮助人们更好地理解了演化与调控CRISPR 系统的机制,发现了用于下一代酶的细胞与基因疗法,并构建了工具以加速识别 DNA 中的异常,例如人类致病性变异。我们隶属于 Anthropic 的生命科学组织,与从事药物发现以及训练 Claude 掌握生物学和化学知识的团队并肩工作。
我们的实验室位于湾区,看起来就像一个典型的分子生物学实验室。我们开展的研究仅涉及生物安全风险等级中较低的层级(BSL-1 和 BSL-2),我们不处理能够感染人类的病原体。所有实验室工作均由人类科学家完成。尽管我们曾尝试通过 Model Hardware Standard 等举措利用 AI 加速实验室工作,但这种方法不太适合我们分子生物学研究中涉及的那种临时性工作流程。
我们如何工作
我们的许多工作流程都涉及让 Claude 在大量与功能未知的蛋白质相关的 DNA 序列集合中进行搜索。一个典型的模式始于对某个给定蛋白质家族的调研。Claude 阅读相关文献,并从公开数据中复现已有结果以检验其方法。随后,它寻找不符合任何已描述系统的家族成员或基因组邻居,并为每个候选对象撰写一份简短、人类可读的报告,提出一种功能并描述支持其主张的证据。
在后续分析中,Claude 会批判性地评估证据——通常大多数候选对象会在这个阶段被淘汰。一次调研可能最终留下一个值得测试的候选对象,也可能一个都不剩。
当一个候选对象通过了我们的审查,我们会在实验室中对其进行测试,在标准实验室菌株中表达该蛋白质,并对其进行生化和结构表征,Claude 则帮助解读数据。我们的工作在 Claude Science 和 Claude Code 中完成,这些工具任何科学家都可以使用,有时还会使用我们自己搭建的、可协调多个并行运行的 Claude 会话的框架。
由于 Claude 产生假设的效率极高,这些假设本身已成为我们的研究对象。单次行动就能产出数百到数千份候选报告,我们一直在追问:我们认为值得测试的提案与搁置一旁的提案之间,区别究竟在哪里。我们学到的东西会反馈到给 Claude 的指令中,教会它模仿我们自己的科学品味。
Claude 找到了 ART
在过去几年里,研究人员发现了更多逆转录酶(RT),其中大多数存在于细菌中,作为免疫系统的一部分发挥作用。几乎所有 RT 家族都是通过基因组分析或基因组挖掘发现的,这需要研究人员在序列数据库中搜索无人表征过的基因,注意到那些不寻常的基因,并弄清楚它们的功能。
Claude 智能体收集了超过 200,000 个 RT,筛选出 3,500 个新的候选系统,并进一步缩小到 20 个最有说服力的候选对象,对其进行分析以生成人类可读的报告。对于一位专家科学家来说,这类分析可能需要数周到数月的工作。
在研究过程中,Claude 注意到一个不寻常的 RT 家族,决定对其进行更详细的检查。在梳理 RT 附近的原始 DNA 序列时,该智能体惊呼:“[RT 旁边的 DNA] 太壮观了:我肉眼就能看到一个串联重复阵列……那是一个类似 CRISPR 的……重复阵列?!”

随后,它的推进方式与人类科学家面对潜在发现时大致相同。它统计了重复序列的数量并测量了它们的间距,将布局与已知的 RT 系统进行比较,并在文献中搜索该模式此前的任何报道。经过彻底分析后,它确信自己发现了一个新的生物系统,并提交了一份报告供人类审查。
它发现的系统名为 ART,主要存在于噬菌体中,由三部分组成:RT、其旁边的一个伙伴基因,以及一长串均匀间隔的 DNA 重复序列。这种重复序列的排布类似于 CRISPR 阵列,后者储存着一组不同的 RNA 序列库,使 CRISPR-Cas 系统成为可编程的生物技术工具。我们的初步实验表明,ART 阵列同样被表达为一组不同的短 RNA,这暗示该系统可能存在类似的机制。
目前正在进行进一步的实验以确定 ART 的工作机制,我们分享这些早期发现,是为了向社区展示 Claude 能够自主检测异常并推动分析,从而开启生物学发现。
你可以在我们的技术报告中找到更多细节(此处)。
与我们合作
我们希望这项工作能向更广泛的科学界展示 AI 驱动假设生成的价值,我们也希望与其他科学家合作,将这一方法扩展到基因组学及其他领域的广泛问题中。如果你有研究问题的提案,我们期待你的来信。
We’re introducing a new life sciences research group and laboratory at Anthropic. Our focus is on fundamental biology research using Claude: exploring datasets of DNA to identify uncharacterized protein families, generating hypotheses at scale, and testing them through experiments in the lab. This post introduces the team behind this work and shares early results in which Claude discovered a novel enzyme system with properties reminiscent of CRISPR, with only high-level direction from our scientists.
Many discoveries that have revolutionized biology and medicine started with a scientist noticing something odd in the staggering diversity of molecular machines found in nature. Restriction enzymes, proteins that cut DNA at specific short sequences, were found in bacterial immune systems, where they destroy the DNA of invading viruses. Researchers realized they could use these enzymes to cut DNA at chosen places and splice genes from one organism into another, which launched the biotechnology industry. Taq polymerase, an enzyme that copies DNA at high temperatures, was identified in a bacterium in a Yellowstone hotspring. It became the basis for PCR, the DNA-copying method used in much of modern diagnostics. CRISPR was first noticed as an unusual repeat sequence in the DNA of certain bacteria, and is now the foundation of gene editing-based medicines.
In the spring of 2026, we formed a research group to see whether general AI models can systematize and accelerate such discoveries. We believe that this acceleration will come from establishing a new way of doing biology research, in which agents collaborate with humans in every step of the process. Developing this new way of working required that we build our own lab and a single team working on everything from training Claude in biology to running experiments in the lab.
Today, we’re sharing early results from one of our first research programs, in which Claude autonomously discovered a novel enzyme system that is associated with an array of DNA repeats, a pattern reminiscent of CRISPR. Although we don’t yet know its function, the system that Claude discovered has a set of characteristics that have only ever been found together in a handful of other systems, all of which are programmable and perform operations like cutting, copying, and pasting DNA. Beyond CRISPR, which has already transformed science and medicine, several other such systems are now in development as promising tools.
The system that Claude found is based on a reverse transcriptase (RT), enzymes that copy RNA into DNA. While this underlying RT, found in a jumbo phage, had been identified in previous studies, Claude appears to be the first to notice the system’s defining features—an associated array of non-coding DNA sequences and an additional accessory protein of unknown function.
After reviewing the pre-print, Feng Zhang, one of the pioneers of CRISPR genome editing and a professor at MIT and the Broad Institute said:
This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation. I hope this work encourages more scientists to explore how AI can support their research.
We gave Claude a prompt to search through a massive database of DNA sequences for interesting new examples of RTs. Our involvement was limited to the initial prompt and the lab work, while Claude agents combed through the database, investigated the distinct RT families, and used their own judgement to identify interesting candidates. After 21 hours spent searching this data by roughly 950 agents using 210 million tokens, one of the agents spotted something remarkable: a repeating pattern of DNA sequences that occurs next to the gene for an odd-looking RT. After further analysis and testing in our lab, we recognized that this pattern marked a previously uncharacterized enzyme system found in bacteriophages (the viruses that infect bacteria) that we call array-associated reverse transcriptases (ART).
Our work to understand the primary function of ARTs is ongoing. However, we think it is important to share such findings early, both to demonstrate Claude’s capabilities and to give the broader community insight into what we’re working on. We have released a pre-print (here) that discusses this in more detail.
About our lab
We are a team of scientists who have spent our careers exploring unusual proteins, and specialize in using computational approaches to systematically read DNA, interpret its evolution, and pick out biological systems for further characterization. Our research prior to joining Anthropic has helped to better understand the evolution and regulation of CRISPR systems, discover new enzymes for next-generation cell and gene therapies, and build tools for accelerating the identification of anomalies in DNA, such as human pathogenic variants. We are part of Anthropic’s life sciences organization, alongside teams whose work includes drug discovery, and training Claude in biology and chemistry.
Our lab, located in the Bay Area, looks like a typical molecular biology lab. We do research that involves only the lower-levels of the biosafety risk level (BSL-1 and BSL-2) and we do not handle pathogens that can infect humans. All of the lab work is performed by human scientists. Although we’ve experimented with using AI to accelerate lab work with initiatives like the Model Hardware Standard, this approach is less conducive to the sort of ad hoc workflows that are involved in our molecular biology research.
How we work
Many of our workflows involve having Claude search through the vast collection of DNA sequences associated with proteins without a known function. One typical pattern begins with a survey of a given protein family. Claude reads the relevant literature and reproduces the established results from public data to check its methods. It then searches for family members or genomic neighbors that fit no described system, and writes a short, human-readable report for each candidate that proposes a function and describes the evidence supporting its claims. In follow-up analyses, Claude critically evaluates the evidence—typically most candidates are eliminated at this stage. A survey may end with a single candidate worth testing, or with none.
When a candidate survives our review, we test it in the laboratory, expressing the protein in standard laboratory strains and characterizing it biochemically and structurally, with Claude helping to interpret the data. We do our work in Claude Science and Claude Code, the same tools available to any scientist, and sometimes with a harness of our own that coordinates many Claude sessions running in parallel.
Because Claude produces hypotheses so prolifically, the hypotheses themselves have become an object of study for us. With hundreds to thousands of candidate reports from a single campaign, we have been asking what distinguishes the proposals we judge worth testing from those we set aside. What we learn goes back into the instructions we give Claude and teaches it to mimic our own scientific taste.
Claude finds ART
In the past few years, researchers have discovered many more reverse transcriptases (RTs), most of them in bacteria, where they act as part of the immune system. Nearly all RT families were found by genomic analysis, or genome mining, which requires researchers to search sequence databases for genes that no one has characterized, notice the unusual ones, and work out what they do.
Claude agents gathered over 200,000 RTs, picked out 3,500 new candidate systems, and narrowed those to the 20 most-compelling candidates that they analyzed to produce human-readable reports. For an expert scientist, this type of analysis can take weeks to months of work.
During the course of its research, Claude noticed an unusual RT family and decided to examine it in greater detail. While combing through the raw DNA sequence near the RT, the agent exclaimed: “[The DNA next to the RT] is spectacular: I can see by eye a tandem repeat array … that's a CRISPR-like … repeat array?!”

It then proceeded much as a human scientist would when faced with a potential discovery. It counted the repeats and measured their spacing, compared the layout with the known RT systems, and searched the literature for any previous report of the pattern. After a thorough analysis it was convinced that it had found a new biological system, and filed a report for human review.
The system it found, ART, is found mainly in bacteriophages and consists of three parts: the RT, a partner gene beside it, and a long array of evenly spaced DNA repeat sequences. The repeat layout resembles a CRISPR array, which holds a bank of different RNA sequences that make CRISPR-Cas systems programmable biotechnological tools. Our first experiments show that the ART array is also expressed as a set of distinct short RNAs, suggesting that something analogous may be at play for this system.
Further experiments are underway to determine how ART works, and we are sharing these early findings to show the community that Claude can autonomously detect anomalies and drive analyses to initiate biological discoveries.
You can find more detail in our technical report (here).
Work with us
We hope this work demonstrates the value of AI-driven hypothesis generation to the wider scientific community, and we would like to work with other scientists to extend this approach to a broad range of problems, in genomics and in other fields. If you have a proposal for a research question, we would like to hear from you.