Thomas Wolf 团队发布开源模型 Carbon-A,直接在 DNA 中定位基因,并据此构建覆盖 22617 个物种、共 5.66 亿候选基因的 Carbon Annotation Database。模型与数据库一并开放,团队与 Active Site 和 UCSD 合作在猫、鸡和拟南芥等物种中完成部分新基因的湿实验验证,并找到 239 个参考注释中缺失的 RNA 证据基因。
Scientists have sequenced the genomes of thousands of species. But for most of them, nobody knows *where the genes are*
Today we're releasing Carbon-A, an open model that finds genes directly in DNA. We used it to create a database of 566 million candidate genes across 22,617 species, from fungi to mammals (that we are sharing as well)
Why bother? A few reasons
1. Elephants rarely get cancer. Part of the explanation turned up in their genes: extra copies of a tumor-suppressor gene humans have only one of. You can't ask that kind of question about a species until someone, or some tool, has found its genes.
2. The first GLP-1 drug was based on a peptide from Gila monster venom. Nobody would have put the Gila monster on a priority list. Same for wild relatives of crops, which carry resistance to drought and disease. Many of these species have never been annotated.
3. Most of what we know about genes comes from about a dozen model species, like mice, flies and yeast. That focus worked remarkably well, but annotation pipelines still lean heavily on comparisons with them, which makes genes unique to other species easy to miss, and those can be the most interesting ones.
Carbon-A belongs to a newer family of tools that read DNA directly. It learned from known genomes, but it doesn't need a close relative to read a new one.
4. Even familiar genomes still have gaps. With our partners at Active Site and UCSD, we found RNA evidence for 239 genes missing from the reference annotations of species as common as cats, chickens, hamsters and a lab plant.
5. AlphaFold can predict the shape of a protein, but only once someone has found the gene that makes it. In a species with no annotated genes, it has nothing to work with.
Some notes on safety:
Carbon-A doesn't design DNA or predict what a gene does. It marks where genes are in DNA. This is one of the reasons we think releasing it openly is the right call.
More annotations also help health research: many species that carry or cause disease are among those covered, and work on controlling the diseases they spread often starts from their genes.
Model and database and more details in Georgia's thread 👇
Today, we're releasing the Carbon Annotation Database and the model behind it, Carbon-A. We've used this model to discover 566.34 million new gene candidates across 22,617 species, thousands of which have never been studied like this before. We wet-lab validated a number of these new genes in well-studied species, such as in cats, chicken, and arabidopsis (plant), with our partners at Active Site and @UCSD, with more extensive wet lab results on previously-unstudied genomes coming soon.在 X 查看被引用的帖子
来源:Thomas Wolf · x.com