# Basecamp Research 获 1.4 亿美元融资，用全球生物样本训练 EDEN 生物模型设计抗生素与基因治疗工具

- 来源：The Decoder：AI News（RSS）
- 作者：Maximilian Schreiner
- 发布时间：2026-09-23 21:52
- AIHOT 分数：56
- AIHOT 链接：https://aihot.news/items/cmue5z5pa0o5uroghdq72lty5
- 原文链接：https://the-decoder.com/inside-basecamp-research-the-ai-startup-turning-evolution-into-training-data

## AI 摘要

Basecamp Research 获 1.4 亿美元融资，投资方包括 S32、Nvidia、Anthropic 的 Anthology Fund 和 NATO Innovation Fund，用于开发生物 AI 模型 EDEN 并推进体内细胞疗法进入临床。

## 正文

Basecamp Research / GPT-Image-2 prompted by THE DECODER

Basecamp Research has raised $140 million from investors including Nvidia and Anthropic's Anthology Fund. The London company trains AI models on genetic material from rainforests, oceans, and hot springs to design antibiotics and tools for cell therapies. In an interview with THE DECODER, CTO Phil Lorenz explains why biology is a far bigger problem for AI than language, and why good scores on paper don't guarantee good molecules.

According to the company, investor S32 led the round. Nvidia, Anthropic's Anthology Fund, the NATO Innovation Fund, and Redalpine also took part, among others.

Founded in 2020, Basecamp plans to use the money to keep developing its biological AI models, called EDEN, and to push its own therapy candidates toward clinical development. Cell therapies come first. These treatments genetically modify a patient's cells so they can fight cancer, for example. Basecamp wants to make that change directly inside the body.

To get there, the company looks to microorganisms from habitats researchers have barely studied. The models are supposed to learn patterns from their genomes that can be put to medical use.

"Given the complexity of biology, we just need orders of magnitude more data," Phil Lorenz, Basecamp's CTO, tells THE DECODER. What matters most, he says, is pairing that massive data collection with smaller, targeted experiments.

There's still a lot of development work ahead before any of this becomes a treatment. So far, Basecamp has lab results and data from mice. That doesn't show whether its approach can produce safe and effective therapies in humans.

Why public genome databases can't carry biological AI

The Wall Street Journal recently reported that pharma companies are pouring billions into AI, but solid evidence of noticeably higher success rates in clinical development is still scarce. Lorenz doesn't see that as a reason to doubt the technology. Scaling delivered big gains for language models very quickly. Biology is making progress too, he says, but the big open questions, like speeding up clinical trials, are far harder. As a problem, biology is "way, way, way more complicated and way larger."

The data is also fundamentally different. Language models learn from huge text collections, while pharma companies often work with small, siloed datasets from individual studies.

Lorenz uses an estimate from Epoch AI to show how wide the gap is. It puts the upper limit of all existing words at roughly five quadrillion tokens. Basecamp estimates the number of nucleotides on Earth at 10 to the power of 37. "If you took a stack of cards, like poker cards, of 10 to the power of 37 cards, that stack would surround the observable universe a million times," Lorenz says. You don't need a dataset that big, he says, but you do need one far bigger than what exists today.

Public genome databases also paint an incomplete picture of nature. According to Basecamp's research paper on its BaseData database, about 68 percent of the sequence volume in the Sequence Read Archive comes from just five species. Humans alone account for about 54 percent.

Basecamp Research works with researchers in dozens of countries to collect samples on site. | Image: Basecamp Research

That focus makes sense for medical research. For a model meant to learn as many different biological processes as possible, though, Basecamp sees it as a limitation, since many other life forms barely show up in the data. "If you were to train LLM only on newspaper articles from 1975, it would be a really, really bad model," Lorenz said in a case study from Microsoft, the company's cloud partner. "That is kind of where we are in biology."

So Basecamp collects its own samples with local research partners, from rainforest soil, volcanic ground, and the deep sea off Antarctica. According to the latest funding announcement, the network now spans more than 30 countries and all seven continents. Microsoft puts the number of participating organizations at 208 across 31 countries.

At the time of the interview, the dataset held about 15 trillion tokens, according to Lorenz. That already puts it in the same range as the text datasets used to train models like Claude or GPT. Here, though, a token is a single DNA building block. The AI doesn't process words like a chatbot does. It processes strings of characters that describe genetic material.

According to Microsoft, the first EDEN generation trained on 9.7 trillion DNA building blocks from more than a million newly sequenced species. Basecamp ran the training with Microsoft researchers on Azure, and the compute reportedly matched GPT-4's. OpenAI has never officially disclosed that figure, however.

Over the next year and a half, the dataset is supposed to grow roughly a hundredfold, passing one quadrillion tokens. The framework for that is the Trillion Gene Atlas announced in March, in which Basecamp, Anthropic, Nvidia, PacBio, and Ultima Genomics plan to assemble genetic data on the scale of one trillion genes.

In its latest funding announcement, Basecamp already calls the Trillion Gene Atlas the training foundation for its EDEN models. It doesn't say how far the planned expansion has gotten.

Bacterial arms races offer a blueprint for new drugs

Lorenz sees evolution, which drives all biological processes, as the link between environmental samples and medicine. Whether a bacterium in a hot spring adapts to heat or a cancer cell spreads, similar selective pressures are at work. The genomes of humans, mammals, and nearly all vertebrates are largely sequenced, while the rest of biodiversity has barely been cataloged. Many drugs also originally came from plants, bacteria, and fungi.

He points to competing microorganisms as a concrete example. Some of the most powerful therapeutic molecules come from "biological warfare" between organisms, he says. When a bacterial species tries to dominate an ecosystem, it sometimes develops substances that inhibit or kill rival species, securing food and habitat for itself. This happens billions of times in countless places on Earth, especially where nutrients are scarce. In medicine, such substances can be useful as antibiotics.

In the Microsoft case study, Lorenz names phages as another example. These viruses infect bacteria and inject their DNA into the cells. That back-and-forth produced tools that can be used for gene editing.

Basecamp wants to learn from a wide range of these evolved solutions. The models aren't just supposed to rediscover known compounds. They're meant to design new candidates from the patterns they've learned.

To do this, the company records the chemical, physical, and ecological conditions at each site along with the genetic material. It also focuses on long, continuous stretches of DNA. These show which genes sit next to each other and might work together. Short, isolated fragments often lose that context.

This figure from a 2024 paper compares the protein diversity captured in public databases with Basecamp's holdings at the time. | Image: Vince, Gowers, and McGibbon, GEN Biotechnology, CC BY 4.0

An AI-designed antibiotic clears its first animal test

Lorenz sees EDEN's antibiotic designs as early proof of medical value. "We prompt on a pathogen and then the model designs an antibiotic that kills it," he says.

In a research paper on the EDEN model family that hasn't been peer-reviewed yet, the authors report on a small, curated set of antimicrobial peptides, short protein chains that can attack bacteria. According to the paper, 97 percent of the tested candidates showed activity in the lab.

That rate only applies to the tested selection, not to anything the model might design. Basecamp says it didn't stop at test tube experiments, though.

A company announcement from June describes tests with a candidate called EDEN-7. In mice infected with multidrug-resistant bacteria, it reportedly worked about as well as a last-resort antibiotic, the kind of drug reserved for hard-to-treat infections.

According to Basecamp, the model generated the candidate directly, with no rounds of tweaking before the test. The work was done with researchers at the University of Pennsylvania led by Cesar de la Fuente.

Lorenz puts a lot of weight on this step from design to experiment. Basecamp isn't just about collecting data and training models, he says. What matters is whether the molecules actually work.

Besides antibiotics and the insertion tools described below, Basecamp also had EDEN generate a synthetic gut microbiome of about 9,000 bacterial species, according to Microsoft. To check the results, the team brought in UC Berkeley microbiome researcher Jill Banfield, who set strict criteria and is now a co-author of the related paper. "It is good to ask a skeptical human and not just AI," Lorenz said.

Basecamp's therapy bet rests on inserting genes inside the body

For its own therapy work, Basecamp is using EDEN differently for now. The model is supposed to design biological tools that insert larger stretches of DNA at chosen spots in the human genome.

These include large serine recombinases, enzymes that can rejoin DNA. Basecamp wants to design them so they insert therapeutically useful genetic information into cells at precise locations.

The approach targets a core problem in gene therapy. Many inherited diseases stem from different mutations depending on the patient. Instead of fixing each one, the enzymes would insert a healthy copy of the gene regardless of the specific mutation. These insertion tools come from that same conflict between phages and bacteria, and they have to be reprogrammed before they can work in the human genome.

In the EDEN paper, the authors report several working enzyme candidates for every disease-relevant target region they studied. Half of the generated serine recombinases were active in human cells.

One possible use is CAR-T cells, genetically modified immune cells designed to recognize and attack cancer cells. In an announcement from January, Basecamp said primary human T cells modified this way, meaning immune cells taken directly from donors, cleared more than 90 percent of tumor cells in the lab.

With the new funding, the company wants to develop the approach into a treatment that works directly inside the body. That would simplify the complex manufacturing of cell therapies, which can currently cost hundreds of thousands of dollars per patient, according to the funding announcement.

A genetic edit that works in the lab doesn't answer how the necessary components get safely to the right cells in the body, though. Arc Institute co-founder Patrick Hsu raised this point in a Financial Times report from May 2025. He named delivery, possible toxic effects, and the cells' defense responses as open challenges.

In the long run, the principle behind the antibiotic designs is supposed to extend to diseases in general. "Our ambition is to make biology programmable. You prompt on a disease and out comes a molecule that will address that," Lorenz told Microsoft. He admits the company is still far from that goal. "These molecules are great, they are active, they work. But for them to go into the clinic, there is a lot more that needs to happen."

The therapy pipeline Basecamp published on its website shows how long the road still is. As of September 2026, none of its six programs has moved past lead optimization, the stage where promising candidates get refined. Preclinical studies come next, which are required before applying to test in humans, followed by clinical trials.

Basecamp's published therapy pipeline shows none of its six programs has left lead optimization yet. | Image: Basecamp Research

Four programs are in lead optimization. They include in vivo CAR-T cell therapies for blood cancer and for an unnamed autoimmune disease, a liver gene therapy for the inherited metabolic disorder phenylketonuria (PKU), and antimicrobial peptides against resistant pathogens. All of the cell and gene therapy programs rely on large serine recombinases. A CAR-T therapy for solid tumors and peptides for diabetes are still in early research.

The overview doesn't say which program will go into human trials first, or when. Basecamp didn't initially answer THE DECODER's questions on this, or on its preclinical results, their independent review, and comparison data with other AI systems. The company said it would respond later.

Why a third of Basecamp's GPUs go to reinforcement learning

For Lorenz, the data from nature is a foundation, not a replacement for medical experiments. At its lab in Boston, Basecamp generates data from experiments with T cells, a type of human immune cell, among other work.

These experiments don't produce billions of tokens. They usually yield hundreds to thousands of data points. In return, they answer narrow questions, like whether a genetic change works in a specific cell type.

Basecamp wants to use these results to steer an already broadly trained model toward specific tasks. That includes extra training rounds with selected examples as well as reinforcement learning, where the model gets adjusted based on scores for its outputs.

"Right now a third of our GPUs are reserved for reinforcement learning experiments," Lorenz says, referring to the graphics processors the company trains its models on.

He doesn't say what measurable gains these experiments have delivered so far. But the way Basecamp splits its compute shows it isn't betting only on feeding pretraining more and more data. Reinforcement learning has driven much of the recent performance gains in language models like GPT-6.

Basecamp also wants to broaden the data itself. According to Lorenz, the first EDEN model family was a pure DNA model. DNA encodes proteins, which in turn fold into structures. That means a DNA model already picks up many signals you'd otherwise need separate protein or structure models to capture. The next generation is supposed to handle more complex, more clinical tasks, like the safety and toxicology assessments mentioned earlier, and process other kinds of data beyond DNA sequences. Lorenz won't say which ones.

Better benchmark scores don't mean better molecules

"The more data [we] collect, the more I think we need to collect even more data," Lorenz says. Basecamp has grown its dataset roughly tenfold each year so far. With every new training run, he says, the team sees "performance increases, but never a performance plateau."

Data quality matters too, according to Lorenz. Basecamp puts a lot of effort into processing samples, from extraction to sequencing to assembly. Over the past six months, the company trained identical model architectures under otherwise identical conditions on its own data and on public data. Its own data produced a steeper scaling curve. The gap was small for small models and grew as the models got bigger.

Lorenz measures this with perplexity, which tracks how well a model predicts the next characters in a sequence. Lower is better. To reach a perplexity of two, Basecamp's data saves several hundred thousand to millions of GPU hours, he says. For a perplexity of 1.5, the extrapolation points to a difference of billions of GPU hours. That figure is a projection, not a measured result.

The basic principles of the scaling laws known from language models also hold in biology, Lorenz says. The specific ratios differ by model type, though, for example between DNA and protein models.

At the same time, he warns against relying on technical metrics alone. Basecamp compared several model architectures, including StripedHyena, the architecture behind the Arc Institute's Evo model, and Llama. StripedHyena sometimes reached lower perplexity, but Llama did better on the biological tasks that followed.

That's why Basecamp went with Llama, Lorenz says. The company also regularly checks training checkpoints to see how well the models handle biological tasks. As an example of a capability that only shows up clearly as models grow, he points to predicting possible immune reactions to a molecule. It's weaker in the smallest EDEN variant. He didn't share specific numbers.

"Not to overpromise what AI can or can't do, but just do the experiments, figure out what it does," Lorenz says. The team is often surprised itself by which properties emerge with scale.

Synthetic training data is a lower priority for now. According to Lorenz, it only adds more points in regions of sequence space that are already known and doesn't open up new ones. Basecamp mainly wants to capture biodiversity nobody has seen before. In internal tests, models trained on synthetic data sometimes produced working proteins, but they never achieved real generalization to new tasks.

Chatbots become the front door to biological models

Running biological models is supposed to stop being a job only for specialists. Basecamp has made selected EDEN features for antibiotic design and picking potential vaccine targets available through Claude and Claude Science.

Researchers can describe their task in plain language. Claude then taps EDEN to suggest candidates for further study, for example. For vaccines, the first question is which parts of a pathogen could serve as targets for a protective immune response.

Lorenz says Basecamp won't release the models openly. He cites safety concerns and agreements with data partners, and he sees the two as closely linked. Many partners prefer a model whose commercial use earns them a share over an open-source model whose use nobody can track. The Arc Institute, by contrast, released its Evo model openly and left pathogen data out of training.

Basecamp also removes certain virus data from its training material, among other steps, and screens generated sequences against databases of known pathogens. These measures are meant to lower the risk of misuse. That doesn't prove they fully prevent problematic outputs, though.

Who gets paid for nature's data is still up for debate

The data strategy also depends on whether countries and local institutions are willing to work with Basecamp. The company doesn't start sampling without prior informed consent from landowners and from local and national governments, Lorenz says.

"Every single token that EDEN has been pretrained on, we can trace back to the exact geographic location and the actual consent that was given to us," he says.

Partners shouldn't have to wait until a drug possibly reaches the market years later to benefit. Lorenz points to hiring local scientists and investing in sequencing technology and labs meant to build a local "bioeconomy." The BaseData paper also describes licensing payments split according to each partner's share of the training data used.

According to that paper, Basecamp had made payments to 52 beneficiaries in 19 countries by the end of 2024. In the fastest case, nine months passed between sampling and the first payment.

Lorenz, who worked with open databases back when he was in academia, doesn't see this model as an alternative to public data but as a parallel system. Consent and benefit-sharing build trust, he says, and that gives partners a reason to scale data collection to the level Basecamp needs for its models.

Whether that share is enough is disputed. The Financial Times reported in 2025 that the payments amount to one percent of revenue. Researcher Jim Thomas told the paper he appreciated that the data is traceable but criticized the size of the payments. The data also drives up the company's value, he argued, and the communities it comes from don't get a fair cut of that.

Aurelie Dingom of Cameroon's environment ministry also said higher shares would be fair. Basecamp's founders pointed the paper to contractual safeguards and said long-term partnerships are in the company's own business interest.

Basecamp now has another $140 million for its next step in medical development. Its experiments so far show that EDEN can design biologically active candidates. Upcoming studies will have to show whether those candidates turn into safe, effective treatments that are easier to manufacture.
