HypoEvolve:用遗传算法让多智能体 LLM 发现科学假设

HuggingFace Daily Papers(社区热门论文)·2026-09-14 08:00·3天前
AI 导读

HypoEvolve 通过代际遗传算法协调多个专业 LLM 智能体,对假设种群进行组合、修订与保留,使协作过程对假设质量的影响可直接检验。在 34 种癌症类型上,该方法在 DepMap 和 Open Targets 两项外部指标上均超越六个基线,DepMap 选择性达 0.171,最强基线为 0.115。相对单次生成的提升还可泛化到未见过的癌症类型。

HuggingFace Daily Papers(社区热门论文)
41AI 编辑部评分,满分 100

HypoEvolve:用遗传算法让多智能体 LLM 发现科学假设

2026-09-14 08:00· 3天前
AI 导读

HypoEvolve 通过代际遗传算法协调多个专业 LLM 智能体,对假设种群进行组合、修订与保留,使协作过程对假设质量的影响可直接检验。在 34 种癌症类型上,该方法在 DepMap 和 Open Targets 两项外部指标上均超越六个基线,DepMap 选择性达 0.171,最强基线为 0.115。相对单次生成的提升还可泛化到未见过的癌症类型。

Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality remains an open question. Answering this question requires separating the effects of agents' scientific capabilities from those of their collaboration. A framework must therefore preserve agents' scientific roles and support rules for combining, revising, and retaining hypotheses.

Building on this view, we introduce HypoEvolve, which makes collaboration explicit through successive updates to a hypothesis population. Specifically, we propose a generational genetic algorithm to coordinate specialized large language model (LLM) agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability. Each generation specifies how scientific judgments and new proposals reshape the population, making collaboration effects on hypothesis quality directly testable. Moreover, we design our evaluation around scientifically meaningful hypotheses that explain how a proposed intervention could work.

Drug repurposing links these explanations to target-level biological claims assessed against external evidence. Specifically, we adapt DepMap and Open Targets into complementary external measures grounded in experimental, genetic, and clinical evidence. Across 34 cancer types, HypoEvolve achieves the highest scores against six baselines on both measures. DepMap selectivity reaches 0.171, versus 0.115 for the strongest baseline. Gains over single-pass generation also generalize to held-out cancer types. HypoEvolve advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org