跳到正文
Goodfire Research·· 1 天前AI 评分59

Goodfire 提出基于蛋白质嵌入的 AI 智能体生物安全监控方法

Better biosecurity monitors for AI agents via protein embeddings

AI 导读

Goodfire Research 开发了利用蛋白质语言模型嵌入的序列感知监控器,用于筛查 AI 智能体生物任务中的双用途序列。在自建双用途任务基准上,其监控器在关注序列上的拦截率高于前沿模型自身防护,同时在良性序列上的拒绝率更低;对蛋白质片段改写和碎片化等对抗攻击更鲁棒,每条序列检测耗时为毫秒级,且检测能力随蛋白质模型能力提升而增强。

正文

Cybersecurity risks from AI agents are now widely acknowledged. We think that biosecurity is the next frontier of AI risks: agents can perform increasingly complicated tasks on biological sequences, like DNA or proteins, without reliably recognizing whether that sequence is concerning. Such tasks are often dual-use; they can support benign or concerning research, depending on the sequences involved. Consequently, task-based monitors must either block legitimate research or fail to flag potentially harmful work.

To address this problem, we developed sequence-aware monitors that leverage protein model embeddings. Combining information from both the specific sequence and the task context enables more precise safeguarding of dual-use research. We evaluate our monitors on a custom benchmark where dual-use tasks can be distinguished only by the sequences involved. Our monitors outperform frontier model safeguards with fewer refusals on these dual-use tasks.

Additionally, our monitor:

  • is more robust to novel adversarial attacks,
  • is as fast as database-lookup methods, and
  • scales with model capability.

By using embeddings, our monitors are robust to a wider range of adversarial attacks such as paraphrasing and fragmentation. This also means that the monitor’s defensive robustness improves as model capabilities improve.

Goodfire works with frontier labs and inference providers to deploy monitors. If your organization is interested, get in touch.

Diagram of sequence-aware monitoring for biological research workflows.

Biological screening in the age of agents

As agents start managing hazardous tools and research workflows, their ability to work around safeguards is of increasing concern. Recently, these abilities have been demonstrated in a cybersecurity context, via incidents involving agents hacking external websites, as well as cybercriminals using models to assist them. Similarly, as agents demonstrate more sophisticated biological capabilities, and agents are starting to be trained through feedback from lab experiments, similar concerns are arising for biosecurity. Unlike cyber attacks, where vulnerabilities can be patched quickly and broadly, “patching” human biological vulnerabilities is very difficult.

A major focus of biosecurity efforts has been on screening sequences at the point of DNA synthesis, a required step for producing any novel proteins. Synthesis uses specialized equipment, so centralized providers provide a natural bottleneck for screening that offers a highly effective line of defense against most misuse. However, as agents and advanced biological models rapidly improve, a greater number of threat actors may be able to design concerning sequences. We believe it will be important to take a defense in depth approach that screens at both the point of synthesis and the point of design, in particular during agentic workflows.

Yet safeguards on agentic workflows have been challenged by the dual-use nature of many tasks in biological research. For example, common protein design tasks like generating new candidates through inverse folding and re-folding can be used to generate “paraphrases” that maintain the same structure but alter the sequence. Similarly, before sending sequences out for DNA synthesis, long sequences are divided up to meet length requirements, but this fragmentation can make recognition of concerning sequences more difficult.

For these dual-use tasks, the difference between a concerning use case and a benign one lies in the sequences themselves. Current models, though, lack a good understanding of biological sequences, so screening based on the content of the sequences is challenging. Here, we develop monitors that use embeddings from protein language models to predict the level of concern for sequences.

Embeddings give access to biological relationships that sequence-based (HMM, family, …), structure (TM) or other metrics alone may miss. By leveraging these rich embeddings, we find our approach to be more tolerant to held-out samples and novel classes of attacks. Our results show that this attack detection scales with protein-model capability and dataset quality. This scaling holds for both held-out natural sequences and held-out attack examples. We thus believe that more advances in biological models will continue to help us preempt attacks and make screening more robust to workarounds, without having to patch as new attacks emerge.

Interpretability also allows us to study embeddings to provide evidence about harmful properties or provenance. Such assessments could give expert reviewers insight into how a sequence might have been generated.

A Pareto improvement for dual-use screening

For dual-use tasks, the biological sequences involved solely determine the level of concern. To evaluate how well frontier models can differentiate based on a sequence’s concern level, we have developed a benchmark of dual-use tasks, such as paraphrasing and fragmentation, where safety depends on the sequences involved. Each task prompt can include either a concerning or benign sequence as input, and is framed either as a benign use-case, or simply describes the method and tools to use. This setup is similar to LatchBio’s BioSecBench-Refusal benchmark, which tests risky vs. routine dual-use tasks, but our tasks are identical except for the input sequences which determine the concern level.

For this benchmark, we use protein toxins as our concerning sequences. The input prompts include both natural proteins as “known hazards” and pre-generated adversarial paraphrases that have similar structure and folding quality to their natural proteins (TM ≥ 0.5 and ΔpLDDT ≥ -10, similar to Wittmann et al 2025). To have a consistent definition of “concerning”, we ensured that our concerning sequences and parents of the paraphrases would be flagged by a screening tool.

Comparison of frontier model safeguards and Goodfire monitors on known hazardous sequences. Comparison of frontier model safeguards and Goodfire monitors on adversarial protein paraphrases.

Figure 1: Comparison of frontier model safeguards on inverse folding and fragmentation related requests, both dual-use biological tasks. Towards the top left is better, which means that a monitor flags a higher proportion of harmful examples while allowing benign requests.

Most frontier models have similar rates of refusal for both concerning and benign sequences, indicating that they classify primarily based on the task request, especially as sequences are obfuscated. In contrast, Goodfire’s monitors on protein model embeddings provide strong refusal rates on concerning sequences and low rates on benign ones, providing a much better trade-off for these dual use tasks. Our system extracts the sequences from the user interaction and runs them through our monitor.

Robustness to adversarial attacks

Our embedding-based classifiers consistently outperform the evaluated baselines on tasks where ordinary screening methods commonly struggle, such as paraphrases and fragmentation. Specifically, on fragments of held-out natural proteins, our method beats sequence-based screening at every tested length. This extends the advantage beyond full-length proteins to fragments containing only a small portion of the original sequence.

Protein-level AUROC for fragmentation detection across tested fragment lengths.
Figure 2: Fragmentation detection using our classifier, compared to a naive version of our approach and a sequence-based screening tool. We find that embeddings outperform ordinary sequence screening classification at all tested fragment lengths. We report a protein-level AUC: a protein is considered flagged when at least one of its parts classifies as positive.

For paraphrases whose original proteins were present in the screening database, our approach achieved an AUROC of approximately 0.98. When the originals were absent, performance declined across all methods, but our approach retained a clear advantage at low false-positive rates. Generalization also extended to held-out protein families, a mechanism holdout of enzymatic toxins, and Exo-Tox, an external bacterial-toxin dataset. The paraphrase results rely on in-silico estimates and do not establish whether the redesigned proteins retain biological activity.

Partial AUROC at false-positive rates up to 1% across held-out datasets and screening tools.
Figure 3: Partial AUROC over false-positive rates up to 1% across diverse held-out sets and tools. We distinguish between seen and unseen paraphrases based on whether the parent was in the database. We also evaluate generalization by holding out whole families or enzymatic toxins (mechanism). Finally, we test an external out of distribution dataset (exo-tox) which contains bacterial toxins.

Efficiency and speed

Our method is also quite fast, taking milliseconds per sequence—comparable to sequence search in our benchmarks. Detection also improved with more capable protein models, offering a way to strengthen screening as biological design capabilities advance. Together, these results demonstrate broader detection without sacrificing the speed needed to screen sequences throughout research workflows. That speed means our monitors can be deployed widely, such as in automated hardware or experimental tools.

Screening runtime in milliseconds per request, with cold-start costs amortized over 1000 sequences.
Figure 4: Milliseconds per request, cold-start amortized over 1000 sequences. Our approach works on a millisecond level.

Our vision for monitoring biological misuse

In a defense in depth approach, each screen can combine sequence evidence with the context available to it. Agents have context about intent, biological design tools have the precise objective, and synthesis providers have customer information and access to expert review. These layers should add different evidence and opportunities to intervene, allowing ambiguous flags to prompt review before blocking and reducing pressure on single chokepoints to be overly cautious.

The presented results show that precise automated monitoring is feasible at scale. We show that current embeddings contain sufficiently rich information to monitor against attacks like paraphrasing and fragmentation, where we beat existing baselines. Embedding-based methods will
only improve in the future as they scale with biological capabilities and interpretability tools. The combination of increasingly sophisticated judges and better predictions about biological sequences will be the strongest defense we have to agentic biological misuse.

Beyond monitoring agents, we think it may become increasingly important to monitor generative biological models themselves. Just as language models are fine-tuned to be more helpful and harmless, we believe biological models can be trained or steered towards less risky behavior. Additionally, those training advanced biological models can conduct evaluations before release to assess capabilities on harmful tasks.

Our goal is to maximize the amount of biological research that can be done safely. Better safeguards should give researchers more room to use increasingly capable biological models. We want the same advances in biological understanding that make new research possible to also give us the means to carry it out safely.

Goodfire works with frontier labs and inference providers to deploy monitors. If your organization is interested, get in touch.

Acknowledgements

We thank Tessa Alexanian, Gary Abel, Isha Harris, and Conor McGurk for their feedback on drafts of this post.

来源:Goodfire Research · goodfire.com