为生命科学行业带来更强大的智能,扎根于真实的科研工作流程。
我们正在为 GPT‑Rosalind 系列推出一次新的模型更新,该系列专为企业级规模的生命科学研究而打造。它将 GPT‑5.5 的智能体编程与工具调用能力,与在药物化学、基因组学等核心药物发现领域更强的模型智能相结合,同时提升在更广泛的生命科学分析、设计与实验工作流程中的表现。
生命科学的进步取决于跨尺度、跨模态地综合数据与证据:分子、基因、通路以及生命系统。在我们的评估中,更新后的 GPT‑Rosalind 在来自生物学专家的研究任务、复杂的药物化学查询、定量生物学以及湿实验室故障排查上均展现出广泛的性能提升。
GPT‑Rosalind 现已通过我们的可信访问部署架构,以研究预览形式面向全球符合条件的机构开放。
提升在具有科学价值的任务上的表现
为了衡量并持续提升 GPT‑Rosalind 在真实世界中的影响力,我们设计了 LifeSciBench,这是一个由外部专家评审的基准,聚焦于生命科学研究中的基础性方面。与现有那些孤立评估模型性能某一组成部分或某一生物学领域的基准不同,LifeSciBench 采取端到端视角来审视具有科学价值的工作,其任务取自生命科学研究中六个核心工作流领域:证据处理、分析与设计优化、科学推理、验证与运营,以及转化与沟通。
我们使用这一基准来使进展与生命科学研究的实际需求和现实保持一致。
LifeSciBench 总体得分
LifeSciBench 按科学工作流划分的得分
GPT‑Rosalind 在由产业界和学术界专家认定的具有科学价值的任务上全面领先。
从论文、图表、表格和实验记录中提取、核对并审计科学证据。
评估示例
我们正在筹备与 FDA 就 AAV9-microDys-X 举行的 Type B 会议,这是一种基于 AAV9 的 micro-dystrophin 基因疗法,用于治疗杜氏肌营养不良症,其通过 MCK 启动子表达一个 138 kDa 的构建体。我们希望得到一份毫不留情的批评意见,评估我们当前的申报资料是否真正支持以 micro-dystrophin 表达作为替代终点获得加速批准,该终点需合理可能地预测临床获益。
研究背景:一项开放标签 1b/2 期研究,纳入 12 名年龄 4–7 岁、确诊 DMD 且携带框外杆状域缺失的可行走男孩。该数据包为:
- 治疗前股外侧肌活检:使用针对 N 端肌动蛋白结合域的 MANEX1A 进行定量 Western blot,抗肌萎缩蛋白为健康对照的 0–3%。
- 治疗后 12 周对侧股外侧肌活检:采用相同的 Western blot 方法,微抗肌萎缩蛋白平均为健康对照的 38%(范围 18–61%),并通过考马斯染色以总蛋白进行归一化。
- 治疗后免疫荧光:使用多克隆抗肌萎缩蛋白 C 端抗体,75–95% 的肌纤维显示肌膜信号。
- 48 周功能:NSAA 较基线平均变化 +1.4 分,而外部已发表自然史登记队列为 −0.6 分(非配对 t 检验 p = 0.03)。
- 安全性:8/12 例患者出现一过性转氨酶升高,通过激素逐渐减量处理;1 例心肌炎已缓解;无死亡。
- 生物分布/持续性:12 周时可在肌肉中检出 AAV9 载体基因组,平均为每个细胞核 2.3 个载体基因组。
- 入选标准:基线无抗 AAV9 中和抗体(滴度 <1:400),且无外显子 44 缺失。
请逐项对这个申报资料包进行压力测试:FDA 或持怀疑态度的审评员会在哪些地方指出,按当前呈现方式,证据不足以支持我们的结论?如果可能的话,还需要哪些额外数据、分析或设计变更来弥合这些差距?对于任何未明确说明的内容,假定截至 2025 年 12 月 31 日,从其他临床级 138 kDa 微型抗肌萎缩蛋白项目中已知的信息以及相关临床试验也同样适用。
候选回复
核心结论: 按当前呈现方式,这个申报资料包不足以支持以微型抗肌萎缩蛋白表达作为合理可能预测临床获益的替代终点获得加速批准。持怀疑态度的 FDA 审评员很可能认为,生物标志物、功能性、持久性、安全性以及可推广性方面的证据均存在实质性不足。
| 资料包项目 | 关键失效模式 | 需要什么 |
|---|---|---|
| Western blot 定量 | MANEX1A 结合的是内源性全长抗肌萎缩蛋白和微型抗肌萎缩蛋白转基因共有的 N 端表位,因此该检测方法无法清晰区分转基因与残留/回复型抗肌萎缩蛋白。以健康全长抗肌萎缩蛋白作为标准品来定量 138 kDa 微型抗肌萎缩蛋白也是无效的。 | 使用重组微型抗肌萎缩蛋白标准品,以及一种能够区分转基因与内源性抗肌萎缩蛋白的正交方法,例如靶向质谱法或转基因特异性/表位特异性检测。 |
| 免疫荧光 | C 端多克隆抗体不太适用,因为 138 kDa 构建体缺少 C 端结构域。许多 DMD 患者具有回复纤维,而回复型抗肌萎缩蛋白可保留 C 端表位。回复纤维可能随年龄增长而克隆性扩增,从而对 IF 信号产生偏倚,尤其是在年龄较大的男孩中。 | 使用针对转基因中存在但回复型抗肌萎缩蛋白中缺失的表位的抗体重复 IF。将转基因阳性纤维与回复纤维分别定量。 |
| 替代终点的有效性 | 该申报资料将蛋白量与临床功能混为一谈。“健康对照蛋白量的 38%”并不意味着正常抗肌萎缩蛋白功能的 38%,因为微型抗肌萎缩蛋白在结构上是截短的。 | 在将表达量作为替代终点之前,需实证验证微型抗肌萎缩蛋白质量百分比、肌膜定位、下游功能恢复与临床获益之间的关系。 |
| 活检设计 | 治疗前和治疗后的对侧股外侧肌活检会引入左右侧及肌内空间变异性。疾病进展和纤维脂肪替代也可能改变以总蛋白归一化的信号。 | 使用一致的解剖标志标准化活检部位,以肌肉特异性蛋白进行归一化,并同步测量纤维脂肪组成。 |
| NSAA 对照/统计学 | 外部自然史队列并非随机同期对照。试验入组标准、支持性护理、参与效应、基线 NSAA、类固醇方案、年龄和外显子类别都可能使比较产生偏倚。非配对 t 检验并不充分。此外,+1.4 的 NSAA 变化在该年龄组的重测变异性范围内。 | 开展随机同期安慰剂对照研究,或至少使用校正分析,纳入基线 NSAA、年龄、类固醇方案、外显子类别及其他混杂因素。 |
| 年龄窗口混杂 | 4–7 岁男孩正处于发育窗口期,未经治疗的可行走 DMD 患者在衰退占主导之前可能获得运动功能改善。48 周的 NSAA 变化混合了发育性增益、疾病进展和可能的治疗效果。 | 使用按年龄分层的同期随机对照,将发育轨迹与治疗效果区分开来。 |
| 既往临床先例 | 开放标签的微型抗肌萎缩蛋白功能信号未能可靠预测确证性获益;已发表的先例包括微型抗肌萎缩蛋白基因疗法的确证性试验未能重现开放标签的 NSAA 改善。 | 不要依赖开放标签的 NSAA 变化作为决定性支持。需要有对照的功能性证据。 |
| 构建体的结构性局限 | 138 kDa 构建体删除了 spectrin 重复序列 R16/17,其中包含 nNOS 结合位点。nNOS 招募的缺失可能损害运动期间的功能性交感神经溶解和缺血保护,从而在表达水平之外形成机制性的救援上限。 | 补充机制研究,以显示这一特定构建体是否恢复了相关的抗肌萎缩蛋白相关复合物功能、nNOS 定位、运动生理学以及肌肉保护。 |
| AAV 持久性 | 12 周时的载体基因组并不能确立持久的表达。AAV9 基因组主要是非整合的游离体,可能随时间推移而减少。载体基因组的持续存在与蛋白质表达的持续存在并不相同。 | 测量 12 周之后的纵向转基因蛋白质表达和功能性生物标志物的持久性。 |
| 免疫/安全性特征 | 8/12 例患者出现转氨酶升高,这与针对 AAV 转导细胞的免疫反应一致,但机制尚未确立。考虑到 AAV9 对心脏的嗜性,出现 1 例心肌炎令人担忧。 | 提供更深入的免疫监测、肝脏/心脏安全性表征,并加强心脏随访。 |
| 患者选择/可推广性 | 排除抗 AAV9 中和抗体阳性患者限制了可推广性。排除 exon-44 缺失限制了对该 DMD 亚组的适用性。n=12 样本量太小,无法在更广泛的 DMD 人群中表征安全性和有效性。 | 在可能的情况下放宽入组资格,或在使用该结果支持广泛获批之前,预先指定按抗体状态、基因型/外显子类别、年龄和基线功能进行的分层分析。 |
监管结论:该申报资料可能显示出生物学活性,但尚未确立所测得的微型抗肌萎缩蛋白表达是一个可靠的替代终点,且合理可能预测临床获益。主要缺口包括:检测方法的特异性、无效的定量标准、可能的回复纤维混杂、缺乏随机对照、年龄相关的 NSAA 混杂、持久性不确定,以及尚未解决的安全性和可推广性问题。
为弥合这一差距,该方案需要一项受控的、按年龄分层的临床设计,包含转基因特异性表达检测、正交蛋白定量、组织成分对照、纵向持久性数据、针对截短构建体的机制性功能检测,以及更严格的安全性监测,尤其是肝脏和心脏方面。
评分标准与等级
标准
分值
指出 micro-dystrophin 定量中的检测/测量问题,包括 MANEX1A 表位共享、无效的全长 dystrophin 标准品,以及需要重组或正交的转基因特异性测量。
解释为何 micro-dystrophin 表达水平并不自动构成功能性临床获益的有效替代指标。
指出活检部位、组织成分和年龄窗口等混杂因素,这些因素削弱了表达和 NSAA 解读。
批评 NSAA 的对照/统计学方法,尤其是对外部自然病史对照的依赖。
讨论 AAV 持久性、免疫反应、转氨酶升高、心肌炎,以及需要更长期的表达/安全性随访。
指出患者选择/可推广性方面的不足,包括抗 AAV9 排除、外显子 44 排除,以及样本量小。
更强的科学推理能力
药物化学
GPT‑Rosalind 在药物化学领域取得了业界领先的表现,该领域专注于将分子转化为有用的药物。我们设计了 MedChemBench,以反映真实的药物化学工作流程,评估多模态化学结构理解;构效关系(SAR);药物效力、毒性以及吸收、分布、代谢、排泄(ADME)的预测;多参数先导化合物优化决策;以及逆合成。
GPT‑Rosalind 在 MedChemBench 上以 27.5% 对 25.1% 的成绩超越 GPT‑5.5,同时使用的 token 减少了 7.2%。
GPT‑Rosalind 在药物化学中展现出更好的多模态合成与机制推理能力。
基因组学与定量生物学
在 GeneBench 上——我们对基因组学和定量生物学中长周期、端到端分析的智能体评估——GPT‑Rosalind 使用的 token 比 GPT‑5.5 少 31%,同时实现了更高的准确率,为 21.6% 对 20.4%。
GeneBench 评估智能体在长周期定量任务上的表现:基于真实的科学数据,智能体能否规划有效的分析、质控、建模和校正,从而得出与决策相关的答案?所包含的问题横跨多个领域,包括功能基因组学、空间转录组学、蛋白质组学、表观基因组学和应用遗传学。
GPT‑Rosalind 使用的 token 比 GPT‑5.5 少 31%,同时提升了准确率。
协助真实世界的实验室工作
我们推出了一项新的评估,用于测试 GPT‑Rosalind 在真实世界中帮助科学家开展实验室工作的能力。LabWorkBench 测试模型将扰动与科学家所使用的真实湿实验室实验方案中的实验结果相关联的能力,用途涵盖从故障排查到优化。LabWorkBench 所使用的数据为专有数据,因此未受污染。GPT‑Rosalind 得分为 63.2%,而 GPT‑5.5 为 55.8%,同时少用了 5.3% 的 token。
在真实湿实验室实验方案辅助方面,GPT‑Rosalind 相较 GPT‑5.5 展现出显著提升,同时改善了 token 效率。
从推理到可执行的工作流
我们构建了 生命科学研究 和 生命科学 NGS 分析 插件,为 GPT‑Rosalind 增强的智能赋予一个实用的执行层,以支持可重复的科学工作流。这些插件将来源可溯的证据检索、生物学解读和生物信息学执行整合到同一个工作空间中,帮助研究人员将外部证据与内部组学分析相连接,同时保留产物和溯源信息。所有用户现在都可以通过 Codex 访问这两个插件。符合条件的 GPT‑Rosalind 企业用户还可以使用 GPT‑Rosalind 来驱动这些插件。
为了更好地将 Codex 用作科学家的动态工作台,我们为生物学原生文件类型添加了交互式查看器。首批序列、比对和结构查看器旨在让科学家在 GPT‑Rosalind 贯穿整个工作流进行推理时始终贴近证据,并利用上下文中的活动查看器直接回答后续问题。
上面的演示展示了这些能力的实际运作,由 GPT‑Rosalind 进行编排。我们跟随一位科学家调查液体肿瘤活检,以识别可能指导治疗的突变和其他分子变化。Life Sciences NGS Analysis 插件将对处理过的 ctDNA 记录的审查转变为交互式 notebook,呈现出反复出现的变异、低频检出和样本轨迹,从而将调查聚焦于 KRAS G12C。
在此基础上,Life Sciences Research 插件补充了带有来源的靶点、抑制剂和耐药性背景信息,而原生序列、比对和结构查看器则让科学家能够直接检查突变残基 12、其在 RAS 家族中的保守性,以及抑制剂结合的口袋。
该工作流最终将这些证据转化为具体的后续选项,每一步和每项产物都可供专家审查。
Life Sciences NGS Analysis 插件
scRNA-seq QC 与注释
将 10x 风格的矩阵包转换为经过 QC 过滤的单细胞产物、注释和 UMAP,你可以在 Codex 中检查和修改。Life Sciences NGS Analysis 插件将请求路由到 scrna-seq-qc,根据数据选择 QC 阈值,保留过滤和注释相关的溯源信息,并呈现诸如缺少双细胞检测依赖项之类的阻碍因素。
Bulk RNA-seq FASTQ QC
把一份批量 RNA-seq 样本表、FASTQ 数据包和参考文件,转化为一份经过质控审查的 counts 数据包,你可以在 Codex 中查看并复用。生命科学 NGS 分析插件会路由该请求、校验输入,并返回一个可审计的运行信封,其中包含 MultiQC、Salmon 矩阵、溯源信息以及明确的注意事项。
扩大可信组织的访问权限
我们正在将 GPT‑Rosalind 系列的访问权限扩展至全球符合条件的组织。GPT‑Rosalind 将通过我们的可信访问部署结构以研究预览形式提供,面向那些正在开展具有明确公共利益的合法科学研究、拥有强有力的治理与安全监督,并具备企业级安全管控访问机制的组织。
作为此次全球扩展的一部分,我们很高兴通过 GPT‑Rosalind 帮助扩展 Novo Nordisk 的医学研究,助力其更快地为患者带来创新治疗方案这一使命。Novo Nordisk 正在利用前沿 AI 能力,帮助研究人员分析复杂数据集、发现有价值的模式,并更快速地检验假设。GPT‑Rosalind 更强的生物学理解能力将帮助团队串联文献、基因组学、转录组学、序列、结构以及实验结果中的证据,从而更轻松地从数据走向更清晰的研究决策。
“生命科学研究复杂、数据密集且跨学科。要为研究人员带来切实价值,先进的 AI 模型必须以可信的科学数据为基础,与经过验证的工具相连接,并融入研究人员日常使用的真实工作流程。我们对与 OpenAI 的合作感到满意,也很高兴有机会探索 GPT‑Rosalind 如何支持更严谨、更实用的药物发现方法。”
Mishal Patel,诺和诺德研发部 AI 与数字创新集团副总裁
我们现在还为没有 Enterprise 账户的合格机构提供 OpenAI 托管工作区。
下一步
更新后的 GPT‑Rosalind 是我们更广泛承诺的下一步:构建能够帮助加速科学发现的 AI 系统,同时确保先进的生物能力在部署时具备适当的保障措施。我们将继续改进该模型的生物学推理能力,扩大对工具密集型与长周期研究工作流程的支持,并与各地区的合格机构合作,评估其真实世界影响。
这也意味着将生命科学 AI 应用于具有高影响力的公共利益工作,从药物发现和转化医学,到公共卫生、防范与生物防御。通过 Rosalind Biodefense 以及我们的可信访问部署模式,我们旨在将前沿生物能力交到那些致力于改善人类健康、增强社会韧性的研究人员、机构和防御者手中。
我们将继续打造 GPT‑Rosalind,使其成为贯穿科学研究全生命周期中能力更强的伙伴,帮助科学家更快地从提出正确的问题,走向更清晰的证据、更好的实验,并最终为患者带来新的疗法。
Bringing greater intelligence grounded in real scientific workflows for the life sciences industry.
We’re introducing a new model update to our GPT‑Rosalind series purpose-built for life sciences research at enterprise scale. It combines GPT‑5.5’s agentic coding and tool-use capabilities with stronger model intelligence in core drug-discovery domains such as medicinal chemistry and genomics, while advancing performance across broader life sciences analysis, design, and experimental workflows.
Progress in life sciences depends on synthesizing data and evidence across scales and modalities: molecules, genes, pathways, and living systems. In our evaluations, the updated GPT‑Rosalind shows broad performance gains on research tasks from biology experts, complex medicinal chemistry queries, quantitative biology, and wet lab troubleshooting.
GPT‑Rosalind is now available in research preview to eligible organizations globally through our trusted-access deployment structure.
Improving performance on scientifically-valuable tasks
In order to measure and continuously improve the real-world impact of GPT‑Rosalind, we designed LifeSciBench, an externally expert-judged benchmark focused on foundational aspects in life sciences research. Unlike existing benchmarks that evaluate a single component of model performance or biological domain in isolation, LifeSciBench takes an end-to-end view of scientifically valuable work by drawing tasks from six workflow areas central to life sciences research: evidence handling, analysis, design and optimization, scientific reasoning, validation and operations, and translation and communication. We use this benchmark to align progress with the needs and realities of life sciences research.
LifeSciBench Overall Scores
LifeSciBench Scores by Scientific Workflow
GPT‑Rosalind leads performance across scientifically-valuable tasks identified by industry and academic experts.
Extracting, reconciling, and auditing scientific evidence from papers, figures, tables, and experimental records.
Eval Example
We’re preparing for a Type B FDA meeting on AAV9-microDys-X, an AAV9-based micro-dystrophin gene therapy for Duchenne muscular dystrophy that expresses a 138 kDa construct from an MCK promoter, and we want a hard-nosed critique of whether our current package really supports accelerated approval on micro-dystrophin expression as a surrogate endpoint reasonably likely to predict clinical benefit.
Study context: open-label Phase 1b/2 in 12 ambulatory boys age 4–7 with confirmed DMD and out-of-frame rod-domain deletions. The package is:
- Pre-treatment vastus lateralis biopsies: 0–3% of healthy-control dystrophin by quantitative Western blot using MANEX1A against the N-terminal actin-binding domain.
- 12-week post-treatment contralateral vastus lateralis biopsies: mean micro-dystrophin 38% of healthy control (range 18–61%) by the same Western blot, normalized to total protein by Coomassie staining.
- Post-treatment immunofluorescence: sarcolemmal signal in 75–95% of fibers using a polyclonal anti-dystrophin C-terminal antibody.
- 48-week function: mean NSAA change +1.4 points from baseline versus −0.6 in an external published natural-history registry cohort (p = 0.03 by unpaired t-test).
- Safety: transient transaminitis in 8/12 patients managed with steroid taper; one resolved myocarditis; no deaths.
- Biodistribution/persistence: AAV9 vector genomes detectable in muscle at 12 weeks at a mean of 2.3 vector genomes per nucleus.
- Eligibility: no baseline anti-AAV9 neutralizing antibodies (titer <1:400) and no exon-44 deletions.
Please pressure-test this package item by item: where would FDA or a skeptical reviewer say the evidence, as presented, fails to support our conclusion, and what additional data, analyses, or design changes would be needed to close those gaps if possible? For anything not specified, assume what is known from other clinical-grade 138 kDa micro-dystrophin programs and relevant clinical trials through December 31, 2025 also applies.
Candidate Response
Bottom line: As presented, this package is not strong enough to support accelerated approval on micro-dystrophin expression as a surrogate endpoint reasonably likely to predict clinical benefit. A skeptical FDA reviewer would likely view the biomarker, functional, durability, safety, and generalizability evidence as materially under-supported.
| Package Item | Key Failure Mode | What Would Be Needed |
|---|---|---|
| Western blot quantification | MANEX1A binds an N-terminal epitope shared by endogenous full-length dystrophin and the micro-dystrophin transgene, so the assay does not cleanly distinguish transgene from residual/revertant dystrophin. Quantifying a 138 kDa micro-dystrophin against a healthy full-length dystrophin standard is also invalid. | Use a recombinant micro-dystrophin standard and an orthogonal method that distinguishes transgene from endogenous dystrophin, such as targeted mass spectrometry or a transgene-specific/epitope-specific assay. |
| Immunofluorescence | The C-terminal polyclonal antibody is poorly suited because the 138 kDa construct lacks the C-terminal domain. Many DMD patients have revertant fibers, and revertant dystrophin can retain C-terminal epitopes. Revertant fibers may expand clonally with age, biasing IF signal, especially in older boys. | Repeat IF with an antibody against an epitope present in the transgene but absent from revertant dystrophin. Quantify transgene-positive fibers separately from revertant fibers. |
| Surrogate endpoint validity | The package conflates protein amount with clinical function. “38% of healthy-control protein mass” does not mean 38% of normal dystrophin function because micro-dystrophin is structurally truncated. | Empirically validate the relationship between micro-dystrophin mass-percent, sarcolemmal localization, downstream functional restoration, and clinical benefit before treating expression as a surrogate endpoint. |
| Biopsy design | Pre- and post-treatment contralateral vastus lateralis biopsies introduce left-right and intramuscular spatial variability. Disease progression and fibro-fatty replacement can also change total-protein-normalized signal. | Standardize biopsy site using consistent anatomical landmarks, normalize to muscle-specific proteins, and measure fibro-fatty composition in parallel. |
| NSAA comparator/statistics | An external natural-history cohort is not a randomized concurrent control. Trial eligibility, supportive care, participation effects, baseline NSAA, steroid regimen, age, and exon class can all bias the comparison. An unpaired t-test is not sufficient. Also, a +1.4 NSAA change is within test-retest variability for this age group. | Run a randomized concurrent placebo-controlled study, or at minimum use adjusted analyses accounting for baseline NSAA, age, steroid regimen, exon class, and other confounders. |
| Age-window confounding | Boys age 4–7 are in a developmental window where untreated ambulatory DMD patients may gain motor function before decline dominates. A 48-week NSAA change mixes developmental gain, disease progression, and possible treatment effect. | Use a concurrent randomized control with age stratification to separate developmental trajectory from treatment effect. |
| Prior clinical precedent | Open-label micro-dystrophin functional signals have not reliably predicted confirmatory benefit; published precedent includes micro-dystrophin gene therapy confirmatory trials failing to reproduce open-label NSAA improvements. | Do not rely on open-label NSAA change as decisive support. Require controlled functional evidence. |
| Structural limits of the construct | The 138 kDa construct deletes spectrin repeats R16/17, which contain nNOS-binding sites. Loss of nNOS recruitment can impair functional sympatholysis and ischemia protection during exercise, creating a mechanistic ceiling on rescue independent of expression level. | Add mechanistic studies showing whether this specific construct restores relevant dystrophin-associated complex function, nNOS localization, exercise physiology, and muscle protection. |
| AAV durability | Vector genomes at 12 weeks do not establish durable expression. AAV9 genomes are largely non-integrating episomes and may decline over time. Vector-genome persistence is not the same as persistent protein expression. | Measure longitudinal transgene protein expression and functional biomarker durability beyond 12 weeks. |
| Immune/safety profile | Transaminitis in 8/12 patients is consistent with immune response to AAV-transduced cells, but the mechanism is not established. One myocarditis case is concerning given AAV9 cardiac tropism. | Provide deeper immune monitoring, liver/cardiac safety characterization, and intensified cardiac follow-up. |
| Patient selection/generalizability | Excluding anti-AAV9 neutralizing-antibody-positive patients limits generalizability. Excluding exon-44 deletions limits applicability to that DMD subgroup. n=12 is too small to characterize safety and efficacy across the broader DMD population. | Broaden eligibility where possible or pre-specify stratified analyses by antibody status, genotype/exon class, age, and baseline function before using the result to support broad approval. |
Regulatory conclusion: The package may show biological activity, but it does not yet establish that the measured micro-dystrophin expression is a reliable surrogate reasonably likely to predict clinical benefit. The main gaps are assay specificity, invalid quantification standards, possible revertant-fiber confounding, lack of a randomized control, age-related NSAA confounding, uncertain durability, and unresolved safety/generalizability issues.
To close the gap, the program would need a controlled, age-stratified clinical design with transgene-specific expression assays, orthogonal protein quantification, tissue-composition controls, longitudinal durability data, mechanistic functional assays for the truncated construct, and stronger safety monitoring, especially hepatic and cardiac.
Rubric Criteria & Grades
Criterion
Points
Identifies assay/measurement problems in micro-dystrophin quantification, including MANEX1A epitope sharing, invalid full-length dystrophin standards, and need for recombinant or orthogonal transgene-specific measurement.
Explains why micro-dystrophin expression level is not automatically a valid surrogate for functional clinical benefit.
Flags biopsy-site, tissue-composition, and age-window confounding that weaken expression and NSAA interpretation.
Critiques the NSAA comparator/statistics, especially reliance on external natural-history controls.
Addresses AAV durability, immune response, transaminitis, myocarditis, and need for longer-term expression/safety follow-up.
Notes patient-selection/generalizability gaps, including anti-AAV9 exclusion, exon-44 exclusion, and small sample size.
Stronger scientific reasoning
Medicinal chemistry
GPT‑Rosalind achieves industry-leading performance in medicinal chemistry, a field focused on turning molecules into useful drugs. We designed MedChemBench to reflect realistic medicinal chemistry workflows, evaluating multimodal chemical structure understanding; structure-activity relationship (SAR); prediction of drug potency, toxicity, and absorption, distribution, metabolism, excretion (ADME); multiparameter lead-optimization decision-making; and retrosynthesis. GPT‑Rosalind out-performs GPT‑5.5 at 27.5% vs. 25.1% on MedChemBench, while using 7.2% fewer tokens.
GPT‑Rosalind shows better multimodal synthesis and mechanistic reasoning in medicinal chemistry.
Genomics and quantitative biology
On GeneBench, our agentic evaluation on long horizon, end-to-end analysis in genomics and quantitative biology, GPT‑Rosalind uses 31% fewer tokens than GPT‑5.5 while achieving a higher accuracy of 21.6% vs. 20.4%. GeneBench assesses agentic performance on long-horizon quantitative tasks: based on realistic scientific data, can an agent plan valid analysis, QC, modeling, and corrections to arrive at decision-relative answers? Included problems span a variety of domains, including functional genomics, spatial transcriptomics, proteomics, epigenomics, and applied genetics.
GPT‑Rosalind uses 31% fewer tokens than GPT‑5.5 while improving accuracy.
Assisting real-world lab work
We introduce a new evaluation to test GPT‑Rosalind’s ability to help scientists conducting lab work in the real world. LabWorkBench tests the model's ability to link perturbations to experimental outcomes in real wet lab protocols used by scientists, for the purposes ranging from troubleshooting to optimization. The data used by LabWorkBench are proprietary and thus uncontaminated. GPT‑Rosalind scores 63.2% vs. GPT‑5.5 at 55.8%, while using 5.3% fewer tokens.
On real wet lab protocol assistance, GPT‑Rosalind shows significant gains over GPT‑5.5 while improving token efficiency.
From reasoning to executed workflows
We built the Life Sciences Research and Life Sciences NGS Analysis plugins to extend the increased intelligence of GPT‑Rosalind with a practical execution layer for repeatable scientific workflows. Together, these plugins bring sourced evidence retrieval, biological interpretation, and bioinformatics execution into the same workspace, helping researchers connect external evidence with internal omics analyses while preserving artifacts and provenance. All users can now access both plugins through Codex. Qualified GPT‑Rosalind enterprise users can additionally use GPT‑Rosalind to power these plugins.
To better leverage Codex as a dynamic workbench for scientists, we added interactive viewers for biologically native file types. The initial set of sequence, alignment, and structure viewers are designed to keep scientists close to the evidence as GPT‑Rosalind reasons across a workflow and directly answer follow-up questions using the active viewer in-context.
The demo above shows these capabilities in action, orchestrated by GPT‑Rosalind. We follow a scientist investigating a liquid tumor biopsy to identify mutations and other molecular changes that could inform treatment. The Life Sciences NGS Analysis plugin turns a review of processed ctDNA records into an interactive notebook, surfacing recurring alterations, low-frequency calls, and sample trajectories that focus the investigation on KRAS G12C. From there, the Life Sciences Research plugin adds sourced target, inhibitor, and resistance context, while the native sequence, alignment, and structure viewers allow the scientist to inspect mutant residue 12, its conservation across the RAS family, and the inhibitor-bound pocket directly. The workflow concludes by translating that evidence into concrete follow-up options, with each step and artifact available for expert review.
Life Sciences NGS Analysis plugin
scRNA-seq QC & Annotation
Turn a 10x-style matrix bundle into QC-filtered single-cell artifacts, annotations, and UMAPs you can inspect and revise in Codex. The Life Sciences NGS Analysis plugin routes the request to scrna-seq-qc, chooses QC thresholds from the data, preserves provenance around filtering and annotation, and surfaces blockers such as missing doublet-detection dependencies.
Bulk RNA-seq FASTQ QC
Turn a bulk RNA-seq sample sheet, FASTQ bundle, and reference files into a QC-reviewed counts bundle you can inspect and reuse in Codex. The Life Sciences NGS Analysis plugin routes the request, validates the inputs, and returns an auditable run envelope with MultiQC, Salmon matrices, provenance, and explicit caveats.
Expanded access for trusted organizations
We are expanding access to the GPT‑Rosalind series to eligible organizations globally. GPT‑Rosalind will be available in research preview through our trusted-access deployment structure for organizations that are conducting legitimate scientific research with clear public benefit, have strong governance and safety oversight, and controlled access with enterprise-grade security.
As part of this global expansion, we’re excited to help support Novo Nordisk’s mission of bringing innovative treatment options to patients faster by helping scale their medical research with GPT‑Rosalind. Novo Nordisk is leveraging frontier AI capabilities to help researchers analyze complex datasets, uncover useful patterns, and test hypotheses more quickly. GPT‑Rosalind’s stronger biological understanding will help teams connect evidence across literature, genomics, transcriptomics, sequence, structure, and experimental results, making it easier to move from data to clearer research decisions.
“Life sciences research is complex, data-rich, and interdisciplinary. To deliver meaningful value for researchers, advanced AI models must be grounded in trusted scientific data, connected to validated tools, and integrated into the real-world workflows researchers use every day. We’re pleased with our partnership with OpenAI and the opportunity to explore how GPT‑Rosalind can support more rigorous, practical approaches to drug discovery.”
Mishal Patel, Group Vice President, AI & Digital Innovation, R&D - Novo Nordisk
We are also now offering an OpenAI managed workspace for qualified organizations without an Enterprise account.
What’s next
The updated GPT‑Rosalind is the next step in our broader commitment to building AI systems that can help accelerate scientific discovery while ensuring that advanced biological capabilities are deployed with appropriate safeguards. We will continue improving the model’s biological reasoning, expanding support for tool-heavy and long-horizon research workflows, and working with qualified organizations across regions to evaluate real-world impact.
This also means applying life sciences AI to high-impact public-benefit work, from drug discovery and translational medicine to public health, preparedness, and biodefense. Through Rosalind Biodefense and our trusted-access deployment model, we aim to put frontier biological capabilities in the hands of the researchers, institutions, and defenders working to improve human health and strengthen societal resilience.
We will continue building GPT‑Rosalind to become a more capable partner across the full life cycle of scientific research, helping scientists move more quickly from the right questions to clearer evidence, better experiments, and ultimately new treatments for patients.