23:59
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
OpenAI 发布内部研究加速报告:已达成自动化研究实习生目标,推进 2028 年 3 月自动化 AI 研究员OpenAI 发文披露自动化研究进展,宣布已达成去年秋天设定的今年 9 月拥有自动化研究实习生(可在人类指导下完成耗时数天的明确研究任务)的目标,并计划在 2028 年 3 月前造出自动化 AI 研究员。
另有 6 家信源报道X:Rohan Paul (@rohanpaul_ai)X:Kim (@kimmonismus)Hacker News 热门(buzzing.cc 中文翻译)X:Noam Brown (@polynoamial)IT之家(RSS)The Decoder:AI News(RSS)
推荐理由:OpenAI 以内部数据披露 coding agent 对研究工作的实际影响和 RSI 进展,读者可以据此了解前沿实验室的自动化研究现状。
21:33
The Decoder:AI News(RSS)精选
OpenAI 发布 GPT-6 Astra 提示词指南,含 slop 词屏蔽清单OpenAI 在模型文档中说明 GPT-6 Astra 相比 GPT-5.6 Sol 更常提出澄清问题、对上下文更敏感,并给出让模型更主动、审计 AGENTS.md 等技能文件、控制写作风格、约束子智能体委派和测试规模的提示词建议。
推荐理由:原文汇总了 OpenAI 官方文档中针对 GPT-6 Astra 的提示词建议和 slop 词屏蔽清单,开发者可直接迁移到自己的提示词写法。
19:40
实测GPT-6 Astra:速度、前端与代码能力对比GPT-5.6 Sol的全面升级GPT-6 Astra正式向所有订阅用户推送,作者实测后认为其综合能力追平Claude Fable 5,且额度100%可用。相比GPT-5.6 Sol,速度明显提升,大型系统审查从数小时缩短到约10分钟,代码扫描找出大量此前未发现的性能问题并2小时完成修复;前端3D生成和审美大幅强化,写作在白描和用词上更好但仍缺中文留白感。
推荐理由:作者实测了GPT-6 Astra在速度、前端生成、代码深度和写作上的具体变化,并给出可迁移的AGENT.md简化思路。
04:29
OpenAI 发布 GPT-6 Astra,ARC-AGI 3 得分 99.9%OpenAI 的 GPT-6 Astra 今日起向部分组织推出,随后面向 ChatGPT Plus、Pro、Business、Enterprise 用户开放,API 定价为每百万输入 $10、每百万输出 $50,与 Claude Fable 5/5.1 持平。
另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Sherwin Wu(@sherwinwu)X:Greg Brockman (@gdb)
推荐理由:作者汇总了 GPT-6 Astra 的定价、ARC-AGI 3 与安全基准成绩及第三方对比,读者可据此了解它相对 Claude Fable 的定位。
04:07
Sam Altman@sama精选 OpenAI 发布 GPT-6 AstraGPT-6 Astra is here.We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building.We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more.It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait.It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.译Sam Altman 宣布 GPT-6 Astra 发布,称其为计算机使用、专业工作、科学、编码、网络安全等领域全球最佳模型。官方表示为确保该能力级别所需的安全与对齐标准而多花了些时间,并公布 FrontierMath Tier 4 得分 98%、ARC-AGI 3 得分 99.9%、ExploitBench 得分 100%。
推荐理由:官方公告给出三项基准成绩和安全方面的说明,读者可以据此判断模型在科学、编码等场景的定位。
03:57
Artificial Analysis@ArtificialAnlys精选 Artificial Analysis 评测 GPT-6 Astra:编码智能体追平 Fable 5 但价格涨至 2.5 倍GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher pricesPricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes.We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase.Artificial Analysis Coding Agent Index - key takeaways:➤ Rivals top models: In Codex, GPT-6 Astra scores 67 in the Index - approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the Index with a score of 70.➤ 70% more token efficient than GPT-5.6 Sol: GPT-6 Astra sees a substantial improvement in token efficiency, using one third of the tokens compared to GPT-5.6 Sol (max) in the Codex harness, and one fifth of the tokens of Claude Opus 5 (xhigh). Various effort levels of the model occupy the Pareto frontier of token efficiency.➤ Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score.Artificial Analysis Intelligence Index - key takeaways:➤ Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max).➤ ~10% fewer output tokens, offset by price increase: GPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in token use at max effort compared to GPT-5.6 Sol. However, due to the 2.5x increase in price, the model is 75% more expensive per task than its predecessor at max effort.➤ Hallucinates half as much as GPT-5.6 Sol: GPT-6 Astra sees a large jump in AA-Omniscience, our knowledge and hallucination benchmark. This is driven by a significant decrease in hallucination rate from 92% to 51% at max effort. Unlike some models, this improvement does not come at the cost of accuracy - Astra increased accuracy by 4 points at the same time.➤ ~80 point gain in AA-Briefcase Elo: GPT-6 Astra improves ~80 points in AA-Briefcase, our frontier long-horizon knowledge work evaluation. Models are tested on multi-week projects, with many linked tasks and thousands of source files. Astra sees a significant increase in both rubric scores and Analytical Quality Elo in AA-Briefcase compared to its predecessor. In the other direction, we observe a reduction in Presentation Quality Elo, where GPT-5.6 Sol (max) still leads all models.➤ Mixed progress on other evaluations: The model sees a 6 point gain in Humanity’s Last Exam, a long-standing evaluation with emphasis on mathematics, science, and humanities. This is offset by a drop of ~80 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI’s dataset measuring economically valuable tasks across 44 occupations. We also observe 2-3 point regressions on other evaluations across a mix of capabilities, including reductions in τ³-Banking (customer support), SciCode (Python problems in a scientific domain), and AA-LCR (long context reasoning over large documents).Congratulations @OpenAI and @sama on the launch!译Artificial Analysis 发布 GPT-6 Astra 评测,其 Coding Agent Index 得分 67,约等于 Claude Opus 5 和 Fable 5,且成本不到 Fable 5 的一半;token 效率比 GPT-5.6 Sol (max) 高约 70%。
推荐理由:Artificial Analysis 以双指数实测数据拆解 GPT-6 Astra 的编码效率收益与涨价抵消逻辑,读者可据此评估换用成本。
03:48
Mark Chen@markchen90精选 OpenAI 发布 GPT-6 Astra,主打 Computer Use 与 Agent 对齐进展GPT-6 Astra is here! This is a big moment for our research team - years of work on pretraining, reinforcement learning, and post-training have come together in our most capable and aligned model yet. It can build and test software, work across apps on your computer, and even help you take a crack at open scientific problems!Capabilities that felt like grand challenges a few years ago have become tools people can actually use. One example is Computer Use - if you’ve tried this before and felt like it was too slow or not good enough, I encourage you to give it another shot. We’ve come a long way since Operator, and it “just works” now.We’re also asking these systems to act on your behalf for more consequential work. Agents needs to stay aligned with your goals and values, think transparently, and respond to oversight even when tasks become difficult. We’ve made substantial progress on these behaviors in Astra, alongside stronger monitoring that can stop potentially unauthorized actions. That work is part of what makes this release possible.I think alignment is one of the most important research frontiers in AI, and it remains far from solved. Our ability to understand and align models has to keep pace with model capabilities. We want to give people more room to think, build, and discover with increasingly powerful tools that remain *under their control*.Huge thanks to the researchers and teams who got us here. There’s a lot more work ahead, and I’m incredibly excited about what we can make possible in the near future!译OpenAI 首席研究官 Mark Chen 宣布 GPT-6 Astra 发布,称其为团队多年预训练、强化学习和后训练工作的成果,是迄今能力最强、对齐最好的模型。OpenAI: 这就是 GPT-6 Astra。 你在电脑上能做的任何事,Astra 都能为你完成,而且速度很快。
推荐理由:OpenAI 首席研究官亲述 GPT-6 Astra 的能力来源与对齐进展,读者可以了解 Computer Use 和 Agent 监督方面的改进。
03:32
The Decoder:AI News(RSS)精选
OpenAI 发布 GPT-6 Astra,并首次将其列为 Preparedness Framework 下关键级网络安全模型OpenAI 发布其最强模型 GPT-6 Astra,总裁 Greg Brockman 称其可能已接近 AGI,并以 Welcome to the AGI era 结束发布。
另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)X:Sherwin Wu(@sherwinwu)
推荐理由:原文汇总了 Astra 的基准成绩、定价和 Preparedness Framework 分级,读者可以据此比较它相对前代和竞品的实际变化。
03:10
Rohan Paul@rohanpaul_ai精选 OpenAI 发布 GPT-6 Astra,先向受限网络安全客户开放OpenAI is rolling GPT-6 Astra out first to restricted cybersecurity customers, API access and AWS expected to follow over the coming days.In a press briefing with reporters OpenAI President Greg Brockman suggested this may be the model later remembered as AGI. his closing line, "Welcome to the AGI era."The rollout is deliberately narrow: Daybreak cybersecurity customers get first access, with Plus, Pro, Business, Enterprise, API and AWS availability expected over the following days.OpenAI says Astra improves computer use, coding, financial modeling and finished professional work such as spreadsheets and presentations, alongside stronger results on several reasoning and cybersecurity benchmarks.At $10 per million input tokens and $50 per million output tokens, Astra costs 2.5 times GPT-5.6 Sol's current API rates. It's the same pricing as Anthropic's Feble 5.1.Brockman suggested the industry may eventually move beyond token-based pricing. He argued that tokens are not directly comparable between AI companies or across different model families, so “the price per task is what matters.”Astra scored 100% on ExploitBench and found two zero-day V8 vulnerabilities in an internal evaluation, although those results used Daybreak Blue access rather than the default production configuration.That capability led OpenAI to classify Astra as its first Critical-cyber model and restrict its most advanced cybersecurity access to vetted users.译OpenAI 发布 GPT-6 Astra,首先向经过审核的 Daybreak 网络安全客户开放,Plus、Pro、Business、Enterprise、API 和 AWS 将在未来数日内跟进。另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)X:Sherwin Wu(@sherwinwu)
推荐理由:原文给出定价、能力范围和网络安全受限开放策略,读者可以了解 GPT-6 Astra 与前代及竞品的价格对比。
03:01
OpenAI 发布 GPT-6 Astra,称已进入 AGI 时代OpenAI 发布下一代的旗舰模型 GPT-6 Astra,称其为能力上的世代跃升,Greg Brockman 表示现在可能已进入 AGI 时代。
推荐理由:报道把 GPT-6 Astra 的能力宣称、AGI 判断与 Hugging Face 事件后的安全安排放在一起,读者可对照看待发布时机的商业与信任背景。
02:29
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
OpenAI 发布 GPT-6 Astra:多项基准刷新纪录, cybersecurity 能力达 Critical 阈值OpenAI 发布新一代模型 GPT-6 Astra,称其在计算机使用、软件工程、科学和网络安全等方向达到 SOTA。
另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)X:Sherwin Wu(@sherwinwu)
推荐理由:官方发布给出多项评测数字、定价和可用渠道,读者可以据此比较它相对前代和竞品的能力与成本变化。
19:29
Hugging Face 发布开源工具 funes,为编码智能体提供可本地持有的记忆层Hugging Face 发布开源工具 funes,为 Claude Code、Codex、pi、Hermes 等编码智能体提供本地记忆层,把已有会话记录索引成 Lance 数据集,一条 funes add 命令即可让 Agent 自主召回原始出处(Agent、时间戳、会话、轮次)。
推荐理由:原文给出本地记忆层的检索机制、跨 Agent 复用方式,以及召回比压缩和交接更省成本的基准数字,方法可直接复用。
02:29
GitHub Copilot 如何在不牺牲任务质量的前提下降低 AI 编码成本GitHub 工程师 Erik Kristensen 分享了 Copilot 降本的四项改动:选择性压缩工具输出、移除 view 工具行号前缀(线下推理成本降约 5%,线上用户日均推理成本降约 3%)、压缩 task-tool 提示词(每轮省约 1300 token,每活跃小时归一化成本降 2.9%)、后台任务完成后直接交付结果(AI Credits 用量降约 2.3%)。
推荐理由:GitHub 用四项改动说明压缩 token 为何要看整个任务而非单次调用,并给出可复用的评估方法与踩坑教训。