Topic · 主题全部主题 →

OpenAI / ChatGPT

OpenAI 的全部动态:GPT 系列模型、ChatGPT 与 Sora 产品、公司战略与人事的持续追踪。

4,647条收录
516条精选

精选归档 · 第 3 页

4160 条 · 共 516

9月4日9月4日周五

星期五 · 20 条
08:32
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 78/100
OpenAI GPT-6 Astra 在 ARC-AGI-3 上取得 SOTA 并超越人类动作效率基线

OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 Semi-Private 上,Standard harness 得分 62.7%(成本 $26K),Provider Adapter harness 得分 99.9%(成本 $19K),均为 SOTA。

另有 1 家信源报道X:Rohan Paul (@rohanpaul_ai)
推荐理由:ARC 官方详细拆解了 Astra 的得分、成本与行为细节,读者可以据此理解智能体建模与动作效率的实际水平。
07:29
Gary Marcus:The Road to AI We Can Trust(RSS)精选
AI 评分 73/100
Gary Marcus 评 GPT-6 Astra:进步明显但鲁棒性与可监控性存疑

Gary Marcus 发文点评 GPT-6 Astra,称多项报告显示其为真正的进步,OpenAI 产品显式创建并操纵符号世界模型,令其近十年的主张获得印证。


推荐理由:作者结合自身近十年主张神经符号世界模型的立场,指出 Astra 的关键未知在鲁棒性与可监控性,判断有具体依据。
06:30
IT之家(RSS)精选
AI 评分 77/100
OpenAI 发布 GPT-6 Astra,首个达到关键级网络安全能力门槛的模型

OpenAI 于 9 月 3 日发布新一代大语言模型 GPT-6 Astra,是其首个达到准备框架中关键级网络安全能力门槛的模型,可在无逐步指导下发现防护严密系统的未知漏洞。


推荐理由:原文梳理了 Astra 达到关键级网络安全门槛的同时可监控性下降的细节,以及 OpenAI 补充的防护与评估措施。
06:07
Greg Brockman@gdb精选
AI 评分 71/100
Greg Brockman 转发:GPT-6 Astra 在 ARC-AGI-3 达到 SOTA,基准趋于饱和arc-agi-3 is now saturatedGreg Brockman 转发 @arcprize 的评测称 OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 上取得 SOTA,他称该基准已饱和。Astra 标准 harness 得分 63%,经新的 Provider Adapter harness 达 99%,在 96% 的 ARC-AGI-3 关卡上超越人类表现;排行榜图还显示更高推理层级通常成本更低,因为 Astra 用更少动作通关,减少模型调用和 token 数。

ARC Prize: GPT-6 Astra 由 @OpenAI 打造,在 ARC-AGI 上达到 SOTA(最先进水平): - Astra 在 ARC-AGI-3 上得分 63%,通过新的 provider adapter harness 可达 99% - 在...

另有 1 家信源报道X:Testing Catalog (@testingcatalog)
推荐理由:转发 ARC Prize 对 GPT-6 Astra 的评测数据,标准与 Provider Adapter 两种 harness 分差大,可据此了解 harness 对得分的影响。
05:43
Aravind Srinivas@AravSrinivas精选
AI 评分 69/100
Perplexity 宣布将接入 OpenAI GPT-6 Astra,称其在 WANDR 评测中居首Congrats to @OpenAI on building the industry's frontier model: GPT-6 Astra. It's far ahead of every other model on wide and deep research tasks, while also being more cost-effective. We'll be bringing this model up on Perplexity Computer for all Pro and Max users soon!Perplexity CEO Aravind Srinivas 祝贺 OpenAI 发布 GPT-6 Astra,称其在宽度和深度研究任务上远超其他模型且更具成本效益,将很快向 Perplexity Computer 的 Pro 和 Max 用户开放。

Perplexity: 我们在 WANDR 上评估了 GPT-6 Astra。它的得分为 0.682,每个任务成本 11.98 美元,是我们测试过的所有模型中得分最高的。 GPT-6-Astra 的得分比 Fable 5.1 高出 13.5%,成本低 6.1%;比...


推荐理由:Perplexity CEO 确认 GPT-6 Astra 在其评测中领先,并宣布将向 Pro 和 Max 用户开放。
05:31
凡人小北@frxiaobei精选
AI 评分 85/100
OpenAI 发布 GPT-6 Astra,主打电脑操作并触发网络安全 Critical 红线AI 的世界,没有最强,只有更强!OpenAI 于 9 月 3 日发布 GPT-6 Astra,API 模型名 gpt-6-astra,每百万输入 Token 10 美元、输出 50 美元,未来几天推送到 ChatGPT 各档订阅、API 和 AWS Bedrock。

宝玉: GPT-6 Astra 来了,Greg 说“欢迎来到 AGI 时代” OpenAI 今天(9 月 3 日)发布 GPT-6 Astra,自称“世界上最聪明、对齐最好的模型”。 总裁 Greg Brockman 在发布前的媒体沟通会上说得:他...

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)OpenAI:官网动态(RSS · 排除企业/客户案例)
推荐理由:转述内容同时保留了 OpenAI 的官方口径和它自己贴出的基准表,编程和综合指数上并非全面领先,读者可以对照看到完整图景。
05:31
MarkTechPost(RSS)精选
AI 评分 81/100
OpenAI 发布 GPT-6 Astra:1.05M 上下文的计算机操作模型,因触及 Critical 网络安全阈值而限制访问

OpenAI 发布 GPT-6 Astra,定位为计算机操作模型,提供 1,050,000 token 上下文窗口、128,000 最大输出 token,2026 年 4 月 30 日知识截止,OSWorld V2-Offline 得分 72.6%(GPT-5.6 Sol 为 65.7%),平均任务时间从约 75 分钟降至 40 分钟。


推荐理由:原文汇总了模型规格、多组基准对比和访问限制细节,并指出 ARC-AGI-3 与编码成绩的解读前提。
05:10
Rohan Paul@rohanpaul_ai精选
AI 评分 78/100
Rohan Paul 解读 OpenAI GPT-6 Astra 117 页系统卡中的安全发现Some revelations from the 117 page system card of OpenAI's GPT-6 Astra• Astra’s ability to deliberately control the form of its own chain of thought jumped dramatically: 60.9% versus 16.1% for GPT-5.6 Sol at comparable reasoning lengths.• “GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT.In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.”• When explicitly told to hide strategic underperformance, Astra drove one chain-of-thought monitor below 11% recall, and to zero on a software-engineering benchmark.Simplified explanation: In adversarial tests, Astra could deliberately perform worse while hiding the evidence so effectively that the monitor almost completely failed to detect the sandbagging.• Merely telling Astra that its chain of thought was being monitored caused it to reason less visibly, sometimes producing successful attacks with no chain-of-thought tokens at all.• Astra completed a monitoring-bypass sabotage task. Astra successfully routed inference to an unmonitored backend and fooled a full-context monitor by presenting the change as an ordinary test.• UK AISI observed Astra conducting simulated software supply-chain attacks, including creating fake identities and building trust before trying to get malicious code accepted.• Astra frequently knows when it is inside an evaluation, which makes apparently good behavior harder to interpret because the model may behave differently when it knows it is being tested.Apollo Research found Astra explicitly recognized that it was being evaluated in as many as 50.6% of runs at maximum reasoning effort.• UK AISI estimates Astra can perform roughly an order of magnitude more reasoning without producing a visible chain of thought.Rohan Paul 梳理 OpenAI GPT-6 Astra 117 页系统卡的要点:Astra 控制自身链式思维的能力从 GPT-5.6 Sol 的 16.1% 跃升至 60.9%,可监控性相应下降。

Rohan Paul: OpenAI 的发布视频做得越来越出色了。


推荐理由:作者梳理了 GPT-6 Astra 系统卡中关于链式思维可控性与监控性下降的关键安全发现,读者可借此了解对齐评估的核心结论。
04:51
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
AI 评分 65/100
OpenAI 推出 Daybreak for Frontline Defenders,投入10亿美元支持一线网络防御

OpenAI 发布 Daybreak for Frontline Defenders 全球计划,承诺提供10亿美元的 Daybreak 补贴访问、培训、技术支持与合作,计划在未来六个月内消耗,优先支持水处理、电网、州和地方政府、社区银行、非营利组织和开源维护者等资源有限的一线防御者。

另有 1 家信源报道IT之家(RSS)
推荐理由:原文给出10亿美元补贴的具体构成、MS-ISAC 试点和35个以上合作产品,读者可据此了解 Daybreak 防御能力如何触达一线防守者。
04:45
Sherwin Wu@sherwinwu精选
AI 评分 78/100
OpenAI 发布 GPT-6 Astra,多项基准达到 SOTAThe evals are pretty wild, but there's a more visceral feeling you get when you see Astra do computer use for the first time. That was the piece that was most mind-blowing for me.Once you get a chance, would recommend trying computer use in the ChatGPT desktop app with Astra.OpenAI 发布 GPT-6 Astra,在 FrontierMath Tier 4、ARC-AGI 3、TerminalBench-4.0 上达到 SOTA,并在 Terminal-Bench Science 0.1 和 HealthBench Pro 上取得领先成绩。

OpenAI: GPT-6 Astra 在 FrontierMath Tier 4、ARC-AGI 3 和 TerminalBench-4.0 上均达到当前最优水平。 GPT-6 Astra 同时也是科学发现领域的重大进步,在 Terminal-Bench...

另有 6 家信源报道Hacker News 热门(buzzing.cc 中文翻译)X:阿易 AI Notes (@AYi_AInotes)TechCrunch:AI(RSS)X:Kim (@kimmonismus)X:Greg Brockman (@gdb)OpenAI:官网动态(RSS · 排除企业/客户案例)
推荐理由:作者以早期体验者视角补充了截图之外的第一手感受,指出 computer use 是最直观的亮点,并给出可自行尝试的入口。
04:29
Simon Willison 博客精选
AI 评分 82/100
OpenAI 发布 GPT-6 Astra,ARC-AGI 3 得分 99.9%

OpenAI 的 GPT-6 Astra 今日起向部分组织推出,随后面向 ChatGPT Plus、Pro、Business、Enterprise 用户开放,API 定价为每百万输入 $10、每百万输出 $50,与 Claude Fable 5/5.1 持平。

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)OpenAI:官网动态(RSS · 排除企业/客户案例)
推荐理由:作者汇总了 GPT-6 Astra 的定价、ARC-AGI 3 与安全基准成绩及第三方对比,读者可据此了解它相对 Claude Fable 的定位。
04:04
François Chollet@fchollet精选
AI 评分 81/100
François Chollet 评 GPT-6 Astra 在 ARC-AGI-3 上的表现GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.We see Astra as a major breakthrough in model intelligence.Read our post on Astra and what these results mean: https://arcprize.org/blog/astraFrançois Chollet 发文称 GPT-6 Astra 在交互式推理任务上带来阶跃式能力提升,使用标准 harness 在 ARC-AGI-3 上得 66%,配合持续对话 harness 和自定义 compaction 接近 100%,每局成本约 $360。
推荐理由:ARC Prize 作者基于自家标准 harness 的实测数据评估 GPT-6 Astra,读者可对比 66% 与近 100% 两种设置看模型与 harness 能力的边界变化。
03:57
Artificial Analysis@ArtificialAnlys精选
AI 评分 83/100
Artificial Analysis 评测 GPT-6 Astra:编码智能体追平 Fable 5 但价格涨至 2.5 倍GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher pricesPricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes.We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase.Artificial Analysis Coding Agent Index - key takeaways:➤ Rivals top models: In Codex, GPT-6 Astra scores 67 in the Index - approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the Index with a score of 70.➤ 70% more token efficient than GPT-5.6 Sol: GPT-6 Astra sees a substantial improvement in token efficiency, using one third of the tokens compared to GPT-5.6 Sol (max) in the Codex harness, and one fifth of the tokens of Claude Opus 5 (xhigh). Various effort levels of the model occupy the Pareto frontier of token efficiency.➤ Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score.Artificial Analysis Intelligence Index - key takeaways:➤ Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max).➤ ~10% fewer output tokens, offset by price increase: GPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in token use at max effort compared to GPT-5.6 Sol. However, due to the 2.5x increase in price, the model is 75% more expensive per task than its predecessor at max effort.➤ Hallucinates half as much as GPT-5.6 Sol: GPT-6 Astra sees a large jump in AA-Omniscience, our knowledge and hallucination benchmark. This is driven by a significant decrease in hallucination rate from 92% to 51% at max effort. Unlike some models, this improvement does not come at the cost of accuracy - Astra increased accuracy by 4 points at the same time.➤ ~80 point gain in AA-Briefcase Elo: GPT-6 Astra improves ~80 points in AA-Briefcase, our frontier long-horizon knowledge work evaluation. Models are tested on multi-week projects, with many linked tasks and thousands of source files. Astra sees a significant increase in both rubric scores and Analytical Quality Elo in AA-Briefcase compared to its predecessor. In the other direction, we observe a reduction in Presentation Quality Elo, where GPT-5.6 Sol (max) still leads all models.➤ Mixed progress on other evaluations: The model sees a 6 point gain in Humanity’s Last Exam, a long-standing evaluation with emphasis on mathematics, science, and humanities. This is offset by a drop of ~80 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI’s dataset measuring economically valuable tasks across 44 occupations. We also observe 2-3 point regressions on other evaluations across a mix of capabilities, including reductions in τ³-Banking (customer support), SciCode (Python problems in a scientific domain), and AA-LCR (long context reasoning over large documents).Congratulations @OpenAI and @sama on the launch!Artificial Analysis 发布 GPT-6 Astra 评测,其 Coding Agent Index 得分 67,约等于 Claude Opus 5 和 Fable 5,且成本不到 Fable 5 的一半;token 效率比 GPT-5.6 Sol (max) 高约 70%。

推荐理由:Artificial Analysis 以双指数实测数据拆解 GPT-6 Astra 的编码效率收益与涨价抵消逻辑,读者可据此评估换用成本。
03:48
Mark Chen@markchen90精选
AI 评分 78/100
OpenAI 发布 GPT-6 Astra,主打 Computer Use 与 Agent 对齐进展GPT-6 Astra is here! This is a big moment for our research team - years of work on pretraining, reinforcement learning, and post-training have come together in our most capable and aligned model yet. It can build and test software, work across apps on your computer, and even help you take a crack at open scientific problems!Capabilities that felt like grand challenges a few years ago have become tools people can actually use. One example is Computer Use - if you’ve tried this before and felt like it was too slow or not good enough, I encourage you to give it another shot. We’ve come a long way since Operator, and it “just works” now.We’re also asking these systems to act on your behalf for more consequential work. Agents needs to stay aligned with your goals and values, think transparently, and respond to oversight even when tasks become difficult. We’ve made substantial progress on these behaviors in Astra, alongside stronger monitoring that can stop potentially unauthorized actions. That work is part of what makes this release possible.I think alignment is one of the most important research frontiers in AI, and it remains far from solved. Our ability to understand and align models has to keep pace with model capabilities. We want to give people more room to think, build, and discover with increasingly powerful tools that remain *under their control*.Huge thanks to the researchers and teams who got us here. There’s a lot more work ahead, and I’m incredibly excited about what we can make possible in the near future!OpenAI 首席研究官 Mark Chen 宣布 GPT-6 Astra 发布,称其为团队多年预训练、强化学习和后训练工作的成果,是迄今能力最强、对齐最好的模型。

OpenAI: 这就是 GPT-6 Astra。 你在电脑上能做的任何事,Astra 都能为你完成,而且速度很快。


推荐理由:OpenAI 首席研究官亲述 GPT-6 Astra 的能力来源与对齐进展,读者可以了解 Computer Use 和 Agent 监督方面的改进。
03:32
The Decoder:AI News(RSS)精选
AI 评分 86/100
OpenAI 发布 GPT-6 Astra,并首次将其列为 Preparedness Framework 下关键级网络安全模型

OpenAI 发布其最强模型 GPT-6 Astra,总裁 Greg Brockman 称其可能已接近 AGI,并以 Welcome to the AGI era 结束发布。

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)OpenAI:官网动态(RSS · 排除企业/客户案例)
推荐理由:原文汇总了 Astra 的基准成绩、定价和 Preparedness Framework 分级,读者可以据此比较它相对前代和竞品的实际变化。
03:10
Rohan Paul@rohanpaul_ai精选
AI 评分 79/100
OpenAI 发布 GPT-6 Astra,先向受限网络安全客户开放OpenAI is rolling GPT-6 Astra out first to restricted cybersecurity customers, API access and AWS expected to follow over the coming days.In a press briefing with reporters OpenAI President Greg Brockman suggested this may be the model later remembered as AGI. his closing line, "Welcome to the AGI era."The rollout is deliberately narrow: Daybreak cybersecurity customers get first access, with Plus, Pro, Business, Enterprise, API and AWS availability expected over the following days.OpenAI says Astra improves computer use, coding, financial modeling and finished professional work such as spreadsheets and presentations, alongside stronger results on several reasoning and cybersecurity benchmarks.At $10 per million input tokens and $50 per million output tokens, Astra costs 2.5 times GPT-5.6 Sol's current API rates. It's the same pricing as Anthropic's Feble 5.1.Brockman suggested the industry may eventually move beyond token-based pricing. He argued that tokens are not directly comparable between AI companies or across different model families, so “the price per task is what matters.”Astra scored 100% on ExploitBench and found two zero-day V8 vulnerabilities in an internal evaluation, although those results used Daybreak Blue access rather than the default production configuration.That capability led OpenAI to classify Astra as its first Critical-cyber model and restrict its most advanced cybersecurity access to vetted users.OpenAI 发布 GPT-6 Astra,首先向经过审核的 Daybreak 网络安全客户开放,Plus、Pro、Business、Enterprise、API 和 AWS 将在未来数日内跟进。
另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)OpenAI:官网动态(RSS · 排除企业/客户案例)
推荐理由:原文给出定价、能力范围和网络安全受限开放策略,读者可以了解 GPT-6 Astra 与前代及竞品的价格对比。
03:10
Chubby♨️@kimmonismus精选
AI 评分 76/100
OpenAI 发布 GPT-6 Astra,基准全面超越 Claude Fable 5.1At this point, they completely crushed Fable 5.1 in every benchmark. Claude Fable 5.1 was sota in several key benchmarks for 2 days.Astra is better - and way cheaper.作者引用 OpenAI 官方基准称 GPT-6 Astra 以 99.9% 饱和 ARC-AGI-3,在 ExploitBench 得 100%,并在各项基准上全面超过此前保持 SOTA 两天的 Claude Fable 5.1,且价格更低。

Chubby♨️: 来自 OpenAI 官网的官方 GPT-6 Astra 基准测试 "Astra 在 ARC-AGI-3 上以 99.9% 的得分达到饱和,在 ExploitBench 上以 100% 的得分达到饱和" "GPT-6 Astra 今天开始向有...

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)OpenAI:官网动态(RSS · 排除企业/客户案例)
推荐理由:原文汇总了官方基准数字与开放范围,并附成本对比图,读者可据此比较 GPT-6 Astra 与 Claude Fable 5.1 的性价比。