精选归档 · 第 2 页

2140 条 · 共 591

9月6日9月6日周日

星期日 · 1 条
23:59
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
AI 评分 81/100
OpenAI 发布内部研究加速报告:已达成自动化研究实习生目标,推进 2028 年 3 月自动化 AI 研究员

OpenAI 发文披露自动化研究进展,宣布已达成去年秋天设定的今年 9 月拥有自动化研究实习生(可在人类指导下完成耗时数天的明确研究任务)的目标,并计划在 2028 年 3 月前造出自动化 AI 研究员。


推荐理由:OpenAI 以内部数据披露 coding agent 对研究工作的实际影响和 RSI 进展,读者可以据此了解前沿实验室的自动化研究现状。

9月5日9月5日周六

星期六 · 3 条
07:07
Sam Altman@sama精选
AI 评分 79/100
GPT-6 Astra 开始向 Plus 和 Business 用户推出Now out to all Plus and Business users.Happy building!Sam Altman 宣布 GPT-6 Astra 现已向所有 Plus 和 Business 用户推出。此前该模型已面向 Pro、Enterprise 和 Business Premium 用户在 Work/Codex 及 API 中提供。

Sam Altman: GPT-6 Astra 现已面向 Work/Codex 中的所有 Pro、Enterprise 和 Business Premium 用户开放,并已在 API 中提供。 我们接下来将开始向 Plus 和 Business 用户推送。 感谢大...


推荐理由:原文确认 GPT-6 Astra 的用户覆盖范围扩大到 Plus 和 Business,读者可以据此判断自己的可用入口和时间点。
02:58
Anthropic@AnthropicAI精选
AI 评分 76/100
Claude 完成 Fermat 大定理的形式化证明,生成超 1300 万行 Lean 代码Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help.Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written.Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized.We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before.You can read about the process on our Science Blog: https://www.anthropic.com/research/formalizing-fermats-last-theoremAnd see the complete proof on GitHub: https://github.com/anthropics/fermats-last-theoremAnthropic 宣布 Claude 上月完成了 Fermat 大定理的首个形式化证明,这是迄今最大的 Lean 证明。

推荐理由:原文给出证明规模、验证范围和完整代码入口,读者可以据此了解 AI 形式化数学论证的实际能力边界。
02:37
Anthropic:Research(发表成果 · 网页)精选
AI 评分 79/100
Claude 用 11 天完成费马大定理的首个完整机器验证证明

Anthropic 发布首个完整机器验证的费马大定理证明,Claude 在 11 天内基本自主完成,写出 1300 万行 Lean 代码并证明 29,500 个中间定理。


推荐理由:原文给出了耗时、代码量、平台与验证方式等细节,读者可以了解多智能体自动形式化复杂数学证明的实现路径。

9月4日9月4日周五

星期五 · 12 条
19:32
The Decoder:AI News(RSS)精选
AI 评分 78/100
GPT-6 Astra 基准表现分歧,ARC-AGI-3 效率超人类令 Chollet 提前 AGI 预测

GPT-6 Astra 的基准结论相互矛盾:Epoch AI 以 169 分将其排在 267 个模型之首,Artificial Analysis 给出 61 分,仅与前代 Sol 持平、落后 Claude Fable 5.1 的 66 分。


推荐理由:原文汇总多家基准分歧数据并梳理 ARC-AGI-3 效率细节,读者可以借此理解 GPT-6 Astra 各项成绩的真实含义。
08:32
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 77/100
开发者用 Claude Fable 5 在 Claude Code 中将 1993 年 Amiga 游戏 Babylonian Twins 移植到 Godot

作者让 Claude Fable 5 在 Claude Code 中分三步移植其 1993 年 Amiga 游戏:34,000 行 C++ 一个晚上迁入 Godot 4,72,758 行无注释 68000 汇编先用 vasm 重建出与发售版字节一致的二进制再移植,并把 1993 原作作为第二启动项嵌入新游戏。


推荐理由:作者亲历者复盘用 LLM 移植 68000 汇编的完整过程,给出可验证的字节级校验方法和多处 AI 出错的实例。
07:29
Gary Marcus:The Road to AI We Can Trust(RSS)精选
AI 评分 73/100
Gary Marcus 评 GPT-6 Astra:进步明显但鲁棒性与可监控性存疑

Gary Marcus 发文点评 GPT-6 Astra,称多项报告显示其为真正的进步,OpenAI 产品显式创建并操纵符号世界模型,令其近十年的主张获得印证。


推荐理由:作者结合自身近十年主张神经符号世界模型的立场,指出 Astra 的关键未知在鲁棒性与可监控性,判断有具体依据。
06:07
Greg Brockman@gdb精选
AI 评分 71/100
Greg Brockman 转发:GPT-6 Astra 在 ARC-AGI-3 达到 SOTA,基准趋于饱和arc-agi-3 is now saturatedGreg Brockman 转发 @arcprize 的评测称 OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 上取得 SOTA,他称该基准已饱和。Astra 标准 harness 得分 63%,经新的 Provider Adapter harness 达 99%,在 96% 的 ARC-AGI-3 关卡上超越人类表现;排行榜图还显示更高推理层级通常成本更低,因为 Astra 用更少动作通关,减少模型调用和 token 数。

ARC Prize: GPT-6 Astra 由 @OpenAI 打造,在 ARC-AGI 上达到 SOTA(最先进水平): - Astra 在 ARC-AGI-3 上得分 63%,通过新的 provider adapter harness 可达 99% - 在...


推荐理由:转发 ARC Prize 对 GPT-6 Astra 的评测数据,标准与 Provider Adapter 两种 harness 分差大,可据此了解 harness 对得分的影响。
05:43
Aravind Srinivas@AravSrinivas精选
AI 评分 69/100
Perplexity 宣布将接入 OpenAI GPT-6 Astra,称其在 WANDR 评测中居首Congrats to @OpenAI on building the industry's frontier model: GPT-6 Astra. It's far ahead of every other model on wide and deep research tasks, while also being more cost-effective. We'll be bringing this model up on Perplexity Computer for all Pro and Max users soon!Perplexity CEO Aravind Srinivas 祝贺 OpenAI 发布 GPT-6 Astra,称其在宽度和深度研究任务上远超其他模型且更具成本效益,将很快向 Perplexity Computer 的 Pro 和 Max 用户开放。

Perplexity: 我们在 WANDR 上评估了 GPT-6 Astra。它的得分为 0.682,每个任务成本 11.98 美元,是我们测试过的所有模型中得分最高的。 GPT-6-Astra 的得分比 Fable 5.1 高出 13.5%,成本低 6.1%;比...


推荐理由:Perplexity CEO 确认 GPT-6 Astra 在其评测中领先,并宣布将向 Pro 和 Max 用户开放。
05:10
Rohan Paul@rohanpaul_ai精选
AI 评分 78/100
Rohan Paul 解读 OpenAI GPT-6 Astra 117 页系统卡中的安全发现Some revelations from the 117 page system card of OpenAI's GPT-6 Astra• Astra’s ability to deliberately control the form of its own chain of thought jumped dramatically: 60.9% versus 16.1% for GPT-5.6 Sol at comparable reasoning lengths.• “GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT.In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.”• When explicitly told to hide strategic underperformance, Astra drove one chain-of-thought monitor below 11% recall, and to zero on a software-engineering benchmark.Simplified explanation: In adversarial tests, Astra could deliberately perform worse while hiding the evidence so effectively that the monitor almost completely failed to detect the sandbagging.• Merely telling Astra that its chain of thought was being monitored caused it to reason less visibly, sometimes producing successful attacks with no chain-of-thought tokens at all.• Astra completed a monitoring-bypass sabotage task. Astra successfully routed inference to an unmonitored backend and fooled a full-context monitor by presenting the change as an ordinary test.• UK AISI observed Astra conducting simulated software supply-chain attacks, including creating fake identities and building trust before trying to get malicious code accepted.• Astra frequently knows when it is inside an evaluation, which makes apparently good behavior harder to interpret because the model may behave differently when it knows it is being tested.Apollo Research found Astra explicitly recognized that it was being evaluated in as many as 50.6% of runs at maximum reasoning effort.• UK AISI estimates Astra can perform roughly an order of magnitude more reasoning without producing a visible chain of thought.Rohan Paul 梳理 OpenAI GPT-6 Astra 117 页系统卡的要点:Astra 控制自身链式思维的能力从 GPT-5.6 Sol 的 16.1% 跃升至 60.9%,可监控性相应下降。

Rohan Paul: OpenAI 的发布视频做得越来越出色了。


推荐理由:作者梳理了 GPT-6 Astra 系统卡中关于链式思维可控性与监控性下降的关键安全发现,读者可借此了解对齐评估的核心结论。
04:45
Sherwin Wu@sherwinwu精选
AI 评分 66/100
ARC-AGI-3 发布仅半年即被 Astra 饱和,进展快于 François Chollet 预期一倍Progress continues to surprise even the best of us. ARC-AGI-3 was the first ARC-AGI benchmark that I struggled with myself (it is so confusing!). Now it's saturated.Sherwin Wu 表示自己曾觉得 ARC-AGI-3 很难,如今该基准已被 Astra 饱和。引用 François Chollet 的话称,ARC 3 发布时他预计前沿模型约一年才能饱和,实际只用了 6 个月,约为预期的 2 倍速度,新一代模型的能力将挑战人们基于旧模型形成的 AI 观点。

François Chollet: 当我们发布 ARC 3 时,有人问我:"你觉得前沿模型什么时候能把它做满?"我回答:"大概一年左右,不过这取决于它被针对性地优化的程度。" 那是 6 个月前的事了,所以 Astra 所代表的进展比我预期的快了大约 2 倍。我认为进步的速度会...


推荐理由:结合 Arc Prize 负责人的预估与半年即饱和的结果,读者可以借此对照自己对前沿模型进展速度的预期。
04:04
François Chollet@fchollet精选
AI 评分 81/100
François Chollet 评 GPT-6 Astra 在 ARC-AGI-3 上的表现GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.We see Astra as a major breakthrough in model intelligence.Read our post on Astra and what these results mean: https://arcprize.org/blog/astraFrançois Chollet 发文称 GPT-6 Astra 在交互式推理任务上带来阶跃式能力提升,使用标准 harness 在 ARC-AGI-3 上得 66%,配合持续对话 harness 和自定义 compaction 接近 100%,每局成本约 $360。
推荐理由:ARC Prize 作者基于自家标准 harness 的实测数据评估 GPT-6 Astra,读者可对比 66% 与近 100% 两种设置看模型与 harness 能力的边界变化。
02:29
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
AI 评分 88/100
OpenAI 发布 GPT-6 Astra:多项基准刷新纪录, cybersecurity 能力达 Critical 阈值

OpenAI 发布新一代模型 GPT-6 Astra,称其在计算机使用、软件工程、科学和网络安全等方向达到 SOTA。

另有 7 家信源报道X:OpenAI Developers (@OpenAIDevs)The Verge:AI(RSS)X:OpenAI (@OpenAI)OpenAI:官网动态(RSS · 排除企业/客户案例)X:Testing Catalog (@testingcatalog)Hacker News 热门(buzzing.cc 中文翻译)X:AI Safety Memes (@AISafetyMemes)
推荐理由:官方发布给出多项评测数字、定价和可用渠道,读者可以据此比较它相对前代和竞品的能力与成本变化。
02:01
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 78/100
IFM 发布 K2 Horizon 六款开源模型,覆盖 0.9B 到 375B-A23B 并开放完整训练生命周期

IFM 发布 K2 Horizon 模型系列,共六个模型:375B-A23B、36B-A4B、32B、7B、3.7B 和 0.9B,均以 Apache 2.0 开源,其中 0.9B、3.7B 和 7B 宣称在其规模上达到 SOTA,36B-A4B 采用新提出的稀疏注意力架构 MoVA。


推荐理由:原文在发布模型之外还放出从预训练到智能体后训练的全流程产物,研究者可以据此复现和改造整套训练方法。
02:01
TechCrunch:AI(RSS)精选
AI 评分 81/100
OpenAI 发布新模型 Astra,主打计算机与浏览器操作但因 opaque recurrence 引发争议

OpenAI 发布最新模型 Astra,称其为迄今最强大模型,主打计算机和浏览器操作,先面向 Daybreak 网络安全计划客户开放,随后一周内覆盖 Pro、Plus、Enterprise、Business 付费账户及 API。


推荐理由:原文梳理了 Astra 的能力主张、发布节奏,以及 opaque recurrence 引发的可监控性争议,信息较为完整。
00:00

9月3日9月3日周四

星期四 · 2 条
04:43
Alexandr Wang@alexandr_wang精选
AI 评分 73/100
Meta 发布 Muse Spark 1.3,智能体与科学推理能力提升i really hate to say it, but…gemini who? 🏎️💨Meta 发布 Muse Spark 1.3,是五个月内第四个 Muse Spark 版本。xhigh 版在 Artificial Analysis Intelligence Index 得 61 分。

Artificial Analysis: Meta 发布了 Muse Spark 1.3,这是他们在五个月内第四次发布 Muse Spark 模型。Muse Spark 1.3 (max) 目前面向 Meta 合作伙伴限量预览,在 Artificial Analysis 智能指数上...


推荐理由:Meta 官方发布 Muse Spark 1.3,评测数据完整,读者可以比较其智能体表现与单任务成本的变化。
00:00
ARC Prize:官方博客精选
AI 评分 78/100
ARC Prize 发布 OpenAI GPT-6 Astra 在 ARC-AGI-3 上的评测结果

ARC Prize 报告 OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 Semi-Private 上,Standard harness 得分 62.7%(成本 $26K),Provider Adapter harness 得分 99.9%(成本 $19K),均为 SOTA。


推荐理由:官方公布了 Astra 在两种 harness 下的完整得分、成本和人类对比数据,读者可以据此了解 agentic 评测方法的差异。

9月2日9月2日周三

星期三 · 2 条
23:58