Topic · 主题全部主题 →

AI 编码

AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。

5,049条收录
580条精选

近期焦点

近 14 天 · 按多源报道热度

最新精选

120 条 · 共 580

9月6日

星期日 · 2 条
23:59
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
AI 评分 81/100
OpenAI 发布内部研究加速报告:已达成自动化研究实习生目标,推进 2028 年 3 月自动化 AI 研究员

OpenAI 发文披露自动化研究进展,宣布已达成去年秋天设定的今年 9 月拥有自动化研究实习生(可在人类指导下完成耗时数天的明确研究任务)的目标,并计划在 2028 年 3 月前造出自动化 AI 研究员。

另有 6 家信源报道X:Rohan Paul (@rohanpaul_ai)X:Kim (@kimmonismus)Hacker News 热门(buzzing.cc 中文翻译)X:Noam Brown (@polynoamial)IT之家(RSS)The Decoder:AI News(RSS)
推荐理由:OpenAI 以内部数据披露 coding agent 对研究工作的实际影响和 RSI 进展,读者可以据此了解前沿实验室的自动化研究现状。
05:59
🚨 AI News | TestingCatalog@testingcatalog精选
AI 评分 76/100
OpenAI GPT-6 Astra 登顶 Code Arena WebDev 榜首,领先 Claude Fable 5.1 达 35 分GPT-6 Astra from OpenAI claimed a top spot on WebDev Arena, surpassing recently released Claude Fable 5.1 by 35 points.“It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken, which matches the latest Claude model pricing.”Are we set to hit 2k by the end of this year?GPT-6 Astra (Max) 以 1797 分登顶 Code Arena: WebDev,领先第 2 名 Claude Fable 5.1 (Max) 35 分、第 3 名 Claude Opus 5 (Max) 1688 分。

Arena.ai: 真实世界的结果已经出炉。Code Arena 上出现了新的第一名--GPT-6 Astra (Max)! 它还重塑了帕累托前沿,成为 $40/Mtoken 价位上性能最强的模型,这与最新的 Claude 模型定价持平。 OpenAI 的 G...

另有 1 家信源报道X:Rohan Paul (@rohanpaul_ai)
推荐理由:榜单数据给出了与 Claude 系列的具体分差和同价位对比,读者可以据此评估新模型的实际编码位置。

9月5日

星期六 · 5 条
21:33
The Decoder:AI News(RSS)精选
AI 评分 78/100
OpenAI 发布 GPT-6 Astra 提示词指南,含 slop 词屏蔽清单

OpenAI 在模型文档中说明 GPT-6 Astra 相比 GPT-5.6 Sol 更常提出澄清问题、对上下文更敏感,并给出让模型更主动、审计 AGENTS.md 等技能文件、控制写作风格、约束子智能体委派和测试规模的提示词建议。


推荐理由:原文汇总了 OpenAI 官方文档中针对 GPT-6 Astra 的提示词建议和 slop 词屏蔽清单,开发者可直接迁移到自己的提示词写法。
19:40
公众号:数字生命卡兹克精选
AI 评分 77/100
实测GPT-6 Astra:速度、前端与代码能力对比GPT-5.6 Sol的全面升级

GPT-6 Astra正式向所有订阅用户推送,作者实测后认为其综合能力追平Claude Fable 5,且额度100%可用。相比GPT-5.6 Sol,速度明显提升,大型系统审查从数小时缩短到约10分钟,代码扫描找出大量此前未发现的性能问题并2小时完成修复;前端3D生成和审美大幅强化,写作在白描和用词上更好但仍缺中文留白感。


推荐理由:作者实测了GPT-6 Astra在速度、前端生成、代码深度和写作上的具体变化,并给出可迁移的AGENT.md简化思路。
07:07
Sam Altman@sama精选
AI 评分 79/100
GPT-6 Astra 开始向 Plus 和 Business 用户推出Now out to all Plus and Business users.Happy building!Sam Altman 宣布 GPT-6 Astra 现已向所有 Plus 和 Business 用户推出。此前该模型已面向 Pro、Enterprise 和 Business Premium 用户在 Work/Codex 及 API 中提供。

Sam Altman: GPT-6 Astra 现已面向 Work/Codex 中的所有 Pro、Enterprise 和 Business Premium 用户开放,并已在 API 中提供。 我们接下来将开始向 Plus 和 Business 用户推送。 感谢大...


推荐理由:原文确认 GPT-6 Astra 的用户覆盖范围扩大到 Plus 和 Business,读者可以据此判断自己的可用入口和时间点。
03:59
OpenAI@OpenAI精选
AI 评分 81/100
OpenAI 发布 GPT-6 Astra,面向 Pro、Enterprise 和 Business Premium 用户开放GPT-6 Astra is now available to all Pro, Enterprise, and Business Premium users in ChatGPT Work and Codex. It's also live in the API.It might take a few days to roll out to our Plus and Business users. Thank you for your patience.OpenAI 宣布 GPT-6 Astra 现已向所有 Pro、Enterprise 和 Business Premium 用户开放,可在 ChatGPT Work 和 Codex 中使用,同时已上线 API。Plus 和 Business 用户的推送可能需要几天时间。另有 7 家信源报道X:OpenRouter (@OpenRouter)IT之家(RSS)X:Greg Brockman (@gdb)Hacker News 热门(buzzing.cc 中文翻译)X:OpenAI Developers (@OpenAIDevs)X:Testing Catalog (@testingcatalog)X:Sam Altman (@sama)
推荐理由:官方宣布 GPT-6 Astra 上线范围与渠道,Plus 和 Business 用户还需等待几天,读者可据此确认自己能否用上。
00:29
GitHub Blog精选
AI 评分 68/100
GitHub 发布 Project HydraFusion 研究预览,用多模型运行时编排降低 Copilot 成本

GitHub 推出 Project HydraFusion 研究预览,通过运行时多模型编排,在 Single、Cascade、Critique 三种执行模式间为每个任务选择工作流,以平衡质量、成本和延迟。

另有 1 家信源报道MarkTechPost(RSS)
推荐理由:原文给出三种执行模式和三个基准的质量与成本数据,读者可以据此评估多模型编排对编码工作流的影响。

9月4日

星期五 · 11 条
08:32
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 77/100
开发者用 Claude Fable 5 在 Claude Code 中将 1993 年 Amiga 游戏 Babylonian Twins 移植到 Godot

作者让 Claude Fable 5 在 Claude Code 中分三步移植其 1993 年 Amiga 游戏:34,000 行 C++ 一个晚上迁入 Godot 4,72,758 行无注释 68000 汇编先用 vasm 重建出与发售版字节一致的二进制再移植,并把 1993 原作作为第二启动项嵌入新游戏。


推荐理由:作者亲历者复盘用 LLM 移植 68000 汇编的完整过程,给出可验证的字节级校验方法和多处 AI 出错的实例。
05:31
凡人小北@frxiaobei精选
AI 评分 85/100
OpenAI 发布 GPT-6 Astra,主打电脑操作并触发网络安全 Critical 红线AI 的世界,没有最强,只有更强!OpenAI 于 9 月 3 日发布 GPT-6 Astra,API 模型名 gpt-6-astra,每百万输入 Token 10 美元、输出 50 美元,未来几天推送到 ChatGPT 各档订阅、API 和 AWS Bedrock。

宝玉: GPT-6 Astra 来了,Greg 说“欢迎来到 AGI 时代” OpenAI 今天(9 月 3 日)发布 GPT-6 Astra,自称“世界上最聪明、对齐最好的模型”。 总裁 Greg Brockman 在发布前的媒体沟通会上说得:他...

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)X:Sherwin Wu(@sherwinwu)
推荐理由:转述内容同时保留了 OpenAI 的官方口径和它自己贴出的基准表,编程和综合指数上并非全面领先,读者可以对照看到完整图景。
04:29
Simon Willison 博客精选
AI 评分 82/100
OpenAI 发布 GPT-6 Astra,ARC-AGI 3 得分 99.9%

OpenAI 的 GPT-6 Astra 今日起向部分组织推出,随后面向 ChatGPT Plus、Pro、Business、Enterprise 用户开放,API 定价为每百万输入 $10、每百万输出 $50,与 Claude Fable 5/5.1 持平。

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Sherwin Wu(@sherwinwu)X:Greg Brockman (@gdb)
推荐理由:作者汇总了 GPT-6 Astra 的定价、ARC-AGI 3 与安全基准成绩及第三方对比,读者可据此了解它相对 Claude Fable 的定位。
03:57
Artificial Analysis@ArtificialAnlys精选
AI 评分 83/100
Artificial Analysis 评测 GPT-6 Astra:编码智能体追平 Fable 5 但价格涨至 2.5 倍GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher pricesPricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes.We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase.Artificial Analysis Coding Agent Index - key takeaways:➤ Rivals top models: In Codex, GPT-6 Astra scores 67 in the Index - approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the Index with a score of 70.➤ 70% more token efficient than GPT-5.6 Sol: GPT-6 Astra sees a substantial improvement in token efficiency, using one third of the tokens compared to GPT-5.6 Sol (max) in the Codex harness, and one fifth of the tokens of Claude Opus 5 (xhigh). Various effort levels of the model occupy the Pareto frontier of token efficiency.➤ Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score.Artificial Analysis Intelligence Index - key takeaways:➤ Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max).➤ ~10% fewer output tokens, offset by price increase: GPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in token use at max effort compared to GPT-5.6 Sol. However, due to the 2.5x increase in price, the model is 75% more expensive per task than its predecessor at max effort.➤ Hallucinates half as much as GPT-5.6 Sol: GPT-6 Astra sees a large jump in AA-Omniscience, our knowledge and hallucination benchmark. This is driven by a significant decrease in hallucination rate from 92% to 51% at max effort. Unlike some models, this improvement does not come at the cost of accuracy - Astra increased accuracy by 4 points at the same time.➤ ~80 point gain in AA-Briefcase Elo: GPT-6 Astra improves ~80 points in AA-Briefcase, our frontier long-horizon knowledge work evaluation. Models are tested on multi-week projects, with many linked tasks and thousands of source files. Astra sees a significant increase in both rubric scores and Analytical Quality Elo in AA-Briefcase compared to its predecessor. In the other direction, we observe a reduction in Presentation Quality Elo, where GPT-5.6 Sol (max) still leads all models.➤ Mixed progress on other evaluations: The model sees a 6 point gain in Humanity’s Last Exam, a long-standing evaluation with emphasis on mathematics, science, and humanities. This is offset by a drop of ~80 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI’s dataset measuring economically valuable tasks across 44 occupations. We also observe 2-3 point regressions on other evaluations across a mix of capabilities, including reductions in τ³-Banking (customer support), SciCode (Python problems in a scientific domain), and AA-LCR (long context reasoning over large documents).Congratulations @OpenAI and @sama on the launch!Artificial Analysis 发布 GPT-6 Astra 评测,其 Coding Agent Index 得分 67,约等于 Claude Opus 5 和 Fable 5,且成本不到 Fable 5 的一半;token 效率比 GPT-5.6 Sol (max) 高约 70%。

推荐理由:Artificial Analysis 以双指数实测数据拆解 GPT-6 Astra 的编码效率收益与涨价抵消逻辑,读者可据此评估换用成本。
03:48
Mark Chen@markchen90精选
AI 评分 78/100
OpenAI 发布 GPT-6 Astra,主打 Computer Use 与 Agent 对齐进展GPT-6 Astra is here! This is a big moment for our research team - years of work on pretraining, reinforcement learning, and post-training have come together in our most capable and aligned model yet. It can build and test software, work across apps on your computer, and even help you take a crack at open scientific problems!Capabilities that felt like grand challenges a few years ago have become tools people can actually use. One example is Computer Use - if you’ve tried this before and felt like it was too slow or not good enough, I encourage you to give it another shot. We’ve come a long way since Operator, and it “just works” now.We’re also asking these systems to act on your behalf for more consequential work. Agents needs to stay aligned with your goals and values, think transparently, and respond to oversight even when tasks become difficult. We’ve made substantial progress on these behaviors in Astra, alongside stronger monitoring that can stop potentially unauthorized actions. That work is part of what makes this release possible.I think alignment is one of the most important research frontiers in AI, and it remains far from solved. Our ability to understand and align models has to keep pace with model capabilities. We want to give people more room to think, build, and discover with increasingly powerful tools that remain *under their control*.Huge thanks to the researchers and teams who got us here. There’s a lot more work ahead, and I’m incredibly excited about what we can make possible in the near future!OpenAI 首席研究官 Mark Chen 宣布 GPT-6 Astra 发布,称其为团队多年预训练、强化学习和后训练工作的成果,是迄今能力最强、对齐最好的模型。

OpenAI: 这就是 GPT-6 Astra。 你在电脑上能做的任何事,Astra 都能为你完成,而且速度很快。


推荐理由:OpenAI 首席研究官亲述 GPT-6 Astra 的能力来源与对齐进展,读者可以了解 Computer Use 和 Agent 监督方面的改进。
03:32
The Decoder:AI News(RSS)精选
AI 评分 86/100
OpenAI 发布 GPT-6 Astra,并首次将其列为 Preparedness Framework 下关键级网络安全模型

OpenAI 发布其最强模型 GPT-6 Astra,总裁 Greg Brockman 称其可能已接近 AGI,并以 Welcome to the AGI era 结束发布。

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)X:Sherwin Wu(@sherwinwu)
推荐理由:原文汇总了 Astra 的基准成绩、定价和 Preparedness Framework 分级,读者可以据此比较它相对前代和竞品的实际变化。
03:10
Rohan Paul@rohanpaul_ai精选
AI 评分 79/100
OpenAI 发布 GPT-6 Astra,先向受限网络安全客户开放OpenAI is rolling GPT-6 Astra out first to restricted cybersecurity customers, API access and AWS expected to follow over the coming days.In a press briefing with reporters OpenAI President Greg Brockman suggested this may be the model later remembered as AGI. his closing line, "Welcome to the AGI era."The rollout is deliberately narrow: Daybreak cybersecurity customers get first access, with Plus, Pro, Business, Enterprise, API and AWS availability expected over the following days.OpenAI says Astra improves computer use, coding, financial modeling and finished professional work such as spreadsheets and presentations, alongside stronger results on several reasoning and cybersecurity benchmarks.At $10 per million input tokens and $50 per million output tokens, Astra costs 2.5 times GPT-5.6 Sol's current API rates. It's the same pricing as Anthropic's Feble 5.1.Brockman suggested the industry may eventually move beyond token-based pricing. He argued that tokens are not directly comparable between AI companies or across different model families, so “the price per task is what matters.”Astra scored 100% on ExploitBench and found two zero-day V8 vulnerabilities in an internal evaluation, although those results used Daybreak Blue access rather than the default production configuration.That capability led OpenAI to classify Astra as its first Critical-cyber model and restrict its most advanced cybersecurity access to vetted users.OpenAI 发布 GPT-6 Astra,首先向经过审核的 Daybreak 网络安全客户开放,Plus、Pro、Business、Enterprise、API 和 AWS 将在未来数日内跟进。
另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)X:Sherwin Wu(@sherwinwu)
推荐理由:原文给出定价、能力范围和网络安全受限开放策略,读者可以了解 GPT-6 Astra 与前代及竞品的价格对比。
03:10
Chubby♨️@kimmonismus精选
AI 评分 76/100
OpenAI 发布 GPT-6 Astra,基准全面超越 Claude Fable 5.1At this point, they completely crushed Fable 5.1 in every benchmark. Claude Fable 5.1 was sota in several key benchmarks for 2 days.Astra is better - and way cheaper.作者引用 OpenAI 官方基准称 GPT-6 Astra 以 99.9% 饱和 ARC-AGI-3,在 ExploitBench 得 100%,并在各项基准上全面超过此前保持 SOTA 两天的 Claude Fable 5.1,且价格更低。

Chubby♨️: 来自 OpenAI 官网的官方 GPT-6 Astra 基准测试 "Astra 在 ARC-AGI-3 上以 99.9% 的得分达到饱和,在 ExploitBench 上以 100% 的得分达到饱和" "GPT-6 Astra 今天开始向有...

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Sherwin Wu(@sherwinwu)X:Greg Brockman (@gdb)
推荐理由:原文汇总了官方基准数字与开放范围,并附成本对比图,读者可据此比较 GPT-6 Astra 与 Claude Fable 5.1 的性价比。
03:01
The Verge:AI(RSS)精选
AI 评分 83/100
OpenAI 发布 GPT-6 Astra,称已进入 AGI 时代

OpenAI 发布下一代的旗舰模型 GPT-6 Astra,称其为能力上的世代跃升,Greg Brockman 表示现在可能已进入 AGI 时代。


推荐理由:报道把 GPT-6 Astra 的能力宣称、AGI 判断与 Hugging Face 事件后的安全安排放在一起,读者可对照看待发布时机的商业与信任背景。
02:29
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
AI 评分 88/100
OpenAI 发布 GPT-6 Astra:多项基准刷新纪录, cybersecurity 能力达 Critical 阈值

OpenAI 发布新一代模型 GPT-6 Astra,称其在计算机使用、软件工程、科学和网络安全等方向达到 SOTA。

另有 6 家信源报道TechCrunch:AI(RSS)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:阿易 AI Notes (@AYi_AInotes)X:Greg Brockman (@gdb)X:Sherwin Wu(@sherwinwu)
推荐理由:官方发布给出多项评测数字、定价和可用渠道,读者可以据此比较它相对前代和竞品的能力与成本变化。

9月3日

星期四 · 2 条
19:29
Hugging Face:Blog(RSS)精选
AI 评分 72/100
Hugging Face 发布开源工具 funes,为编码智能体提供可本地持有的记忆层

Hugging Face 发布开源工具 funes,为 Claude Code、Codex、pi、Hermes 等编码智能体提供本地记忆层,把已有会话记录索引成 Lance 数据集,一条 funes add 命令即可让 Agent 自主召回原始出处(Agent、时间戳、会话、轮次)。


推荐理由:原文给出本地记忆层的检索机制、跨 Agent 复用方式,以及召回比压缩和交接更省成本的基准数字,方法可直接复用。
02:29
GitHub Blog精选
AI 评分 61/100
GitHub Copilot 如何在不牺牲任务质量的前提下降低 AI 编码成本

GitHub 工程师 Erik Kristensen 分享了 Copilot 降本的四项改动:选择性压缩工具输出、移除 view 工具行号前缀(线下推理成本降约 5%,线上用户日均推理成本降约 3%)、压缩 task-tool 提示词(每轮省约 1300 token,每活跃小时归一化成本降 2.9%)、后台任务完成后直接交付结果(AI Credits 用量降约 2.3%)。


推荐理由:GitHub 用四项改动说明压缩 token 为何要看整个任务而非单次调用,并给出可复用的评估方法与踩坑教训。