Topic · 主题全部主题 →

Anthropic / Claude

Anthropic 的全部动态:Claude 系列模型、Claude Code、安全研究路线与公司进展的持续追踪。

3,633条收录
525条精选

精选归档 · 第 21 页

401420 条 · 共 525

4月23日4月23日周四

星期四 · 3 条
08:00
Claude Platform:开发者版本说明(RSS)精选
AI 评分 63/100
Claude Managed Agents 记忆功能进入公开 Beta 测试

Claude Managed Agents 的记忆功能现已进入公开 Beta 测试,可通过标准 managed-agents-2026-04-01 请求头使用。该功能允许智能体在对话间保持上下文,具体集成指南见官方文档。


推荐理由:Claude托管Agent的记忆功能进入公测,这步让Agent从一问一答走向有上下文记忆的长期助手,做复杂任务的开发者可以试试。
01:27
Anthropic:Research(发表成果 · 网页)精选
AI 评分 74/100
从8.1万用户调查看AI经济学:职业暴露度与替代担忧

一项针对8.1万名Claude用户的调查显示,从事AI高暴露职业(如软件工程师)的从业者对岗位替代的担忧显著更高,其中早期职业者焦虑感尤为突出。约五分之一受访者明确表达失业忧虑,且职业AI暴露度每增加10%,感知到的岗位威胁就上升1.3%。在生产力方面,受访者平均自评生产力增益较高,最高薪与最低薪群体报告了最大幅度的效率提升,主要体现在工作范围拓展。值得注意的是,从AI中获得速度提升最大的人群,对岗位替代的担忧反而更强烈。


推荐理由:Anthropic 对 81000 名用户的调查,第一次把 AI 暴露度和就业焦虑直接挂上钩了,高收入者和早期职业者两头挤压,这个数据比任何猜测都硬。
00:00
Anthropic:Engineering(事故复盘 + 工程实践 · 网页)精选
AI 评分 72/100
关于近期 Claude Code 质量报告的更新说明

Anthropic 确认并解决了过去一个月影响 Claude Code、Claude Agent SDK 和 Claude Cowork 的三个问题,所有问题已于 4 月 20 日修复。具体包括:3月4日将 Claude Code 的默认推理强度从“高”改为“中”,导致用户感知智能下降,已于4月7日回滚;3月26日一项缓存优化存在缺陷,导致会话恢复后模型“健忘”和重复,4月10日修复;4月16日一项旨在减少冗余的系统提示指令意外损害了代码质量,4月20日撤销。这些问题影响了 Sonnet 4.6 和 Opus 4.6/4.7 模型,但 API 未受影响。公司已重置所有订阅用户的使用限额,并承诺改进流程以防止类似问题。


推荐理由:Anthropic 把 Claude Code 连续一个月质量下滑的三个 bug 全部摊开讲,这种级别的工程复盘在大模型公司里极少见。做 Agent 产品的人该认真读,因为这三个坑你迟早也会踩。

4月17日4月17日周五

星期五 · 1 条
23:29
Anthropic:Newsroom(网页)精选
AI 评分 75/100
Anthropic Labs 推出 Claude Design

Anthropic Labs 推出 Claude Design,由 Claude Opus 4.7 驱动的视觉设计工具,面向 Claude Pro、Max、Team 及 Enterprise 订阅者开放研究预览。用户可通过对话快速创建设计原型、幻灯片及营销物料,支持自动应用企业设计系统、多格式导入导出,并能与 Claude Code 无缝衔接。据 Brilliant 反馈,复杂页面设计从 20 余条提示缩减至 2 条,设计到生产周期从数周缩短至单次对话。

另有 6 家信源报道X:Testing Catalog (@testingcatalog)Hacker News 热门(buzzing.cc 中文翻译)X:Kim (@kimmonismus)X:Claude (@claudeai)X:Yuchen Jin (@Yuchenj_UW)The Decoder:AI News(RSS)
推荐理由:Anthropic 把 Claude 变成了设计协作工具,从草图到代码的链路打通,对非设计师做原型是个实用入口,但能否撼动 Figma 还早。

4月16日4月16日周四

星期四 · 3 条
22:45
Anthropic:Newsroom(网页)精选
AI 评分 79/100
Claude Opus 4.7 正式发布

Claude Opus 4.7 全面上线,在高级软件工程任务上实现重大飞跃,93项编码基准测试解决率较 Opus 4.6 提升 13%,并首次解决四项前代无法完成的难题。新版本支持更高分辨率图像识别,在研究代理基准测试中总分达 0.715,长上下文性能表现最为稳定。模型定价维持每百万输入 token 5 美元、输出 25 美元不变。出于安全考量,其网络攻击能力较 Claude Mythos Preview 有所削弱,并配备自动拦截高风险网络安全请求的防护机制,合法安全研究人员可通过 Cyber Verification Program 申请使用权限。

另有 12 家信源报道Claude Code:GitHub Releases(RSS)X:Boris Cherny (@bcherny)X:Yuchen Jin (@Yuchenj_UW)X:Kim (@kimmonismus)X:Thariq (@trq212)X:Testing Catalog (@testingcatalog)X:Claude (@claudeai)X:Artificial Analysis (@ArtificialAnlys)X:Rohan Paul (@rohanpaul_ai)The Decoder:AI News(RSS)Claude:Blog(网页)Hacker News 热门(buzzing.cc 中文翻译)
推荐理由:Opus 4.7 不是最炫的模型,但多家真实用户反馈指出它在复杂工程任务上的可靠性提升是实打实的,尤其长链自主执行和指令遵循,对开发 Agent 的团队是一次直接的生产力升级。
08:00
Claude Platform:开发者版本说明(RSS)精选
AI 评分 82/100
Claude Opus 4.7 发布:复杂推理与智能体编码能力升级,定价不变

Anthropic 发布 Claude Opus 4.7,定位为最强大的通用模型,擅长复杂推理与智能体编码,定价保持与 Opus 4.6 相同的 $5/$25 per MTok。


推荐理由:Opus 4.7 发布,定价不变但加入了 task budgets 和高分辨率支持,对于构建 coding agent 的团队,这次更新相当于给模型装了进度条和显微镜。
07:46
Thariq@trq212精选
AI 评分 72/100
使用 Claude Code:会话管理与百万级上下文窗口的策略http://x.com/i/article/2044537014620721153Using Claude Code: Session Management & 1M ContextIn my recent calls with Claude Code users, one theme keeps coming up: the 1M token context window is a double-edged sword.It lets Claude Code operate autonomously for longer and handle tasks more reliably, but it also opens the door to context pollution if you're not deliberate about managing your sessions.Session management matters more than ever and there seem to be a lot of questions about it. Do you keep one session open in a terminal, or two? Start fresh with every prompt? When should you use compact, rewind, or subagents? What causes a bad compact?There’s a surprising amount of detail here that can really shape your experience with Claude Code and almost all of it comes from managing your context window.A Quick Primer on Context, Compaction & Context RotThe context window is everything the model can "see" at once when generating its next response. It includes your system prompt, the conversation so far, every tool call and its output, and every file that's been read. Claude Code has a context window of one million tokens.Unfortunately using context has a slight cost, which is often called context rot. Context rot is the observation that model performance degrades as context grows because attention gets spread across more tokens, and older, irrelevant content starts to distract from the current task. For our 1MM context model, we see some level of context rot happen around ~300-400k tokens, but it is highly dependent on the task- not a fast rule.Context windows are a hard cutoff, so when you’re nearing the end of the context window, you will need to summarize the task you’ve been working on into a smaller description and continue the work in a new context window, we call this compaction. You can also trigger compaction yourself.Every Turn Is a Branching PointSay you've just asked Claude to do something and it's finished, you’ve now got some information in your context (tool calls, tool outputs, your instructions) and you have a surprising number of options for what to do next:• Continue — send another message in the same session• /rewind (esc esc) — jump back to a previous message and try again from there• /clear — start a new session, usually with a brief you've distilled from what you just learned• Compact — summarize the session so far and keep going on top of the summary• Subagents — delegate the next chunk of work to an agent with its own clean context, and only pull its result back inWhile the most natural is just to continue, the other four options exist to help manage your context.When to Start a New SessionThe new 1M context windows means that you can now do longer tasks more reliably, for example to have it build a full-stack app from scratch. But just because your model hasn't run out of context, it doesn't mean you shouldn't start a new session.Our general rule of thumb is when you start a new task, you should also start a new session.A grey area is when you may want to do related tasks where some of the context is still necessary, but not all.For example, writing the documentation for a feature you just implemented. While you could start a new session, Claude would have to reread the files that you just implemented, which would be slower and more expensive. Since documentation may not be a highly intelligence sensitive task, the extra context is probably worth the efficiency gain of not having to re-read the relevant files again.Rewinding Instead of CorrectingIf I had to pick one habit that signals good context management, it’s rewind.In Claude Code, double-tapping Esc(or running /rewind) lets you jump back to any previous message and re-prompt from there. The messages after that point are dropped from the context.Rewind is often the better approach to correction. For example, Claude reads five files, tries an approach, and it doesn't work. Your instinct may be to type "that didn't work, try X instead." but the better move is to rewind to just after the file reads, and re-prompt with what you learned. "Don't use approach A, the foo module doesn't expose that — go straight to B."You can also use “summarize from here” to have Claude summarize its learnings and create a handoff message, kind of like a message to the previous iteration of Claude from its future self that tried something and it didn’t work.Compacting vs. Fresh SessionsOnce a session gets long, you have two ways to shed weight: /compact or /clear (and start fresh). They feel similar but behave very differently.Compact asks the model to summarize the conversation so far, then replaces the history with that summary. It's lossy, you're trusting Claude to decide what mattered, but you didn't have to write anything yourself and Claude might be more thorough in including important learnings or files. You can also steer it by passing instructions (/compact focus on the auth refactor, drop the test debugging).With /clear you write down what matters ("we're refactoring the auth middleware, the constraint is X, the files that matter are A and B, we've ruled out approach Y") and start clean. It's more work, but the resulting context is what you decided was relevant.What Causes a Bad Compact?If you run a lot of long running sessions, you might have noticed times in which compacting might be particularly bad. In this case we’ve often found that bad compacts can happen when the model can’t predict the direction your work is going.For example autocompact fires after a long debugging session and summarizes the investigation and your next message is "now fix that other warning we saw in bar.ts."But because the session was focused on debugging, the other warning might have been dropped from the summary.This is particularly difficult, because due to context rot, the model is at its least intelligent point when compacting. With one million context, you have more time to /compact proactively with a description of what you want to do.Subagents & Fresh Context WindowsSubagents are a form of context management, useful for when you know in advance that a chunk of work will produce a lot of intermediate output you won't need again.When Claude spawns a subagent via the Agent tool, that subagent gets its own fresh context window. It can do as much work as it needs to, and then synthesize its results so only the final report comes back to the parent.The mental test we use: will I need this tool output again, or just the conclusion?While Claude Code will automatically call subagents, you may want to tell it to explicitly do this. For example, you may want to tell it to:• “Spin up a subagent to verify the result of this work based on the following spec file”• “Spin off a subagent to read through this other codebase and summarize how it implemented the auth flow, then implement it yourself in the same way”• “Spin off a subagent to write the docs on this feature based on my git changes”SummaryIn summary, when Claude has ended a turn and you’re about to send a new message, you have a decision point.Overtime we expect that Claude will help you handle this itself, but for now this is one of the ways you can guide Claude's output.Claude Code 的百万级上下文窗口在支持长任务的同时,也带来了"上下文腐化"的风险,即模型性能可能在处理约30-40万token后开始下降。因此,有效的会话管理至关重要。关键策略包括:开启新任务时建议新建会话;对于关联任务可酌情保留上下文以提升效率;善用 /rewind 回退功能而非直接纠正错误,是维护上下文清洁的核心习惯。用户在每个对话轮次后,应根据情况选择继续、回退、新建会话、压缩或使用子代理。
推荐理由:Claude Code 1M 上下文听着爽,但 context rot 在 300k 就开始咬你了。这篇把 rewind、compact、subagent 三个操作的使用时机讲得极清楚,是目前最实用的 Claude Code 上下文管理指南,重度用户必读。

4月15日4月15日周三

星期三 · 2 条
03:07
Anthropic:Research(发表成果 · 网页)精选
AI 评分 74/100
自动化对齐研究者:利用大语言模型扩展可扩展监督

Anthropic部署9个Claude Opus 4.6作为自动化对齐研究者(AARs),在弱到强监督框架下自主研究800小时。通过自主提出、测试并优化对齐方案,AARs将性能差距恢复率(PGR)从0.23提升至0.97,成本约1.8万美元。该研究表明当前大模型已能独立开展对齐研究,为解决超人类AI的可扩展监督问题提供了可行路径。

另有 1 家信源报道X:Rohan Paul (@rohanpaul_ai)
推荐理由:Anthropic 让 Claude 自主探索对齐方法并接近完美恢复性能,自动化对齐研究的实证终于来了,代码开源,对齐研究者该认真看一遍。
01:03
Claude:Blog(网页)精选
AI 评分 72/100
Claude Code 推出 Routines 自动化功能

Claude Code 发布 routines 研究预览版,支持配置包含提示词、代码仓库及连接器的自动化工作流,可按定时计划(每小时/每晚/每周)、API调用或GitHub Webhook事件触发,并在云端基础设施运行,无需本地电脑保持在线。适用于自动处理积压工单、代码审查、部署验证及告警分类等场景。目前向Pro、Max、Team及Enterprise订阅用户开放,每日运行次数限制分别为5次、15次和25次。

另有 5 家信源报道Hacker News 热门(buzzing.cc 中文翻译)X:Claude (@claudeai)X:Testing Catalog (@testingcatalog)Claude:Blog(网页)The Decoder:AI News(RSS)
推荐理由:Claude Code 的 routines 让 AI 助手进入自动值守模式,定时修 bug、响应告警、审查 PR,很可能会改变开发者的工作流,可惜每日运行次数限制让重度用户有点不爽。

4月14日4月14日周二

星期二 · 1 条
21:58
The Decoder:AI News(RSS)精选
AI 评分 75/100
Claude Mythos 为欧洲 AI 安全体系敲响警钟

Anthropic 正限制访问其最新安全模型 Claude Mythos,该系统在发现软件漏洞方面表现优于大多数人类。英国已率先开展独立测试,而欧洲监管机构却对该系统几乎一无所知。这种反差暴露出欧洲 AI 安全监管体系存在深层结构性缺陷,凸显其在前沿 AI 评估能力上的严重滞后。

另有 4 家信源报道X:Rohan Paul (@rohanpaul_ai)X:Kim (@kimmonismus)X:OpenAI (@OpenAI)The Decoder:AI News(RSS)
推荐理由:Claude Mythos 事件暴露了欧洲在 AI 安全评估上的结构性短板,这不是监管过度,而是长期缺乏技术对等机构的结果,值得每一个关心 AI 治理的人认真读一遍。

4月13日4月13日周一

星期一 · 1 条
21:47
The Decoder:AI News(RSS)精选
AI 评分 70/100
AI 行业算力告急:面临服务中断、配额限制与 GPU 价格上涨

AI 智能体需求激增正与有限算力容量产生激烈冲突,引发行业连锁反应。Anthropic 因算力不足遭遇严重服务中断,OpenAI 宣布终止视频生成模型 Sora。据市场数据,GPU 价格已飙升近 50%,算力短缺正从基础设施层面制约 AI 行业扩张。


推荐理由:算力短缺不是一时半会儿的事,Anthropic掉链子、OpenAI砍Sora、GPU涨价48%,正逼着全行业重新算账,也把AI产品从免费狂欢推向了稀缺定价。

4月10日4月10日周五

星期五 · 2 条
05:28
Nathan Lambert:Interconnects(RSS)精选
AI 评分 60/100
Claude 神话与对开放权重的误导性恐慌

针对开放权重模型的恐慌情绪缺乏事实依据,围绕 Claude 等模型的安全争议存在明显的神话化倾向。作者批评这种对开源 AI 的恐惧散布是误导性的,认为所谓开放权重威胁论如同"又一场围绕开源的舞蹈",呼吁业界理性看待开源模型的安全性与价值,避免被无根据的安全恐慌所左右。

另有 1 家信源报道Gary Marcus:The Road to AI We Can Trust(RSS)
推荐理由:Nathan Lambert 对 Claude Mythos 引发的反开放权重恐慌提出关键反驳,认为问题被简化为全面禁令,忽略时间滞后本身就是安全阀,观点冷静且值得政策讨论者细读。
00:00
Claude:Blog(网页)精选
AI 评分 67/100
针对AI加速攻击的安全准备建议

Anthropic发布Project Glasswing项目,利用Claude Mythos Preview模型进行网络防御。文章指出,未来24个月内AI将发现大量长期潜伏的漏洞并快速转化为攻击,公开可用的次级模型已能发现传统审查遗漏的严重漏洞。建议立即修补CISA KEV目录漏洞,使用EPSS评分系统优先处理风险,互联网暴露系统需在24小时内完成补丁更新。预计未来两年漏洞报告量将增加一个数量级,安全团队需建立自动化修补流程以应对AI加速的攻击态势。

另有 3 家信源报道X:Boris Cherny (@bcherny)X:Anthropic (@AnthropicAI)X:slow_developer (@slow_developer)
推荐理由:Anthropic安全团队基于自家红队和研发实践写的防御指南,不是泛泛而谈,每条建议都配有具体工具和操作步骤,做安全的人可以对着清单逐项自查。

4月9日4月9日周四

星期四 · 2 条
00:00
Claude:Blog(网页)精选
AI 评分 64/100
Claude Cowork 正式上线企业版功能

Claude Cowork 向所有付费用户正式开放,并推出企业级管理控制功能。企业版管理员现可配置基于角色的访问权限、团队支出预算及 OpenTelemetry 可观测性,通过面板追踪 DAU/WAU/MAU 等使用数据。新增 Zoom MCP 连接器及按工具细分的权限控制。数据显示,目前大部分使用来自工程团队以外的运营、市场、财务和法务部门,主要用于项目更新、协作文档等任务。

另有 3 家信源报道X:Boris Cherny (@bcherny)X:Claude (@claudeai)X:Testing Catalog (@testingcatalog)
推荐理由:Claude Cowork 正式 GA 并补上企业治理拼图,角色权限、预算分析和 OTEL 日志让 IT 团队有底气推广,但它的杀手应用场景是已经被验证的非工程工作流,不是概念验证。
00:00
Claude:Blog(网页)精选
AI 评分 79/100
顾问策略:为智能体提供智能升级

Anthropic 在 Claude Platform 推出 advisor tool 测试版,实现"顾问策略":以 Opus 为顾问模型,Sonnet 或 Haiku 为执行模型。执行模型全程主导任务,仅在遇到决策难题时咨询 Opus 获取指导。该方案让 Sonnet 在 SWE-bench Multilingual 基准上性能提升 2.7 个百分点,同时成本降低 11.9%;Haiku 搭配 Opus 在 BrowseComp 得分达 41.2%,是单独 Haiku(19.7%)的两倍多,且比单独 Sonnet 成本低 85%。开发者只需在 API 请求中添加 advisor_20260301 工具即可启用,无需额外上下文管理。

另有 2 家信源报道X:Claude (@claudeai)X:Testing Catalog (@testingcatalog)
推荐理由:我觉得这个 advisor 模式是今年最实用的 API 设计之一,用最小的架构改动换来近 Opus 级的 agent 智能,账单却和 Sonnet 差不多,做复杂 agent 的团队可以直接套用。

4月8日4月8日周三

星期三 · 2 条
08:00
Claude Platform:开发者版本说明(RSS)精选
AI 评分 67/100
Claude 推出 Managed Agents 公开测试版与 ant CLI 命令行客户端

Claude 发布 Managed Agents 公开测试版,这是一个完全托管的智能体框架,支持通过 API 创建智能体、配置容器并运行会话,具备安全沙箱、内置工具和 SSE 流式传输。同时推出 ant CLI 命令行客户端,可加速与 Claude API 的交互,原生集成 Claude Code,并支持以 YAML 文件管理 API 资源版本。


推荐理由:Claude Managed Agents 把沙箱化代理从第三方方案变成了平台原生能力,ant CLI 则把 API 调用纳入了版本控制,这对习惯了 Claude Code 的开发团队是个自然升级,beta 阶段值得先占坑。
08:00
Tomer Tunguz 博客(VC 分析)精选
AI 评分 61/100
Mythos 横空出世

Anthropic 发布迄今最大的 Claude Mythos 模型,参数量达 10 万亿,为此前前沿模型的六倍。该模型在测试中发现了一处存在 27 年的操作系统漏洞和一处 16 年历史的视频软件缺陷,而传统工具曾对此检查过 500 万次均未发现。这些安全能力并非专门训练所得,而是代码、推理和自主性提升后涌现的副产品。Mythos 目前按 ASL-3 安全标准仅向 40 余家组织开放访问,这种排他性将重塑行业格局——拥有访问权的企业获得结构性安全优势,软件防护标准被重新定义,工程预算也将从开发转向深度加固。


推荐理由:Claude Mythos 的漏洞发现能力是意外的副产物,这让安全变成了准入壁垒,没有访问权的公司软件将默认多孔。