Topic · 主题全部主题 →

AI 编码

AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。

5,063条收录
582条精选

精选归档 · 第 7 页

121140 条 · 共 582

7月11日7月11日周六

星期六 · 2 条
09:20
Claude Code:GitHub Releases(RSS)精选
AI 评分 56/100
Claude Code v2.1.207 发布

Claude Code v2.1.207 发布。Auto 模式在 Bedrock、Vertex AI 和 Foundry 上无需 CLAUDE_CODE_ENABLE_AUTO_MODE 即可使用,可通过 disableAutoMode 设置关闭。修复了流式响应中包含超长列表、表格、段落或代码块时终端冻结和按键延迟的问题;修复了非交互式运行中远程托管设置被永久记录为已同意而未显示安全同意对话框的问题;修复了自动更新程序每次发布时覆盖 ~/.local/bin/claude 自定义启动脚本或符号链接的问题。Bedrock、Vertex 和 Claude Platform on AWS 默认切换为 Claude Opus 4.8。Auto 模式不再从 .claude/settings.local.json 读取 autoMode,改为使用 ~/.claude/settings.json。修复了 Windows 上 AWS 凭证解析卡住时无限挂起的问题,60 秒超时保护现在生效。


推荐理由:对使用 Bedrock、Vertex 和 Foundry 的开发者来说,自动模式默认开放是最实用的变化,加上模型默认升到 Opus 4.8,是个值得升级的稳定版。

7月10日7月10日周五

星期五 · 5 条
10:13
Claude Code:GitHub Releases(RSS)精选
AI 评分 62/100
Claude Code v2.1.206 发布

Claude Code v2.1.206 发布,主要更新包括:为 /cd 命令添加目录路径建议;新增 /doctor 检查以建议修剪 CLAUDE.md 文件中模型可从代码库推导的内容;/commit-push-pr 现在自动允许 git push 到仓库配置的推送远程仓库;/login 支持 Anthropic 运营的公共网关端点;后台智能体在更新后自动升级。修复了过期登录导致所有模型报错、claude --resume--continue 在启动时无键盘响应、MCP 服务器忽略 request_timeout_ms 配置、OAuth MCP 服务器需手动重新认证、/model 选择器价格显示错误、桌面会话卡在“运行中”状态、Windows 上键盘输入被忽略等多项问题。


推荐理由:Claude Code 这次更新修复了一串体验问题,/doctor 检查能帮团队清理冗余的项目文件,后台静默升级也减少了等待时间,日常使用的开发者更新一下不会亏。
06:21
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 71/100
Cognition 推出 SWE-1.7,接近 GPT 5.5 与 Opus 智能水平

Cognition 发布迄今最强模型 SWE-1.7,基于 Kimi K2.7 基座训练,通过强化学习管线改进(基础设施、训练稳定性、数据质量、长程任务技术)实现前沿智能水平并大幅降低成本。在 FrontierCode 1.1 Main 基准上达 42.3%(Kimi K2.7 Code 为 30.1%,GPT-5.5 为 43.0%,Opus 4.8 为 46.5%),Terminal-Bench 2.1 达 81.5%,SWE-Bench Multilingual 达 77.8%。模型针对长周期异步任务优化,现已在 Devin(Web、桌面、CLI)通过 Cerebras 以 1000 TPS 提供。


推荐理由:Cognition 把 RL 训练的编码模型推到了和 GPT-5.5 叫板的水平,成本还低得多,做代码 agent 的值得认真看看他们的训练稳定性、多集群等思路。
06:21
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 71/100
Bun 被 Anthropic 收购后用 Rust 重写,月下载超 2200 万

Bun 于 2025 年 12 月被 Anthropic 收购,作者使用预发布版 Claude Fable 5 进行了大量 Rust 重写。Bun 最初用 Zig 在一年内构建,如今 CLI 月下载超 2200 万,被 Claude Code 等采用。广泛功能带来稳定性挑战,v1.3.14 修复了多项 use-after-free、内存泄漏等 bug。团队通过 ASAN、Fuzzilli 模糊测试等系统性预防,并借助 Rust 的内存安全特性减少此类缺陷。

另有 1 家信源报道IT之家(RSS)
推荐理由:Bun 创始人把 54 万行 Zig 用 Claude 11 天重写为 Rust,对抗式审查和动态工作流的细节是近期最值得看的 AI 辅助工程实战复盘,做基础设施的可以认真读。
01:12
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
AI 评分 85/100
GPT-5.6 系列模型发布:Sol、Terra、Luna

OpenAI 推出 GPT-5.6 系列,包括旗舰 Sol、平衡型 Terra 和成本最优 Luna。在 Agents' Last Exam 上,Sol 以 53.6 分超越 Claude Fable 5(自适应推理)13.1 分;中等推理时以约四分之一成本领先 11.4 分。Terra 和 Luna 以约十六分之一成本超越 Fable 5。在 Artificial Analysis Coding Agent Index 上,Sol 以 80 分创 SOTA,高于 Fable 5 的 2.8 分,使用不到一半输出 token、一半时间,成本低约三分之一。引入 Programmatic Tool Calling 减少中间数据往返,ultra 模式协调多个智能体加速复杂任务。模型配备迄今最强安全防护,经人类红队和自动化测试验证。

另有 17 家信源报道X:Greg Brockman (@gdb)X:Jason Liu (@jxnlco)X:OpenRouter (@OpenRouter)X:OpenAI (@OpenAI)The Decoder:AI News(RSS)X:OpenAI Developers (@OpenAIDevs)OpenAI:官网动态(RSS · 排除企业/客户案例)Hacker News 热门(buzzing.cc 中文翻译)公众号:数字生命卡兹克X:Kim (@kimmonismus)X:Charlie Holtz(@charlieholtz)X:Rohan Paul (@rohanpaul_ai)TechCrunch:AI(RSS)X:Testing Catalog (@testingcatalog)X:Sam Altman (@sama)MarkTechPost(RSS)X:Noam Brown (@polynoamial)
推荐理由:GPT-5.6 在编码、安全、知识工作上全面领先,且 token 效率大幅优化,是我今年见过的真正能改变 agent 开发成本结构的发布。程序化工具调用和 ultra 模式值得立即试用。

7月9日7月9日周四

星期四 · 5 条
22:30
AI at Meta@AIatMeta精选
AI 评分 65/100
Meta 发布 Muse Spark 1.1 模型Straight from @finkd — Muse Spark 1.1 is live.来自 @finkd 的消息 - Muse Spark 1.1 已上线。

Mark Zuckerberg: (1) 今天我们发布了 Muse Spark 1.1--一款价格极低但能力强大的智能体与编程模型。该模型已通过我们全新的 Meta Model API 以及 Meta AI 开放使用。

另有 6 家信源报道X:Artificial Analysis (@ArtificialAnlys)X:Elvis Saravia (@omarsar0, DAIR.AI)The Decoder:AI News(RSS)IT之家(RSS)Hacker News 热门(buzzing.cc 中文翻译)Simon Willison 博客
推荐理由:Muse Spark 1.1 主打 agent 和 coding 能力,同时压到极低价,这可能是 Meta 低价模型策略的正式开场。虽然这次没给具体数据,但后续模型、API 的布局值得关注。
15:16
IT之家(RSS)精选
AI 评分 77/100
官方支招两种AI方案:Claude Fable 5搭配Sonnet 5省token

Anthropic官方建议将Claude Fable 5用作规划层、Sonnet 5执行任务以降低成本。顾问模式下,Sonnet 5主执行,仅需额外指导时调用Fable 5;SWE-bench Pro测试显示相比完全用Fable 5可达92%性能,成本仅63%。协调者模式下,Fable 5充当规划者,将子任务分派给多个Sonnet 5工作智能体;BrowseComp基准上达到Fable 5单独运行96%表现,成本为46%。

另有 2 家信源报道X:Claude Devs (@ClaudeDevs)The Decoder:AI News(RSS)
推荐理由:Anthropic 官方亲自下场教你省钱,把 Fable 5 当架构师、Sonnet 5 当码农,能保住九成以上性能同时省下近半成本,用 Claude 开发的人今天就可以在项目里试试。
04:08
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
AI 评分 70/100
OpenAI 审计 SWE-Bench Pro 发现约 30% 的评测任务存在缺陷

OpenAI 对编码评测基准 SWE-Bench Pro 进行详细审计,发现约 30% 的任务存在缺陷。在 731 个任务的公开子集中,前沿模型通过率在八个月内从 23.3% 提升至 80.3%,但数据质量检查显示大量任务存在测试过于严格、提示词描述不足、测试覆盖不全或误导性提示等问题。OpenAI 建议模型开发者仔细审视评测结果,并指出 AI 智能体在规模化数据质量检查中日益增长的实用性。

另有 2 家信源报道X:OpenAI (@OpenAI)The Decoder:AI News(RSS)
推荐理由:OpenAI 自己审计了 SWE-Bench Pro,发现三成任务有缺陷,这个基准给出来的分数可能要打问号,做模型评测和选型的人该认真看看。
02:50
xAI:News(网页)精选
AI 评分 85/100
xAI 发布 Grok 4.5

xAI 今日推出 Grok 4.5,为其最强模型,专为编程、智能体任务和知识工作打造。模型训练于数万块 NVIDIA GB300 GPU,DeepSWE 1.0 得分 62.0%,DeepSWE 1.1 得分 53%,Terminal Bench 2.1 得分 83.3%,SWE Bench Pro 解决率 64.7%。服务速度 80 TPS,token 效率约为 Opus 4.8 (max) 的 4.2 倍,在 Harvey Legal Agent Benchmark 排名第一。定价 $2/百万输入 token、$6/百万输出 token。即日起在 Grok Build、Cursor 和 SpaceXAI API 可用,欧盟地区预计 7 月中旬上线。

另有 7 家信源报道X:Elon Musk (@elonmusk, xAI)Cursor BlogX:Michael Truell (@mntruell)MarkTechPost(RSS)IT之家(RSS)The Decoder:AI News(RSS)Hacker News 热门(buzzing.cc 中文翻译)
推荐理由:Grok 4.5 在编码任务上追平第一梯队,但真正的杀手锏是极致性价比——输出成本只有对手的四分之一,还会倒逼编码模型继续降价。
01:22
ClaudeDevs@ClaudeDevs精选
AI 评分 73/100
Claude Code 的 Model 与 Effort:知道更多 vs. 更加努力http://x.com/i/article/2074606120292020224Model and effort in Claude Code: knowing more vs. trying harderClaude Code gives you two settings that both seem to "make the answer better": the model, and the effort level. But what do these actually do to the output? And how do you know whether to reach for a different model or just change the effort level?It's easy to assume that choosing a larger model like Fable gives you a smarter output than Sonnet, and that a higher effort level just means Claude thinks longer before it answers.The first assumption is true. Our largest models are more capable, according to industry-standard benchmarks.But effort means more than "thinking time." Effort controls how much work Claude does on your request overall. That includes how long it thinks, but also:• how many files it reads;• how much it verifies; and• how far it pushes through a multi-step task before checking in with you.At higher effort, Claude takes more of those actions (read files, run tests, double-check) before it comes back to you. At lower effort, it would rather ask you for more context than spend tokens figuring something out on its own.How model selection worksTo understand what the model setting actually controls, it helps to start at the very beginning, from the moment you press enter.Claude Code assembles your message together with the system prompt, tool definitions, your CLAUDE.md, the conversation history, and any files in context. All of this is sent as one request to the API.The model never sees any of that as plain text, though. The first thing that happens on the server is tokenization: the text gets split into pieces, and each piece is mapped to an integer from a fixed vocabulary the model was trained with. const might map to 1978, await might map to 4293. From here on, your prompt is an array of integers.The model's job is to take that array and predict which token comes next. It does this by computing a probability for every token in its vocabulary and picking from the top. After "const x = await", a well-trained model puts high probability on "fetch" (very likely) and near-zero on "banana" (not likely at all).What turns your input tokens into those probabilities is the weights (also called parameters): billions of numbers organized into large matrices. To predict one token, the model runs your input through those matrices (a long chain of matrix multiplications) and reads the probabilities at the end. The weights are where everything the model "knows" lives.The weights of each model are set during training, and by the time you're sending requests they're read-only. Nothing in your prompt, your CLAUDE.md, or your context changes them. If you've run into the word inference, that's all it means: using the model after training is done, with the weights fixed.Everything Claude knows about TypeScript, popular frameworks, or any other general programming knowledge was encoded into those weights at training time.Your prompt and context can still steer the prediction. Putting your real code in front of Claude is steering, and it works really well. However, this doesn't add anything to the weights themselves.If a library didn't exist when the model was trained, it isn't in the weights. You can put the docs in context and Claude will use them, but that's steering, not teaching. Claude's response is only influenced for that one request, but the underlying model hasn't retained anything.When Claude confidently calls an API that doesn't exist (a hallucination), that's the weights producing a token sequence that looks plausible from training patterns, not a failed lookup.So what does changing the model actually do? It swaps which set of frozen weights handles your request.The model doesn't generate a whole answer at once. It predicts one token, appends it to the sequence, and runs the whole computation again to get the next one. A 200-token response is 200 separate passes through the weights. This loop is where most of your wait time (and your output cost) comes from.The model setting decides which weights handle your request, and it also decides what each output token costs.What it doesn't decide is how many tokens get generated. That number can vary a lot for the same prompt, depending on how much work Claude decides to do.Which is exactly what effort controls.How effort worksWhile Claude Code is working on a task, the tokens it generates fall into a few categories:• Thinking: the reasoning you see streaming before and between actions.• Tool calls: structured blocks naming a tool like Read or Edit and its arguments, which Claude Code then parses and executes.• Text to you: the plan, progress updates, the summary at the end.These are all ordinary output tokens from the same loop, billed at the same rate. Thinking tokens, for example, are generated exactly like the other output tokens and stay in context for the rest of that turn.By the time Claude moves on to writing code, its earlier reasoning is part of the input, just like a file it read.So how does effort change any of this? The effort level is sent to the model as part of the request, right alongside your prompt. The model was trained to understand how to behave at each effort level, and that learned behavior is baked into the frozen weights.When your request arrives, effort is just one more input the model responds to, the same way it responds to your prompt text. It sets how thorough, and how certain, Claude needs to be before it considers the task done. That gets weighed on every turn, and higher confidence takes more tokens to reach.At higher effort levels, Claude often starts by creating a plan, and the effort level influences the depth and breadth of that plan. But the plan isn't frozen in place. As Claude gets results back from its actions, it updates its picture of how much progress it's made and how certain it is of the accumulated result.When step 1 of a three-hypothesis debugging plan finds the bug, "investigate hypotheses 2 and 3" may no longer be necessary. Claude will usually say this explicitly (e.g. "the first check found it, so the remaining checks aren't needed") and skip ahead. You see this happen in Claude Code when task lists get revised mid-run.Higher effort does make Claude more likely to double-check, like verifying the answer it found, or still look into the hypotheses it could have skipped. However, it generally won’t artificially inflate usage on a simple task just because the effort level is turned up. "Overthinking" is something our team specifically watches for during model training as it degrades effectiveness.Picking an effort levelFor most tasks, use the model's default effort level. The default is the level where Claude scales its token usage to what most people would want to spend on a task.Think of effort as a manual override on how hard and how long Claude works. Reach for it deliberately when you have a strong preference for thoroughness or speed based on your domain or the type of work you do, and treat it as a general preference, not a task-by-task decision.One practical note following the launch of Opus 4.8: in our testing, the default effort setting on Opus 4.8 produces better results for about the same amount of tokens as the default effort setting on Opus 4.7 on the same task.What to change when Claude gets it wrongWhen Claude gets something wrong, your first instinct shouldn't be to change a setting. It should be to look at the context you gave it. Is your prompt too vague? Is Claude connected to the right tools? Does it have the right skills?If you're increasing effort on a task that shouldn't need it, the fix is usually upstream: in your context, your CLAUDE.md, or how the task is scoped.But say you've given clear context and Claude still gets it wrong. The question to ask yourself is: did it not try hard enough, or did it not know enough?Model: the problem was too hardPick a larger model when the problem is genuinely hard, like subtle bugs, unfamiliar domains, architecture decisions. A larger model is what you want when the smaller model is confidently wrong no matter how much context you give it.Larger models are also better at handling ambiguity. On smaller models, specific instructions that direct the execution are a better recipe for success.Pick a smaller model when the work is routine: edits you can describe precisely, mechanical changes, questions about code that's already in context. There's no reason to pay for capability the task doesn't need.If Claude had all the pertinent context, clearly tried, and still got it wrong; that's a signal to pick a larger model. And if you're on the larger model and the work has been routine for a while, dropping down will increase speed and typically reduce cost without impacting the quality of the output.Effort: Claude didn't try hard enoughPick a higher effort level if Claude did it wrong by not trying hard enough: skipping a file, not running the tests, or not double-checking its work. This is most relevant if you'd selected an effort level below the model's default.The specialist, the expert, and the generalistOne way I like to think about the two settings is that Fable is a specialist who can handle problems almost no one else has, Opus is the expert, and Sonnet is a really good generalist. The effort level decides how much time any of them spends on your task.Opus at low effort is like getting five minutes with an expert who has deep experience with problems like yours. They bring knowledge that isn't anywhere in your codebase; patterns they've seen before, gotchas they know to check for, the kind of experience you only get from having solved a lot of similar problems. But five minutes means a quick read of your code, not a careful pass through every file.Sonnet at high effort is the generalist with the whole afternoon. They're great at coding, and they'll read everything, run things, double-check their work, and end up understanding your specific code thoroughly.Fable is the specialist you call when everyone else is stuck. Even at low effort, they'll spot the thing no one else would. That recognition is also what you're paying the most for, so it's worth saving it for the tasks that need it.None of these is universally "better". The model setting is roughly how capable; the effort setting is roughly how thorough. Most real tasks need some of both.Effort, model, and token consumptionSo how do model selection, effort, and token consumption all interact? It depends on the task.On routine work at the same effort level, both the larger and smaller models generally get it right. The larger model consumes more tokens with extra verification steps, at a higher per-token price. That's why dropping to the smaller model for routine stretches saves real money at no quality cost.On harder, multi-step work, the equation flips. The smaller model has to grind toward the limit of its ability, burning iterations, while the larger model reaches the same quality bar in fewer steps.You're paying more per token for the larger model, but on tasks that genuinely stretch the smaller one, the total cost per task can come out lower. And more importantly: the larger model can finish tasks the smaller one can't, even at the highest effort settings.This is most pronounced with Fable. On long, multi-step work it pulls furthest ahead. In our testing, it finished jobs Opus and Sonnet can't reach at any effort level. It also costs the most per token, which is the other reason to save it for the work that really needs it.The key point in the graphs above: effort picks how far Claude is willing to travel along the curve. That doesn't mean Claude will need to go that far to finish the task.Lastly, effort shapes token consumption, but it doesn't limit it. The only hard cap in the system is max_tokens, which truncates a response mid-stream when hit, but it's a blunt instrument and mostly relevant to API developers. Softer controls like task budgets or asking Claude to keep it brief in your prompt are more helpful. They're guidance the model is trained to follow (it'll look to wrap up as it gets near the limit) rather than a wall it runs into.Effort changes how much work Claude does. The model changes what Claude knows.When you're unhappy with a result, check the context before you touch either setting: give Claude a clear prompt, the right tools and skills, and a way to verify its own work.If Claude still gets it wrong, ask yourself: did it not know enough, or did it not try hard enough? Not knowing enough is a model problem, not trying hard enough is an effort problem.This article was written by @lydiahallie, member of technical staff on the Claude Code team.Claude Code 的 model 和 effort 两种设置都旨在提升输出,但机制不同。model 越大,模型能力越强(基于行业标准基准测试)。effort 控制 Claude 在请求上的总工作量,包括思考时间、读取文件数、验证程度、多步任务推进深度等。高 effort 时 Claude 会执行更多操作(读文件、跑测试、再检查);低 effort 时更倾向询问上下文。模型选择本质是切换不同的冻结权重集--权重在训练时固定,prompt 和上下文只能引导(steering)而不能改变权重。模型幻觉是权重产生看似合理但错误的 token 序列。
推荐理由:Claude Code 官方这篇把 model 和 effort 的取舍讲得比他处都透,读完就知道什么任务该堆算力、什么任务该降模型省钱。

7月8日7月8日周三

星期三 · 3 条
12:44
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 71/100
AI 审计代理在 Cloudflare CIRCL 中发现 7 个漏洞

zkSecurity 的 AI 审计代理 zkao 持续扫描 Cloudflare 的 CIRCL 密码学库,使用 Opus 4.6 + skills 和 GPT-5.3 + skills 等模型发现并确认了 7 个真实漏洞。其中包括阈值 RSA 中 float64 精度丢失(AI 自评 Critical)和属性基加密(CP-ABE)访问控制完全失效(Critical,由 zkao 自行发现)。所有漏洞已在上游修复,多数在 HackerOne 上获得确认和奖励。AI 生成的候选发现仍需人工验证,但 zkao 已能自动完成大部分验证工作。


推荐理由:zkSecurity用AI扫了Cloudflare的密码学库,挖出7个真实漏洞,从浮点数精度损失到访问控制完全破防。这是AI在密码学审计里第一次证明自己能找到能用的漏洞,不是纸上谈兵。虽然后面发现AI对严重性的判断还很瞎,但整体值得安全从业者一读。
02:12
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 73/100
YC CEO声称每日用AI部署3.7万行代码,开发者审查发现前端代码大量臃肿低效

Y Combinator CEO Garry Tan在X上宣称,他与AI编码代理每天在五个项目中部署37000行代码,并保持连续72天发布记录。波兰开发者Gregorein深入审查Tan网站前端代码,发现大量臃肿与低效问题:页面加载169次请求、总计6.42MB数据(对比Hacker News仅7次12KB);包含28个测试文件、78个未使用的JavaScript控制器、八种格式Logo(含空文件)、未压缩的旧PNG等。Gregorein指出,AI虽能快速生成代码,但质量仍应优先于数量。


推荐理由:Garry Tan 日行三万七千行的神话被代码审查拆穿,这不是节奏问题,是 AI 编码‘有量无质’的典型病征。认为 AI 编码能光速开发的人该冷静一下了。

7月7日7月7日周二

星期二 · 2 条
03:13
ClaudeDevs@ClaudeDevs精选
AI 评分 70/100
Claude Code 团队详解四种智能体循环类型http://x.com/i/article/2074204645845839872Getting started with loopsThere’s a lot of talk right now about "designing loops" instead of prompting your coding agent. If you spend some time on X trying to pin down what a loop actually is, you'll come across multiple different answers.On the Claude Code team, we define loops as agents repeating cycles of work until a stop condition is met. We categorize a few different types of loops based on:• How they are triggered• How they are stopped• What Claude Code primitive is used• What type of task is most appropriate for each.We’ll cover the main loop types, when to use each, and how to maintain code quality while managing token usage. Not all tasks require complex loops; start with the simplest solution and use these patterns selectively.Turn-based loops• Triggered by: A user prompt.• Stop criteria: Claude judges it has completed the task or needs additional context.• Best used for: Shorter tasks that are not part of a regular process or schedule.• Managed usage by: Write specific prompts and improve verification using skills to reduce the number of turns.Every prompt you send starts a manual loop with you directing each turn. Claude gathers context, takes action, checks its work, repeats if needed, and responds. We call this the agentic loop.For example, ask Claude to create a like button. It reads your code, makes the edit, runs the tests, and hands back something it believes works. You then manually check the work, and write the next prompt.You can improve the verification step by encoding your manual steps as a SKILL.md so Claude can check more of its own work, end-to-end. This should include tools or connectors to allow Claude to see, measure or interact with the result. The more quantitative the checks are, the easier it is for Claude to self-verify.For example, in your SKILL.md file you may specify:Goal-based loop (/goal)• Triggered by: A manual prompt in real-time.• Stop criteria: Goal achieved OR maximum number of turns reached.• Best used for: Tasks that have verifiable exit criteria.• Managed usage by: Setting a specific completion criteria and explicit turn caps, “stop after 5 tries.”Sometimes, a single turn is not enough, especially for more complex tasks. Agents do better when they can iterate. You can extend how long Claude keeps iterating by defining what done looks like with /goal.When you define the success criteria, Claude doesn’t have to make a determination on what is “good enough” and end the loop early. Each time Claude tries to stop, an evaluator model checks your condition and sends it back to work until the goal is met or a number of turns you define is reached.This is why deterministic criteria, such as number of tests passed or clearing a certain score threshold, are so effective.For example:Time-based loop (/loop and /schedule)• Triggered by: A specified time interval.• Stop criteria: You cancel it, or the work completes (the PR merges, the queue is empty).• Best used for: For recurring work, or interfacing with external environments / systems.• Managed usage by: Set longer intervals or react based on events rather than time.Some agentic work is recurring: the task stays the same and only the inputs change. For example, summarizing Slack messages every morning. Other work depends on external systems, and a simple way to interface with one is to check it on an interval and react to what changed. For example, a PR which may receive code reviews or fail CI.For these, you can trigger when Claude runs with /loop which re-runs a prompt on an interval. For example:/loop runs on your computer, so if you turn it off, it stops. You can move the loop to the cloud by creating a routine with /schedule.Proactive loops• Triggered by: An event or schedule, with no human in real time.• Stop criteria: Each task exits when its goal is met. The routine itself runs until you turn it off.• Best used for: Recurring streams of well-defined work: bug reports, issue triage, migrations, dependency upgrades, etc.• Managed usage by: Routing routines to smaller, faster models and using the most capable model for judgment calls.The primitives above, along with other Claude Code features like auto mode and dynamic workflows (research preview) can be composed into a loop for long-running work.For example, to handle incoming feedback, you can use:1. /schedule (research preview) to run a routine that checks for new reports1. /goal to define what done looks and skills to document how to verify it1. Dynamic workflows to orchestrate agents that triage each report, fix it, and review the fix1. Auto mode so the routine runs without stopping to ask for permissionPutting it together, a prompt could look like this:Maintaining code qualityThe quality of a loop’s output depends on the system around it. When designing the system:• Keep the codebase itself clean: Claude follows patterns and conventions that already exist in your codebase.• Give Claude a way to verify its own work: Encode what good looks like for you and your team with skills.• Make docs easy to reach: Frameworks and libraries docs have up-to-date best practices.• Use a second agent for code reviews: A reviewer with fresh context is less biased and not influenced by the main agent’s reasoning. You can use the built-in /code-review skill or Code Review for Github.When an individual result doesn’t meet the standard, don’t stop at fixing the individual issue, try to encode it to improve the system for all future iterations.Managing token usageTo manage token usage, loops should have clear boundaries:• Choose the right primitive and model for the job: Smaller tasks don’t need multiple agents or loops. Some tasks can use cheaper and faster models.• Define clear success and stop criteria: Be specific about what done looks like so Claude can arrive at the solution sooner (but not too soon).• Pilot before a large run: Dynamic workflows can spawn hundreds of agents. Gauge usage on a smaller slice of the work first.• Use scripts for deterministic work: Running a script is cheaper than reasoning through the steps. For example, a PDF skill can ship a form-filling script that Claude runs each time, instead of re-deriving the code.• Don’t run routines more often that you need to: Match the interval to how often the thing you’re watching changes• Review usage: The /usage command breaks down recent usage by skills, subagents, and MCPs, /goal with no arguments shows number of turns and token usage so far, /workflows shows each agent’s token usage and you can stop an agent at any time.Getting startedTo summarize:To get started with loops, look at the work you already do. Pick one task where you’re the bottleneck and ask which piece you could hand off: can you write the verification check? Is the goal clear enough? Does the work arrive on a schedule?Once you have an idea, run the loop, observe the results like where it stalls or over-reaches, and don’t be afraid to iterate on it.For more information, read the Claude Code docs on running agents in parallel, as well as the loop, schedule, goal, and dynamic workflows pages.This article was written by @delba_oliveiraClaude Code 团队将"设计循环"定义为智能体重复工作直到满足停止条件,划分四种类型:1)回合循环--手动提示触发,Claude 自判完成,适合短任务,可通过 SKILL.md 提升验证;2)目标循环--/goal 手动触发,达成目标或达最大轮数停止,需确定性完成标准(如测试通过数);3)时间循环--/loop 和 /schedule 按间隔触发,适合同步消息、检查 PR 等重复任务,可云端运行;4)主动循环--事件或计划触发,无人实时参与,每个子任务独立退出。建议从最简单方案开始,选择性使用复杂循环。
推荐理由:Claude Code 团队官方的循环设计指南,把 `/goal`、`/loop` 这些原语讲得很清楚,想从单次提示转向自主代理工作流的开发者可以直接照着搭。
02:20
Claude:Blog(网页)精选
AI 评分 70/100
Claude Fable实地指南:发现你的未知

Claude Fable是第一款要求用户主动澄清未知才能获得高质量工作的模型。与Claude Fable协作是一个在实现前后迭代发现未知的过程。通过将问题分解为已知的已知、已知的未知、未知的已知和未知的未知四类,用户可以借助Claude Fable和Claude Code进行盲点检查、头脑风暴、原型设计、实现笔记记录以及答辩解释,从而高效挖掘并解决深藏于代码库和设计与实现中的潜在问题。


推荐理由:Anthropic 官方分享的 Claude Fable 协作方法论,把「发现未知」拆成盲点扫描、原型、面试等可操作步骤,如果你用 Claude Code 但常觉得代理跑偏,这篇是必读实践指南。

7月6日7月6日周一

星期一 · 1 条
20:07
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 77/100
AI颠覆初级程序员就业市场:斯坦福数据揭示年轻开发者就业锐减19%

斯坦福数字经济实验室基于ADP薪资数据发现,美国22-25岁软件开发人员就业较2022年峰值下降19%,而41-49岁增长14%。入门级岗位招聘减少28%,计算机科学毕业生失业率达6.1%,高于文科专业。核心推手是2024-2025年兴起的智能体编程(Agentic programming)。总程序员就业增长4.4%,但全部来自年长群体。GitHub一年新增3600万账号,80%新用户一周内使用Copilot。编程工作未消失,但“初级程序员”头衔正在消亡。


推荐理由:用斯坦福、BLS、GitHub和App Store的数据把“初级程序员被替代”讲成了确凿事实,更关键是指出编程正从职业变成每个人的基本能力——对所有写代码的人是必读。

7月5日7月5日周日

星期日 · 1 条
22:21
Meituan LongCat@Meituan_LongCat精选
AI 评分 81/100
美团 LongCat-2.0 完全开源(MIT 许可),1.6T MoE 模型开放权重与推理代码🐱 LongCat-2.0 is now fully open-source — MIT licensed, no restrictions.Since our launch a few days ago, the response from the community has been incredible. Thank you for all the feedback, discussions, and interest.Today, we’re releasing the model weights and inference code to everyone. ◆ 1.6T MoE · ~48B active · 1M token context ◆ Agent-native: Integrates directly with Claude Code, OpenClaw, and Hermes Agent ◆ Deployment: Support both GPU and NPU platforms— verified on large-scale domestic clusters📑 Tech Blog: https://longcat.ai/blog/longcat-2.0/ 🤗 HuggingFace: https://huggingface.co/meituan-longcat/LongCat-2.0 💻 GitHub: https://github.com/meituan-longcat/LongCat-2.0 🪄 ModelScope: https://modelscope.ai/collections/meituan-longcat/LongCat-20 👇 Inference Code GPU: https://github.com/sgl-project/sglang/pull/30042 NPU: https://github.com/meituan-longcat/SGLang-FluentLLM/tree/npu美团今日宣布 LongCat-2.0 完全开源(MIT 许可),公开模型权重与推理代码。该模型为 MoE 架构,总参数量 1.6T,每 token 激活约 48B,支持 1M token 上下文。技术亮点包括 LongCat Sparse Attention 高效处理长文本、Zero-Compute Experts 动态激活 33B-56B 零浪费计算、MOPD 按任务路由 Agent/Reasoning/Interaction 三组专家。Benchmark 成绩:Terminal-Bench 2.1 70.8;SWE-bench Pro 59.5(超越 GPT-5.5 的 58.6);SWE-bench Multilingual 77.3;FORTE 73.2;RWSearch 78.8;BrowseComp 79.9。原生集成 Claude Code、OpenClaw、Hermes Agent 等工具,支持 GPU 与 NPU 部署,已在大规模国内集群验证。

Meituan LongCat: 推出 LongCat-2.0 🐱 1.6T 参数 · MoE 架构,约 48B 活跃参数 · 1M 上下文窗口 这是 @OpenRouter 上 Owl Alpha 背后的完整模型--现已可用。 从零开始为智能体编程构建: ◆ LongC...

另有 1 家信源报道MarkTechPost(RSS)
推荐理由:国内大厂首个在 SWE-bench Pro 上超过 GPT-5.5 的开源模型,MIT 协议无任何限制,搞 agentic coding 的团队可以直接用,是个重要转折。

7月4日7月4日周六

星期六 · 1 条
03:44
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 83/100
pxpipe:通过图像化压缩输入token降低Claude Code成本

pxpipe是一个本地代理,将系统提示、工具文档和历史记录等密集文本渲染为PNG图像,利用图像token成本取决于像素尺寸的特性压缩输入token。在Fable 5模型上,约25k文本token压缩为约2.7k图像token,端到端账单降低59–70%。SWE-bench Lite 10个实例全部通过,成本从$54降至$27;SWE-bench Pro 19对测试中18对判定一致,单次请求成本降低约60%。该方法有损(精确ID等需保持文本),默认仅处理claude-fable-5请求,可通过PXPIPE_MODELS变量控制。


推荐理由:pxpipe 通过把大量上下文渲染成图像来降低 token 开销,实测能削减 60-70% 的账单,对重度使用 Claude Code 的开发者很诱人,但它有损,精确值可能读错,适合容错高的编码场景。