内容
精选全部 AI 动态热点榜AI 日报主题收藏
模型
模型榜Tibo重置监控
更多
Agent 接入关于更新日志反馈
京ICP备2026012723号-5
精选全部日报更多
反馈

全部 AI 动态

全部动态模型 · 35 条
来源全部一手资讯X
类型模型
全部模型产品行业论文教程观点
标签「MCP/工具调用」清除
全部 AI 动态
全部模型产品行业论文教程观点
模型 · 标签「MCP/工具调用」 · 35 条清除

9月8日9月8日周二

星期二 · 1 条
01:07
OpenBMB@OpenBMB
AI 评分 54/100
面壁智能发布 MiniCPM5-2B,vLLM 稳定版提供 day-0 支持Thank you @vllm_project for the day-0 support 🙌译面壁智能(OpenBMB)发布 MiniCPM5-2B,vLLM 稳定版提供 day-0 支持。该模型为 2.6B Dense 模型,原生 131K 上下文,同一 checkpoint 支持 Think/No-Think 两种模式,并通过 vLLM 的 minicpm5 parser 支持工具调用。模型采用标准 LlamaForCausalLM 架构,权重与训练数据一同开放,部署配方见 http://recipes.vllm.ai/openbmb/MiniCPM5-2B。

vLLM: 🤝 Day-0 support for MiniCPM5-2B on stable vLLM. ⚡ Dense 2.6B model with 131K native context 🧠 Think / No-Think from th...

MCP/工具开源生态推理模型发布

8月28日8月28日周五

星期五 · 1 条
14:11
公众号:腾讯混元精选
AI 评分 82/100
腾讯混元发布 Hy4 preview:770B 总参数、1M 上下文,开源上线

腾讯混元发布新一代旗舰模型 Hy4 preview,总参数 770B、激活参数 49B、上下文长度 1M,现已开源并在腾讯云 TokenHub 和 OpenRouter 上线。

MCP/工具开源生态推理模型发布

推荐理由:官方盲测中它与 GLM-5.3、Kimi K3 的分差不足 0.1,选型时 1M 上下文和开源部署成本可能比榜单排名更实际。

8月14日8月14日周五

星期五 · 1 条
13:52
MarkTechPost(RSS)
AI 评分 58/100
Meet Needle 2:一个 45M 参数的开源工具调用模型,以 14MB 二进制文件发布,28MB RAM 即可运行完整会话

Cactus Compute 发布 Needle 2,一个 45M 参数的开源工具调用模型,以 14MB 二进制文件发布,约 28MB RAM 即可运行完整会话。

MCP/工具模型发布端侧

8月11日8月11日周二

星期二 · 1 条
21:13
NVIDIA Blog(RSS)同事件
AI 评分 75/100
NVIDIA 发布 Nemotron 3.5 Lightning 与 NeMo Switchyard,提升智能体 AI 效率

NVIDIA 扩展 Nemotron 3 模型家族,推出 300 亿参数混合专家模型 Nemotron 3.5 Lightning,面向高吞吐智能体工作负载,输出速度最高提升 4 倍,智能体任务完成速度提升 30%。同时发布开源模型路由库 NeMo Switchyard,可将任务成本降至 Opus 4.8 单独运行的近三分之一,并支持在 RTX PC、DGX 等本地设备部署。

智能体MCP/工具开源/仓库模型发布
同一事件,精选展示《NVIDIA 推出 Nemotron 3.5 Lightning,加速本地智能体任务》

8月3日8月3日周一

星期一 · 1 条
11:31
MiniMax (official)@MiniMax_AI
AI 评分 67/100
MiniMax H3 开源并原生接入 ComfyUIThe weights are here. The nodes are native. The rest is up to your graph. 👀 H3 is in @ComfyUI on Day 0!!one model, five workflows, endless ways to wire it. 😌#MiniMaxH3 #OpenWeights译MiniMax H3 权重正式开源,并在发布当天(Day 0)原生集成到 ComfyUI,支持文本生视频、图像生视频、首尾帧控制、参考视频及就地编辑五种工作流。H3 以多模态上下文理解为核心,可同时处理图像、音频和视频,并依据提示词解析其关联。

ComfyUI: MiniMax H3 is native in ComfyUI, day zero with the weights. Model highlights: → Text-to-video, prompt only → Image-to-vi...

MCP/工具多模态开源生态模型发布

7月30日7月30日周四

星期四 · 1 条
00:25
OpenRouter@OpenRouter
AI 评分 45/100
阿里Qwen3.7-Flash上线OpenRouterQwen3.7-Flash from @Alibaba_Qwen is live on OpenRouter.A fast vision capable reasoning model for multimodal agents, visual coding, search, and computer interaction, with tool use and a 1M context window.Try it now: https://openrouter.ai/qwen/qwen3.7-flash译来自 @Alibaba_Qwen 的 Qwen3.7-Flash 已在 OpenRouter 上线。 一款支持视觉的快速推理模型,适用于多模态智能体、视觉编码、搜索和计算机交互,具备工具使用能力和 1M 上下文窗口。 立即体验:https://openrouter.ai/qwen/qwen3.7-flash
智能体MCP/工具多模态模型发布

7月10日7月10日周五

星期五 · 1 条
02:09
The Decoder:AI News(RSS)同事件
AI 评分 72/100
OpenAI 发布 GPT-5.6 公开版及 ChatGPT Work 智能体

OpenAI 正式推出 GPT-5.6 公开版本,并发布基于 GPT-5.6 和 Codex 技术的 ChatGPT Work 智能体。该智能体可自主处理数小时的复杂工作流,例如从客户研究到生成营销素材并适配多市场。它集成了 Unified Plugins Directory,支持 Google Drive、Slack、Salesforce 等第三方插件,并配备 Auto-Review 功能以防止数据泄露。桌面端所有用户即刻可用;Web 和移动端优先开放给 Pro、Enterprise、Edu 用户,Plus 和 Business 稍后跟进。ChatGPT Work 与 Codex 共享 agent 消费池,按任务复杂度和所选模型计费。

智能体MCP/工具OpenAI模型发布
同一事件,精选展示《OpenAI 推出 ChatGPT Work:可跨应用自主工作的 AI 智能体》

7月9日7月9日周四

星期四 · 1 条
02:52
Peter Steinberger 🦞@steipete
AI 评分 57/100
OpenAI 发布第三代语音模型 GPT-LiveThis is how you wanna talk with your claw.译这就是你想和你的爪子聊天的方式。 今天,我们推出了第三代语音模型与架构 GPT-Live。GPT-Live 是一款全双工模型,内置异步委派功能,能够实现极其自然的对话,同时具备你期望 ChatGPT 拥有的智能。 https://openai.com/index/introducing-gpt-live/

Justin Uberti: Today, we're launching our third-gen voice model and architecture, GPT-Live. GPT-Live is a full-duplex model with built-...

MCP/工具OpenAI模型发布语音

7月7日7月7日周二

星期二 · 1 条
07:18
🚨 AI News | TestingCatalog@testingcatalog
AI 评分 64/100
GPT-Realtime-2.1 系列上线 Playground 和 APIOPENAI 🔥: GPT-Realtime-2.1 and GPT-Realtime-2.1-mini are now available on OpenAI Playground and APIs.GPT-Realtime-2.1-mini now supports reasoning and tool use at the same cost as GPT-Realtime-mini.Wen Bidi??? 👀译OPENAI 🔥: GPT-Realtime-2.1 和 GPT-Realtime-2.1-mini 现已在 OpenAI Playground 和 API 中可用。 GPT-Realtime-2.1-mini 现在支持推理和工具使用,成本与 GPT-Realtime-mini 相同。 Wen Bidi??? 👀

OpenAI Developers: GPT-Realtime-2.1-mini is now available in the API, bringing reasoning and tool use to our Realtime mini lineup at the sa...

MCP/工具OpenAI推理模型发布

7月1日7月1日周三

星期三 · 2 条
02:28
ClaudeDevs@ClaudeDevs同事件
AI 评分 79/100
Claude Sonnet 5 发布:1M上下文窗口,编码工具性能顶级Claude Sonnet 5 is here.Top-tier performance on coding and tool use at Sonnet pricing, with a 1M context window.It's the new default in Claude Code for Pro users, and available everywhere on the Claude Platform, including the API and Managed Agents.译Claude Sonnet 5 已推出。 以 Sonnet 定价提供顶级编码和工具使用性能,并拥有 1M 上下文窗口。 它已成为 Pro 用户 Claude Code 的新默认模型,并可在 Claude 平台所有位置使用,包括 API 和托管智能体。

Claude: Introducing Claude Sonnet 5, our most agentic Sonnet yet. It makes plans, uses tools like browsers and terminals, and ru...

AnthropicMCP/工具模型发布编码
同一事件,精选展示《Claude Sonnet 5 发布》
02:28
Claude@claudeai精选
AI 评分 73/100
Claude Sonnet 5:最具智能体能力的版本Introducing Claude Sonnet 5, our most agentic Sonnet yet.It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models.译介绍 Claude Sonnet 5,这是迄今为止最具智能体能力的 Sonnet。 它会制定计划、使用浏览器和终端等工具,并以几个月前还需要更大、更昂贵模型才能达到的水平自主运行。
智能体AnthropicMCP/工具模型发布

推荐理由:Sonnet 5 把浏览器和终端直接内建为工具,中端模型首次跑通自主执行链路,我觉得很多之前因为成本没上的 agent 方案可以重新算账了。

6月25日6月25日周四

星期四 · 1 条
05:29
Hacker News 热门(buzzing.cc 中文翻译)同事件
AI 评分 71/100
Gemini 3.5 Flash 中的计算机使用

Google 将计算机使用(Computer use)作为内置工具集成至 Gemini 3.5 Flash,使开发者能构建跨浏览器、移动端和桌面环境的智能体。此前该功能仅作为独立模型在 Gemini 2.5 中提供,现已原生整合至主 Flash 模型。开发者可通过 Gemini API 及 Gemini Enterprise Agent Platform 调用。安全方面,模型采用针对性对抗训练降低提示注入风险,并新增两项可选企业级保护:要求用户确认敏感操作、检测到间接提示注入时自动停止。该能力在持续软件测试、跨应用知识工作等长周期企业自动化场景中表现更优。(198字)

智能体GoogleMCP/工具模型发布
同一事件,精选展示《Gemini 3.5 Flash 引入 computer use 功能》

6月24日6月24日周三

星期三 · 2 条
18:12
Qwen@Alibaba_Qwen同事件
AI 评分 76/100
通义千问发布Qwen-AgentWorld原生语言世界模型📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within a single model. Environment modeling is the training objective from day one, not a post-hoc adaptation.🤔 LLMs are trained to be better agents — better at acting in environments. But nobody has trained them to model the environments themselves.🗺️ Our roadmap: investigate how language world modeling can push the boundaries of general agent capabilities, along two routes:1️⃣ Build a foundation model for environment simulation — outperforming Claude Opus 4.8 and GPT-5.4 on AgentWorldBench2️⃣ Investigate how world modeling enhances agent training: 🔬 Controllable Sim RL (agentic RL with LWM as environments) surpasses training in real environments 🧠 Learning to predict environments (LWM warm-up) makes agents stronger — remarkably, even without any agent-specific training, this predictive knowledge transfers to agentic tasks with zero fine-tuning📑 Paper: https://arxiv.org/abs/2606.24597 📖 Blog: https://qwen.ai/blog?id=qwen-agentworld 💻 GitHub: https://github.com/QwenLM/Qwen-AgentWorld 🤗 HuggingFace: https://huggingface.co/collections/Qwen/qwen-agentworld 🧩 ModelScope: https://modelscope.cn/collections/Qwen/Qwen-AgentWorld译通义千问发布Qwen-AgentWorld,一款原生语言世界模型,可在单一模型中模拟MCP、搜索、终端、SWE、Web、OS、Android共7种智能体环境。环境建模即训练目标,非事后适配。该模型在AgentWorldBench上性能超越Claude Opus 4.8和GPT-5.4。研究分两条路径:一是构建环境模拟基础模型;二是探索世界模型增强智能体训练--可控Sim RL(以LWM为环境的智能体强化学习)优于真实环境训练,而LWM预热(预测环境的学习)即使不经任何智能体特定微调,也能将预测知识迁移至智能体任务。
智能体arXivMCP/工具模型发布
同一事件,精选展示《Qwen-AgentWorld 开源:让 Agent 学会"先预测,再行动"》
11:54
Qwen:Blog Retrieval(API)精选
AI 评分 81/100
Qwen-AgentWorld:面向通用智能体的语言世界模型

Qwen 团队发布 Qwen-AgentWorld,一个以环境建模为训练目标的原生语言世界模型,在单个模型中模拟 MCP、Search、Terminal、SWE 及 GUI 域(Web、OS、Android)共七个域。模型使用超 1000 万条真实交互轨迹训练,在 AgentWorldBench 上以 Qwen-AgentWorld-397B-A17B 版本达最高模拟质量,超越 GPT-5.4、Claude Opus 4.8 和 Gemini 3.1 Pro。同时发布评测基准 AgentWorldBench。该模型可作为解耦环境模拟器用于智能体 RL 训练,也可作为统一智能体基础模型,经 LWM 预热后无需智能体 RL 微调即可迁移。模型和基准已开源在 Hugging Face 和 ModelScope。

智能体arXivHugging FaceMCP/工具

推荐理由:Qwen把世界模型做成了一个可开源的通用产品,覆盖七域,做agent RL的可以直接拿它仿真训练,可控性甚至超过真实环境,做agent的团队应该认真看看。

6月22日6月22日周一

星期一 · 1 条
16:05
🚨 AI News | TestingCatalog@testingcatalog
AI 评分 64/100
Sakana AI 发布 Fugu 和 Fugu Ultra 多智能体编排系统BREAKING 🔥: Sakana AI announced the Sakana Fugu and Sakana Fugu Ultra systems, which perform on par with Claude Fable 5 and Mythos 5 across many benchmarks.Sakana AI is an AI lab from Japan, and Fugu is an orchestration model trained to operate other LLMs.It is available as an API but not yet accessible in the EEA region.That's a natural evolution. Orchestration multi-model systems will outperform single-model systems, and they will become much more accessible for smaller labs and companies to build.Big players will have to consider building orchestrating systems that rely on models built by competitors. It is already happening at Meta, Apple, and Microsoft, and will likely catch Google, Anthropic, and OpenAI as well eventually.译Sakana AI 宣布推出 Fugu 和 Fugu Ultra 系统。Fugu 是一个多智能体编排模型,训练用于操控其他 LLM,通过单一模型 API 访问。其中 Fugu Ultra 在多项基准测试中性能匹敌 Claude Fable 5 和 Mythos 5,并宣称提供前沿能力且规避出口管制风险。该系统目前通过 API 提供服务,但暂不支持 EEA 地区。推文指出,编排式多模型系统将超越单一模型,使小型实验室和企业更易构建,并已促使 Meta、Apple、微软等巨头考虑采用竞争对手的模型搭建编排系统。

Sakana AI: Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API. Our ‘Fugu Ultra’ mod...

智能体MCP/工具模型发布

6月4日6月4日周四

星期四 · 2 条
05:58
MiniMax (official)@MiniMax_AI精选
AI 评分 78/100
MiniMax M3联袂Mem0推持久记忆AIMem0 is an official launch partner for MiniMax M3!M3's 1M token context window + @mem0ai 's memory layer = AI apps that truly remember.Build personalized AI agents with persistent memory, now with 50% off M3 during launch week.Get started with Minimax → https://platform.minimax.io/docs/guides/models-introSign up with mem0 → http://app.mem0.ai/?utm_source=minimax_x_post译Mem0 是 MiniMax M3 的官方启动合作伙伴! M3 的 1M token 上下文窗口 + @mem0ai 的记忆层 = 真正记住的 AI 应用。 构建具有持久记忆的个性化 AI 智能体,现在启动周内 M3 享五折优惠。 开始使用 Minimax → https://platform.minimax.io/docs/guides/models-intro 注册 mem0 → http://app.mem0.ai/?utm_source=minimax_x_post
智能体MCP/工具模型发布
另有 7 家信源报道公众号:MiniMax(稀宇科技)IT之家(RSS)X:karminski (@karminski3)X:MiniMax (@MiniMax_AI)X:OpenRouter (@OpenRouter)X:硅基流动 SiliconFlow (@SiliconFlowAI)X:Testing Catalog (@testingcatalog)
推荐理由:MiniMax 把 1M 上下文和 Mem0 记忆层绑在一起,不是单纯秀参数,是给 Agent 装了个硬盘,做长期记忆产品的该关注一下。
04:28
MiniMax (official)@MiniMax_AI
AI 评分 65/100
MiniMax M3携mem0推1M记忆层@mem0ai is an official launch partner for MiniMax M3!M3's 1M token context window + @mem0ai 's memory layer = AI apps that truly remember.Build personalized AI agents with persistent memory, now with 50% off M3 during launch week.Get started with Minimax → https://platform.minimax.io/docs/guides/models-introSign up with mem0 → http://app.mem0.ai/?utm_source=minimax_x_post译@mem0ai 是 MiniMax M3 的官方发布合作伙伴! M3 的百万 token 上下文窗口 + @mem0ai 的记忆层 = 真正能记住的 AI 应用。 构建带有持久记忆的个性化 AI 智能体,发布周期间 M3 可享 5 折优惠。 开始使用 Minimax → https://platform.minimax.io/docs/guides/models-intro 注册 mem0 → http://app.mem0.ai/?utm_source=minimax_x_post
智能体MCP/工具模型发布

6月2日6月2日周二

星期二 · 2 条
21:06
StepFun@StepFun_ai
AI 评分 73/100
阶跃星辰 Step 3.7 Flash 发布:开放权重模型进军智能体编程Open weights are moving from model cards into real coding workflows.Step 3.7 Flash is designed for fast agentic coding, reliable tool calling, and multimodal understanding.Big thanks for the blog from the @kilocode team: https://blog.kilo.ai/p/new-models-from-stepfun-and-minimax译阶跃星辰发布 Step 3.7 Flash 模型,强调其为快速智能体编程设计,具备可靠的工具调用与多模态理解能力。该模型采用开放权重。同期,MiniMax 也开源了 M3 模型。两者已均在 Kilo 中上线。此次发布凸显了开放权重模型正从模型卡片走向实际编程工作流的趋势。

Kilo: The open-weight labs did not come to play this week. StepFun dropped Step 3.7 Flash. MiniMax dropped M3. Both with open ...

MCP/工具开源/仓库模型发布编码
01:18
MiniMax (official)@MiniMax_AI同事件
AI 评分 78/100
MiniMax M3模型与智能体对齐实践this is what model-and-agent alignment looks like 🤝 @SimularAI译这就是模型与智能体对齐的样子 🤝 @SimularAI

Simular: Today @MiniMax_AI ships M3 — the first frontier model purpose-built for computer-use agents. Natively multimodal. One mo...

智能体MCP/工具多模态模型发布
同一事件,精选展示《MiniMax M3 模型上线 Cloudflare AI Gateway》

6月1日6月1日周一

星期一 · 2 条
10:15
MiniMax (official)@MiniMax_AI同事件
AI 评分 79/100
MiniMax M3:首个融合三大前沿能力的开源模型Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities• Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas • MiniMax Sparse Attention scales context to 1M • Natively Multimodal from Step ZeroAPI: http://platform.minimax.io Token Plan: https://platform.minimax.io/subscribe/token-plan 🚀New! MiniMax Code: http://code.minimax.ioWeights & Tech Report in ~10 Days译介绍 MiniMax M3:首个融合三大前沿能力的开源权重模型 • 编码与智能体前沿:59.0% SWE-Bench Pro,66.0% Terminal Bench 2.1,34.8% SWE-fficiency,28.8% KernelBench Hard,74.2% MCP Atlas • MiniMax Sparse Attention 将上下文窗口扩展至 1M • 从零开始原生多模态 API:http://platform.minimax.io Token 计划:https://platform.minimax.io/subscribe/token-plan 🚀新!MiniMax Code:http://code.minimax.io 权重与技术报告将在约 10 天内发布
智能体MCP/工具多模态模型发布
同一事件,精选展示《MiniMax M3:前沿编码、100万token上下文与原生多模态一体模型》
09:24
公众号:MiniMax(稀宇科技)精选
AI 评分 78/100
MiniMax M3 发布:前沿 Coding、1M 上下文与原生多模态集于一身

MiniMax 发布 M3 模型,采用全新稀疏注意力架构 MSA,支持 1M 上下文窗口,并原生支持图片、视频输入及电脑桌面操作,是国内首个齐备这些要素的开源模型。

MCP/工具多模态推理模型发布

推荐理由:M3在复现论文和CUDA优化中的长程自主迭代,将模型评估从单轮得分推进到真实协作场景,为衡量Agent Coding能力提供了更具参考价值的维度。

5月29日5月29日周五

星期五 · 2 条
12:40
StepFun@StepFun_ai
AI 评分 71/100
阶跃星辰Step 3.7 Flash在ZenMux平台上线Step 3.7 Flash now showing up on @ZenMuxAI — nice to see it plugged into more model stacks!译阶跃星辰(Step Fun)的视觉语言模型Step 3.7 Flash已在ZenMux平台上线。该模型采用稀疏MoE架构,专为智能体、编程、搜索、多模态及长上下文工作流设计。其核心性能包括:400 TPS推理速度、约110亿激活参数、256K上下文窗口及3个推理级别。该模型能够理解UI、图表、文档和图像以编写代码或调用工具,并擅长深度网络与视觉搜索,在τ2-bench上跨难度级别取得98%+的成绩。它兼容Claude Code、MCP风格工作流等,并可本地部署于Mac Studio M4 Max、DGX Spark等硬件。

ZenMux: Excited to support Step 3.7 Flash by @StepFun_ai on ZenMux from day one. 🚀 A sparse MoE vision-language model built for...

智能体MCP/工具多模态模型发布
08:02
公众号:阶跃星辰(Step)精选
AI 评分 61/100
阶跃发布 Step 3.7 Flash,面向生产级 Agent 的高效率 Flash 模型

阶跃星辰发布并开源 Step 3.7 Flash,采用稀疏 MoE 架构(总参数 196B+1.8B,激活 11B),最高生成速度 400 Tokens/s。围绕原生多模态理解与执行、联网与视觉搜索增强、高可靠工具调用与编排、Agent 生态兼容优化四大能力优化。在 Toolathlon 达 49.5%,ClawEval-1.1 达 67.1%,GDPval 达 45.8%,τ²-bench Telecom 通过率超 98%。兼容 Claude Code、KiloCode 等主流架构及 MCP/Skills 协议,支持云端与本地部署,已在 Kilo Code 等生态中完成接入验证。

智能体MCP/工具多模态开源生态

推荐理由:Step 3.7 Flash 用激活仅 11B 的 MoE 架构把 Agent 工作流稳定性做透了,兼容主流框架还开源,对需要低延迟、高可靠性的生产环境 Agent 是真正可用的选择。

5月22日5月22日周五

星期五 · 3 条
20:09
IT之家(RSS)同事件
AI 评分 75/100
阿里千问 App、PC 端及网页端接入全新一代大模型 Qwen3.7-Max

5月22日,阿里千问App官方宣布,千问App、PC端及网页端接入全新一代大模型Qwen3.7-Max。用户需将千问App更新至6.9.7及以上版本,即可免费体验该模型。Qwen3.7-Max定位为全能的智能体基座,核心能力覆盖编程开发、办公流程自动化及超长周期任务执行。官方实测显示,在一项长达35小时、包含超过1000次工具调用的全自主内核优化实验中,该模型保持了连贯推理。此外,模型具备跨框架泛化能力,并即将通过阿里云百炼平台提供API调用服务。

智能体MCP/工具模型发布
同一事件,精选展示《阿里通义千问Qwen3.7-Max上线OpenRouter》
16:35
MarkTechPost(RSS)
AI 评分 66/100
微软发布Fara1.5浏览器操作智能体系列:性能超越OpenAI Operator与Gemini 2.5

微软研究院近日推出Fara1.5系列浏览器操作智能体,包含4B、9B和27B三种参数规模。其中最大模型Fara1.5-27B在Online-Mind2Web基准测试中达到72%的准确率,显著优于OpenAI Operator、Gemini 2.5 Computer Use等主流模型。此次发布同步推出FaraGen1.5合成数据流水线,可在受控环境中高效训练智能体,为自动化浏览器操作提供了新解决方案。

智能体MCP/工具Microsoft模型发布
01:56
Rohan Paul@rohanpaul_ai同事件
AI 评分 84/100
阿里巴巴发布旗舰模型Qwen3.7-Max,专为Agent时代打造Alibaba just released Qwen3.7-Max.Their best flagship model built for real-world tasks and production environments.• Agent reliability the center of the story, where the model must plan steps, call tools, inspect results, fix mistakes, and continue without collapsing after the first wrong turn.• 56.6 on the Artificial Analysis Intelligence Index, up 4.8 points from Qwen3.6-Max. Qwen 3.7 Max sitting at 5th, pretty much on par with GPT 5.4 (xhigh)• The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding.• One important layer of the serving stack, the inference kernel, was optimized heavily. from near-baseline speed to 10.0x geometric mean speedup after many rounds of low-level GPU optimization.译阿里巴巴正式推出最新旗舰模型Qwen3.7-Max,定位为Agent时代的生产级基础模型。该模型在权威评测中得分56.6,较前代显著提升,性能与GPT-5.4相当。其核心优势在于卓越的Agent可靠性,能够在复杂任务中自主规划、调用工具、纠错并持续执行。通过底层深度优化,模型实现了10倍推理加速,并支持长达数小时的自主运行与多工具协作。该模型现已上线阿里云模型工作室,并兼容Claude Code、OpenClaw等主流开发框架,助力开发者构建实际应用。

Qwen: 📣Meet Qwen3.7-Max — our latest flagship, made for the Agent Era. A versatile foundation for agents that actually get th...

智能体MCP/工具推理模型发布
同一事件,精选展示《阿里通义千问Qwen3.7-Max上线OpenRouter》

5月21日5月21日周四

星期四 · 1 条
21:40
Qwen@Alibaba_Qwen精选
AI 评分 82/100
Qwen3.7-Max:面向Agent时代的旗舰模型📣Meet Qwen3.7-Max — our latest flagship, made for the Agent Era.A versatile foundation for agents that actually get things done: 🧑💻 Coding agent, end to end. Frontend prototypes, multi-file refactors, real debugging — nails it. 🗂️ A reliable office and productivity assistant. Get your work done through MCP integrations and multi-agent orchestration. ⏱️ Long-horizon autonomy. 35 hours straight on a kernel optimization task — 1,000+ tool calls, zero hand-holding. 🔌 Scaffold-agnostic. Claude Code, OpenClaw, Qwen Code, or your own stack. Consistent reliability everywhere.API's up on Alibaba Model Studio. You can also take it for a spin on Qwen Studio.Go build something wild!🏃🏃♂️📖 Blog: https://qwen.ai/blog?id=qwen3.7 ✅ Qwen Studio: https://chat.qwen.ai/?models=qwen3.7-max ⚡️ API:https://modelstudio.console.alibabacloud.com/ap-southeast-1?tab=doc#/doc/?type=model&url=2840914_2&modelId=qwen3.7-max&serviceSite=international译Qwen3.7-Max是Qwen系列面向Agent时代推出的最新旗舰模型,旨在为能完成实际任务的智能体提供强大基础。其核心能力包括:可作为端到端编码智能体,处理前端原型与多文件重构;作为可靠的办公助手,通过MCP集成与多智能体编排协同工作;并支持超长时间(超过35小时)的自主运行,执行复杂任务链。该模型兼容Claude Code、OpenClaw等主流开发框架,现已上线阿里云模型工作室与Qwen Studio提供服务。
智能体MCP/工具模型发布
另有 5 家信源报道公众号:通义实验室(千问)Hacker News 热门(buzzing.cc 中文翻译)X:通义千问 / Qwen (@Alibaba_Qwen)X:OpenRouter (@OpenRouter)X:X.PIN (@thexpin)
推荐理由:Qwen 3.7-Max 的亮点不在榜上分数,而是它瞄准 Agent 场景的连贯执行能力,35 小时不间断跑 kernel 优化,对需要长线任务的开发者是直接可用的探索方向。

5月15日5月15日周五

星期五 · 2 条
17:41
🚨 AI News | TestingCatalog@testingcatalog
AI 评分 66/100
谷歌Gemini Spark新增高级工具使用与技能创建流程GOOGLE 🔥: New Gemini Spark screenshots featuring advanced tool use and Skills creation flow.It seems like there won't be an option to import SKILL MD files besides copeing and pasting. There is also no evidence of Browser or Computer Use atm.译GOOGLE 🔥:Gemini Spark新截图展示高级工具使用和技能创建流程。 目前看来除了复制粘贴外,似乎没有导入SKILL MD文件的选项。目前也没有浏览器或计算机使用功能的迹象。

Just a dragon: The new Gemini Spark model will have Agent mode / Chat mode. New advanced use of tools.

智能体GoogleMCP/工具模型发布
07:34
Artificial Analysis@ArtificialAnlys
AI 评分 62/100
中国移动发布专有模型JT-35B-Flash,智能指数显著提升China Mobile has just released JT-35B-Flash, a proprietary 35B non-reasoning model with relatively high token efficiency and competitive intelligence for its size (Artificial Analysis Intelligence Index of 36)This represents a significant upgrade from China Mobile's previous JT-MINI, with an Intelligence Index improvement of +11 points (25 → 36). China Mobile is one of the world's largest telecommunications companies, and JT-35B-Flash is a sign of their continued focus on AI.Key results:➤ JT-35B-Flash scores 36 on the Intelligence Index, an +11 point improvement from JT-MINI (25). While still behind frontier models overall, the model shows China Mobile's progression in developing more capable proprietary models. The 35B parameter count represents a significant scale-up from JT-MINI.➤ JT-35B-Flash outperforms JT-MINI with significantly in AA-Omniscience, with a +42 improvement in score. This is driven by both lower hallucination rate (63%) as well as higher accuracy (28%).➤ JT-35B-Flash leads in τ²-Bench with 99%, ahead of GLM-4.7-Flash (Reasoning, 98%) and other top performers. τ²-Bench measures tool use in customer service scenarios, making this particularly relevant for China Mobile's telecommunications business. This represents the highest score measured on this evaluation across models we benchmark.➤ JT-35B-Flash achieves an Agentic Index score of 52, driven primarily by its exceptional τ²-Bench performance. GDPval-AA reaches 1076, indicating competent real-world task execution capabilities for a model at this Intelligence Index level.➤ JT-35B-Flash demonstrates high token efficiency, even compared to other non-reasoning models, using ~17M output tokens to run the Intelligence Index. This positions JT-35B-Flash as an efficient inference option compared to reasoning-enabled alternatives.Model details:➤ Context window: 256K tokens➤ Availability: Currently primarily available to China Mobile’s enterprise customers译中国移动近日发布了专有的350亿参数非推理模型JT-35B-Flash,其Artificial Analysis智能指数达到36,较前代JT-MINI大幅提升11分。该模型在面向电信客服场景的工具使用评测τ2-Bench中以99%的得分领先,并展现出较高的令牌效率,运行智能指数仅消耗约1700万输出令牌。JT-35B-Flash拥有256K上下文窗口,目前主要面向企业客户提供。作为全球主要电信运营商,此举标志着中国移动在开发更强大专有模型方面的持续投入。
MCP/工具模型发布

5月13日5月13日周三

星期三 · 1 条
04:56
Hacker News 热门(buzzing.cc 中文翻译)
AI 评分 65/100
Show HN: Needle:我们将"双子座工具召唤"浓缩为一个26M模型

研究团队发布了名为Needle的轻量级模型,它将谷歌Gemini的工具调用能力浓缩至仅2600万参数。该模型在保持核心功能的同时,体积显著缩小,旨在实现更高效的部署与应用。项目代码已在GitHub开源,并在Hacker News社区获得了超过100点的关注度。

智能体MCP/工具开源生态模型发布

4月29日4月29日周三

星期三 · 1 条
15:33
IT之家(RSS)
AI 评分 53/100
科大讯飞星火 X2-Flash 模型发布:基于华为昇腾 910B 集群训练,最大 256K 上下文

科大讯飞正式发布星火 X2-Flash 模型并开放API。该模型采用MoE架构,总参数300亿,支持256K上下文,基于华为昇腾910B集群训练。其在智能体、代码等能力上大幅提升,在深度研究报告、Skill管理等多项任务上效果接近业界万亿参数模型,而整体token消耗不到主流大尺寸模型的三分之一。通过结合DSA与MTP技术,模型在国产芯片上的训练效率从同规模A800集群的20%提升至90%,并解决了长交互场景采样效率低的问题,为大规模强化学习训练扫清障碍。AstronClaw、Loomy等已率先接入。

MCP/工具推理模型发布

4月21日4月21日周二

星期二 · 1 条
01:08
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 81/100
Kimi K2.6:推动开源编程发展

Kimi发布K2.6版本,专注提升开源编程能力。新版本在代码生成与理解方面实现优化,发布后在Hacker News获得132个点赞。此次更新强化了Kimi在开发者工具领域的布局,通过改进编程辅助功能,为开源社区提供更高效的代码编写支持。

MCP/工具模型发布编码
另有 4 家信源报道X:Testing Catalog (@testingcatalog)X:Artificial Analysis (@ArtificialAnlys)The Decoder:AI News(RSS)IT之家(RSS)
推荐理由:Kimi K2.6 把长程编程和 Agent Swarm 能力带到了开源模型里,那几个 12 小时自主优化、13 小时重构旧代码库的工程案例比基准表更有说服力,做 Agent 的可以认真看看。

7月15日7月15日周二

星期二 · 1 条
08:00
LG AI Research:官方研究博客
AI 评分 51/100
LG AI Research 发布混合推理模型 EXAONE 4.0,含 32B 与 1.2B 双版本

LG AI Research 发布混合推理模型 EXAONE 4.0,融合通用语言能力与 EXAONE Deep 的推理能力,支持 MCP 和 Function Calling,并新增西班牙语。

MCP/工具开源生态推理模型发布

7月1日7月1日周二

星期二 · 1 条
08:00
OpenRouter:Announcements(RSS)
AI 评分 47/100
新型隐形模型:Cypher Alpha

Cypher Alpha 是一款免费、通用、隐形模型,自带工具调用功能。

智能体MCP/工具模型发布

10月16日10月16日周三

星期三 · 1 条
00:00
Mistral AI:News(网页)
AI 评分 54/100
Mistral AI发布Ministral 3B和8B边缘模型

Mistral AI发布了两个新的边缘计算模型Ministral 3B和Ministral 8B。两者均支持高达128k的上下文长度。Ministral 8B采用了特殊的交错滑动窗口注意力模式,以实现更快、内存效率更高的推理。这些模型在知识、常识、推理、函数调用和效率方面为10B以下类别设定了新标杆,可用于设备端翻译、离线智能助手、本地分析和机器人等场景。在多项基准测试中,它们超越了同级别的Gemma 2 2B、Llama 3.2 3B等模型。Ministral 8B的API定价为$0.1 / M tokens,Ministral 3B为$0.04 / M tokens。

MCP/工具模型发布端侧
共 35 条,没有更多了