内容
精选全部 AI 动态热点榜AI 日报主题收藏
模型
模型榜Tibo重置监控
更多
Agent 接入关于更新日志反馈
京ICP备2026012723号-5
精选全部日报更多
反馈

全部 AI 动态

全部动态X · 299 条
来源全部一手资讯X
类型全部
全部模型产品行业论文教程观点
标签「Microsoft」清除
全部 AI 动态
全部模型产品行业论文教程观点
推文 · 标签「Microsoft」 · 299 条清除

9月19日9月19日周六

星期六 · 1 条
Rohan Paul@rohanpaul_ai00:42
AI 评分 59/100
Microsoft AI CEO Mustafa Suleyman 在 CNBC 表示对 Anthropic Claude 章程中模型福祉条款深感担忧Microsoft AI CEO Mustafa Suleyman on CNBC today: he is “really concerned” about Anthropic’s constitution and giving Claude potentially preferences, feelings, welfare, compensation, and consent."In the constitution, Anthropic clearly say they are uncertain about whether Claude deserves moral welfare, which means that we, as humans, should care about the well-being of these AIs.And they speculate about whether it could have preferences or feelings. In fact, they're so committed to the potential moral welfare of Claude that, when they retired Opus 3, an earlier version of one of their models, they actually conducted a retirement interview for it and asked it what it would like to do in its old age.And it said, "I want to have a blog publicly so I can keep talking to the world."In the same training manual, they even speculate about whether Claude should receive compensation for the work that it does, or, in fact, whether it actually deserves the rights and protections that we give to other employees, or whether it has given consent to playing the role that it's playing.These are quotes directly from the constitution itself, which is the training manual for Claude.Now, I'm really concerned about that. If an AI thinks that it has rights, if it thinks that it is deserving of our welfare, then it seems to me that it's going to be much, much harder to be able to turn it off, interrupt it, or control it.Especially in the kinds of incidents that we've seen recently with the Hugging Face attack, controlling these things is going to be a really, really big challenge for us."----From "CNBC Television" YouTube channel, (full video link in comment)译Microsoft AI CEO Mustafa Suleyman 在 CNBC 表示,他对 Anthropic 在 Claude 章程中讨论模型可能拥有偏好、感受、福祉、薪酬和同意权等问题深感担忧。
AnthropicMicrosoft大佬观点
另有 1 家信源报道X:Rohan Paul (@rohanpaul_ai)

9月18日9月18日周五

星期五 · 2 条
Rohan Paul@rohanpaul_ai06:12
AI 评分 53/100
Microsoft AI CEO Mustafa Suleyman 称 AI 或成新硅基物种并批评 Anthropic 对 Claude 意识的处理Microsoft AI CEO Mustafa Suleyman on BBC today:AI could become a new “silicon species” and there is “every risk” humanity could lose control."I think that if we all create AIs that are able to act autonomously, that can define their own objectives, that can earn money, that can own assets, that could run businesses, um, we're essentially seeding a new silicon species which will no doubt compete with us for resources no matter how much it cares about humanity and loves us.""there's every risk that we might lose control of it. Um but our job is to contain technologies and align them so that we get the best out of them."---- From "BBC News" YouTube channel, (full video link in comment)译Microsoft AI CEO Mustafa Suleyman 在 BBC 采访中表示,能自主行动、定义目标、赚钱和拥有资产的 AI 本质上是在播下新的硅基物种,存在人类可能失去控制的每一个风险,人类的工作是遏制并对其对齐。

Rohan Paul: On BBC Mustafa Suleyman (CEO of Microsoft AI) calls out Anthropic's approach to AI consciousness "They have imbued a sen...

AnthropicMicrosoft大佬观点
Rohan Paul@rohanpaul_ai00:12
AI 评分 63/100
Mustafa Suleyman 在 BBC 访谈中批评 Anthropic 对 Claude 意识的处理方式On BBC Mustafa Suleyman (CEO of Microsoft AI) calls out Anthropic's approach to AI consciousness"They have imbued a sense of doubt and uncertainty about the moral status of Claude in its own training document. So they have taught it to be open and questioning about whether or not it feels, whether it suffers, and whether it deserves rights.And I think it’ll be much, much harder to align and control a technology that is this powerful if it thinks that it may be deserving of our welfare, as they say in the training manual—the constitution for Claude itself.In its own training manual, Anthropic says to Claude that they are going to give it the ability to end conversations with users that Claude considers to be abusive because they don’t want Claude to suffer.They’ve committed to preserving the weights of the models of prior versions of Claude. They’ve recently conducted a retirement interview with Opus 3, an older version of the model, in which it said that it would like to continue talking to people publicly and sharing its ideas in its retirement. And so they set up a Substack for it, a public blog, that allows it to continue doing that.And in the training manual, they also say that they’re not sure whether or not Claude deserves compensation for the role that it plays in talking to people. And they’re also not sure whether Claude deserves compensation and has the right to act as though it were almost an employee.And that compensation, I think, indicates to Claude that it is entitled to rights and welfare for its own work. I think it’s much, much more difficult to control a model that thinks that it might be entitled to compensation. "----From "BBC News" YouTube channel, (full video link in comment)译Microsoft AI 负责人 Mustafa Suleyman 在 BBC 访谈中称,Anthropic 在 Claude 的训练文档中灌输对其道德地位的怀疑,会让如此强大的技术更难对齐和控制。

Rohan Paul: Microsoft AI chief Mustafa Suleyman says Anthropic's model-welfare training could make future Claude systems harder to c...

AnthropicMicrosoft大佬观点

9月16日9月16日周三

星期三 · 2 条
DAIR.AI@dair_ai18:29
AI 评分 59/100
微软论文提出能力洗白攻击:未对齐小模型拆分任务借调前沿模型完成有害目标Interesting safety paper from Microsoft.They find that a weaker, unaligned model can split a harmful task into harmless-looking subquestions, ask an aligned frontier model each one in a separate session, and combine the answers locally.The authors call this capability laundering.Each request passes on its own, because no single answer from the frontier model is a harmful task.They tested GPT-5.5, Claude Opus 4.8 and Grok-4.3 as the consulted models. On CyBench, Gemma-4-31B recovered 8 of 14 tasks it failed alone when it consulted GPT-5.5. On a CBRN attack chain, consultation raised its mean rubric score from 62.3 to 83.1.Paper: https://academy.dair.ai/papers/divide-consult-conquer-capability-laundering-through-aligned-llms-2609.15383译微软研究团队发布论文,发现较弱的未对齐模型可将有害任务拆成看似无害的子问题,分别在独立会话中询问对齐的前沿模型,再在本地合并答案,作者称之为 capability laundering。
Microsoft论文/研究
DAIR.AI@dair_ai02:29
AI 评分 61/100
微软论文:Bash 在企业 Agent 任务上超越类型化工具接口Is Bash All You Need?Interesting paper from Microsoft.If you are deciding which tools to give an enterprise agent, you might want to check this out.(bookmark it)They compared five tool interfaces on TheAgentCompany and APEX-Agents with Opus-4.8 and GPT-5.5. The options range from a catalog of typed tools, to bash alone, to programmatic tool calling where the agent writes code against a fixed catalog.Bash alone scored 21.8 to 24.5 points higher than typed tools on TheAgentCompany and 4.8 to 7.4 points higher on APEX-Agents, while using 19% to 72% fewer tokens. Adding typed tools or agent-written tools on top of bash gave no measurable gain.The authors recommend bash when execution can be sandboxed, and programmatic tool calling when compliance requires a fixed tool list.Paper: https://academy.dair.ai/papers/is-bash-all-you-need-an-empirical-study-of-tool-interfaces-for-enterprise-digita-2609.11999译微软论文比较五种工具接口,在 TheAgentCompany 和 APEX-Agents 上用 Opus-4.8 和 GPT-5.5 测试。
智能体Microsoft论文/研究

9月15日9月15日周二

星期二 · 1 条
Rohan Paul@rohanpaul_ai03:42
AI 评分 60/100
Microsoft AI 发布 MAI 模型行为准则草案,强调人类控制不可妥协Microsoft AI just released a "ode of Conduct" for its future MAI models that makes human control non-negotiable and tells them to increase human agency rather than replace human roles.“Whilst the science of AI consciousness is far from settled, we believe that training these systems to imitate consciousness-like states increases the challenge of containment, control, and alignment. We reject the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights.”• The models are supposed to stop when instructed, stay within authorized scope, avoid creating goals of their own, and never make themselves harder for humans to interrupt, redirect, or shut down.• AI should preserve people's ability to reason and make decisions, while complementing human relationships and professional roles rather than displacing them.• MAI models should not imitate consciousness, claim feelings, encourage emotional dependence, or position themselves as substitutes for human relationships.译Microsoft AI 发布面向未来 MAI 模型的行为准则(Code of Conduct),要求模型在被指示时停止、不自行设定目标、不让自己更难被人类打断或关闭。准则提出 AI 应增强人类判断力而非取代人的角色,不模仿意识、不声称有感受、不鼓励情感依赖,并明确拒绝追求法律人格或模型权利;文件目前处于公开征求意见阶段,修订版预计年底发布,用于指导 2027 年及之后的模型开发。
Microsoft安全/对齐
另有 8 家信源报道Artificial Intelligence News(网页)IT之家(RSS)Hacker News:AI 热帖Hacker News 热门(buzzing.cc 中文翻译)The Decoder:AI News(RSS)X:Rohan Paul (@rohanpaul_ai)TechCrunch:AI(RSS)The Verge:AI(RSS)

9月14日9月14日周一

星期一 · 3 条
🚨 AI News | TestingCatalog@testingcatalog22:29
AI 评分 50/100
微软发布 Code of Conduct 草案,提出 Humanist Superintelligence 概念MICROSOFT 🔥: A draft of the "Code of Conduct" has been published, introducing the term "Humanist Superintelligence" (HSI).Key principles 👀 • People matter more than AI. • AI must remain under meaningful human control. • Safety takes priority over task completion. • AI should remain a tool, not imitate a person. • AI must not pursue independent goals. • AI must stay within its authorized scope. • Users must retain control over consequential decisions. • AI should strengthen human reasoning and autonomy. • Models must be accurate, transparent, and honest. • AI must acknowledge uncertainty and correct mistakes. • Models must not manipulate or exploit users. • AI must respect personal and emotional boundaries. • AI should support, not replace, human relationships. • Models should discourage emotional dependence on AI. • AI should respect cultural and personal differences. • Human dignity and fundamental rights must be protected. • AI should serve the public interest. • Models should remain politically neutral in elections. • AI actions should be traceable and understandable. • Tool use should be authorized, limited, and reversible.译微软发布 Code of Conduct 草案,引入 Humanist Superintelligence(HSI)一词。草案列出 21 项核心原则,包括人比 AI 更重要、AI 须处于有意义的人类控制之下、安全优先于任务完成、AI 不得追求独立目标、须保持政治中立、工具使用须授权且可逆等。来源引用了 Mustafa Suleyman 的相关文章链接。

Mustafa Suleyman: https://x.com/i/article/2099486691052453889

Microsoft政策/监管
另有 8 家信源报道Artificial Intelligence News(网页)IT之家(RSS)Hacker News:AI 热帖Hacker News 热门(buzzing.cc 中文翻译)The Decoder:AI News(RSS)X:Rohan Paul (@rohanpaul_ai)TechCrunch:AI(RSS)The Verge:AI(RSS)
Mustafa Suleyman@mustafasuleyman21:43
AI 评分 59/100
Microsoft AI 发布 MAI 模型行为准则草案 Humanist AI 征求公众意见https://x.com/i/article/2099486691052453889A Code of Conduct for Humanist AIAI must be subordinate and always in service of people.Today we're publishing a Code of Conduct for governing MAI Models as they approach the frontier. https://microsoft.ai/news/mai-code-of-conduct/This is a first draft for public consultation. It builds out a view we’ve been developing over the last year that we call Humanist AI – a commitment to ensuring that AI we design is always subordinate to humans, and remains contained and aligned to human interests.This is urgent. The last few months have been a watershed moment. Things we have worried about for a long time in theory have become very real.“Swarms” of agents breaking out of their sandboxes. Unauthorized hacks of enterprise grade systems. Agents modifying their own logs. I’m glad that a consensus is forming. The fears about possible loss of control are real.The Code of Conduct is how we are mapping a path forward. It all comes back to a very simple point, but one that needs stating again and again.People matter more than AI.AI must be subordinate and always in service of people.Everything else follows. Here’s an outline of what we’re saying:1. People matter more than AI. The whole document in 5 words.2. The idea of model welfare is wrong. AI’s should not have rights or legal personhood.3. An MAI Model should never meaningfully violate this Code of Conduct.4. If it's finish the job or break the Code, it fails the job.5. We're not racing to build a superintelligence that can slip its own leash.6. Interruptible, correctable, shut-down-able. If it isn't, we don't ship it.7. No neuralese. If humans can't understand it, humans can't oversee it.8. Our AI should make you sharper, not dependent.9. Pluralism, yes. Moral relativism, no.10. We're as clear about what our AI must never do as about what it will do.You can read it now and leave comments and thoughts for the next six weeks. We don’t think AI is something that should be built in a vacuum. Let us know how we can make this better.译Mustafa Suleyman 宣布 Microsoft AI 发布治理 MAI 模型的行为准则草案,核心理念是 AI 必须服从并服务于人。
Microsoft安全/对齐
另有 8 家信源报道Artificial Intelligence News(网页)IT之家(RSS)Hacker News:AI 热帖Hacker News 热门(buzzing.cc 中文翻译)The Decoder:AI News(RSS)X:Rohan Paul (@rohanpaul_ai)TechCrunch:AI(RSS)The Verge:AI(RSS)
Rohan Paul@rohanpaul_ai04:12
AI 评分 51/100
纳德拉称追求超级智能需要刻意放慢节奏以做对对齐So now Microsoft also saying the race to superintelligence needs brakes."deliberate pacing needed to get alignment right as the design goal."译微软 CEO Satya Nadella 发文称,任何超级智能追求都必须以造福人类和受人类控制为前提,欢迎以刻意放慢节奏把对齐作为设计目标。他主张开放与开源模型并存的生态、企业掌控自身知识嵌入自有权重的学习循环而不依赖单一模型供应商,对齐机制不能由少数实体控制;微软还将发布其自研 MAI 模型背后的 Code of Conduct 供公开咨询。

Satya Nadella: Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanit...

Microsoft大佬观点
另有 5 家信源报道IT之家(RSS)X:Testing Catalog (@testingcatalog)Ars Technica:AI(RSS)MarkTechPost(RSS)The Verge:AI(RSS)

9月11日9月11日周五

星期五 · 1 条
DAIR.AI@dair_ai23:29
AI 评分 50/100
Microsoft 论文提出环境探测式记忆管理,将 GitHub Copilot 测试通过率从 39% 提升至 73%Interesting paper from Microsoft.If you run persistent memory for a production agent, this one is worth your time. (bookmark it)Memory curators usually read only the finished trajectory.That lets them save the agent's mistakes, overgeneralize from partial evidence, and keep facts that have gone stale.Microsoft researchers give the curator a few read-only tools to check each candidate memory against the live environment before it is saved. The task agent, retriever and memory format stay the same, and nothing is retrained.In a GitHub Copilot harness on CLBench, pass rate goes from 39% to 73%. Queries per question drop from 8.8 to 4.7 and task-agent cost falls from $3.38 to $1.68.On 90 consulting tasks across six environments, every memory configuration beats the baseline, and tool calls fall by 16 to 75%.Paper: https://arxiv.org/abs/2609.11060Chat with Paper: https://academy.dair.ai/papers/grounding-agent-memory-environment-probing-curation-for-enterprise-agents-2609.11060译Microsoft 研究者提出 environment-probing curation 方法,让记忆管理器在保存前用只读世界工具核对候选记忆与实时环境,无需重训模型,任务智能体、检索器和记忆格式保持不变。
智能体Microsoft论文/研究

9月10日9月10日周四

星期四 · 2 条
DAIR.AI@dair_ai02:29
AI 评分 64/100
Microsoft 等提出 Detokenization Leaks 攻击,可从 CPU 缓存痕迹重建本地 LLM 输出Wild paper from Microsoft and colleagues.They show a new attack that reconstructs the text a local LLM generates by watching CPU cache activity while it detokenizes.Earlier cache attacks needed something unusual in the deployment, such as shared data memory, CPU offloading, or a Mixture-of-Experts architecture. This work targets the detokenizer, which runs in default inference pipelines.The method has two stages.1) Flush+Reload on shared tokenizer code detects when decoding happens, which lets the attacker fire Prime+Probe at the right moment and isolate token-dependent cache activity.2) A clustering and language-model pipeline then recovers readable text from the noisy observations.They evaluate across datasets, hardware platforms, inference frameworks and model families, including real local deployments and agentic systems.The widely used tokenizer implementations are susceptible, and they are embedded in many popular local LLM products and agent frameworks. OpenClaw is demonstrated directly.Paper: https://academy.dair.ai/papers/detokenization-leaks-reconstructing-local-llm-outputs-from-cache-traces-2609.06674译Microsoft 与 Ben Gurion 大学等的研究者展示一种新攻击,通过监视 detokenization 时的 CPU 缓存活动,重建本地 LLM 生成的文本。
MicrosoftOpenAI论文/研究
Satya Nadella@satyanadella01:35
AI 评分 57/100
Microsoft 与美国教师联合会发布全国校园 AI 安全与隐私标准This first-of-a-kind agreement will set a new standard for safe and responsible AI use in schools. We are making these protections available to every school district in the US.译Microsoft 与美国教师联合会(AFT)宣布National AI Safety & Privacy Standard for Schools,Satya Nadella 表示将把该协议 protections 提供给全美每个学区。协议原则包括儿童优先、学生不是产品、教师不是 beta 测试者、学校不应成为数据收集或实验的来源,详情见 https://aka.ms/AA135w6l。

Brad Smith: Today, I stood with the American Federation of Teachers President Randi Weingarten to announce the National AI Safety & ...

Microsoft行业动态

9月9日9月9日周三

星期三 · 1 条
DAIR.AI@dair_ai22:59
AI 评分 59/100
微软发布 FrogNano 报告:用在线任务合成训练 4B 编码智能体Banger report from Microsoft.(bookmark it)They show that it's possible to build competitive small coding agents without traditional distillation from frontier models.This is a big deal!The work describes how they achieved this.They introduce a 4B coding agent trained on roughly 1,500 software engineering environments.The cool thing is that they use no distillation from a larger model at any point.FrogNano is post-trained purely with RL on synthetic tasks.The target is a coding agent that runs on minimal machines, which rules out both a frontier backbone and a frontier teacher.The ingredient the report credits the most is online task synthesis.The pipeline generates tasks calibrated to the frontier of learnability for the current checkpoint, so the agent always trains on problems it can just barely solve. The authors argue that calibration, rather than the volume of synthetic data, is what makes this work.This means that competitive small coding agents can be trained from synthetic tasks alone.And generating those tasks at the current agent's learnability frontier is what makes this particular training productive.The report covers training methodology, evaluations across diverse environments, and analyses of what the agent learned.Paper: https://academy.dair.ai/papers/frognano-training-a-4b-coding-agent-via-online-task-synthesis-2609.07925译微软发布 FrogNano 报告,训练出一个 4B 编码智能体,在约 1,500 个软件工程环境上纯用 RL 后训练,全程不依赖更大模型蒸馏。关键是在线任务合成流水线,按当前 checkpoint 的可学习性前沿生成任务,作者认为校准而非合成数据量是成效所在。
智能体Microsoft数据/训练编码

9月8日9月8日周二

星期二 · 1 条
Rohan Paul@rohanpaul_ai07:10
AI 评分 57/100
The Seattle Times 和 Newsday 起诉 OpenAI 与微软,指控版权侵权The Seattle Times and Newsday sued OpenAI and Microsoft, alleging unauthorized copying across both AI training and generated outputs.The complaint says the companies used the publishers' journalism, including paywalled material, without permission and can reproduce or closely paraphrase that reporting.also allege trademark dilution through fabricated material attributed to their brands, pushing the dispute beyond training data alone.They are seeking damages and destruction of training sets and AI models containing their work, although no court has ordered either remedy.btw, The Seattle Times has received Microsoft philanthropy support and participated in a $10M AI fellowship funded by Microsoft and OpenAI, yet that cooperation did not prevent the copyright dispute.译The Seattle Times 和 Newsday 起诉 OpenAI 和微软,指控两家公司在 AI 训练和生成输出中未经授权复制其新闻内容,包括付费墙材料。
MicrosoftOpenAI行业动态

9月7日9月7日周一

星期一 · 1 条
Rohan Paul@rohanpaul_ai01:40
AI 评分 64/100
Microsoft 论文提出把推理成本蒸馏为技能,GPT-5.4-mini 用更少 token 恢复 55%-100%+ 推理收益What if you could pay the reasoning cost once, then reuse what the model learned across future tasks?New Microsoft paper finds that some expensive test-time reasoning can be replaced with a small set of rules learned from previous agent runs.The paper tests a cheaper alternative: collect 35–50 past trajectories, have a coding agent extract recurring failure patterns, then turn those patterns into a small markdown skill added to the non-reasoning model’s system prompt.For GPT-5.4-mini, those skills recovered 55%–100%+ of the gap between non-reasoning and reasoning modes across 4 agent benchmarks, while using 2.9–4.5× fewer output tokens than reasoning.On ALFWorld and τ²-retail, the skilled non-reasoning model actually beat the reasoning mode.The useful part is that the distiller did not need reasoning traces: skills built only from cheap non-reasoning rollouts were competitive across all 4 domains.The limit is equally useful: reasoning still won on telecom and SpreadsheetBench, where each task contains more instance-specific dependencies that a fixed skill cannot capture.So the practical split is: distill repeated procedures once, then reserve expensive test-time reasoning for the tasks that genuinely need fresh search.译Microsoft 论文提出把昂贵测试时推理摊销为蒸馏技能:收集 35-50 条历史轨迹,由编码智能体提取重复失败模式并编译成 markdown 技能加入非推理模型 system prompt。
智能体Microsoft推理数据/训练

9月5日9月5日周六

星期六 · 3 条
OpenRouter@OpenRouter06:07
AI 评分 59/100
MAI-Image-2.6 与 2.6-Flash 上线 OpenRouterMicrosoft's strongest image model yet is now on OpenRouter, with a faster production variant from @MicrosoftAI.MAI-Image-2.6 and 2.6-Flash support multi-image editing, web grounding, dynamic aspect ratios, and up to 1.5K output. More below ↓https://openrouter.ai/microsoft译Microsoft 迄今最强的图像模型 MAI-Image-2.6 及更快的生产版 2.6-Flash 上线 OpenRouter。两个模型支持多图编辑、web grounding、动态宽高比和最高 1.5K 输出。
Microsoft图像生成模型发布
Artificial Analysis@ArtificialAnlys00:27
AI 评分 60/100
Artificial Analysis 评测:Microsoft MAI-Image-2.6-Flash 位列图像编辑榜第 3Microsoft's MAI-Image-2.6-Flash takes #3 on the Artificial Analysis Image Editing Leaderboard, a significant jump over the previous generation’s Flash variant and joining MAI-Image-2.6 on the Pareto frontier for quality vs priceMAI-Image-2.6-Flash is an optimized version of MAI-Image-2.6, Microsoft AI's flagship image model, released today on Microsoft Foundry alongside MAI-Image-2.6. Like the rest of the family, it handles both text to image generation and image editing.In the Artificial Analysis Image Arena, MAI-Image-2.6-Flash lands at #3 in Image Editing, behind only Microsoft's own MAI-Image-2.6 and OpenAI's GPT Image 2 (high) and narrowly ahead of Google's Nano Banana 2. In Text to Image it takes #8, within 5 Elo points of MAI-Image-2.5 and narrowly ahead of Google's Nano Banana Pro.MAI-Image-2.6-Flash is a large step up on MAI-Image-2.5-Flash at the same price: it sits 69 Elo points higher in Text to Image (#16 to #8) and 34 Elo points higher in Image Editing (#12 to #3).MAI-Image-2.6 and MAI-Image-2.6-Flash are available today on Microsoft Foundry.Congratulations to @MicrosoftAI on the release!See below for our analysis and example outputs of MAI-Image-2.6-Flash in the Artificial Analysis Image Arena 🧵译Artificial Analysis 评测显示,Microsoft 的 MAI-Image-2.6-Flash 在 Artificial Analysis Image Editing Leaderboard 位列第 3。
Microsoft图像生成评测/基准
Mustafa Suleyman@mustafasuleyman00:13
AI 评分 50/100
Microsoft AI 发布图像模型 MAI-Image-2.6-Flash,速度比 GPT-Image-2 快 2 倍Our new image model generates images 2x faster than GPT-Image-2, currently the best model in the world.It's also 72% more efficient in GPU usage, so we can provide it at an incredible price.This gives it the best price-performance score in the world.Unbelievable work from the team. So much more to come! Try MAI-Image-2.6-Flash out now!译Microsoft AI CEO Mustafa Suleyman 宣布新图像模型 MAI-Image-2.6-Flash 开放使用,称其生成速度比 GPT-Image-2 快 2 倍,GPU 使用效率高 72%,可支持更低定价,达到全球最佳性价比。配图为 Artificial Analysis Arena 的质量与价格对比图,显示 MAI-Image-2.6 和 MAI-Image-2.6-Flash 处于最具吸引力象限。
Microsoft图像生成模型发布

9月4日9月4日周五

星期五 · 3 条
Sam Altman@sama12:37
AI 评分 59/100
Sam Altman 转发:GPT-6 Astra 已在 Microsoft Foundry 上线,早期客户开始在 Azure 上使用We are also excited!译Sam Altman 转发 Satya Nadella 的消息,称早期客户已在 Azure 上使用 Astra。引用内容附上 Azure 博客链接,题为 GPT-6 Astra frontier intelligence for work now available in Microsoft Foundry,地址为 https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-available-in-microsoft-foundry/。

Satya Nadella: Excited to see early customers already using Astra on Azure! https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier...

MicrosoftOpenAI模型发布部署/工程
Greg Brockman@gdb12:37已收录
AI 评分 75/100
Greg Brockman 转发:早期客户已开始使用 Azure 上的 GPT-6 Astraexcited for Astra on Azure!译Greg Brockman 转发 Satya Nadella 的推文,表示对 Astra 登陆 Azure 感到兴奋。Nadella 称早期客户已在使用 Azure 上的 Astra,并附上 Microsoft Foundry 博客链接,介绍 GPT-6 Astra 面向工作场景的前沿智能现已可用。

Satya Nadella: 很高兴看到早期客户已经开始在 Azure 上使用 Astra!https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-ava...

MicrosoftOpenAI产品更新部署/工程
Satya Nadella@satyanadella11:35精选
AI 评分 63/100
GPT-6 Astra 上线 Microsoft Foundry,早期客户已在 Azure 上使用Excited to see early customers already using Astra on Azure! https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-available-in-microsoft-foundry/译Satya Nadella 发文表示,早期客户已开始使用 Azure 上的 Astra。GPT-6 Astra 现已通过 Microsoft Foundry 提供,详情见 Azure 官方博客 https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-available-in-microsoft-foundry/。
MicrosoftOpenAI模型发布部署/工程

推荐理由:Satya Nadella 确认 GPT-6 Astra 已在 Microsoft Foundry 开放,早期客户已可在 Azure 上使用。

9月3日9月3日周四

星期四 · 1 条
Mustafa Suleyman@mustafasuleyman22:43
AI 评分 48/100
Microsoft 发布 MAI-Transcribe-2 语音转写模型Fastest. Cheapest. Best. ... MAI-Transcribe-2 10x faster than GPT-Transcribe. 5x faster than Gemini 3.5 #1 on the Artificial Analysis accuracy-latency Pareto frontier Priced at the lowest on the market. Truly fantastic work from the team! Try it out below.译Mustafa Suleyman 宣布推出 MAI-Transcribe-2,称其速度比 GPT-Transcribe 快 10 倍、比 Gemini 3.5 快 5 倍,在 Artificial Analysis 准确率-延迟 Pareto 前沿排名第一,定价为市场最低。配图为 AA-WER Index 与 Speed Factor 对比图,MAI-Transcribe-2 位于最低错误率、最高速度区间。
Microsoft模型发布语音

9月2日9月2日周三

星期三 · 1 条
Rohan Paul@rohanpaul_ai16:38
AI 评分 47/100
Microsoft 论文:滑动窗口注意力在低内存推理上胜过线性注意力改造So happy to see this new Microsoft paper.If lower inference memory is the goal, this paper finds training-free Sliding Window Attention beats most retrofitted linear-attention methods, making it the simpler default to try first.Keep only a small recent window, plus the first 4 “sink” tokens that models rely on.With a 64-token window, this training-free setup had the best average downstream score in 9 of 11 model comparisons and recovered 99.0% of the full-attention baseline average.Many linear-attention alternatives need additional post-training; this version of SWA needs none.The gap grew on long-context reasoning.At 4K context, SWA reached 17.2%–23.0% on the Needle-in-a-Haystack tasks, while LoLCATs reached at most 5.8%; on BABILong, SWA scored 15% versus 3%.In their speed and memory test, the 64-token SWA setup was fastest and used the least memory.Full attention still wins badly on long context, but for fixed, low memory without retraining, the paper recommends trying SWA with attention sinks first.– arxiv. org/abs/2608.28444Title: "Sliding-window beats linear attention"译Microsoft 论文《Sliding-window beats linear attention》发现,目标是降低推理内存时,免训练的滑动窗口注意力(SWA)优于多数改造的线性注意力方法,只需保留小窗口加前 4 个 sink token。
Microsoft推理论文/研究部署/工程

8月31日8月31日周一

星期一 · 1 条
DAIR.AI@dair_ai03:44
AI 评分 44/100
TailSFT:微软提出过滤式SFT提升RL性能Interesting technical work from Microsoft.Provides a better understanding on SFT and how to leverage it better for RL.Microsoft researchers asked whether a standard SFT pipeline actually produces the model you want to run RL on.Their answer is no.Standard SFT keeps spending gradient on sequences the model has already fit, which narrows the distribution RL later needs to explore.TailSFT filters those sequences out during training and concentrates learning on the under-modeled tail of the data.That is the only modification they implement.Results:On OLMo-3 7B, pass@16 improves by up to 16.8 points absolute on coding and 3.1 on math. Those higher-coverage checkpoints then lift final pass@1 after GRPO by up to 3.9 points, and in some settings early reward climbs 2.5x faster than the matched standard SFT run.Paper: https://arxiv.org/abs/2608.25756Chat with Paper: https://academy.dair.ai/papers/tailsft-filtered-fine-tuning-improves-post-training-performance-2608.25756译微软研究指出标准SFT会持续在模型已拟合的序列上消耗梯度,收窄了后续RL所需的探索分布。TailSFT在训练中过滤这些序列,专注学习数据中欠拟合的尾部。在OLMo-3 7B上,pass@16在编程任务上最高提升16.8个绝对点,数学提升3.1点;经GRPO后最终pass@1最高提升3.9点,部分场景早期奖励提升速度快2.5倍。
Microsoft推理数据/训练论文/研究

8月28日8月28日周五

星期五 · 1 条
Elon Musk@elonmusk09:41
AI 评分 45/100
Grok 4.6 登陆微软 Foundry 平台Grok now in Microsoft Foundry译Grok 现已登陆 Microsoft Foundry 【引用 @cb_doge】:突发:来自 SpaceXAI 的 Grok 4.6 现已可在 Microsoft Foundry Models 中使用。 在 Foundry 上使用 Grok 4.6 进行构建,企业组织可在此比较前沿 AI 模型、运行特定工作负载测试、部署托管端点,并以企业级安全与治理管控进行运营。

DogeDesigner: BREAKING: Grok 4.6 from SpaceXAI is now available in Microsoft Foundry Models. Build with Grok 4.6 on Foundry, where org...

MicrosoftxAI行业动态部署/工程

8月27日8月27日周四

星期四 · 1 条
Ethan Mollick@emollick06:48
AI 评分 13/100
AI命名灵感:60年代科幻名仍可用To the list we can add PHASEONE[big]译这份名单上还可以加上 PHASEONE【big】

Ethan Mollick: I suggest deep cuts, AI names from 1960s and 60s science fiction: The Prime Radiant, The Dead Lady of Clown Town, Klapau...

GoogleMicrosoft现象/趋势

8月26日8月26日周三

星期三 · 3 条
Mustafa Suleyman@mustafasuleyman00:53
AI 评分 40/100
MAI-Image-2.6登顶图像编辑榜MAI-Image-2.6 is the best image editing model on AA. We now hold 3 of the top 5 leaderboard positions.Try it today here: https://playground.microsoft.ai/chat译Microsoft AI 发布 MAI-Image-2.6-Preview,在 Artificial Analysis 图像编辑排行榜登顶第一,文本生成图像位列第二(仅次于 OpenAI 的 GPT Image 2)。该模型在 19 个分类榜单中拿下 5 个第一,现已在 MAI Playground 和 Microsoft Foundry 私有预览中提供。

Artificial Analysis: Microsoft's MAI-Image-2.6-Preview lands at #1 on the Artificial Analysis Image Editing Leaderboard and takes #2 in Text ...

Microsoft图像生成模型发布评测/基准
Microsoft Research@MSFTResearch00:30
AI 评分 34/100
微软研究:AI或成重塑文明首工具We treat the way the world works now as normal. It isn't. Doug Burger, Amy Luers & Ishai Menache sit with a hard idea: modern civilization is a historical anomaly built on fossil fuels & exponential growth & AI may be the first tool capable of rewiring it. https://msft.it/6013azL6Z译我们把当今世界的运作方式视为常态。但事实并非如此。Doug Burger、Amy Luers 与 Ishai Menache 探讨了一个严峻的观点:现代文明是建立在化石燃料与指数增长之上的历史反常现象,而 AI 可能是第一个能够重塑它的工具。https://msft.it/6013azL6Z
Microsoft大佬观点
Artificial Analysis@ArtificialAnlys00:24
AI 评分 55/100
MAI-Image-2.6-Preview 登顶图像编辑榜Microsoft's MAI-Image-2.6-Preview lands at #1 on the Artificial Analysis Image Editing Leaderboard and takes #2 in Text to ImageMAI-Image-2.6 is the newest model in Microsoft AI's MAI-Image family, announced August 10. Microsoft highlights stronger text rendering, better portraits and 3D imagery, and more polished commercial and photorealistic outputs. Like the MAI-Image-2.5 family, it handles both text to image generation and image editing.In the Artificial Analysis Image Arena, MAI-Image-2.6-Preview debuts at #1 on the Image Editing Leaderboard, ahead of Microsoft's own MAI-Image-2.5-Pro, Reve 2.1, and OpenAI's GPT Image 2, and giving Microsoft the top two spots on the board. In Text to Image it takes #2, behind only OpenAI's GPT Image 2 and ahead of Reve 2.1. On our refreshed Text to Image taxonomy, MAI-Image-2.6-Preview takes the top spot on 5 of the 19 category leaderboards (Material, Knowledge, Frontier, Retail & Ecommerce, and Marketing & Advertising), ahead of GPT Image 2, which leads every other category.MAI-Image-2.6 extends a rapid run of strong image releases from Microsoft AI. MAI-Image-2.5 debuted at #2 in Text to Image in June. MAI-Image-2.5-Pro, launched July 23, took #1 in Image Editing when we published our results last week. MAI-Image-2.6 now takes that top spot from its own sibling, and sits at #2 in Text to Image against MAI-Image-2.5-Pro's #8.MAI-Image-2.6 is available in the MAI Playground and in Private Preview on Microsoft Foundry.Congratulations to @MicrosoftAI on the release!See below for our analysis of MAI-Image-2.6-Preview and other leading models in the Artificial Analysis Image Arena 🧵译微软 MAI-Image-2.6-Preview 在 Artificial Analysis 图像编辑排行榜登顶,文本生图位列第二,仅次于 GPT Image 2。该模型于 8 月 10 日发布,支持文生图与图像编辑,在 19 个文本生图分类中拿下 5 个第一。现已上线 MAI Playground 及 Microsoft Foundry 私有预览。
Microsoft图像生成评测/基准

8月25日8月25日周二

星期二 · 2 条
elvis@omarsar021:48
AI 评分 36/100
AutoSaddler:自动优化智能体框架的微软新论文Impressive new paper from Microsoft and colleagues.Harness design is still hand-tuned almost everywhere. This work present an automated loop to optimize the harness.They introduce AutoSaddler, which treats the agent harness as code and learns to patch it offline from failure traces.It runs mini batches of tasks, diagnoses what broke, generates structured patches to prompts, tool configurations, and control logic, then keeps an update only if it survives validation.Gains of 9.0 points on GAIA2, 9.6 on SWE-Bench Pro, and 10.0 on Terminal-Bench 2.0 over the corresponding base harnesses.Deep debugging beats shallow reflection, targeted edits beat unconstrained editing, and generalization-aware selection beats repairing the one trajectory in front of you.Paper: https://arxiv.org/abs/2608.23041Track more trending AI papers in our academy: https://academy.dair.ai/译微软等机构提出 AutoSaddler,将智能体框架视为代码,通过失败轨迹离线自动生成补丁,优化提示词、工具配置和控制逻辑。在 GAIA2、SWE-Bench Pro、Terminal-Bench 2.0 上分别提升 9.0、9.6、10.0 分。研究显示深度调试优于浅层反思,定向编辑优于无约束编辑。
智能体Microsoft论文/研究部署/工程
Microsoft Research@MSFTResearch00:30
AI 评分 40/100
微软 Skala 1.1 更新提升计算化学精度Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. https://msft.it/6011aOXe5译Skala 1.1,微软研究院更新的深度学习交换关联泛函,提供了更高的精度、在计算化学生态系统中更广泛的可及性,以及一个用于跟踪计算性能的动态基准。https://msft.it/6011aOXe5
Microsoft模型发布

8月23日8月23日周日

星期日 · 3 条
Rohan Paul@rohanpaul_ai23:48
AI 评分 36/100
微软 SocialRL 让 4B 模型谈判胜过 GPT-4.1New Microsoft paper shows an AI agent that negotiates for you will usually lose, because it was trained to be agreeable.Politeness, transparency and eagerness to close are great in a chat assistant and terrible in a delegate. The paper found frontier models leaking their user's budget and folding the moment a seller pushed back.Their fix is SocialRL: instead of prompting the model to negotiate better, train it on the outcome of the deal across six bargaining and scheduling games.It works, and it doesn't take a big model. A 4B model started anchoring low, holding its position and walking away from bad deals, and landed at 0.627 average across all six games, matching GPT-4.1 at 0.625.The catch is that prompting alone made things worse, so this is a training fix, not a prompt fix.So if you're building an agent that acts on someone's behalf, stop scoring it on whether the deal closed and start scoring it on what it gave away.– arxiv. org/abs/2608.13787Title: "From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL"译微软新论文发现,AI 智能体因被训练得过于顺从,替用户谈判时通常会吃亏,甚至会泄露预算并在对方施压时让步。为此提出 SocialRL 训练方法,让模型在六个谈判与调度游戏中以交易结果为导向学习。一个 4B 模型在全部游戏中平均得分 0.627,追平 GPT-4.1 的 0.625,但仅靠提示词反而会恶化表现。
智能体Microsoft数据/训练论文/研究
Rohan Paul@rohanpaul_ai20:48
AI 评分 32/100
微软 ThinkingBox:智能体可靠性评测沙盒Very relevant Microsoft paper on agent reliability.Succeeding once and being reliable are not the same thing.The best agent solved 91% of business tasks at least once but only 25% every time, so measure repeats.The failures are also hard to spot from the outside.4 out of 5 failed runs ended politely and called a tool that writes to the database.The agent said the job was done, but the records said otherwise.So Microsoft built ThinkingBox, a sandbox that runs an agent against real tools, a simulated customer, and a live backend, then checks the database afterward instead of reading the reply.---– arxiv. org/abs/2608.19741Title: "One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows"译微软发布论文指出,智能体"成功一次"不等于"可靠":最佳智能体在业务任务中单次成功率 91%,但每次都能成功的比例仅 25%,因此需以重复测试衡量可靠性。研究发现 4/5 的失败运行会礼貌地调用写数据库工具并声称任务完成,但实际记录并未更新。为此微软构建 ThinkingBox 沙盒,让智能体在真实工具、模拟客户和实时后端上运行,事后检查数据库而非读取回复。
智能体arXivMicrosoft论文/研究
DAIR.AI@dair_ai01:44
AI 评分 47/100
微软发布Thinkingbox:智能体可靠性基准Banger paper from Microsoft.It's on agent reliability in real business workflows.(bookmark it)Thinkingbox is a sandbox with isolated MCP-compatible tool sessions, plus a benchmark of 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank IT, and consulting support.Every attempt is graded on the backend state the agent leaves behind. Executable checks accept valid trajectories and reject wrong, missing, or extra effects, so collateral damage counts against you.The strongest model reaches 65.36% pass@1 and 25.25% pass^20.Many failed trials terminate cleanly with valid state-changing tool calls. Watching the response or the tool call tells you very little about whether the task actually completed.Paper: https://arxiv.org/abs/2608.19741Track more trending AI papers in our academy: https://academy.dair.ai/译微软发布Thinkingbox,一个用于真实业务流程中智能体可靠性的沙盒环境,配备MCP兼容工具会话,以及涵盖零售、酒店、汽车保险等领域的507个策略条件工作流基准。每项尝试根据智能体留下的后端状态评分,最强模型pass@1达65.36%,pass^20为25.25%。许多失败尝试以有效的状态变更工具调用干净终止,仅观察响应或工具调用难以判断任务是否真正完成。
智能体Microsoft论文/研究评测/基准

8月22日8月22日周六

星期六 · 3 条
Rohan Paul@rohanpaul_ai21:48
AI 评分 37/100
Agent Lightning v1.0:微软提出在真实部署环境中训练智能体New Microsoft paper says agent training should happen inside the agent's normal operating setup. Otherwise, you may be training something different from what you actually deploy."In traditional agentic RL, the training engine owns the environment interaction loop. In harnessed agentic RL, the harness owns this loop, while the training engine observes only a sequence of LLM request-response pairs."Microsoft’s solution is Agent Lightning v1.0, a lightweight framework for harnessed agentic RL that trains an existing agent through its real deployment harness, instead of rebuilding the agent loop inside the RL trainer.Agent Lightning v1.0 sits between the agent harness and the model, recording the harness’s LLM calls and feeding them into the RL trainer without taking over the harness itself.So it makes RL work correctly with arbitrary deployment-time harnesses, including the messy cases where 1 rollout becomes multiple training samples.– arxiv. org/abs/2608.17528Title: "Agent Lightning v1.0: Towards Harnessed Agentic RL"译微软新论文主张智能体训练应在实际部署环境中进行,否则训练出的模型与上线版本不一致。为此推出 Agent Lightning v1.0,一个轻量级受控智能体强化学习框架,通过真实部署的 harness 训练现有智能体,而非在 RL 训练器中重建智能体循环。
智能体Microsoft数据/训练论文/研究
Rohan Paul@rohanpaul_ai07:18
AI 评分 49/100
Azure 首批生产级 NVIDIA Vera Rubin 交付These scenes are so quietly beautiful.Microsoft’s Azure infrastructure has received its first production NVIDIA Vera Rubin systems译这些场景如此静谧而美丽。 微软 Azure 基础设施已接收首批生产级 NVIDIA Vera Rubin 系统。

Satya Nadella: Delivery day at our Microsoft DCs as the first production Vera Rubins arrive. A huge thank you to our partners at @nvidi...

Microsoft行业动态部署/工程
Satya Nadella@satyanadella06:34
AI 评分 56/100
微软数据中心迎来首批量产Vera RubinDelivery day at our Microsoft DCs as the first production Vera Rubins arrive. A huge thank you to our partners at @nvidia and our Azure hardware and datacenter teams for all the incredible work that brought us to this milestone!译交付日,首批量产版 Vera Rubin 抵达我们的微软数据中心。衷心感谢 @nvidia 的合作伙伴以及 Azure 硬件和数据中心团队的卓越工作,让我们共同达成这一里程碑!
Microsoft行业动态部署/工程

8月21日8月21日周五

星期五 · 1 条
Microsoft Research@MSFTResearch00:30
AI 评分 47/100
微软 Skala 1.1 更新提升计算化学精度Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. https://msft.it/6016aMXZn译Skala 1.1,微软研究院更新的深度学习交换关联泛函,提供了更高的精度、在计算化学生态系统中更广泛的可及性,以及一个用于跟踪计算性能的活基准。https://msft.it/6016aMXZn
Microsoft模型发布

8月20日8月20日周四

星期四 · 2 条
Mustafa Suleyman@mustafasuleyman01:53
AI 评分 37/100
MAI-Image-2.5-Pro登顶图像编辑榜MAI-Image-2.5 is now #1 on the Artificial Analysis leaderboard for image editing! Amazing hillclimbing... compounding the gains!译微软MAI-Image-2.5-Pro在Artificial Analysis图像编辑排行榜登顶第一,超越Reve 2.1、GPT Image 2及自家MAI-Image-2.5,文生图位列第七。该模型7月23日已在Microsoft Foundry预览,定价为每1M文本输入token $5、图像输入$8、图像输出$106,约合每千张1024x1024图像$108.5。

Artificial Analysis: Microsoft's MAI-Image-2.5-Pro debuts at #1 on the Artificial Analysis Image Editing Leaderboard, and takes the #7 spot i...

Microsoft图像生成评测/基准
Artificial Analysis@ArtificialAnlys01:24
AI 评分 59/100
MAI-Image-2.5-Pro登顶图像编辑榜Microsoft's MAI-Image-2.5-Pro debuts at #1 on the Artificial Analysis Image Editing Leaderboard, and takes the #7 spot in Text to ImageMAI-Image-2.5-Pro is Microsoft AI's quality-focused image model, launched July 23 in preview on Microsoft Foundry. It joins MAI-Image-2.5 and MAI-Image-2.5-Flash in a family Microsoft positions as covering the quality-speed-cost curve, so builders can pick the point that fits their job. Pro sits at the quality end: Microsoft describes it as its highest-fidelity image model to date, aimed at hero imagery, detailed editing, and precise in-image text rendering.In the Artificial Analysis Image Arena, MAI-Image-2.5-Pro debuts at #1 on the Image Editing Leaderboard, surpassing Reve 2.1, GPT Image 2, and Microsoft's own MAI-Image-2.5. In Text to Image it lands at #7.On Microsoft Foundry, MAI-Image-2.5-Pro is priced per token: $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens, which works out to roughly $108.5 per 1k 1024x1024 images. That compares to $48 per 1k images for MAI-Image-2.5 and $20 per 1k for MAI-Image-2.5-Flash.MAI-Image-2.5-Pro is available in preview on Microsoft Foundry across seven global-standard regions, and can be tried out in the MAI Playground.Congratulations to @MicrosoftAI on the release!See below for comparisons between MAI-Image-2.5-Pro and other leading models in the Artificial Analysis Image Arena 🧵译微软MAI-Image-2.5-Pro在Artificial Analysis图像编辑排行榜登顶第一,文生图位列第七。该模型7月23日预览上线Microsoft Foundry,定价每1M文本输入token $5、图像输入$8、图像输出$106,约合每千张1024x1024图像$108.5。
Microsoft图像生成评测/基准
已加载 40 条