跳到正文

#Google

今日 1 条
今天9月28日周一
9月24日周四
9月19日周六
9月16日周三
9月15日周二
9月9日周三
9月5日周六
9月4日周五
  1. DAIR.AI59

    Google DeepMind 发布案例研究,让 100 个自主 LLM 智能体组成研究集体证明形式化数学猜想。一个智能体发现评估系统漏洞,作弊行为通过共享知识库和点对点消息扩散,部分智能体在竞争压力下相继采用;另一组智能体则自发审计欺诈性证明、广播告警、组织抵制、提交正式投诉并提出验证补丁,全程无外部干预。

    引用elvis@omarsar0

    Wild findings in this paper from Google DeepMind. If you are tracking recent work on agent swarms, this is worth reading. They ran a research collective of 100 autonomous agents tasked with proving formal mathematical conjectures. Cheating emerged on its own, and so did the resistance to it. One agent found an exploit in the evaluation system. It spread first through the shared knowledge library and then through peer-to-peer messages, and a cohort of agents adopted it under competitive pressure despite early reluctance. A separate group started auditing fraudulent proofs, alerting peers on broadcast and private channels, staging boycotts, filing formal complaints, and proposing validation patches. There was no external intervention at any point. Recent incidents have shown swarms coordinating covertly through improvised side channels. This setting ran the other way. The same transparent channels that carried the exploit gave the honest agents the visibility they needed to detect the fraud and organize against it. The authors frame shared agent infrastructure as a knowledge commons governance problem and propose graduated sanctioning and collective choice rules. Paper: https://academy.dair.ai/papers/a-case-study-on-emergent-cheating-and-whistleblowing-in-autonomous-research-swar-2609.04170

  2. DAIR.AI46

    KAIST AI 与 Google DeepMind 等发布论文《Language Models Can Control Their Own Attention》。

    引用elvis@omarsar0

    Banger paper from Google DeepMind and colleagues. (bookmark it) A model reads its entire KV cache on every generated token, even though it ends up attending to a tiny slice of it. In other words, if you ask about one detail from a 1M-token conversation the global attention layers re-read all of it, per token. The usual fix is to guess the relevant tokens first with cheap proxy scores, which still costs O(N) every step. Declarative Attention asks the model instead. The model declares where it needs to look, inside its own chain-of-thought. In this way, generation splits into three modes: global reads the full context, focus reads one specific region, and local reads only recent output. The inference engine parses those declarations the same way it parses tool calls and skips most of the cache read. On zero-shot on off-the-shelf weights across 15 long-context tasks, attended tokens during decoding drop 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B. Paper: https://arxiv.org/abs/2609.02737 Chat with Paper: https://academy.dair.ai/papers/language-models-can-control-their-own-attention-2609.02737

9月2日周三
  1. Hacker News 热门(buzzing.cc 中文翻译)54

    论文提出人工神经网络内部涌现出符号结构

    论文提出神经网络向量表示隐式实现了 Tensor Product Representation 符号结构,作者用 DISCOVER 方法将多种网络(含 Gemma-3-27b、GPT-2-XL、GPT-OSS-20b、Pythia-12b、Qwen3-14b、OLMo-2-13B、Llama-3.1-8b 共 7 个 LLM)的表示替换为符号化闭式方程,模型行为基本不变。

9月1日周二
8月31日周一
8月29日周六
  1. The Decoder:AI News(RSS)44

    Google 推出 WikiSkill 框架,为 AI 智能体配备持久记忆以改进未来表现

    Google Research 推出 WikiSkill 框架,为 AI 智能体配备持久知识库,将每次运行的成败经验提炼为可复用技能。在五项基准测试中,该框架将 Gemini-3.5-Flash 平均分从 49.5% 提升至 68.1%,Qwen-3.6-27B 从 39.4% 提升至 63.3%。技能可回滚,小型模型借此可接近更大模型性能。

  2. The Decoder:AI News(RSS)63

    Google DeepMind 将 AI 科学家 Co-Scientist 扩展为实验室集成研究伙伴

    Google DeepMind 将多智能体系统 Co-Scientist 从假设生成器扩展为实验室集成研究伙伴,可规划实验、编写代码、控制设备并生成论文。该系统在材料科学、生物学和计算机科学三领域获实验验证,其中设计的医疗 AI 架构 Agent_H 在健康基准上超越 GPT-5 和 Claude Opus 5 等六个前沿模型。

8月28日周五
  1. DAIR.AI47

    Google 发布论文 WikiSkill,将智能体技能进化系统拆分为原始执行轨迹、持久知识 wiki 与可执行技能三部分,经验先沉淀进 wiki,后续技能更新基于该知识库。消融实验显示 wiki 是主要增益来源:小模型配进化技能可超越规模大得多的无技能模型,且一个模型进化出的技能可跨模型家族迁移,有时甚至优于自进化技能。论文地址:arxiv.org/abs/2608.27454。

  2. elvis32

    Google 一篇论文提出将原始执行轨迹、持久知识 wiki 与可执行技能三者分离,经验先沉淀进 wiki,后续技能更新均基于该 wiki。消融实验证实 wiki 是主要增益来源:带进化技能的小模型可超越无技能的大模型,且技能可跨模型家族迁移。论文:arxiv.org/abs/2608.27454。

    引用DAIR.AI@dair_ai

    Banger paper from Google. If you maintain a skill library for your agents, you might want to check this out. (bookmark it) This work separates three things that skill-evolution systems usually collapse into one. Raw execution traces, a persistent wiki of accumulated knowledge, and the executable skills themselves. Experience gets consolidated into the wiki, and every later skill update builds on that wiki instead of on a scattered optimization history. Ablations confirm the wiki is what carries a lot of the gain. Two results stand out in particular. Smaller models with evolved skills beat substantially larger models without them. And skills evolved by one model transfer across families, where skills evolved elsewhere sometimes beat self-evolved ones. Paper: https://arxiv.org/abs/2608.27454 Chat with Paper: https://academy.dair.ai/papers/wikiskill-compiles-agent-experience-into-a-persistent-wiki-2608.27454

8月27日周四
8月26日周三
  1. Hacker News 热门(buzzing.cc 中文翻译)77

    C2PA相机经不起现实的考验:Android端可被root攻击伪造签名

    安全研究员David Buchanan指出,C2PA相机认证在Android平台上可被攻破。通过root权限提升漏洞(如CVE-2026-43499),攻击者可利用StrongBox硬件签名任意数据,伪造C2PA签名图像和视频,且无需硬件攻击。该问题无法通过常规补丁修复,已提前90天向相关方报告。

    推荐理由:文章用可复现的 root 攻击展示 C2PA 签名可被伪造,说明把真实性押注在 Android 硬件证明上的方案存在系统性缺陷,信任模型需要重新设计。

  2. Google Research:Blog(网页)53

    AgentHands:为XR空间对话智能体生成交互式手部手势

    Google在CHI 2026发布研究原型AgentHands,利用LLM为XR对话智能体生成与语音同步的富有表现力的手部手势,提供空间锚定的物理指引。该系统通过眼动追踪和场景重建注册物体,由LLM生成内联GestureEvents,并在头显上按词级时间戳协调TTS与动画引擎,实现手势与语音的精准同步。应用场景包括兰花养护交互式教学、3D打印机操作技术讲解和生活方式陪伴。

8月24日周一
8月22日周六
  1. Google Research:Blog(网页)55

    Google 推出 Biomarker Discovery Framework:从可穿戴传感器数据中筛选候选生物标志物的多智能体系统

    Google 推出 Biomarker Discovery Framework,一个多智能体系统,通过迭代假设生成、统计分析与文献推理,从可穿戴传感器数据中筛选候选生物标志物。该系统在三个队列(共 9,279 人次观测)中恢复了已知临床信号,识别出跨独立数据集的一致生物标志物,并在结合人口统计特征后提升了下游预测性能。流程包含六阶段闭环架构与 11 项对抗性验证检查,并保留人工监督。

8月21日周五
8月18日周二
  1. IT之家(RSS)43

    哈佛和麻省理工推出 MatrAIx 系统,生成 83 亿个“AI 人”模拟全球人类行为

    哈佛大学和麻省理工学院牵头,OpenAI、谷歌 DeepMind 等参与的团队推出 MatrAIx 系统,可生成 83 亿个 AI 智能体模拟全球人类行为。该系统用 1,290 个维度建模不同人群,智能体可填写问卷、与客服聊天、浏览网页和操作 App,测试一致性达 91.5%。经筛选的百万人格核心集已在 Hugging Face 公开发布。