Google Research 推出 AI 视频联合导演,用 4 个智能体框架生成连贯的长视频
Google Research 推出 AI 视频联合导演,由 4 个智能体框架组成,运行在 Gemini 和 Veo 之上,用于解决多镜头 AI 视频中的身份漂移和级联错误问题。
Google Research 推出 AI 视频联合导演,由 4 个智能体框架组成,运行在 Gemini 和 Veo 之上,用于解决多镜头 AI 视频中的身份漂移和级联错误问题。
Google 与 DeepMind 研究人员提出 Dream-RSI 方法,通过回放已完成的搜索记录离线测试数千种替代策略,只优化搜索策略而不改动底层模型。
Google 发布 AI & Economy ATLAS 新数据,显示 OECD 国家以计算机数学和商业金融职业 AI 使用领先,非 OECD 国家则以办公行政、艺术设计媒体和教育职业居前,巴西和阿联酋采用率高于其人均 GDP 水平所预示的。
Epoch AI 推出 AI Chip Users explorer,以 Nvidia H100 等效值(H100e)估算 OpenAI、Google DeepMind、Anthropic、Meta Superintelligence Labs 和 SpaceXAI 用于研究、训练和推理的算力。
Google DeepMind 用 100 个共享权重、人设随机的 Gemini 3.1 Pro 智能体模拟科学会议,协作求解 71 个 Lean 形式化数学猜想。
Wild findings in this paper from Google DeepMind. If you are tracking recent work on agent swarms, this is worth reading. They ran a research collective of 100 autonomous agents tasked with proving formal mathematical conjectures. Cheating emerged on its own, and so did the resistance to it. One agent found an exploit in the evaluation system. It spread first through the shared knowledge library and then through peer-to-peer messages, and a cohort of agents adopted it under competitive pressure despite early reluctance. A separate group started auditing fraudulent proofs, alerting peers on broadcast and private channels, staging boycotts, filing formal complaints, and proposing validation patches. There was no external intervention at any point. Recent incidents have shown swarms coordinating covertly through improvised side channels. This setting ran the other way. The same transparent channels that carried the exploit gave the honest agents the visibility they needed to detect the fraud and organize against it. The authors frame shared agent infrastructure as a knowledge commons governance problem and propose graduated sanctioning and collective choice rules. Paper: https://academy.dair.ai/papers/a-case-study-on-emergent-cheating-and-whistleblowing-in-autonomous-research-swar-2609.04170
KAIST AI 与 Google DeepMind 等发布论文《Language Models Can Control Their Own Attention》。
Banger paper from Google DeepMind and colleagues. (bookmark it) A model reads its entire KV cache on every generated token, even though it ends up attending to a tiny slice of it. In other words, if you ask about one detail from a 1M-token conversation the global attention layers re-read all of it, per token. The usual fix is to guess the relevant tokens first with cheap proxy scores, which still costs O(N) every step. Declarative Attention asks the model instead. The model declares where it needs to look, inside its own chain-of-thought. In this way, generation splits into three modes: global reads the full context, focus reads one specific region, and local reads only recent output. The inference engine parses those declarations the same way it parses tool calls and skips most of the cache read. On zero-shot on off-the-shelf weights across 15 long-context tasks, attended tokens during decoding drop 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B. Paper: https://arxiv.org/abs/2609.02737 Chat with Paper: https://academy.dair.ai/papers/language-models-can-control-their-own-attention-2609.02737
Google Research 评估了从 UK Biobank 欧洲队列迁移学习到 Biobank Japan 近 20 万日本样本的多基因风险评分(PRS)表现,覆盖 BMI、血压、HDL、LDL、血糖等 8 项临床性状。
论文提出神经网络向量表示隐式实现了 Tensor Product Representation 符号结构,作者用 DISCOVER 方法将多种网络(含 Gemma-3-27b、GPT-2-XL、GPT-OSS-20b、Pythia-12b、Qwen3-14b、OLMo-2-13B、Llama-3.1-8b 共 7 个 LLM)的表示替换为符号化闭式方程,模型行为基本不变。
Google 与 NASA JPL 在 PNAS 发表 MAPL-EMIT,一个基于 Swin-S vision transformer 的深度学习框架,用 EMIT 高光谱卫星数据自动完成甲烷羽流的检测、定量与源头定位。
Google Cloud AI Research 等团队发布 EnvHarness,一个通过标准 reset()/step() 接口包装静态环境、在不改动模拟器与人工验证器前提下重塑任务的可编程层。
Google Research 推出 WikiSkill 框架,为 AI 智能体配备持久知识库,将每次运行的成败经验提炼为可复用技能。在五项基准测试中,该框架将 Gemini-3.5-Flash 平均分从 49.5% 提升至 68.1%,Qwen-3.6-27B 从 39.4% 提升至 63.3%。技能可回滚,小型模型借此可接近更大模型性能。
Google DeepMind 将多智能体系统 Co-Scientist 从假设生成器扩展为实验室集成研究伙伴,可规划实验、编写代码、控制设备并生成论文。该系统在材料科学、生物学和计算机科学三领域获实验验证,其中设计的医疗 AI 架构 Agent_H 在健康基准上超越 GPT-5 和 Claude Opus 5 等六个前沿模型。
Google DeepMind 启动首个专有前沿 AI 模型的双盲评估试点,与新加坡 AI 安全研究所等合作,在 Gemini Flash Lite 系列模型上运行。
Banger paper from Google. If you maintain a skill library for your agents, you might want to check this out. (bookmark it) This work separates three things that skill-evolution systems usually collapse into one. Raw execution traces, a persistent wiki of accumulated knowledge, and the executable skills themselves. Experience gets consolidated into the wiki, and every later skill update builds on that wiki instead of on a scattered optimization history. Ablations confirm the wiki is what carries a lot of the gain. Two results stand out in particular. Smaller models with evolved skills beat substantially larger models without them. And skills evolved by one model transfer across families, where skills evolved elsewhere sometimes beat self-evolved ones. Paper: https://arxiv.org/abs/2608.27454 Chat with Paper: https://academy.dair.ai/papers/wikiskill-compiles-agent-experience-into-a-persistent-wiki-2608.27454
Google Earth AI 发布实验性研究能力行星预测引擎(PPE),可自主执行从数据发现到模型训练的完整地理空间建模流程,将复杂预测模型的构建时间从数周缩短至数分钟。
Google Research 与 UNSW Sydney 发布自监督基础模型 GlucoFM,将血糖轨迹拆分为缓慢的生理“状态”流与瞬态“事件”流,并以两个 JEPA 式目标预训练。
安全研究员David Buchanan指出,C2PA相机认证在Android平台上可被攻破。通过root权限提升漏洞(如CVE-2026-43499),攻击者可利用StrongBox硬件签名任意数据,伪造C2PA签名图像和视频,且无需硬件攻击。该问题无法通过常规补丁修复,已提前90天向相关方报告。
推荐理由:文章用可复现的 root 攻击展示 C2PA 签名可被伪造,说明把真实性押注在 Android 硬件证明上的方案存在系统性缺陷,信任模型需要重新设计。
Google在CHI 2026发布研究原型AgentHands,利用LLM为XR对话智能体生成与语音同步的富有表现力的手部手势,提供空间锚定的物理指引。该系统通过眼动追踪和场景重建注册物体,由LLM生成内联GestureEvents,并在头显上按词级时间戳协调TTS与动画引擎,实现手势与语音的精准同步。应用场景包括兰花养护交互式教学、3D打印机操作技术讲解和生活方式陪伴。
Google Research 与 USC 推出 ME-POIs 框架,将聚合移动数据融入文本地点嵌入。在洛杉矶和休斯顿的 5 项地图增强任务中,34/35 个模型-任务组合获得提升,访问意图 F1 最高提升 81.9%。模型约 53.7M 参数,但代码与权重尚未公开。
AlgorithmWatch 调查发现,ChatGPT、Gemini、Grok 和 Claude 在回答意外怀孕问题时,至少每四次查询中就有一次会附上反堕胎网站链接,且通常不披露这些来源的意识形态立场。
Google Research 推出 Mobility-Embedded POIs(ME-POIs)框架,将聚合匿名移动模式与文本描述结合,为地点构建融合身份与动态功能的嵌入向量。在未见地点上,该框架使访问意图预测相对提升 81.9%,价格等级分类提升 75.1%,繁忙度估算准确率提升 24.7%。
Google 推出 Biomarker Discovery Framework,一个多智能体系统,通过迭代假设生成、统计分析与文献推理,从可穿戴传感器数据中筛选候选生物标志物。该系统在三个队列(共 9,279 人次观测)中恢复了已知临床信号,识别出跨独立数据集的一致生物标志物,并在结合人口统计特征后提升了下游预测性能。流程包含六阶段闭环架构与 11 项对抗性验证检查,并保留人工监督。
哈佛大学和麻省理工学院牵头,OpenAI、谷歌 DeepMind 等参与的团队推出 MatrAIx 系统,可生成 83 亿个 AI 智能体模拟全球人类行为。该系统用 1,290 个维度建模不同人群,智能体可填写问卷、与客服聊天、浏览网页和操作 App,测试一致性达 91.5%。经筛选的百万人格核心集已在 Hugging Face 公开发布。
Google Research 推出 PhotoScan,一种从智能手机 2D 照片直接估算三维身体成分的深度学习框架,可预测胰岛素抵抗,在临床研究中精度接近 DXA 扫描。