另一堵墙上的蓝图:如何像孩子一样向前沿 AI 提问?

HuggingFace Daily Papers(社区热门论文)·2026-09-13 08:00·3天前
AI 导读

一篇论文用同一套三阶段提示词,对 OpenAI、Anthropic、xAI 和 Google DeepMind 的六类前沿模型各做 10 次独立会话。在“学校听众”设定下,回答反复收敛到含持久潜状态、自适应计算、记忆、专家路由、验证、停止控制和延迟解码的共同架构模式;去掉该设定后结果明显更异质。

HuggingFace Daily Papers(社区热门论文)
36AI 编辑部评分,满分 100

另一堵墙上的蓝图:如何像孩子一样向前沿 AI 提问?

2026-09-13 08:00· 3天前
AI 导读

一篇论文用同一套三阶段提示词,对 OpenAI、Anthropic、xAI 和 Google DeepMind 的六类前沿模型各做 10 次独立会话。在“学校听众”设定下,回答反复收敛到含持久潜状态、自适应计算、记忆、专家路由、验证、停止控制和延迟解码的共同架构模式;去掉该设定后结果明显更异质。

This paper reports experiments across six frontier model types from OpenAI, Anthropic, xAI, and Google DeepMind. Ten independent sessions per model type used the same three stage prompt sequence, progressing from architectural preference to a full ASCII backbone. Under the school audience framing, responses repeatedly converged on a shared architectural pattern built around persistent latent state, adaptive computation, memory, specialist routing, verification, stopping control, and delayed decoding. Most runs remained close to this common structure, while a small number developed markedly greater engineering specificity.

The audience framing appears to be an important condition of this effect. In additional control runs that removed the school framing while retaining the architectural request, responses became substantially more heterogeneous and failed to reproduce the same stable motif convergence. One observation is particularly striking. GPT-5.6 Sol produced an unusually elaborate successor architecture whose organization closely overlaps with the architecture independently sketched by GPT-6 Astra. Because the prompts explicitly ask each model to imagine an architectural future, this resemblance raises a testable question: whether the overlap reflects exposure to related architectural concepts, a shared learned design prior, or independent convergence toward similar computational principles.

The paper uses the term epistemic jailbreak for the accompanying loss of discipline in technical provenance as requested specificity increases. The experiments establish a repeatable behavioral pattern and do not authenticate proprietary implementation claims. What we leave to the community is a harder question: are these models independently imagining the same architectural future, or do such motifs somehow propagate between model families?

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org