A*-Thought-V2:用 LLM 几何动力学实现高效潜在推理

HuggingFace Daily Papers(社区热门论文)·2026-09-08 01:56·1天前
AI 导读

A*-Thought-V2 将思维链建模为隐藏状态轨迹,通过显式-隐式交错潜在架构替代硬删除,用 3D PCA 空间中的方向角判断推理步骤去留。在 Qwen3.5-9B 和 Qwen3.6-27B 上,该方法平均准确率最高提升 2.6%,响应长度缩短近一半,Accuracy per Computation Unit 提升 2.29 倍,预处理和训练时间分别减少 94.6% 和最高 80.3%。

HuggingFace Daily Papers(社区热门论文)
31AI 编辑部评分,满分 100

A*-Thought-V2:用 LLM 几何动力学实现高效潜在推理

2026-09-08 01:56· 1天前
AI 导读

A*-Thought-V2 将思维链建模为隐藏状态轨迹,通过显式-隐式交错潜在架构替代硬删除,用 3D PCA 空间中的方向角判断推理步骤去留。在 Qwen3.5-9B 和 Qwen3.6-27B 上,该方法平均准确率最高提升 2.6%,响应长度缩短近一半,Accuracy per Computation Unit 提升 2.29 倍,预处理和训练时间分别减少 94.6% 和最高 80.3%。

Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture. After projecting question, step, and solution representations into a 3D PCA space, it measures alignment between each local transition and global question-to-solution direction. Aligned steps remain explicit text, whereas deviating steps are compressed into continuous latent tokens. Directional angles capture both local semantics and reasoning dynamics: small angles indicate direct execution and answer formation, while large angles more frequently involve checking, correction, and branch exploration; their temporal variation reveals exploration, convergence, and refinement stages. To train this architecture, we introduce stepwise embedding forcing, which pools each redundant step into a single latent embedding, and label forcing, which supervises that latent token with a soft multi-modal vocabulary distribution instead of a hard one-hot label. Experiments on Qwen3.5-9B and Qwen3.6-27B across six in-domain and out-of-domain benchmarks show that A*-Thought-V2 improves average accuracy by up to 2.6% while reducing response length by up to half, increasing Accuracy per Computation Unit by 2.29times, and reducing preprocessing and training time by 94.6% and up to 80.3%, respectively. Representation analyses suggest that latent states form a compact region distinct from textual states, while higher entropy at latent-token positions reflects broader soft targets that encourage richer step-level feature learning.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org