跳到正文
ARC Prize· @arcprize · X·· 6 天前AI 评分43
AI 导读

ARC Prize 通过 Baseten 测试 Qwen3.8-27B,发现其自带聊天模板会按推理档位注入不同指令:low 要求思考简短聚焦,xhigh 要求仔细思考、检查假设并考虑替代方案,medium 则不加任何指令。这可能解释了 medium 档得分偏低——这些设置改变的是模型解题的指令方式,而非单纯扩大思考预算。

正文

Chat models commonly include a chat template that formats messages and can add instructions before they reach the model.

We tested Qwen3.8-27B through @Baseten, which uses Qwen3.8-27B's own chat template. It tells low to keep its thinking brief and focused, and xhigh to think carefully, check assumptions, and consider alternatives. Medium adds neither instruction, although thinking remains enabled.

That difference may help explain medium's lower scores: these settings change how the model is instructed to approach a problem, rather than simply giving it a larger thinking budget.

View Qwen3.8-27B's chat template: https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/chat_template.jinja#L51

来源:ARC Prize · x.com