ReFigBench:编码智能体框架影响科学图表重建结果

elvis · @omarsar0 · X·2026-09-23 03:50·4小时前
AI 导读

ReFigBench 基准要求编码智能体将真实 arXiv 概览图重建为可编辑 PowerPoint 幻灯片,覆盖 GPT、Claude、MiMo、MiniMax 四个模型家族的十种配置。

elvis@omarsar0
44AI 编辑部评分,满分 100

ReFigBench:编码智能体框架影响科学图表重建结果

2026-09-23 03:50· 4小时前
AI 导读

ReFigBench 基准要求编码智能体将真实 arXiv 概览图重建为可编辑 PowerPoint 幻灯片,覆盖 GPT、Claude、MiMo、MiniMax 四个模型家族的十种配置。

Impressive paper showing how much the harness changes a coding agent's results.

Harnesses do play a huge role in what you are getting out of the models.

GPT-5.5 was run inside both Claude Code and Codex on the same 1,000 tasks. With a specialized PowerPoint workflow, it improved inside one harness and got worse inside the other.

The harness also changed scores when the prompt was identical.

ReFigBench asks coding agents to rebuild real arXiv overview figures as editable PowerPoint slides that keep the text, layout and connections.

It covers ten configurations across the GPT, Claude, MiMo and MiniMax families, scored by artifact checks, two families of LLM judges and blinded human comparisons.

Perception is still the main bottleneck. The specialized workflow removed native connectors in every configuration, yet human judges still preferred its renderings in most matchups.

Paper: https://arxiv.org/abs/2609.18844

Chat with Paper: https://academy.dair.ai/papers/refigbench-benchmarking-scientific-figure-reconstruction-as-editable-powerpoint-2609.18844

来源:elvis· x.com