Impressive paper showing how much the harness changes a coding agent's results.
Harnesses do play a huge role in what you are getting out of the models.
GPT-5.5 was run inside both Claude Code and Codex on the same 1,000 tasks. With a specialized PowerPoint workflow, it improved inside one harness and got worse inside the other.
The harness also changed scores when the prompt was identical.
ReFigBench asks coding agents to rebuild real arXiv overview figures as editable PowerPoint slides that keep the text, layout and connections.
It covers ten configurations across the GPT, Claude, MiMo and MiniMax families, scored by artifact checks, two families of LLM judges and blinded human comparisons.
Perception is still the main bottleneck. The specialized workflow removed native connectors in every configuration, yet human judges still preferred its renderings in most matchups.
Paper: https://arxiv.org/abs/2609.18844
Chat with Paper: https://academy.dair.ai/papers/refigbench-benchmarking-scientific-figure-reconstruction-as-editable-powerpoint-2609.18844