A benchmark for difficult scientific discovery from @tianqiao_chen now shows GPT-6 Astra improving across all six capabilities.
But fixing its own mistakes improved the least, Repair improved far less than the other dimensions.
For AI agents, a correct final answer can hide a weak investigation.
An agent might reach the right result after ignoring contradictory feedback. That behavior could fail on the next task, even though the current answer passes its check.
@Apodex_AI released TRACES, a new benchmark for testing AI systems on difficult scientific problems where the correct answer may not already be known.
TRACES separately measures Tools, Repair, Alternatives, Coherence, Evidence, and Scope instead of reducing discovery capability to one score.
Instead of scoring only the final answer, TRACES also evaluates how the AI works through the problem, including its tool use, error correction, evidence, and reasoning process.