Mercor 研究:AI 在会计任务上速度和准确率超过持证会计师,但仍无法独立结账
AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision
Mercor 的研究让 12 名平均经验五年半的持证 CPA 完成简化的 APEX Accounting Benchmark 任务,当前模型在同类结构化记账任务上已几乎无错,而 18 个月前最佳模型还低于会计师约 37% 的平均分。
AI models are faster, more accurate, and far cheaper than accountants at structured bookkeeping tasks. That's the finding of a study by Mercor in which 12 licensed CPAs with an average of five and a half years of experience worked through simplified tasks from the APEX Accounting Benchmark. Eighteen months ago, the best models still scored below the accountants' average of about 37 percent. Today, they solve the same tasks almost flawlessly.

The full APEX Accounting benchmark is much bigger, with 160 tasks across 10 simulated companies, built by more than 40 professionals who average 11 years of experience. Claude Opus 5.5 currently leads with 61.8 percent of grading criteria met, followed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent. Still, Mercor says no model fully solved almost 60 percent of the tasks. AI models can't close the books without oversight yet.
Mercor also admits the study's tasks test exactly what AI does best, which is hunting down details and following instructions precisely. The study left out key parts of the job, such as talking with clients, checking in with colleagues, and drawing on context built up over years. Mercor says that's why accountants can't be replaced, though it expects major productivity gains across the industry.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
来源:The Decoder:AI News · the-decoder.com