市场正在告诉我们:借助 AI,我们的生产力应当提升 3 倍。
如果这种生产力提升,仅仅意味着 AI 一天工作 24 小时,而人类只工作 8 小时呢?
OpenAI 公布了其“3 倍”说法背后的数学依据。8 月中旬,其研究人员每工作 8 小时,就对应记录到 3.14 个智能体工作日¹。典型的研究人员会并行运行 4 个智能体²。
这种机器班次伴随着工业级的代价。3 月底,OpenAI 研究人员的推理费用中位数约为每天 14 美元。到 8 月中旬,这笔账单已攀升至每天 600 美元以上:不到五个月激增了 40 倍²。在最高端,第 90 百分位的研究人员每天消耗超过 7,000 美元,年化运行率高达 250 万美元。
按每个工位每年 250 万美元计算,推理的成本行为类似于重型工厂设备。但它有一个财务上的独特之处:它纯粹是运营支出(OPEX)。
汽车工厂用资本支出(CAPEX)购买焊接机器人。它们通过夜班来摊销机械成本——无论使用与否,这些机械都在折旧。AI 系统颠覆了这一算法。推理是按量计费的运营支出。由于没有实体工具,也没有夜班工资,公司可以仅凭纯可变成本让机器通宵运行。
3.14 个工作日的比例,并不意味着思考能力提升了三倍。它意味着一 名工程师监督三个班次的机器运行,而自己只清醒值守一个班次。
然而,与汽车自动焊接机器人不同,这条数字化装配线的缺陷率极高。过去六个月中,成功完成的四到八小时任务里,超过一半仍需要人工干预;实验室坦承,“整体进展速度很可能赶不上这些具体指标的增长速度。”2
算力支出激增40倍,换来的只是三倍的工作时长产出。但由于主管仍需处理超过一半的运行任务,工程师的日常工作从创造性架构设计,变成了巡视车间、清理机器卡顿。
既然缺陷率如此之高,为什么还要让机器彻夜运行?答案是恐惧与野心。
如果你的同行全天候部署四个智能体,那么下线就意味着落后。由雇主买单的超能力带来的兴奋感令人沉醉。当你在别人的资产负债表上获得一支不知疲倦的数字劳动力队伍时,你绝不会关停这座工厂。
四十年来,程序员只需要一台 MacBook 和八小时轮班。而如今,一位顶尖的 OpenAI 研究员同时指挥着四个并行智能体,每年烧掉250万美元的算力,还要花一上午修复前一晚遗留的机器错误。这解释了当下软件工程领域正在蔓延的那种无声的挫败感。3
市场听到3倍生产力,便期待创造性的奇迹。工程师却困于清理彻夜运行的机器人带来的50%废品率。4 市场称之为生产力的3倍跃升。而一位 CFO 只会说,这是在为第二班和第三班付费。就目前而言,这就是一台永不眠机器的真实代价。真正的问题在于,第二班和第三班的产出何时能开始超过第一班。
-
OpenAI 报告的数据是 3.1 个智能体工作日;我们四舍五入取 3.14,纯粹是为了讽刺——因为 3.14 个智能体工作日对 1 个人类工作日的比率,恰好是一个“派”(π),而不是一个分子。 ↩︎
-
OpenAI:研究加速:OpenAI 内部视角 ↩︎ ↩︎ ↩︎
-
Stack Overflow 开发者调查:弥合 AI 信任鸿沟:84% 的开发者使用 AI 工具,但信任度已降至 29%,其中 66% 的人指出代码“几乎正确,但又不完全正确”,45% 的人报告调试 AI 生成的代码比手动编写代码耗时更长。 ↩︎
-
产出率的数学:一个 8 小时的人类班次留下 16 个小时的夜间时段(相当于机器额外运行两个班次)。鉴于 OpenAI 披露超过一半的 4 至 8 小时任务需要人工干预,自主产出率约为 50%。两个机器班次按 50% 的产出率计算,相当于一个有效班次的成品产出。这样交付的工作量约为 2 倍,同时记录的原始班次运行时间是 3 倍。 ↩︎
The market is telling us that we should be 3x more productive with AI.
What if that productivity gain is just an AI working 24 hours a day while a human works eight?
OpenAI published the math behind its 3x claim. In mid-August, its research staff logged 3.14 agent-workdays1 for every 8-hour human shift.2 The typical researcher ran four agents in parallel.
That machine shift comes with an industrial price tag. In late March, the median OpenAI researcher spent $14 a day on inference. By mid-August, that bill climbed past $600 a day : a 40-fold surge in under five months.2 At the top end, the 90th percentile researcher burns through more than $7,000 a day, an annualized run-rate of $2.5m.
At $2.5m a year per seat, inference behaves like heavy factory tooling. But it comes with a financial twist : it is pure OPEX.
Auto plants buy welding robots with capex. They run night shifts to amortize machinery that depreciates whether used or idle. AI systems invert that math. Inference is metered operating expense. With no physical tooling & no graveyard-shift wages, a company can run machines overnight on pure variable cost.
The 3.14 workday ratio is not three times smarter thinking. It is one engineer supervising three shifts of machine runtime while only being awake for one.
Yet unlike an auto welding robot, this digital assembly line has a massive defect rate. Over half of the successful four-to-eight-hour tasks in the last six months still needed human intervention ; the lab is candid that “the overall pace of progress likely won’t keep pace with these specific metrics.”2
A 40-fold surge in compute spend bought three times the work-hours. But with a supervisor still untangling more than half the runs, the engineer’s day shifts from creative architecture to walking the plant floor & clearing machine jams.
Why run the machines through the night if the defect rate is so high? Fear & ambition.
If your peers field four agents around the clock, logging off is falling behind. The rush of a superpower paid for by your employer is intoxicating. When you get a tireless digital workforce on someone else’s balance sheet, you never turn the factory off.
For forty years, a programmer needed only a MacBook & an eight-hour shift. Today, a top OpenAI researcher commands four parallel agents, burns through $2.5m a year in compute, & spends the morning fixing machine errors from the night before. This explains the quiet frustration spreading across software engineering today.3
The market hears 3x productivity & expects creative miracles. The engineer gets stuck untangling a 50% scrap rate from robots that ran all night.4 The market calls it a 3x leap in productivity. A CFO would just call it paying for a second & third shift. For now, that is the honest price of a machine that never sleeps. The real question is when the second & third shifts start to out-yield the first.
-
OpenAI reports 3.1 agent-workdays; we round to 3.14 for the irony, since a ratio of 3.14 agent-workdays to one human workday is, fittingly, a pie, not a numerator. ↩︎
-
OpenAI: Research acceleration : The view inside OpenAI ↩︎ ↩︎ ↩︎
-
Stack Overflow Developer Survey : Closing the AI Trust Gap : 84% of developers use AI tools, but trust has fallen to 29%, with 66% citing code that is “almost right, but not quite” & 45% reporting that debugging AI-generated code takes more time than writing it manually. ↩︎
-
The yield math : an 8-hour human shift leaves 16 overnight hours (two extra shifts of machine runtime). With OpenAI disclosing that more than half of 4-to-8-hour tasks require human intervention, the autonomous yield is ~50%. Two machine shifts at 50% yield equal one effective shift of finished output. That yields ~2x delivered work while logging 3x the raw shift runtime. ↩︎