对智能的一种定义是样本效率——也就是说,在某个给定领域里,你需要看过多少数据才能流畅且胜任地运作。过去几年里,我们在训练样本效率上是否真的取得了很大进展,这一点并不清楚——看起来更像是我们大幅拓宽并改进了数据分布。
AI 一直在变好的主要方式,是通过加入更多且更好的数据,并扩展算力来首先开发这些数据。显然,RL 是实现这一点的主要方式。你可以把 RL 看作一种合成数据生成——你投入大量算力去对抗一个验证器,以找出“好的”数据。然后你训练模型去预测这些正确的 rollout,方式很像你训练它去预测互联网文本中的下一个词。
要让这个过程奏效,模型必须至少先验地具备一定概率来预判出正确解法,这也是为什么你还需要在每一个你希望模型胜任的领域和技能上,拥有数量惊人的人类专家轨迹。
这种人类专家数据有多么任务特定、多么量身定制,很难被夸大。如果你想获得一些直观感受,可以去读读Mercor或 Surge 网站上的一些职位描述。那里有文字专家的招聘信息,负责把遗留文档转换成精美的 Word 文件;还有法律专家,负责撰写逼真的并购尽职调查或证券申报文件;还有管理咨询顾问,负责撰写模板化市场研究,以及几十种其他特定类别。
而且不仅数据必须如此具有领域特异性,数据量还必须如此庞大!每一项技能都对应着至少数百名人类专家,他们在生成示例补全、撰写评分标准,并解释自己的思维链。这个生产专家标注(以及让这些被精心编目的技能得以凝结成形的 RL 环境)的数据产业每年能赚取数十亿美元的收入,很快就会达到数百亿美元,这是有原因的。
想象一下,如果你要学会打磨一个 Word 文件,需要上几十年的课程、同时有数百位教授授课、还要完成数百万道练习任务。即便这样,任务数量的差距仍然低估了真正的鸿沟——模型必须反复打磨它们那数量远为庞大的任务,而且每一项任务都难得多。人类学生可能把一道教科书习题练上一两遍,而 GRPO 会让模型对每项任务生成数百到数千次 rollout。我们正在打造某种弗兰肯斯坦式的怪物,由十亿块精心构造的示例嫁接缝合而成。
Epoch 最近报告称,开源模型仅落后最先进的闭源模型 4 个月。我认为,开源模型和此前的落后者之所以能相对容易地在几个月内追上前沿,原因在于数据才是进步的真正驱动力。而数据可以很容易地从公开 API 中蒸馏出来,超参数、训练技巧和架构上的微观优化则不能——如果后者才是推动大部分进步的因素,那么追赶的难度会比我们观察到的更大。
人们很容易忘记这些模型训练所用的数据量有多大,以及它比我们人类一生中所见的数据多出多少。我们把这些 AI 看作一个闪烁着各种能力的星系,但在它们的中心,肉眼不可见地、将所有星座维系在一起的,是一个难以想象的巨大数据黑洞。
幕间:比较人类与 AI 的样本效率
如果一个人平均每小时听到和看到约 2,000 个词,那么从出生到成年,他们会看到约 2 亿个 token。相比之下,前沿模型训练所用的数据在 10 万亿到 100 万亿 token 之间。这接近一百万倍的差距。
一个人可以在几小时内学会遥操作任何随机的人形机器人或机械臂。机器人技术之所以还没有成为一个十万亿美元级的产业、拥有无穷无尽的 Unitree G1 大军在世界上做各种有用的工作,原因就在于我们的 AI 学习效率远低于人类,即使我们已经收集了数百万小时的演示数据,也不足以让它们执行复杂的、开放式的任务。
一个青少年大约练习 20 小时就能学会开车。即使把他约 16 年积累的物理直觉也算作相关训练数据,那也至少比 Waymo 和 Tesla 训练其自动驾驶汽车模型所需的数据量少了 3-4 个数量级。
我想回应一些对这种比较的常见反对意见:
数十亿年的进化就是我们的预训练,所以把我们一生中所见到的如此之少的数据,与这些冷启动的大语言模型必须从中学习的数据相比,是不公平的。
我们的基因组是 3GB,其中大约 1-2% 是蛋白质编码。这根本不足以存储据称经过预训练的模型参数(前沿模型的规模达到数 TB)。更贴切的类比大概是:进化找到了正确的超参数和损失函数(旁注:我曾与 Adam Marblestone 做过一期有趣的播客,他在其中主张,损失函数才是进化更重大的发现),但相当于参数训练的过程仍然在一生之中持续进行,并被编码在一生中构建起来的大脑神经连接图谱里。
即便我们能够把预训练一个基础模型所需的数万亿 token 解释为是在追赶进化,这仍然无法解释为什么边际能力需要如此多的数据——一旦你接受了教育,你不需要 100 位不同的教授来学习一门新的编程语言,但 AI(即便已经过预训练)却需要。
这些比较并未计入我们一生中所见到的多模态数据。如果把所有这些感官信息都算进去,我们从出生到成年大概处于 10s 到 100s of billions of tokens 的量级。
被隔绝于这类感官信息的盲人/聋人可能缺乏相应感官的能力,但仍然拥有与其他人相同的通用智能。这表明,这数十亿的感官 token 并不是真正让人类变聪明的东西。
事实上,只能通过手语和阅读(而非听觉)交流的聋人,其摄入的语言 token 量远低于我们此前计算的 2 亿,而即便如此,也足以让他们成为完全通用的智能。
缩放定律告诉我们,更大的模型样本效率更高。人脑有 100T 个突触——如果每个突触约为 1 个参数,而当前前沿模型大约为 5T 参数,那么也许再扩大一两个数量级的参数规模,我们就能达到人类水平的样本效率。
缩放定律方程的工作原理是,参数项和数据项被独立地加到损失上。如果你有一个以计算最优方式训练的模型,假设你问:如果我只想最大化样本效率、使用更少的数据——而我会不惜投入任意多的参数来实现这一点,那会怎样?用 Chinchilla 缩放定律论文中的常数(即便换成不同的常数,结果的性质也不会改变),即使你把参数数量增加到无穷大,也只能把保持相同损失所需的数据量减少约 10 倍。
人类的样本效率比这些模型高出数千到数百万倍。当前模型的缩放根本无法弥补这一差距。这确实表明,人类处在一条完全不同的缩放曲线上。
样本效率重要吗?
但你可能会问,样本效率为什么重要?实验室有两个总体目标:自动化白领工作,以及自动化 AI 研究本身。人类水平的样本效率对其中任何一个来说是必需的吗?
白领工作的赌注在于,软件工程师、分析师或会计师所做的常见任务是,嗯,常见的。而我们可以通过 RL 和 SFT 相当容易地把常见任务带入分布之中。这些 AI 实验室的收入曲线表明,即使我们无法复现人类的样本效率,把任务带入分布之中也能带来巨大的价值。
是的,训练 AI 来完成这些任务,效率远低于训练人类。但那又怎样?人类的寿命不允许这些模型所经历的训练在数量和广度上达到那样的程度。如果你作为一个人,有某种奇怪的学习障碍,需要读遍 Github 上每一个公开仓库才能成为一名合格的开发者,那么把你训练出来就是不合理的。
你在教育的早期阶段就会靠社会保障过活了,而且即便你训练完成,你也一次只能做一个项目。但 AI 可以通过一次灌入吉瓦级的训练来学会这些技能。而它们所学到的东西可以摊销到数十亿次会话中,所以我们可以在训练它们时荒谬地低效,却仍然大幅盈利。
白领员工需要做多少“分布外”的思考,而这些是你根本无法提前训练的?这其实更多是关于不同工作性质的问题,而不是关于 AI 研究的问题。而且这也取决于具体工作——有些工作足够机械、足够可预测,早在现代 AI 时代之前就已经被自动化了,比如银行柜员或旅行社代理。
还有一些工作则需要每天处理与数据分布相距甚远的问题。即便是软件工程(AI 被认为会最先取代的工作)也是其中之一。我敢打赌,2028 年对人类软件工程师的总体需求会比现在更多,很大程度上是因为 AI 的互补性投入。
实验室对后一类工作的计划是,先自动化 AI 研究,然后让自动化的 AI 研究员来解决这个样本效率问题。所以问题就变成了:那些不具备人类水平样本效率的 AI,是否仍然能够解决通往类人智能与学习之路上的剩余研究问题。
这个问题我会在未来的博客文章中讨论——我认为人们目前对智能爆炸的思考方式相当笨拙。要么人们完全否定 AI 加速 AI 进展的可能性,要么他们就只是假设最后会凭空冒出一个神。人们并没有在推理,从 LLM 开始,极其迅速的进展究竟会是什么样子。
One definition of intelligence is sample efficiency - that is to say, how much data do you need to see in a given domain in order to operate fluently and competently. It’s not clear that we’ve actually made much progress on training sample efficiency over the last few years - it seems like more so we’ve dramatically widened and improved the data distribution.
The main way that AIs have been getting better is from adding more and better data, and scaling the compute to develop that data in the first place. Obviously RL is the main way that has happened. You can think of RL as a kind of synthetic data generation - you dump a lot of compute against a verifier in order to find the “good” data. Then you train your model to predict these correct rollouts, much in the same way that you might train it to predict the next word in internet text.
For this process to work, the model must have at least prior some probability to anticipate the correct solution, which is why you also need mind-stretching amounts of human expert trajectories in every single field and skill you want the model to be competent at.
It’s hard to overstate how task specific and bespoke this human expert data is. If you want to get some intuition, go read some job descriptions at Mercor or Surge’s websites. There are listings for a word specialists who will convert legacy documents into polished Word files, and legal experts who will write realistic M&A diligences or securities filings, and management consultants who will write up template market research, and dozens more other particular categories.
And it is not only that the data have to be so domain specific, but there has to be so much of it! Each skill corresponds to at least hundreds of human experts who are generating example completions, writing rubrics, and explaining their chain of thought. There’s a reason that the data industry producing these expert labels (and the RL environments in which their meticulously catalogued skills can congeal) is earning billions a year in revenue, soon deca-billions.
Imagine if it took a couple decades worth of courses with hundreds of concurrent professors and millions of practice tasks for you to learn how to polish a word file. Even the task count difference understates the gap - the models have to grind their far more numerous tasks each far harder. Whereas a human student might practice a textbook problem once or twice, GRPO has the model generate hundreds to thousands of rollouts per task. We are building some Frankenstein’s monster, with a billion grafts of carefully constructed examples sewn together.
Epoch recently reported that open models only lag state-of-the-art closed models by 4 months. I think the reason it is relatively easy for open source and previous laggards to catch up to within months of the frontier is that data is the real driver of progress. And data can be easily distilled from public APIs, whereas hyper-parameters and training tricks and architectural micro-optimizations cannot - if the latter were driving most of progress, then catching up would be harder than we are observing it to be.
It is easy to forget how much data these models are trained on, and how much more it is than what we humans see in our lifetimes. We see these AIs as a galaxy glittering with capabilities, but at their center, invisible to the naked eye, holding all the constellations together, is an unimaginably massive black hole of data.
Intermission: Comparing human and AI sample efficiency
If a person hears and sees on average ~2,000 words an hour, then from birth to adulthood, they’ll see ~200 millions tokens. By contrast, frontier models are trained on somewhere between 10s to 100s of trillions of tokens. That is close to a million fold difference.
A person can learn to teleoperate any random humanoid or robot arm within hours. The reason robotics isn’t already a deca-trillion dollar industry, with a endless army of Unitree G1s doing all kinds of useful work in world, is that our AIs learn so much less efficiently than humans, and even the millions of hours of demonstrations we’ve collected is not enough to allow them to perform complex, open ended tasks.
A teenager can learn to drive a car with about 20 hours of practice. Even if you include their ~16 years of accumulated physical intuition as relevant training data, that is at least 3-4 orders of magnitude less than the amount of data Waymo and Tesla have needed to train their self-driving car models.
I wanna deal with some common objections to this kind of comparison:
Many billions of years of evolution is our pre-training, so it’s unfair to compare how little data we see simply within our lifetime to what these cold-started LLMs have to learn from.
Our genome is 3GB, about 1-2% protein coding. That is just not enough space to store the model parameters that are supposedly pretrained (frontier models are terabytes sized). The closer analogy is probably that evolution has found the right hyperparameters and loss functions (Sidenote: I had an interesting podcast with Adam Marblestone where he argued that the loss functions were the more significant find from evolution), but that the equivalent of parameter training is still happening within lifetime, and is encoded in the map of neural connections in the brain built up over a lifetime.
Even if it were the case that we can explain away the trillions of tokens required to pretrain a base model as catching up to evolution, it doesn’t explain why the marginal capabilities take so much data - once you have been educated, you don’t need 100 different professors to learn a new programming language, but the AIs (even once pretrained) do.
These comparisons are not including the multimodal data we see in our lifetimes. If you include all this sensory information, we’re probably in the 10s to 100s of billions of tokens range from birth to adulthood
Blind/deaf people who are cut off from this kind of sensory information might lack faculty with the relevant sense, but still have the same general intelligence as everyone else. Which suggests that all these billions of sensory tokens are not really the thing making humans smart.
In fact, deaf people who can only communicate via sign language and reading (and not from hearing) are ingesting far less than the 200 million language tokens we calculated earlier, and even this is sufficient for them to be fully general intelligences.
Scaling laws tell us that bigger models are more sample efficient. The human brain is 100T synapses - if each synapse is ~1 parameter, and frontier models are currently roughly ~5T parameters, then maybe we could achieve human-level sample efficiency with another order of magnitude or two of parameter scaling.
The way the scaling law equations work is that parameter and data terms are added to the loss independently. If you have a model that is trained compute optimally, and suppose you ask, well what if I just wanna maximize sample efficiency and use less data - and I’ll throw in as many parameters as it takes to make that happen. With the constants from the Chinchilla scaling laws paper (and the nature of the result wouldn’t change even with different constants), even if you increased the number of parameters by infinity, that would only decrease by a factor of ~10 the amount of data you need in order to keep the same loss. Humans are somewhere between thousands to millions of times more sample efficient than these models. Scaling of current models simply can’t make up for that discrepancy. This really does suggest that humans are on a different scaling curve altogether.
Does sample efficiency matter?
But you might ask, why does sample efficiency matter? The labs have two overarching objectives: automate white collar work, and automate AI research itself. Is human level sample efficiency necessary for either?
The bet with white collar work is the common tasks that a software engineer or analyst or accountant does are, well common. And we can bring common tasks into distribution quite easily through RL and SFT. The revenue curves of these AI labs suggest that there is enormous value from bringing tasks into distribution, even if we don’t replicate human sample efficiency.
Yes it is far more inefficient to train AIs to do these tasks than it is to train humans. But so what? Human lifespan does not allow for the quantity and breath of training these models experience. If you as a human had some weird learning disability where you needed to read through every public repository on Github before you could be a competent developer, it would not make sense to train you up. You’d be on Social Security by the early stages of your education, and even once you were trained, you could work on only one project at a time. But AIs can learn these skills by firehosing gigawatts of training at a time. And what they learn can be amortized across billions of sessions, so we can be ludicrously inefficient in training them and still be wildly in the green.
How much “out-of-distribution” thinking do white collar employees need to do that you simply can’t train for in advance? Well this is more a question about the nature of different jobs rather than a question about AI research. And also depends on the job - some jobs are mechanical and predictable enough that they were automated long before the modern era of AI, for example bank tellers or travel agents. And there are other jobs which require dealing on a daily basis with problems that are quite distant from the data distribution. Even software engineering (the jobs AIs are supposed to take first) is one such. I would be willing to bet that there’s overall more demand for human software engineers in 2028 than there is now, largely due to the complementary input of AI.
The labs’ plan for these later kinds of jobs is to first automate AI research, and then have the automated AI researchers solve this sample efficiency problem. So then the question is, can AIs, which do not have human-level sample efficiency, nonetheless solve the remaining research problems on the way to human-like intelligence and learning.
That question I’ll address in a future blog post - I think the way that people currently think about an intelligence explosion is pretty clumsy. Either people dismiss the possibility of AIs speeding up AI progress altogether, or they just assume that God pops out the other end. People are not reasoning about what extremely rapid progress, but starting with LLMs, looks like.