Terence Tao:当下大语言模型的数学原理简单,难在预测能力表现

Rohan Paul · @rohanpaul_ai · X·2026-09-14 05:42·1小时前
AI 导读

Terence Tao 表示,今天大语言模型的训练与运行主要用线性代数、矩阵乘法和少量微积分,本科生即可掌握,构建与运行这些模型的方法是清楚的。真正的谜团是为何模型在部分任务上表现好、另一些上失败且无法提前预测;原因之一是自然文本介于纯噪声与完全结构化数据之间,对应数学理论薄弱,类似物理学在中观尺度上的困境。内容来自 @Briankeating 的 YouTube 频道视频。

Rohan Paul@rohanpaul_ai
50AI 编辑部评分,满分 100

Terence Tao:当下大语言模型的数学原理简单,难在预测能力表现

2026-09-14 05:42· 1小时前
AI 导读

Terence Tao 表示,今天大语言模型的训练与运行主要用线性代数、矩阵乘法和少量微积分,本科生即可掌握,构建与运行这些模型的方法是清楚的。真正的谜团是为何模型在部分任务上表现好、另一些上失败且无法提前预测;原因之一是自然文本介于纯噪声与完全结构化数据之间,对应数学理论薄弱,类似物理学在中观尺度上的困境。内容来自 @Briankeating 的 YouTube 频道视频。

Terence Tao: The math behind today’s LLMs is actually simple. Training and running them mostly uses linear algebra, matrix multiplication, and a bit of calculus, material an undergraduate can handle. We understand how to build and operate these models.

The real mystery is why they work so well on some tasks and fail on others, and why we cannot predict that in advance. We lack good rules for forecasting performance across tasks, so progress is largely empirical.

A key reason is the nature of real-world data. Pure noise is well understood, perfectly structured data is well understood, but natural text sits in between, partly structured and partly random. Mathematics for that middle regime is thin, similar to how physics struggles at meso-scales between atoms and continua.

Because of this gap, we can describe the mechanisms but cannot yet explain capability jumps or give reliable task-level predictions. That mismatch, simple machinery versus hard-to-predict behavior, is the core puzzle.

----

Video from Prof @Briankeating YT Channel (Link in comment)

来源:Rohan Paul· x.com