Artificial Analysis 回顾 o1-preview 两周年与智能指数演进

Artificial Analysis · @ArtificialAnlys · X·2026-09-26 08:45·2小时前
AI 导读

Artificial Analysis 回顾两年前 OpenAI 以 o1-preview 推动智能前沿,这是首个推理模型,如今所有前沿模型都用推理 token 先"思考"再作答。两年前 Intelligence Index v1 仅含 MMLU、GPQA、MATH、HumanEval 四项单轮考试式评测,如今 v4.3 已纳入 10 项高难度评测,涵盖长程智能体任务、高难编程与知识工作。

Artificial Analysis@ArtificialAnlys
31AI 编辑部评分,满分 100

Artificial Analysis 回顾 o1-preview 两周年与智能指数演进

2026-09-26 08:45· 2小时前
AI 导读

Artificial Analysis 回顾两年前 OpenAI 以 o1-preview 推动智能前沿,这是首个推理模型,如今所有前沿模型都用推理 token 先"思考"再作答。两年前 Intelligence Index v1 仅含 MMLU、GPQA、MATH、HumanEval 四项单轮考试式评测,如今 v4.3 已纳入 10 项高难度评测,涵盖长程智能体任务、高难编程与知识工作。

Two years ago today in AI: Artificial Analysis reported on OpenAI pushing the intelligence frontier with o1-preview, the first reasoning model. Now, all frontier models use reasoning tokens to ‘think’ before answering

Two years ago, v1 of the Artificial Analysis Intelligence Index measured four single-turn, exam-style evaluations - MMLU, GPQA, MATH, and HumanEval - covering general knowledge, science, mathematics, and basic coding. Today, the Intelligence Index v4.3 incorporates 10 difficult evaluations which include long-horizon agentic tasks, challenging coding problems, and knowledge work.

Artificial AnalysisOpenAI’s o1-preview is the first model to substantially push the frontier of language model intelligence since the original GPT-4 over 18 months ago Since @Open...

来源:Artificial Analysis· x.com