关于 AI 写作的批评有很多,但大多数都集中在更具创意性、更具个人风格的内容上,比如这个博客。这些批评——包括我自己写的那篇——常常认为,好的写作之所以好,是因为它带有鲜明的个人风格、有观点、有需要传达出来的深刻人文表达,或者是一种你能透过所选文字窥见的思考过程。随着 LLM 作为工具(而非对话助手)变得越来越精细,我认为在让模型产出鼓舞人心的写作这一目标上,我们实际上是在倒退。
另一面则是非虚构写作。填充性、复制性的文本曾是 LLM 真正有用的能力之一(Sam Altman 在最近一期播客中谈到了很多关于 GPT-3 早期业务的内容)。过去看起来,这方面的任何缺陷主要都归结于模型整体智能水平的不足,或者其他训练问题,而所有非虚构和解释性文本最终都会被快速进步的浪潮最终所碾压。在过去几年里,我把这些模型当作写作助手来使用,它们确实变得好了一些,但值得反思的是,是什么在阻碍它们更进一步。
模型在长篇非虚构写作上的停滞不前,对于那些指望模型在不久的将来自主解决宏大的开放性科学问题的人来说,应该敲响警钟。如今的模型在组织和有力呈现其领域内一些最成熟的科学内容时都颇为吃力。这似乎是一个自然的先决条件——我们应该期望模型在能够独立解决广泛、开放性问题之前先掌握这一点。在这一点得到解决之前,LLM 在科学领域的进展看起来更像是摘取低垂的果实、连接跨领域的遥远关联,而非任何革命性的洞见。
对于一个对 AI 进展非常乐观的人来说,这是一个颇有争议的观点,尤其是在 Anthropic 发布博客文章、宣称 Claude 在著名的黎曼猜想上取得进展的当天写下这些。科学问题范围极其广泛,我不认为当前 AI 模型的覆盖范围有人们想象的那么广。
组织知识本身就是一种压缩。这种压缩是产生洞见所必需的。如今的大语言模型在长篇非虚构写作中增加了熵值,我看不出这种模式如何能无限地自我叠加。它们将依赖人类充当某种向导。
我仍然对这些狭窄科学领域(比如我们在数学领域看到的极端进展)向持续、更广泛进步的转化感到非常乐观——大语言模型是科学家们使用过的最强大的助手。但我首先需要解释,观察模型在处理这类基于事实、低层级知识问题的写作时,为何让我看到了令人惊讶的泛化能力缺失。
为了提供更多背景,我刚刚完成了一本后训练教科书的写作,《基于人类反馈的强化学习》(可在Manning或Amazon购买)。我在许多方面借助了大语言模型,从处理方程式的 LaTeX 格式、进行大量文字编辑,到为 TikZ(LaTeX 中)或 Python 等编程语言创建图表。[
为什么模型在写作质量上停滞不前?
我原本以为,模型在非虚构写作方面会取得远远更多的进展。考虑到 2024 年的情况,我几乎以为自己 2026 年出版一本非虚构类书籍会显得很落伍。如今,一些在写作能力上最负盛名的模型已经相当老旧了,例如 OpenAI 的大型GPT 4.5和 Moonshot 的Kimi K2。在这些模型发布前后,它们在编码和数学等其他任务上的表现已经从尚可跃升到了超人水平。也许一个更接近但仍不完美的比较是,模型在搜索和研究类任务上如何从毫无能力发展到表现尚可。大多数其他技能领域的进步速度都非常陡峭,但写好文章似乎与其中大多数技能都不太相关。我不认为写作只是被忽视了,而是它本身就极具挑战性,并且缺乏可以针对性干预的优质训练数据。
确实有一些唾手可得的方法可以让 AI 模型更擅长写作——比如像 Claude Code 这样的专用工具框架、提示词,以及能让模型在输出上花费更多推理 token 的训练环境,但我认为这些不会对能力产生倍增式的影响。写好文章是一项非常困难的任务!我们尚未为这项伟大的智力追求解锁推理时扩展(inference-time scaling),这实在令人遗憾。无论如何,写作似乎与模型所擅长的领域截然不同。1
如今,这些模型在长篇技术写作上似乎确实表现得很糟糕。它们能把单个句子写对,但如果你让它们写完整的一章,结果就会是:夹杂着令人困惑的措辞、组织结构混乱、整体上有点不对劲。它们会在没必要的地方刻意卖弄聪明,过程中还会犯下各种随机的概念性错误。在不久的将来,模型在小错误方面会好得多,尤其是随着模型变得更大——这让它们能容纳更多的世界知识——但我并不认为它们运用这些知识的能力会发生质变。
例如,GPT 系列模型长期以来在发现错别字和小问题方面表现出色。我把接近定稿的书稿以 PDF 形式传给 GPT 5.5 Pro,它在这份 200-300 页的手稿中发现了深层且令人惊讶的细微错别字。
另一方面,Claude 系列模型作为编辑则要有用得多。它们更有品味,往往更能理解任务的心智模型,并且能给出更有意思的建议,帮助打破各种形式的写作瓶颈。
我上面给出的例子都有一个共同的主题。模型知道如何检查每一块内容——在这里通常是一个句子、一个公式或一张图——或者专门针对你卡住的那个具体部分进行修改。凭借这些技能,它们在反复审视各个组成部分并把它们串联起来方面做得并不好,因为它们会在彼此之上层层叠加许多新增内容。这感觉像是一种不可约的复合错误。我们过去在数学和代码中处理过这类错误,但回想起来,RLVR 在减少这类错误方面确实是一个神奇的解决方案。
作为写作者,从当前模型中获取价值
我愿意坦白,我的书里有几句技术解释性文字出自 AI 模型——占比远低于 1%——它们之所以留在那里,是因为我真的很喜欢那些句子。我允许自己考虑在书中纳入一些 AI 生成的 token,因为如果我自己作为真正的专家都觉得那句话正是读者所需要的,那就不像是在作弊。尤其是在编辑过程中,我对每一处都盯得很紧,同时又一直担心在自己手头事务繁多的情况下这本书到底能不能完成,这成了一条极其宝贵的出路。
举个例子,我的编辑留下的一系列问题散布在 LaTeX 文件里,用类似 \editor{} 这样的特定分隔符标记。我会让 Claude Code 定位到每一条评论,打印出前后文,然后告诉我这是简单的笔误修正,还是需要更细致处理的问题。我会先写出回复——也就是要插入的文本——或者先请 Claude 给出建议,然后再动手修改。从智识上说,这是一种非常聚焦的编辑流程,也是改进这本书的一种有趣方式。有时 Claude 建议里的某些措辞最终就被写进了书里。
这无疑是一条滑坡,当我接受几条 AI 建议时,正值我在进行第二次全稿审阅。从情感上讲,这个项目感觉已经完成了,但我还有更多工作要做。走出教材写作这个过程后,我深深感激自己为 Interconnects 写作定下的那条干脆利落的规矩:绝不在内容中使用 AI 输出。
用一种只属于你自己的方式写作要有趣得多——独特的声音,因过程本身而被深深珍视——但撰写一本标准参考书并不是以有趣著称的活动。我理解为什么当人们的大部分写作只是用来填充篇幅的输出、而非通往某个目标的手段时,他们会把 AI 工具当成拐杖。
我写作的动力来自大量书写,是为了学习、为了感受、为了表达。
我在自己的科研工作中也在权衡类似的取舍。AI 模型非常适合处理论文中重复性的部分,比如起草你早已烂熟于心的相关工作或背景章节;但用它来写摘要、引言、实验或结论,就太可惜了。这些部分承载着论文的故事与灵魂——正是在这里,你才能弄明白自己的研究究竟在讲什么。
我确信,因为能让 AI 模型帮我创作和检查非虚构类写作,我创造了多得多的净价值。它们让书写公式变得轻而易举,还能帮忙重构代码仓库、在不同语言之间移植,以及做许多其他事情。一开始这非常有趣,直到后来,我被出版流程的漫长拖得有些疲惫,眼看着这个领域不断向前推进。
举一个例子说明为什么 AI 在这个案例中至关重要:我需要同时在两个地方维护我的书的 Markdown 和 LaTeX 版本,因为读者在网页版上反馈意见,而 Manning 出版社的编辑团队审阅的是一个分叉副本。没有 AI 智能体,同步这两个版本所花的时间很容易就是现在的五倍(而这项任务本身已经耗费了数十个小时)。
与这段经历交织在一起的另一个发现——我在思考智能体时偶然想到的——是使用智能体并不会加快你的理解速度。那种以直觉、品味、本能等形式呈现的理解,才是未来真正有价值的东西。用 AI 做非虚构写作,恰恰会削弱这种成长。更糟的是,如果你本来就不是专家,你根本发现不了它的缺陷。
就我而言,我迫切地想把脑子里的知识倾倒到纸面上,以至于有些时候使用 AI 模型确实是个得力的工具。我写这本书的很大一部分动机,是想为 拒绝采样 或 角色训练 这类重要的后训练方法提供一个单一参考来源,而网上关于这些方法的内容少之又少。
这本教科书在很大程度上是对社区的回馈,能以任何形式完成它本身就是一种巨大的胜利,所以我觉得这样是可以接受的。我确信,如果投入更多人力,我本可以学到更多,成品质量也能略有提升。但决定性的因素是,我感觉这本书在出版之前就会过时——一方面是担心 AI 模型的能力,另一方面是担心这个领域发展太快。
结果证明这个担心其实大错特错?我对最终成果非常满意,而且现在比 2024 年刚开始动笔时更有信心它能经得起时间考验,因为模型在非虚构写作方面远远没有达到宣传中的那种高度。
技术写作将走向何方
模型是不可思议的工具,它们让你能以不同形式表达知识。它们非常适合用来创作创意性的填充内容或背景材料——例如,幻灯片的初稿,其真正的价值在于为教师讲课提供话题切入点——这让任何知识都能从一种媒介转换到另一种媒介。
写作非虚构类或参考型教科书时,有一个微妙而早期的阶段,感觉更接近写一篇像这样高语气的博客文章。在推进早期组织架构、呈现核心骨架的过程中,新的知识被创造出来。这部分需要洞察力,而大语言模型在取代它方面还远远落后。
上面那段和前一节的核心意思是:如果世界上更多专家能用 AI 模型帮他们写一点点书,从而让更多知识分享给世界,我会很高兴。问题在于,今天你只能用 AI 模型节省 10-20% 的精力,而我看不到这个比例在短期内变成多数。
还有社会压力——人们期望大语言模型成为最好的、个性化的教育者,所以他们觉得写书或做教育内容是徒劳的。我认为其中一些观点正在过时,因为最高质量的教育作品存在巨大的匮乏——而且这种匮乏一直存在。AI 擅长把这类内容加工成适合学生的形式,而不是从零开始创造内容。
与此同时,我觉得我们困在一个令人沮丧的局部最优解里:AI 模型总体上会降低人们在非虚构写作上投入的平均精力,但它们本可以促成伟大的表达。愿意开始并坚持到底的人会越来越少。
所以,在未来 2-5 年内,我仍然预期最好的教科书会大量由人类亲手打磨。之后我就不确定了,但考虑到这些模型拥有如此多的知识、以及它们天然倾向于源源不断地输出这些知识,这个时间跨度已经比很多人预测的要长了。
至于能力方面的结论,这些模型在两个场景下表现出色:1)任何真正可验证的领域;2)在获得大量上下文并需要做小幅修改时——比如找 bug、解一道非常具体的数学题或给出反馈——而不是开放式地生成散文。长文写作肯定会先于创意写作败下阵来,但这强烈说明,在问题定义不充分的情况下,模型无法充分表达其全部知识。当我们试图推动模型成为某种“数据中心里的天才”去解决重大科学问题时,这似乎是一个相当根本性的局限。
有一篇很棒的文章,讲的是为什么 LLM 能成为优秀的编辑,但同时也是糟糕的写作者。
这部分与人们使用模型的方式有关。如果你在 Claude Code(或聊天应用里,我相信也一样)中问 Claude Fable 5:“给我写一首关于金鱼的好诗”,模型会很快吐出一个答案。我问模型它是怎么做到的,是否有一个类似草稿纸的东西,在返回好结果之前它会先写上去并迭代更新,它说没有。
它在推理 token 中做了一个极简的计划,然后自回归地生成了一首诗。它完全没有利用推理时扩展(inference-time scaling),也没有把它当作一个困难任务来处理。
这大概就是大多数人使用模型写作的方式,结果平庸也就不足为奇了。要从模型中获取最佳结果,需要非常重度地提示模型,让它们在回答你之前进行大量工作,并在返回文本之前参考其他评判模型的意见。鉴于这些模型目前确实具备一些真实技能,有一种简单的方法可以让长文写作变得更好。
There are a lot of criticisms of AI writing, but most of them are focused on more creative, high-voice writing like this blog. Those — including my own piece — often argue that it is because good writing is high-voice, has a point of view, has a deep human expression that needs to come across, and or a process of thinking that you peek into with the chosen words. As LLMs get more refined as tools, rather than conversational assistants, I think we are actually going backwards on our goals of having models produce inspiring writing.
On the other side of things is non-fiction writing. Filler, copy text was one of the genuinely useful abilities of an LLM (Sam Altman said so much about the early business of GPT-3 on a recent podcast). It has seemed like any flaws here were mostly down to a general lack of intelligence in the models, or some other training issue, and all non-fiction and explanatory text would get obliterated by the rapid pace of progress eventually. Having worked with the models as a writing assistant over the last few years, they’ve gotten a bit better, but it’s worth reflecting on what’s holding them back.
Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future. The models today struggle to organize and compellingly present some of the most established science in their area. This seems like a natural prerequisite that we should expect the models to master before they can solve broad, open-ended problems on their own. Until this is solved, the progress of LLMs for science will look closer to solving low-hanging fruit and merging distant connections across fields, rather than any sort of revolutionary insight.
This is a somewhat controversial take for someone who is very optimistic about AI’s progress, especially writing it on the day that Anthropic published a blog post on Claude making some progress on the famous Riemann Hypothesis. Scientific problems have a vast breadth, and I don’t think current AI models have as much coverage as many think.
Organizing knowledge is a compression. This compression is needed to make insight. Today’s LLMs increase entropy in long-form non-fiction writing, and I don’t see how that can be stacked on top of itself endlessly. They’ll be reliant on humans acting as sort of guides.
I am still very optimistic about translation from these narrow forms of science, like the extreme advancements we’ve seen in math, into consistent, broader progress — LLMs are the most powerful assistants scientists have ever used. I first need to explain how observing the models work on such grounded, low-level knowledge problems in writing makes me see a surprising lack of generalization.
For more context, I just finished writing a post-training textbook, Reinforcement Learning from Human Feedback (buy on Manning or Amazon). I used LLMs in many ways to support this, from helping wrangle LaTeX formatting for equations, doing extensive copyediting, and creating diagrams for programming languages like TikZ (in LaTeX) or Python.
Why have models stagnated in writing quality?
I would’ve expected way more progress on non-fiction writing from the models. I almost thought I would look dumb publishing a non-fiction book in 2026, given how things looked in 2024. Today, some of the most famous models on writing ability are pretty old, examples include OpenAI’s big GPT 4.5 and Moonshot’s Kimi K2. In and around these releases, the models have gone from okay to superhuman at other tasks like coding and mathematics. Maybe a closer, but still imperfect, comparison is how the models went from incapable to decent at search and research tasks. The pace of progress on most other skills is steep, but writing well feels orthogonal to most of them. I do not think writing is just ignored, but rather it’s challenging and lacks good training data to specifically intervene on it.
There is certainly some low-hanging fruit for making AI models better at writing — such as specialized harnesses like Claude Code, prompts, and training environments that make models spend a lot more inference tokens on the output, but I don’t think these will have a multiplicative impact on ability. Writing well is a very hard task! It’s a shame that we haven’t unlocked inference-time scaling for one of the great intellectual pursuits. Regardless, writing seems very different than what the models are good at.1
Today, the models seem genuinely horrible at long-form technical writing. They can get a sentence right, but if you try and get them to write an entire chapter it’ll be a mix of sprinkled with confusing wording, muddled in its organization, and generally a bit off. They try to be too cute where they don’t need to be and in the process make random conceptual errors. The models in the near future will get much better at the small errors, especially as models get bigger — which allows them to hold more world knowledge — but I do not expect their ability to utilize it to transform.
For example, the GPT models have been incredible at finding typos and minor issues for a long time. I passed a near-final draft of my book as a PDF to GPT 5.5 Pro and it found deep, surprising minor typos across the manuscript that is 200-300 pages.
On the other hand, the Claude models have been much more useful as an editor. They have a lot more taste, tend to understand the mental model of the task better, and have more interesting suggestions to unstick the different forms of writer’s block.
The examples I’ve given above all have a sort of consistent theme. The models know how to check every unit of content, in this case usually a sentence or equation or figure, or make one, specific section where you are caught. With these skills, they don’t do a good job revisiting components and stringing them together as they make many additions on top of each other. It feels like a sort of irreducible compounding errors. We used to deal with these errors in math and code, but reflecting on it, RLVR has been a truly magical solution in reducing them.
Getting value out of current models as a writer
I’m willing to share that there are a few technical explanation sentences in my book that came from an AI model — well less than 1% — they’re there because I really loved them. I let myself consider including some AI tokens in the book, as it didn’t feel like cheating if I, as a true expert, felt that the sentence was what the reader needed. Especially in the editing process, where I had a very close eye on things and plenty of concern on if my book would ever be done with all the things I have going on, it was an extremely valuable path forward.
For example, I had a list of questions from my editor interspersed in a LaTeX file with a specific delimiter like \editor{}. I would have Claude Code navigate to each comment, print the context before and after, and let me know if it was an easy typo fix or something more nuanced. I would write a response — the text to insert — or ask Claude for suggestions before fixing it. Intellectually it is a very focusing process of editing, it was a fun way to improve the book. Sometimes phrases from Claude’s suggestions are what made it into the book.
It is definitely a slippery slope and when I accepted a few AI suggestions it was at the point where I was going through my second full-manuscript review. Emotionally the project felt completed but I had more work to do. Coming out of the textbook-writing process I so deeply appreciate the cut and dry rule I have for my writing on Interconnects to never use AI outputs in the content. It is way more fun to write in a way that is only you — high voice, valued so deeply for the process — but writing a standard reference is not really an activity known for being fun. I see why people turn AI tools into a crutch when most of their writing is just an output to fill space, rather than a means to an end. I am motivated to write voluminously to learn, to feel, and to express.
I am working through similar balances in my scientific work too. AI models are great for repetitive pieces of the paper, like drafting a related work or background section that you know by heart, but using them for the abstract, introduction, experiments, or conclusion is a shame. Those are where the story and soul of the work is communicated — it’s where you learn what your research is really about.
I am confident I created a lot more net value by being able to have AI models create and check my non-fiction writing work. They make writing equations trivial, can help refactor the repository, port between languages, and many other things. At the beginning, it was very fun, until I was a bit worn down by the length of the publishing process, watching the field move on.
For an example of why AI was crucial in this case, I had to maintain Markdown and LaTeX versions of my book simultaneously in two spots, as readers gave feedback on the web version and my Manning editorial team reviewed a forked copy. Without AI agents, syncing between the two of them would’ve easily taken me five times as long (and this task took tens of hours already).
Something intertwined with this story, which I stumbled upon when thinking about agents, is how your pace of understanding won’t increase by using agents. That understanding, in the form of intuition, taste, instinct, etc. is what will be valuable in the future. Using AI for non-fiction writing takes away from that progression. Doubly, if you weren’t already an expert you won’t be able to catch its flaws.
In my case, I felt such an urgency to dump the knowledge out of my brain onto the page that there were times that using the AI models was a worthy tool. Much of the motivation of my book was to have a single reference for important post-training methods like rejection sampling or character training, where very little exists on the web.
This textbook was so much of giving back to the community, that it was just such a win to complete it in any form, that I felt it was okay. I would’ve learned more and the product could’ve been marginally improved with more human effort, I am sure. The determining factor was that I felt like the book was going to be aged out by the time it was published, a fear of AI model’s capabilities on one side and how fast the field moves on the other.
This turned out to be really wrong? I’m very happy with the result and I’m more confident in its staying power now than when I started in 2024, as the models have so failed to live up to the hype in non-fiction writing.
Where technical writing goes from here
The models are incredible tools, they let you express knowledge in different forms. They’re wonderful for creating creative filler or background material — e.g. the first draft of slides whose real value is being a talking point for the teacher to lecture over — that let any knowledge be transformed from one medium to another.
There’s some subtle, early phase of writing a non-fiction or reference textbook that feels a bit closer to writing a high-voice blog post like this. When pushing through the early organization and the presentation of the core skeleton new knowledge is created. This is the part that takes insight, and the LLMs are far behind in being able to replace it.
The crux of the above paragraph and preceding section is that I would be happy if more of the world’s experts used AI models to write a tiny bit of their books in order to get more of their knowledge shared with the world. The problem is that you can only use AI models to save 10-20% of the effort today, and I don’t see that percentage becoming the majority anytime soon.
There’s also the social pressure, where people expect LLMs to be the best, personalized educators out there, so they think working on a book or educational content is pointless. I think some of these opinions are aging out, as there’s a massive dearth in the highest quality educational work — and there always has been. AI is great at manipulating said content into the form that suits the student, not creating the content from scratch.
In the meantime I feel that we are stuck in a frustrating local minimum, where AI models are going to on net reduce the average effort spent on non-fiction writing, but they could enable great expression. Fewer people will start and push through.
So, in 2-5 years I still expect the best textbooks to be heavily crafted by the human hand. I’m not sure after then, but that’s longer than many would’ve predicted, given just how much knowledge these models have and their structural propensity to stream it.
As for a conclusion on capabilities, the models are great in two contexts: 1) any truly verifiable domain and 2) when given a ton of context and making a small edit — like finding a bug or solving a very specific math problem or giving feedback — not generating prose in an open-ended manner. Long-form writing will definitely fall before creative writing, but it’s a strong tell that the models are not able to express the full extent of their knowledge in underspecified problems. As we try to push the models to be something like “geniuses in a datacenter” solving grand scientific problems, this seems like a fairly fundamental limitation.
had a great piece on why LLMs make good editors, while being bad writers too.
Part of this is in how people use the models. If you ask Claude Fable 5 in Claude Code (or the chat app, I’m sure too): “write me a great poem about a goldfish,” the model will quickly spew out an answer. I asked the model how it did this, and if it had a sort of scratchpad it wrote to and iteratively updated before returning something good, and it said no. It made a minimal plan in its reasoning tokens and then autoregressively generated a poem. It’s taking no advantage of inference-time scaling or approaching it like a hard task.
This is how most people surely use models for writing, and it’s no surprise the results are mediocre. The way to get the best results out of them would be to prompt the models very heavily, get them to work extensively before answering you, and reference other judge models’ opinions before returning you the text. There’s a simple way to make the long form better, given the models have some genuine skills right now.