2023年12月27日
我热爱阅读技术预测类文献,因为后见之明本身就是未来预测任务的训练数据。例如,这类文献中我最钟爱的作品之一便是阿瑟·C·克拉克的《未来的轮廓》。本文是我对利克莱德1960年(距今64年前!)《人机共生》一文的笔记摘录。在这篇论文中,利克莱德将计算视为一种根本性的智能增强工具。
利克莱德在此论证,从通往完全自动化(AI)的路径来看,“智能增强”(IA)阶段可能只是过渡期,但其持续时间仍足以值得深入思考与探讨。
他引用的那些在当时看来堪称快速进步的狭义AI与AGI(即那个时代的“通用问题求解器”[20])案例,如今已被证实是根本性偏离正确方向、误入歧途的尝试——当时这些工作基于手动编码知识(使用谓词逻辑),并运用逻辑推理规则与搜索来推导结论。如今,AI领域的大多数从业者仅将这些工作视为历史趣闻,它们并非该领域的“主干分支”,而是困在一条死胡同的特性分支。值得注意的是,当今被认为最有前景的方法(大语言模型)在当时不仅完全不具备计算可行性,更因缺乏数字化形式存储的数万亿token训练数据而无法实现。今天是否存在与之对等的技术局限?
空军那份预测“机器将在20年内独立解决具有军事意义的问题”的研究报告,如今看来令人忍俊不禁。有趣的是,“20年之遥”似乎已成为“无从知晓,遥遥无期”的代名词。平心而论,即便在64年后的今天,我也不敢断言我们已实现这一目标。计算机确实大幅提升了态势感知能力,但据我所知,涉及“军事意义”的决策制定仍完全属于人类计算的范畴。
利克利德的一个有趣观察是,在他关于日常计算任务的思维实验中,他大部分的“思考”其实并非真正的思考,而更多是机械、重复、可自动化的数据收集与可视化。正是这一观察让他得出结论:人类与计算机的优势和劣势是互补的;计算机可以承担繁琐的工作,而人类则负责思考性工作。这一范式主导了接下来的64年,直到最近(大约去年),计算机才开始以一种通用、可扩展且影响经济的方式,在“思考”领域取得突破。这种突破并非通过显式、严谨的谓词逻辑,而是通过隐式、软性的统计方式。由此迎来了大语言模型驱动的AI热潮。
接着,利克利德展望了用于智能增强的计算基础设施的未来。我十分欣赏他基于分时技术提出的“思维中心”构想,这在今天或许就是……云计算。不过,某些计算任务已经变得非常廉价,以至于转移到了本地消费级硬件上,例如我的笔记本电脑,就能胜任简单计算、文字处理等工作。虽然利用率很低,但这也没关系。
在“语言问题”一节中,利克利德探讨了如何设计更便于人类使用的编程语言。他提到了FORTRAN这类命令式编程语言,但随后也谈到人类并不擅长给出明确的指令,反而更擅长直接指定目标。或许可以设计出更符合这种原生方式的编程语言,这暗示了声明式编程范式(例如Prolog)。然而,64年后的今天,主流的编程范式在很大程度上仍然保持着简单和命令式的特点。Python或许是当今最流行的编程语言之一,它本质上就是命令式的(一种“改进版的FORTRAN”),但非常人性化,其读写方式类似于伪代码。
关于输入输出,利克利德显然倾向于一种交互模式:人类团队围坐在大屏幕前,与计算机协作共同绘制示意图。显然,利克利德心中所想的,感觉就像是一台大型多人 iPad。我认为这是一个重大的误判。类似的产品确实被制造出来过,但从未真正成为主流的计算范式。相反,在这篇文章之后的几十年里,文本才是王者。显示器在输出端占据了主导地位,但键盘和鼠标(!)在输入端占据了主导地位,并且在 64 年后的今天,很大程度上依然如此。移动计算时代将其改变为触控,但并非以人们想象的方式。利克利德所设想的那种多人可视化环境确实存在(例如 Figma 等?),但它们远非主流的交互形式。这个误判的根源是什么?我认为利克利德将他所熟悉的东西(铅笔和纸)照搬过来,并想象计算会镜像那种界面。而实际上,无论对计算机还是对人来说,更好的界面是键盘和鼠标。
利克利德反复提及计算的军事应用,我想在那个时代,这是人们最关心的话题。我认为,这同样是对计算将如何在社会中应用的一个误判。也许部分原因在于利克利德为政府工作,并且当时这项工作的许多资金可能也来源于此。计算当然已经改善了军事决策,但据我所知,其影响程度远低于我们在企业和消费领域所看到的。
在 I/O 部分,Licklider 还思考了如何让计算机适应人类界面,具体来说就是自动语音识别。在这方面,Licklider 对能力过于乐观,估计需要 5 年时间才能实现。结果呢!!!64 年!!!过去了,虽然语音识别程序已经很多,但它们远未达到足以成为人机交互主导计算范式的水平。事实上,就在两年前 Whisper 发布时,我们所有人都很兴奋。想象一下 Licklider 会如何看待这个现实。即便近年来质量有了显著提升,ASR 也远非完美,仍然会出错,无法很好地处理多人说话场景,并且也没有朝着成为主导输入范式的方向发展。
以我们今天的认知,有哪些“事后诸葛亮”的真相可以告诉当时的 Licklider?
- 关于智能增强将持续很长时间,以及“思维中心”的想法,你的方向是对的。
- 你所知道且正在发展的、用于思考的所有“AI”肯定会有有用的应用,但终将被淘汰。以今天的标准来看,“正确”的方法是你目前无法着手研究的。你首先必须发明互联网,并让计算机变得快得多。而且不是 CPU 那种快法,而是 GPU 那种。但大量用于机械性/重复性工作的计算确实会非常有用——正如你所想象的那样,成为人类大脑的延伸。
- 大多数编程仍然是命令式的,但会变得方便得多。
- 大多数 I/O 是输入端的键盘和鼠标,以及输出端的显示器,并且是单个人与单台计算机的个体行为,尽管它们通过虚拟方式联网。
- 计算的主要应用场景在企业端和消费端,军事用途少得多。
- 语音识别实际上需要 62 年,而不是 5 年,才能达到足以日常使用的质量水平。
当然,有趣的部分在于滑动时间窗口,假设时间上的平移不变性。想象一下你自己对未来的外推。再想象一下它的“事后诸葛亮”。练习留给读者了 :)
本文最初以推文形式发布,随后(非常手动地,我得找个更好的办法)转换成了这篇博文。
Dec 27 2023
I love reading technology prediction documents because the benefit of hindsight is training data for the future prediction task. For example, one of my all-time favorites of this genre is "Profiles of the Future" by Arthur C. Clarke. This post is some of my notes on "Man-Computer Symbiosis" by Licklider 1960 (64 years ago!) where Licklider imagines computing as a fundamentally intelligence amplification tool.
Here, Licklider argues that the period of "intelligence augmentation" (IA) may be transient on the path to full automation (AI), but still long enough to be worth thinking through and about.
His citations for what must have felt like rapid progress in both narrow AI and AGI (of that age, i.e. the "general problem solver" [20]) are today known to be false starts that were off track in a quite fundamental way, at that time based on a manual process of encoding knowledge with predicate logic and using production rules of logic and search to manipulate them into conclusions. Today, most of AI is only aware of all of this work as a historical curiosity, it is not part of the "master branch" of the field, it is stuck in a dead end feature branch. And notably, what is considered today the most promising approach (LLMs) were at that time not only completely computationally inaccessible, but also impossible due to the lack of training data of trillions of tokens in digitized forms. What might be an equivalent of that today?
The study by the Air Force, estimating that machines alone would be doing problem solving of military significance in 20 years time evokes a snicker today. Amusingly, "20 years away" seems to be a kind of codeword for "no idea, long time". Arguably, I'm not sure that we are there even today, 64 years later. Computers do a lot to increase situational awareness, but decision making of "military significance" afaik is still well within the domain of human computation.
An interesting observation from Licklider is that most of his "thinking" in a day-to-day computational task thought experiment is not so much thinking, but more a rote, mechanical, automatable data collection and visualization. It is this observation that leads him to conclude that the strengths and weaknesses of humans and computers are complementary; That computers can do the busy work, and humans can do thinking work. This has been the prevailing paradigm for the next 64 years, and it's only very recently (last ~year) that computers have started to make a dent into "thinking" in a general, scaleable, and economy-impacting way. Not in an explicit, hard, predicate logic way, but in an implicit, soft, statistical way. Hence the LLM-driven AI summer.
Licklider then goes on to imagine the future of the computing infrastructure for intelligence augmentation. I love his vision for a "thinking center" based on time-sharing, which today might be... cloud compute. That said, some computations have also become so cheap that they moved to local consumer hardware, e.g. my laptop, capable of simple calculations, word processing, etc. Heavily underutilized, but it's okay.
In "The Language Problem" section, Licklider talks about the design of programming languages that are more convenient for human use. He cites imperative programming languages such as FORTRAN, but also later talks about how humans are not very good with explicit instructions, and instead are much better at just specifying goals. Maybe programming languages can be made that function more natively in this way, hinting at the declarative programming paradigm (e.g. Prolog). However, the dominant programming paradigm paradigm today, 64 years later, has remained largely simple and imperative. Python may be one of the most popular programming languages today, and it is simply imperative (an "improved FORTRAN"), but very human-friendly, reading and writing similar to pseudo code.
On the subject of I/O, Licklider clearly gravitates to an interaction pattern of a team of humans around a large display, drawing schematics together in cooperation with the computer. Clearly, what Licklider has in mind feels something like a large multiplayer iPad. I feel like this is a major misprediction. Products like it have been made, but have not really taken off as the dominant computing paradigm. Instead, text was king for many decades after this article. Displays became dominant at the output, but keyboard and mouse (!) became dominant at the input, and mostly remain so today, 64 years later. The mobile computing era has changed that to touch, but not in the way that was imagined. Multiplayer visual environments like Licklider imagined do exist (e.g. Figma etc?), but they are nowhere near the dominant form of interaction. What is the source of this misprediction? I think Licklider took what he was familiar with (pencil and paper) and imagined computing as mirroring that interface. When a better interace was the keyboard and mouse, for both computers and people.
Licklider talks again and again about military applications of computing, I suppose that was top of mind in that era. I feel like this is, again, a misprediction about how computing would be used in society. Maybe it was talked about this way in some part because Licklider worked for the government, and perhaps a lot of the funding of this work at the time came from that source. Computing has certainly gone on to improve military decision making, but to my knowledge to a dramatically lower effect than what we see in enterprise and consumer space.
In the I/O section, Licklider also muses about adapting computers to human interfaces, in this case automatic speech recognition. Here, Licklider is significantly over-optimistic on capabilities, estimating 5 years to get it working. Here we are !!! 64 YEARS !!! later, and while speech recognition programs are plentiful, they have not worked nowhere near well enough to make this a dominant computing paradigm of interaction with the computer. Indeed, all of us were excited when just two years ago with the release of Whisper. Imagine what Licklider would think of this reality. And even with the dramatic improvements to the quality recently, ASR is nowhere near perfect, still gets confused, can't handle multiple speakers well, and is not on track to a dominant input paradigm.
What would be the "benefit of hindsight" truths to tell Licklider at this time, with our knowledge today?
- You're on the right track w.r.t. Intelligence Augmentation lasting a long time. And "thinking centers".
- All of "AI" for thinking that you know and is currently developing will cerainly have useful applications, but will become deprecated. The "correct" approach by today's standards are impossible for you to work on. You first have to invent the Internet and make computers a lot faster. And not in a CPU way but in a GPU way. But a lot of computing for the rote/mechanical will indeed be incredibly useful - an extension of the human brain, in the way you imagine.
- Most of programming remains imperative but gets a lot more convenient.
- Most of I/O is keyboard and mouse at I, and display at O, and is an individual affair of a single human with a single computer, though networked together virtually.
- Majority of computing is in enterprise and consumer, much less military.
- Speech Recognition will actually take 62 years instead of 5 to get a good enough quality level for causual use.
The fun part of this, of course, is sliding the window, making the assumption of translation invariance in time. Imagine your own extrapolation of the future. And imagine its hindsight. Exercise left to the reader :)
This article was first published as a tweet and then converted (very manually, I have to find a better way) into this post.