The Big Blob of Compute Hypothesis bilingual 译文
没想到 17 年就有如此具有前瞻性的笔记了,阅读后感觉对我很有帮助
The Big Blob of Compute Hypothesis|大团算力假说
中英对照全文翻译(Bilingual English–Chinese Translation)
| 项目 | 说明 |
|---|---|
| 原文标题 | The Big Blob of Compute Hypothesis |
| 作者与出处 | Dario Amodei,2017 年写于 OpenAI 内部(未正式发表,后经媒体披露) |
| 源文件 | Big_Blob_of_Compute.pdf(共 10 页,纯图片扫描版,无文字层) |
| 文档范围 | 该 PDF 只包含第 1–3 节。正文结束于第 3 节的收尾段落(末句提到“我讨论 AI 安全研究策略(第 6 节)时再回来谈”),PDF 中并未包含作者在开头预告的第 4、5、6 节 |
| 翻译方式 | 先逐页识别扫描图像、逐段转录英文原文,再逐段翻译为中文 |
| 阅读体例 | > 引用块 = 英文原文(含原文的拼写、标点与笔误,一律保留原样);紧随其后的普通段落 = 对应的中文译文;章节标题为中英并列(英文|中文) |
1. Introduction|1. 引言
I’ve considered several times writing up my overall perspective on AGI safety. The main reason I haven’t done so so far is that my view of the problem is pretty high-level -- it’s from a different genre than the kind of “framework-y” agendas that Paul and Eliezer have written, and in fact a big part of it is arguing that such frameworks operate on the wrong level of abstraction (although that said, Paul’s framework is broadly compatible with things that I think, whereas Eliezer’s framework is broadly incompatible). However, then I realized that the core belief behind my safety perspective is also central to my perspective on AGI timelines, so it seemed particularly important to write it up. Some caveats: this is not a precise hypothesis or one I’m super confident in, and I have additional, more specific, reasons for holding both my views on AGI timelines and my views on AGI safety. It’s more like a distillation of my high-level intuitions (informed by experience) about how intelligence probably works, that causes me to expect things to go a certain way by default, and makes me a priori highly skeptical of research directions that don’t confront or don’t seem to fully understand this picture. I’ll describe the view first in terms of just AI itself, and then later go into the (subtle) implications for safety.
我已经好几次考虑过把自己对 AGI 安全的整体看法写成文章。迄今为止没有这么做的,主要原因是,我对这个问题的看法相当高层——它和 Paul、Eliezer 写的那种“框架式”的研究纲领属于不同的体裁,而且其中很大一部分恰恰是在论证:这类框架所处的抽象层次是错的(话虽如此,Paul 的框架与我所想的那些东西大体上是相容的,而 Eliezer 的框架则大体上不相容)。不过后来我意识到,支撑我安全观的那个核心信念,同样也是我关于 AGI 时间线的看法的核心,所以把它写下来就显得格外重要。有几点需要说明:这并不是一个精确的假说,我自己对它也没有特别强的信心,而且我之所以持有这些关于 AGI 时间线和 AGI 安全的看法,还有另外一些更具体的理由。它更像是把我关于智能大概如何运作的高层直觉(这些直觉也受到经验的塑造)提炼出来,正是这些直觉让我默认预期事情会朝某个方向发展,也让我先验地对那些不面对、或似乎没有完全理解这幅图景的研究方向高度怀疑。我会先只就 AI 本身来描述这个看法,然后再谈它(微妙的)安全意涵。
A big division between people who are bullish vs bearish on AGI timelines is the question of whether raw compute or new algorithms are more important to AI progress. I’m definitely closer to the “raw compute” side, but I also think the question is imprecisely stated -- what counts as an algorithm, how exactly does the raw compute have to be applied, and what about things that aren’t either algorithm or compute (like design of environments)? Also, is lots of compute merely sufficient to get to AGI (i.e. it’s one path among several), or is it actually somehow necessary?
在 AGI 时间线上持乐观与悲观看法的人之间,一个重大分歧在于:AI 的进步更依赖原始算力,还是更依赖新算法。我肯定更靠近“原始算力”这一侧,但我也认为这个问题问得不够精确——什么才算算法?原始算力究竟必须以何种方式被使用?还有那些既不属于算法也不属于算力的东西(比如环境的设计)又该怎么算?另外,大量算力仅仅是充分能通向 AGI(也就是说它只是若干条路径之一),还是它其实在某种意义上必要?
My picture of the situation is what I call the Big Blob of Compute (BBOC) hypothesis. It says that creating any given intelligent behavior is mostly about providing a large, minimally structured mass of computational capacity, and then giving it shape and form via interactions with a rich environment and a training process that drives it towards behavior appropriate for the task at hand. These high-level knobs -- what is the training signal (or signals)?, what is the nature of the environment the agent interacts with?, how much experience is it exposed to? -- constitute the most effective ways of controlling how well an intelligent agent performs a given task and how specifically it performs it. Once these knobs are set, then learning to perform the task is mostly about the agent having enough compute and storage capacity to learn the required cognitive skills.
我对这一情形的图景,就是我称之为大团算力(Big Blob of Compute,BBOC)假说的东西。它说的是:要造出任何一种给定的智能行为,主要在于提供一大团只预设最低限度结构的计算能力,然后通过与丰富环境的互动、以及一个把它推向适合当前任务的行为的训练过程,来为它赋予形状与形式。这些高层旋钮——训练信号(或多种训练信号)是什么?智能体与之互动的环境具有怎样的性质?它接触到多少经验?——构成了控制一个智能体在给定任务上表现得多好、以及以多具体的方式完成该任务的最有效手段。一旦这些旋钮设定好,学习完成该任务主要就取决于智能体是否有足够的算力和存储能力来学会所需的认知技能。
A few algorithmic settings do matter, but they tend to fall into a few very simple categories -- they exploit simple symmetries of the world, they help to regularize and condition the compute, or they make the shape of the computational graph smoother and less awkward. More complicated algorithmic ideas rarely help, and in particular supposedly “fundamental” differences in algorithms (I’ll give examples below, like bayesian vs non-bayesian methods, model-based vs model-free RL, or even ML vs non-ML methods) usually end up mostly reindexing or redescribing the same underlying computations, shuffling around exactly the same processing in a kind of shell game, all while giving the surface appearance of doing something substantively different. Also, algorithms which attempt to rigidly prescribe what the compute should do, or how it should be divided up among alleged (according to humans) “pieces” of the task, tend to do particularly poorly, often leading to much worse results compared to an unstructured compute blob of the same size. More broadly, algorithms matter much more in a negative sense than a positive sense -- it’s possible to come up with horrible architectures that totally block the flow of compute, but it’s really hard to beat the unstructured blob by all that much, and it’s also really hard (perhaps impossible) to use algorithmic changes to get the agent to do anything other than what’s implied by its environment and training process.
确实有一些算法上的设定是重要的,但它们往往落入几个非常简单的类别——利用世界的一些简单对称性,帮助对算力进行正则化和条件化,或者让计算图的形状更平滑、更不别扭。更复杂的算法想法很少有用,尤其是那些所谓“根本性”的算法差异(我下面会举例子,比如贝叶斯方法与非贝叶斯方法、基于模型的强化学习与无模型强化学习,甚至 ML 方法与非 ML 方法),最后大多只是对同一批底层计算重新编号或换一种方式描述,用某种掉包戏法把完全一样的处理搬来搬去,与此同时表面上却显得像是在做某种实质不同的事。此外,那些试图硬性规定算力该做什么、或者该如何在人类所谓的任务“部件”之间进行划分的算法,往往表现特别差,比起同等规模的、没有预设结构的一团算力,结果常常要差得多。更广泛地说,算法在负面意义上的重要性远大于正面意义——想出一套彻底阻断算力流动的糟糕架构是可能的,但要想大幅超越没有预设结构的那一团算力真的很难,而且想通过算法上的改动让智能体做出任何超出其环境和训练过程所隐含之事的行为,也同样很难(也许根本不可能)。
This picture informs my view of both AI timelines and AI safety. Since the summary above is somewhat abstract, I’ll first give (section 2) a more detailed formulation of BBOC, and (section 3) examples of where it’s been true or where I predict it’s going to be true. Then I’ll (section 4) describe the perspective on safety that it suggests, and (section 5) describe why that leads to skepticism of MIRI’s HRAD agenda and mild skepticism of “framework”-style AI safety plans in general. Finally, I’ll describe (section 6) what I think a BBOC-compatible approach to safety should look like, and how that relates to what the safety team is currently doing at OpenAI.
这幅图景同时塑造了我对 AI 时间线和 AI 安全的看法。由于上面的概述有些抽象,我会先(第 2 节)给出 BBOC 的更详细表述,然后(第 3 节)举出它曾经成立、或我预测它将会成立的例子。接着我会(第 4 节)描述它暗示的安全视角,并(第 5 节)说明为什么这会导致我对 MIRI 的 HRAD 议程持怀疑态度,并对一般意义上的“框架”式 AI 安全方案持温和的怀疑。最后,我会(第 6 节)描述我认为一种与 BBOC 相容的安全路径应该是什么样子,以及它与 OpenAI 安全团队目前在做的事情有何关联。
2. More Detailed Formulation of BBOC|2. BBOC 的更详细表述
Suppose you want to learn and then repeatedly execute some task, like recognizing images, playing video games, doing science, proving theorems, imitating a human, or even the skill of learning new things quickly. How do you do this? The BBOC hypothesis says that you should worry mostly about the following 7 things; everything else usually has a minor impact or is operating at too low a level of abstraction to be properly general:
假设你想学会某项任务,然后反复执行它——比如识别图像、玩电子游戏、做科学研究、证明定理、模仿人类,甚至是“快速学会新事物”这项技能。你会怎么做?BBOC 假说认为,你主要应该关心下面这 7 件事;其他一切通常只产生次要影响,或者所处的抽象层次太低,不足以真正具备通用性:
- Compute: You need enough computation to learn the task at hand. You need the computation to flow smoothly and freely between observations of the environment and the agent’s actions. It should be relatively unstructured and responsive to the training process; aside from that its exact form is secondary.
1. 算力: 你需要足够的计算量来学会手头的任务。你需要让计算在环境的观测与智能体的动作之间顺畅、自由地流动。它应当相对没有预设结构,并且能对训练过程作出响应;除此之外,它的具体形式是次要的。
- Parameters: The agent needs a mechanism to persistently store information that it learns about the general task during training, and transiently store information that it learns about an instance of the task during execution. This information can be implicit or explicit, directly stored observations or the parameters of a statistical model. (In simple deep neural nets the persistent and transient information are the weights and the activations; for something like evolution it’s the genome of a class of organisms and the synaptic weights of a particular brain.) If you don’t have enough parameters, you won’t learn the task well.
2. 参数: 智能体需要一种机制,在训练期间持久地存储它学到的关于通用任务的信息,并在执行期间暂时地存储它学到的关于某个具体任务实例的信息。这些信息可以是隐式的或显式的,可以是直接存储的观测,也可以是某个统计模型的参数。(在简单的深度神经网络中,持久信息和暂时信息分别是权重和激活值;而对于像进化这样的过程,它们则是一类生物的基因组和某个特定大脑的突触权重。)如果你没有足够的参数,就无法很好地学会任务。
- Quantity of experience: The agent needs enough experience acting in the environment where the task is to be performed. I use the word experience instead of data because static data is a special case of interaction with a dynamic environment. Experience can to some extent be traded off against compute if the agent can internally simulate some aspects of experience.
3. 经验的数量: 智能体需要在任务将要执行的环境中积累足够多的行动经验。我使用“经验”而不是“数据”这个词,是因为静态数据只是与动态环境交互的一种特例。如果智能体能在内部模拟经验的某些方面,那么经验在一定程度上可以与算力相互权衡取舍。
- Distribution of experience: The distribution of experience during learning must be such that the simplest way to do well on the training tasks also does well on an instance of the final task. One way to achieve this is if training tasks instances are drawn i.i.d. from the same distribution as test task instances, but it also works if the training tasks vary in such a way that it takes less bits to find the high-level commonalities between them (say the laws of physics, or some other small description length program) than it does to separately solve each task. Then, even if the test task is very different from the training tasks, the agent will do well as long as the high-level commonalities are preserved. Separately from this, it may be necessary for the training tasks to be simpler than the final task, or to increase smoothly in difficulty.
4. 经验的分布: 学习期间经验的分布必须做到:在训练任务上表现良好的最简单方式,也能在最终任务的某个实例上表现良好。实现这一点的一种方式是让训练任务实例与测试任务实例从同一分布中按 i.i.d. 抽取,但如果训练任务的变化方式使得找出它们之间的高层共性(比如物理定律,或者某个描述长度很短的其他程序)所需的比特数,少于分别求解每个任务所需的比特数,那么同样可行。这样一来,即使测试任务与训练任务非常不同,只要这些高层共性得以保留,智能体就能表现良好。除此之外,训练任务可能有必要比最终任务更简单,或者难度要平滑地递增。
- Normalization and conditioning: Freely flowing numerical computation can very easily have numerical stability or conditioning issues. Many of the “algorithmic” innovations in the last 5 years of AI have actually been simple methods for normalizing and conditioning numerical computation: consider Adam, BatchNormalization, or ResNets. These methods basically just say that numerical quantities that occur in some aspect of the computation should be normalized to 1 or reparameterized as a diff from 1. This simple “flow control” has often outperformed much more sophisticated ideas either within or outside deep learning.
5. 归一化与条件化: 自由流动的数值计算非常容易出现数值稳定性或数值条件问题。过去 5 年 AI 领域许多“算法上”的创新,实际上都是对数值计算进行归一化与条件化的简单方法:想想 Adam、BatchNormalization 或 ResNets 就知道了。这些方法基本上只是在说,计算某个环节中出现的数值量应当被归一化到 1,或者被重新参数化为相对于 1 的偏差。这种简单的“流量控制”常常胜过深度学习内外那些复杂得多的想法。
- Symmetries and shape: Another huge chunk of the most impactful algorithmic innovations has been the use of symmetry to decrease the number of parameters in models, and thus make them learn much faster. Intelligence only works at all because we live in a low entropy world; one way the low entropy manifests is symmetries and sparsity in the physical world. For example, convolutional nets take advantage of spatial translation invariance; RNNs take advantage of time translation invariance (memory architectures might be a future candidate in that they take advantage of general sparsity). We could likely learn the same tasks without these symmetries, but it would take a lot more compute and parameters (which would then eventually discover the symmetries on their own). In a similar vein, the shape of the computation can matter -- we don’t want information to bottleneck at a very small layer, or be forced to compress observations before analyzing them. The shape of the computation should be flexible and sensible. As tasks get more general, symmetries may become less important, as they may not hold across tasks and we may also be able to train meta-learning architectures that experiment with various symmetries at the object level.
6. 对称性与形状: 最具影响力的算法创新中还有很大一部分,是利用对称性来减少模型中的参数数量,从而使模型学得快得多。智能之所以能起作用,完全是因为我们生活在一个低熵世界中;低熵在物理世界中的一种体现方式就是对称性与稀疏性。例如,卷积网络利用了空间平移不变性;RNN 利用了时间平移不变性(记忆结构可能是一个未来的候选,因为它利用了广义的稀疏性)。我们很可能也可以在没有这些对称性的情况下学会同样的任务,但那需要多得多的算力和参数(而这些算力和参数最终会自己发现这些对称性)。类似地,计算的形状也很重要——我们不希望信息在某个非常小的层处形成瓶颈,或者被迫在分析观测之前先压缩它们。计算的形状应当是灵活而合理的。随着任务变得更通用,对称性可能变得不那么重要,因为它们可能无法跨任务成立,而且我们也许能够训练出元学习结构,在对象层面尝试各种不同的对称性。
- Training target: The computation needs to be driven by the correct evaluative mechanism for the task at hand, which can be a loss function, a reward, a prediction error, a GAN loss, a fitness function, a human evaluation, an extended human interaction, or a process for combining any of the above. You will learn whatever the training process asymptotically incentivizes; if you have the wrong training process, you’ll learn the wrong task. This isn’t just a comment on safety: many deep learning papers today learn the wrong task: for example, dialog systems that attempt to match human responses in a conversation (the goal of human conversation isn’t to match the responses of some other human, but to accomplish whatever the human’s goal is in the conversation), non-metalearning systems that train on one task and then apply the results to another (the agent was not trained on the task of transferring to another task, therefore it won’t be optimal at that). Note that the “training target” is not necessarily a single objective function or reward function; it can be a complicated process with several parts that together drive towards some asymptotic condition. For example, the training target in RL from human feedback comes from combining two different learning processes and is something like “learn to act in the environment in such a way that short clips of your actions would be preferred by a human relative to short clips of other actions you could take”.
7. 训练目标: 计算需要由针对手头任务的正确评价机制来驱动,这种机制可以是损失函数、奖励、预测误差、GAN 损失、适应度函数、人类评价、长时间的人机交互,或者把上述任意几种结合起来的过程。训练过程渐进地激励什么,你就会学到什么;如果你的训练过程错了,你就会学错任务。这不仅仅是在谈安全:如今许多深度学习论文都在学错任务:例如,试图在对话中匹配人类回复的对话系统(人类对话的目标并不是匹配另一个人给出的回复,而是达成人类在对话中的目标),以及那些在一个任务上训练、然后把结果应用到另一个任务上的非元学习系统(智能体并没有接受过“迁移到另一个任务”这一任务的训练,因此它在这方面不会是最优的)。需要注意,“训练目标”并不一定是单一的目标函数或奖励函数;它也可以是一个由若干部分组成的复杂过程,这些部分共同把训练推向某种渐近状态。例如,来自人类反馈的强化学习中的训练目标,来自两种不同学习过程的结合,大致相当于“学会以这样的方式在环境中行动:你行动的一小段片段,相比于你可能采取的其他行动的一小段片段,会更受人类偏好”。
BBOC says something like “if you get 1-6 right, you will probably learn the behavior that 7 specifies if it’s learnable in a computationally tractable way. Things not included in 1-7 have a much lower likelihood of being important, although there are exceptions.”
BBOC 大致上是这么说的:“如果你把 1-6 做对了,那么只要 7 所指定的行为能够以计算上可处理的方式学会,你大概率就会学到它。没有包含在 1-7 里的东西,其重要性要低得多,尽管也存在例外。”
A related concept to BBOC (that may help with intuition) is what I call the “snowflake model” of intelligence. If you looked at the fractal structure of a snowflake, you might think that whoever made it did something impossibly intricate and difficult, but that building it piece by piece must somehow be possible because someone did it. In fact, both statements are false: the way to make a snowflake is not to think in terms of its pieces but to know the laws of physics (training target), have enough raw material (compute) and a large enough chamber (parameters), set the temperature, pressure, and humidity correctly (normalization, shape), and wait for long enough (quantity of experience). Furthermore, this is your only way to make snowflakes and your only leverage over their shape; trying to piece together a single one from little bits of ice is basically hopeless. For snowflakes you might call this the “big blob of snow” theory.
与 BBOC 相关的一个概念(可能有助于建立直觉)是我所称的智能的“雪花模型”。如果你去看一片雪花的分形结构,你可能会觉得制作它的人做了某种不可能做到的、极其繁复而困难的事,但既然有人做到了,那么一块一块地把它拼出来必定是可能的。事实上,这两种说法都是错的:制作雪花的方式不是从它的组成部分去思考,而是要知道物理定律(训练目标)、拥有足够的原材料(算力)和足够大的腔室(参数)、正确地设定温度、压强和湿度(归一化、形状),然后等待足够长的时间(经验的数量)。更进一步说,这是你制作雪花的唯一方式,也是你唯一能影响其形状的着力点;试图用一小块一小块冰拼出单独一片雪花基本上是没指望的。对于雪花,你也许可以把它称为“一大团雪”理论。
Ilya Sutskever expresses BBOC as “networks want to learn”.
Ilya Sutskever 把 BBOC 表述为“网络想要学习”。
It’s worth noting some exceptions or apparent exceptions. When training on a narrow version of a task (like playing a single atari game, rather than playing any video game), it’s often possible to hand-engineer a method for that task in particular, leading to apparent large algorithmic improvements and more prescribed structure to the compute, but this comes at the cost of brittleness and poor generalization to broader tasks (e.g. consider the flood of papers proposing small tweaks to RL on atari). This phenomenon of (often unknowingly) overfitting to narrow tasks creates the illusion of successful algorithmic innovation, while also making our agents appear brittle and overcomplicated, and masking the simple relationship between compute and results. I believe this gives a very distorted view of the field for people who mainly pay attention to small tasks and expect breakthroughs to appear there first. By contrast, experience with huge projects seems to cultivate the BBOC intuition[1], and I also believe that BBOC will become more obvious as we move to more general tasks (which will be algorithmically simpler to learn than more narrow tasks).
值得注意的是存在一些例外,或者说表面上的例外。当你在某个任务的窄版本上训练时(比如只玩某一款 Atari 游戏,而不是玩任何电子游戏),往往可以针对那个特定任务手工设计出一种方法,从而带来表面上很大的算法改进,也让算力被规定了更多的结构,但代价是脆弱性以及对更宽泛任务的糟糕泛化(例如,想想那一大批对 Atari 上的 RL 提出小改动的论文)。这种(常常是无意中)对窄任务过拟合的现象,制造出算法创新取得成功的幻觉,同时又让我们的智能体显得脆弱而过度复杂,并且掩盖了算力与结果之间的简单关系。我相信,对于那些主要关注小任务、并期待突破会首先在那里出现的人来说,这会让他们对这个领域产生非常扭曲的看法。相比之下,在巨型项目上的经验似乎会培养出 BBOC 直觉[1:1],我也相信,随着我们转向更通用的任务(这些任务在算法上会比更窄的任务更简单),BBOC 会变得更加显而易见。
An extreme case of the paragraph above is the case of tasks that are so narrow that you don’t need any parameters at all and a human can simply write a short program that solves them (or experiment with a small range of such programs). Examples of this might be sorting, graph navigation, or alpha-beta pruning. Here the method solves the narrow task really well, can be much more or less efficient depending on the algorithm you use, and generalizes perfectly within instances of that narrow task (for example, sorting a larger list of numbers), but has zero ability to generalize outside that task (i.e. responding well when you give it an algorithmic task other than sorting). Basically all AI work that came before ML was of this form -- pre-ML AI was mostly just programming focused on certain selected domains. This history obscures BBOC and makes it look like huge gains from algorithms are possible and that the right algorithms give perfect generalization. In reality, however, generalizing just a bit beyond these single algorithmic tasks requires searching over programs, which will inevitably be heuristic, will inevitably drag in learning/ML/parameters, and if the task is general enough (e.g. solve an arbitrary programming interview question) will require a lot of compute and will follow the normal rules of BBOC.
上一段的一个极端情形是这样一类任务:它们窄到根本不需要任何参数,人类可以直接写一个短程序来解决它们(或者在一小批这类程序里做实验)。这类例子可能是排序、图导航或 alpha-beta 剪枝。在这里,该方法能非常好地解决这个窄任务,其效率可以根据你所用的算法而高得多或低得多,并且在该窄任务的各个实例内部能完美泛化(例如,对更长的数字列表做排序),但它完全没有能力泛化到这个任务之外(也就是说,当你给它一个排序之外的算法任务时,它无法很好地应对)。基本上所有 ML 之前的 AI 工作都属于这种形式——前 ML 时代的 AI 基本上就是针对某些选定领域所做的编程。这段历史遮蔽了 BBOC,让它看起来像是算法能带来巨大收益,而且正确的算法能带来完美的泛化。然而在现实中,只要稍微超出这些单一的算法任务去泛化,就需要在程序空间上做搜索,而这不可避免地是启发式的,不可避免地会引入学习/ML/参数;并且如果任务足够通用(例如解决任意一道编程面试题),就需要大量算力,并遵循 BBOC 的常规规律。
As written, much of BBOC’s description may read like a common sense guide to designing deep learning systems, but most people have not thought carefully about how far this extends or what it implies about intelligence. Many things that people commonly say about deep learning, both within the field and outside it, are strongly contradicted by BBOC, and start to sound silly if you really believe BBOC. Below I give some examples about what matters and what doesn’t in AI training, that helps to make clear how surprising and far-reaching it is.
就目前写下的内容而言,BBOC 的许多描述读起来可能像是一份设计深度学习系统的常识性指南,但大多数人并没有仔细想过它究竟能延伸到多远,或者它对智能意味着什么。人们常说的关于深度学习的许多事情,无论是在这个领域之内还是之外,都与 BBOC 强烈矛盾;如果你真的相信 BBOC,这些话就开始显得很傻。下面我给出一些关于 AI 训练中什么重要、什么不重要的例子,这有助于说明 BBOC 是多么令人惊讶且影响深远。
3. Evidence and Examples of BBOC|3. BBOC 的证据与实例
Here are some things I’ve observed over 3 years of training AI systems and many years of studying the brain, that are either evidence for BBOC, or implications/predictions of BBOC:
以下是我在 3 年训练 AI 系统以及多年研究大脑的过程中观察到的一些事情,它们要么是支持 BBOC 的证据,要么是 BBOC 的含义/预测:
- The classic (and best known) example of BBOC is the use of convolutional neural nets for vision. In the years before 2012 there were elaborate attempts to design edge detectors, contour integration systems, segmentation systems, foreground-background analyzers, pose estimators, and so on, which people then tried to put together to make vision systems. These systems failed because they did not use enough compute (1), did not have enough parameters (2), organized the computation too rigidly (6), and did not have the right objective function (the pieces were trained on e.g. edge detection rather than identification of the final object) (7). CNN’s fixed all these problems. If you keep the elaborate systems but fix (1), (2), and (7), they can do passably on imagenet, but they are still inefficient, complicated, and generalize poorly.
- BBOC 最经典(也最为人熟知)的例子,就是把卷积神经网络用于视觉。在 2012 年之前的那些年里,人们做了大量精心设计的尝试,去构造边缘检测器、轮廓整合系统、分割系统、前景-背景分析器、姿态估计器等等,然后再试图把这些部件拼装成视觉系统。这些系统之所以失败,是因为它们用的算力不够(1),参数不够(2),对计算的组织方式过于死板(6),而且没有正确的目标函数(这些部件是用边缘检测之类的任务训练的,而不是用最终物体的识别任务训练的)(7)。CNN 把这些问题全部解决了。如果你保留那些精心设计的系统,但修好(1)、(2)和(7),它们在 ImageNet 上也能勉强做得不错,但仍然低效、复杂,而且泛化很差。
- Basically my entire experience with training AI systems lines up with the CNN case -- use the simplest architecture and algorithm you can, and get (1)-(7) right, and you’ll do well. If you don’t have enough compute, parameters, or data, nothing can save you. Speech, language modeling, translation, reinforcement learning, generative models, soon robotics -- they all seem to follow this pattern. It takes only a few words to state this fact, but the weight of experience, of training hundreds of models, trying hundreds of things, and seeing what works and what doesn’t, is hard to overstate.
- 基本上,我训练 AI 系统的全部经验都和 CNN 这个例子吻合——用你能找到的最简单的架构和算法,把(1)-(7)都做对,你就会做得很好。如果你的算力、参数或数据不够,那什么也救不了你。语音、语言建模、翻译、强化学习、生成模型,接下来是机器人——它们似乎都遵循这个模式。陈述这个事实只需几句话,但背后经验的分量——训练过数百个模型、尝试过数百种做法、亲眼看到什么行得通什么行不通——再怎么强调都不为过。
- When I was at Google I was involved in a project to see if neural nets are really essential to deep learning. We tried to replace the matrix+nonlinearity in each layer with some other kind of machine learning model (I am not allowed to say exactly what), basically making a “deep
”. We found that these models worked but consistently performed slightly worse than similarly-sized neural nets. Eventually I discovered that any of these custom layers could form a standard neural-net layer with a subset of its parameters if it wanted to, and that is what it had chosen to do; basically it had reparameterized the model we gave it to internally and surreptitiously form a neural net. In other words, the different layer literally caused rearrangement of the same computation that a neural net would have done. On one hand, this tells us that deep neural nets are pretty effective models (other models want to become them), on the other hand, it tells us that other models can work too and there’s nothing magic about “deep learning”; it’s more that certain computations really want to occur and will shape themselves around whatever type of model you give them.
- 我在 Google 的时候,参与过一个项目,想弄清神经网络对深度学习来说是否真的不可或缺。我们尝试把每一层里的矩阵运算+非线性换成另一种机器学习模型(具体是什么我不被允许说),基本上就是做一个“深度〈某种别的机器学习技术〉”。我们发现这些模型能用,但表现始终比规模相近的神经网络略差。最终我发现,这些自定义层中的任何一个,只要它愿意,都能用自己的一部分参数构成一个标准的神经网络层,而它选择做的正是这件事;基本上,它在内部把我们给它的模型重新参数化了,偷偷地变成了一个神经网络。换句话说,这个不同的层实际上只是把神经网络本来会做的同一份计算重新排列了一番。一方面,这告诉我们深度神经网络是相当有效的模型(别的模型都想变成它);另一方面,这也告诉我们别的模型同样能行得通,“深度学习”并没有什么神奇之处;更准确地说,是某些计算本身非常想要发生,并且会围绕你给它的任何一类模型来塑造自己。
- In RL textbooks there is a distinction between model-based RL and model-free RL. The first has an environment model that makes predictions and also a policy, the second has just a policy. In theory model-based RL is a bit like system 2 in that it can think ahead and plan whereas model-free RL is like system 1. In practice model-free RL tends to do some planning internally, and we're starting to suspect that if we simply shaped model-free policies differently, say allowing them to iterate themselves for many steps before taking an action, that they would be able to (implicitly) do planning at least as well as model-based systems. Meanwhile actual model-based RL systems perform relatively poorly, perhaps because we’re artificially forcing computation to take place in 2 disconnected blocks. The point is, the computational activity of planning is important, but whether or not that activity occurs has more to do with the shape of the blob of computation than with the algorithm we’re supposedly using or which neural net is supposed to do what. “Model-free” systems learn a model internally when their environment and objective function require them to, and if it’s sometimes a deficient model that may just be because we don’t tend to shape model-free networks so as to give them enough internal recurrence to implicitly do model-like computations.
- 在强化学习教科书里,基于模型的强化学习和无模型强化学习是有区分的。前者有一个做预测的环境模型,还有一个策略;后者则只有策略。理论上,基于模型的强化学习有点像系统 2,因为它能够前瞻和规划;而无模型强化学习则像系统 1。但在实践中,无模型强化学习往往会在内部做一些规划,而且我们开始怀疑:如果我们只是换一种方式来塑造无模型的策略,比如说允许它们在采取行动之前先自我迭代很多步,那么它们就能够(隐式地)做规划,且至少做得和基于模型的系统一样好。与此同时,实际的基于模型的强化学习系统表现却相对较差,也许是因为我们人为地强迫计算发生在两个互不连接的模块里。关键在于,规划这一计算活动很重要,但这种活动是否会发生,更多地取决于那一团计算的形状,而不是取决于我们自以为在用的算法,或者哪个神经网络应该负责哪件事。“无模型”系统在环境和目标函数要求它们这样做的时候,会在内部学到一个模型;如果它有时学到的是一个有缺陷的模型,那可能只是因为我们通常不会去塑造无模型网络,好给它们足够的内部循环,让它们隐式地做类似模型的计算。
- Similarly, ordinary RL is said to have issues with long time horizons, say planning over 1 million timesteps. Hierarchical RL is supposed to deal with this; the idea is to have goals and subgoals, and different policies set the goals at different timescales. Unfortunately so far hierarchical RL hasn't been able to set goals and subgoals much better than they can be learned implicitly within a normal RL algorithm; there's a tendency for the hierarchy to just collapse into approximately what a single (deeper) neural-net would have done. It's possible we'll eventually make hierarchical RL work, but for now it seems to be yet another example of compute having its own idea about what it wants to do and routing around the structure imposed on it by a supposedly important algorithm. Meanwhile, ordinary RL algorithms are having surprising success on long time horizons when we make the policies really large, sometimes solving tasks with only 1 reward every 10,000 timesteps. We really will eventually need to represent hierarchical structure within our RL policies, but that may happen implicitly rather than through the techniques we've labeled “hierarchical RL”.
- 同样,普通的强化学习据说在长时间跨度上有问题,比如要在 100 万个时间步上做规划。分层强化学习本该解决这个问题;其思路是设定目标和子目标,由不同的策略在不同的时间尺度上设定目标。遗憾的是,到目前为止,分层强化学习设定目标和子目标的能力,并不比在一个普通强化学习算法内部隐式学出这些目标好多少;分层结构有一种倾向,就是最终塌缩成大致等同于单个(更深的)神经网络会做的事情。我们有可能最终让分层强化学习奏效,但就目前而言,它似乎又是一个例子,说明算力对自己想做什么有它自己的主张,并会绕开那个据称很重要的算法强加给它的结构。与此同时,当我们把策略做得非常大时,普通的强化学习算法在长时间跨度上正取得令人惊讶的成功,有时能在每 10,000 个时间步只有 1 个奖励的情况下解出任务。我们最终确实需要在强化学习策略中表示分层结构,但那可能是以隐式的方式发生,而不是通过我们贴上“分层强化学习”标签的那些技术。
- LSTM’s are supposedly better than simple RNN’s because they can store state for longer -- one of the few real algorithmic innovations in deep learning. Yet there’s now evidence that simple RNN’s perform as well as LSTMs when properly tuned; the only advantage of LSTMs is that they work well for a wider range of hyperparameters. In fact, holding constant the number of parameters and the amount of compute, different recurrent architectures (RNN, LSTM, GRU) seem to have eerily similar capacity and performance.
- LSTM 据说比简单的 RNN 更好,因为它们能把状态保存得更久——这是深度学习中为数不多的真正算法创新之一。然而现在有证据表明,只要调参得当,简单的 RNN 表现和 LSTM 一样好;LSTM 唯一的优势在于它能在更宽的超参数范围内都工作良好。事实上,在参数数量和算力保持不变的情况下,不同的循环结构(RNN、LSTM、GRU)似乎具有惊人相似的容量和表现。
- When current neural nets don't generalize well people often talk about this as if it's some kind of algorithmic defect, as if there's some yet-to-be-discovered algorithm somewhere that generalizes drastically better[2]. But my experience is that neural nets generalize better when you simply expose them to a wider distribution of data that forces them to find a general solution rather than just memorizing a few special cases. For example, when training speech models I found that if you trained only on an American accent you would do poorly on many other accents. However if you train on 5 or 6 accents you will immediately do well on an unseen accent. What’s probably going on is that for 1 or 2 accents it's easy to just memorize the peculiarities of each one, whereas as the number increases as it becomes easier (and requires fewer bits) to identify the commonalities between them and store these, which leads to generalization. To summarize: it's not algorithms that lead to generalization, so much as the training setup and environment.
- 当目前的神经网络泛化不好时,人们常常把这件事说成是某种算法上的缺陷,好像某处还存在某个尚未被发现的算法,能大幅提升泛化能力[2:1]。但我的经验是:只要你让神经网络接触更广的数据分布,迫使它去找一个通用的解,而不是仅仅记住少数几个特例,它就会泛化得更好。例如,在训练语音模型时我发现,如果你只用美国口音训练,那么在很多其他口音上都会表现很差。但如果你用 5、6 种口音训练,你立刻就能在一个没见过的口音上表现良好。很可能发生的事情是:在只有 1、2 种口音时,直接记住每种口音的特性很容易;而随着口音数量增加,找出它们之间的共性并把共性存起来会变得更省事(而且需要更少的比特),这就带来了泛化。总结一下:带来泛化的与其说是算法,不如说是训练设置和环境。
- I sometimes hear discussion, in both the academic ML world and in the AI safety world, that deep neural nets are missing something and that Bayesian methods would do better in some way (they’d represent uncertainty, be better calibrated, it would be easier for us to understand their beliefs, etc). But in fact neural nets already do represent uncertainty (with e.g. class probabilities), and could even output a full joint distribution rather than factored probabilities if they wanted to, just like idealized Bayesian methods. The problem is this would take exponentially more parameters, in both the case of neural nets and some idealized Bayesian model; both run into the same practical limitation. People also say that if you had a Bayesian model, you’d understand explicitly what propositions it believes and how much, but this has never made sense to me -- the hard part is figuring out how to carve the world up into “statements”, and this has to be learned by whatever structure, Bayesian or not, you use to hold and fill your model (for example, a Bayesian network or a neural net). Once you’ve done that, the beliefs held in the nodes of a Bayesian network are just as likely to be comprehensible or incomprehensible, calibrated or uncalibrated, as the activations inside a neural network, the latter of which appear for all the world to be holding beliefs and weighing evidence even though they’re not explicitly labeled as doing so. Finally, people also say that Bayesian models are better at generalization and avoiding overfitting. This is true in the ideal limit (since you’re modeling a distribution), but in practice the simplest ad hoc methods that deep learning has improvised to deal with overfitting, like dropout, can be shown to be equivalent, to a Bayesian approximation. Meanwhile, people have actually tried explicitly Bayesian neural nets (with uncertainty on the weights), but they don’t stand out as as clearly superior in generalization or calibration; they behave similarly to methods like dropout, as you’d expect. Once again the key point here is: the algorithms per se matter less than they appear, and give the appearance of something different going on when in fact they often just relabel the same computational process.
- 我有时会在学术界 ML 圈子和 AI 安全圈子里都听到这样的讨论:深度神经网络缺了点什么,而贝叶斯方法会在某些方面做得更好(它们能表示不确定性、校准得更好、我们更容易理解它们的信念,等等)。但事实上神经网络已经能表示不确定性了(例如用类别概率),而且只要它们愿意,甚至可以像理想化的贝叶斯方法那样输出完整的联合分布,而不只是因子化的概率。问题在于,无论对神经网络还是对某个理想化的贝叶斯模型来说,这都会带来指数级增加的参数;两者都会撞上同样的现实局限。人们还说,如果有了贝叶斯模型,你就能明确知道它相信哪些命题、相信到什么程度,但这一点我从来没觉得说得通——真正困难的地方在于如何把世界切分成各种“陈述”,而这件事必须由你用来承载并填充模型的那个结构来学习,无论它是不是贝叶斯的(例如贝叶斯网络或神经网络)。一旦你做到了这一点,贝叶斯网络节点里所持的信念,其可理解或不可理解、校准良好或校准不佳的程度,和神经网络内部的激活值一样;而后者看上去完全像是在持信念、在权衡证据,尽管它们并没有被明确标注为在做这件事。最后,人们还说贝叶斯模型更擅长泛化、更能避免过拟合。这在理想极限下是对的(因为你是在对分布建模),但在实践中,深度学习为应对过拟合而临时想出的最简单的那类特定方法,比如 dropout,可以被证明等价于一种贝叶斯近似。与此同时,人们确实尝试过显式的贝叶斯神经网络(在权重上带有不确定性),但它们在泛化或校准上并没有明显地更胜一筹;它们表现得和 dropout 之类的方法差不多,正如你所预料的那样。这里的关键点再一次是:算法本身的重要性比它们看起来的要低,它们会给人造成一种“有某种不同的事情在发生”的印象,而实际上它们往往只是给同一个计算过程换了个标签。
- People seem very surprised about adversarial examples -- they say it means neural nets don’t really understand the objects they’re classifying, or are brittle in some way. But BBOC isn’t surprised by this: the classifiers weren’t trained on a distribution of experience that included images engineered in this way. In line with BBOC, the current best defense against adversarial examples is to train on them directly. The current second best defense (and the thing I think will ultimately be the solution) is to have a generative model that tells you what is or isn’t a natural image; this works because the generative model was trained with an objective function that tells it to distinguish what is on the data manifold versus what isn't, whereas the classifier had no such objective function. BBOC says you get what you optimize for; what you don't get is a human's idea of what properties the system should have.
- 人们似乎对对抗样本感到非常惊讶——他们说,这意味着神经网络并没有真正理解它们所分类的物体,或者意味着它们在某种意义上是脆弱的。但 BBOC 对此并不惊讶:这些分类器训练时所用的经验分布中,并不包含以这种方式构造出来的图像。与 BBOC 一致,目前对付对抗样本的最佳防御办法就是直接在对抗样本上训练。目前第二好的防御办法(也是我认为最终会成为解决方案的办法)是拥有一个生成模型,由它告诉你什么才是、什么不是自然图像;这之所以有效,是因为这个生成模型训练时所用的目标函数要求它区分什么在数据流形上、什么不在,而分类器没有这样的目标函数。BBOC 说的是:你优化什么,就得到什么;你得不到的,是人所设想的、这个系统应当具备的那些性质。
- News articles occasionally talk about how we need some alternative to backprop, as if we’re missing some key algorithmic insight, some new kind of update rule that will revolutionize everything we do and allow our networks to finally learn in a truly human-like way (“hebbian learning” is often mentioned). It's quite possible we could find a new update rule (in fact I suspect many gradient-like updates would be comparably good) but I don't expect any such change to have anything more than a very modest effect; navigating high-dimensional spaces is hard and empirically we’ve tried a huge number of things that end up doing only modestly better than the simple gradient. Ascribing this almost totemic importance to the update rule is an example of the kind of thinking BBOC is designed to counteract.
- 新闻文章偶尔会谈到我们需要某种替代反向传播的东西,就好像我们缺了某个关键的算法洞见,缺了某种新的更新规则,它将彻底改变我们做的一切,并让我们的网络最终以真正类似人的方式学习(“赫布学习”经常被提到)。我们完全有可能找到一种新的更新规则(事实上我怀疑许多类似梯度的更新会同样好),但我不指望这类改变会带来超过非常有限的影响;在高维空间中寻路本来就很难,而经验上我们已经尝试过大量方法,最终都只比简单的梯度略好一点。把这种近乎图腾般的重要性赋予更新规则,正是 BBOC 旨在抵消的那类思维方式的一个例子。
- For a while people were doing transfer learning by training a policy on a bunch of tasks and then trying to fine tune on a new task; people tried a bunch of complicated variants of this that all worked about equally poorly. This is actually the wrong training procedure (it doesn’t asymptotically lead to fast-adapting policies) and things started working somewhat better when people switched to the right objective function: training an initial policy optimized for its performance after fast adaptation (meaning, learning and evaluation on the task of adapting quickly to a new environment). Again, the detailed algorithm didn’t make much difference; what seems to help is having the right target.
- 有一段时间,人们做迁移学习的方式是:先在一堆任务上训练一个策略,然后尝试在新任务上微调;人们试过一堆复杂的变体,但效果都差不多地差。这实际上是错误的训练流程(它并不会渐进地带来能快速适应的策略),而当人们换用了正确的目标函数后,情况才开始有所好转:训练一个初始策略,并针对它快速适应之后的表现在做优化(也就是说,学习和 在“快速适应新环境”这一任务上的评估)。同样,具体的算法并没有带来多大差别;看起来真正有帮助的是拥有正确的目标。
- There are a lot of algorithms out there that attempt to get better exploration for reinforcement learning, or better intrinsic motivation, or that attempt to explore more safely, etc. A lot of smart people (including me) have worked on these problems, and come up with algorithms that seem to get improvements, but over time these have become a messy and unwieldy patchwork, somewhat reminiscent of the pre-CNN edge detectors. It’s easy to think we’re beyond that because these methods use end-to-end deep RL, but maybe we need to go up one level of abstraction, learning over many agent learning processes how to explore, how to motivate, and how to explore safely. We’ve seen early progress on this type of “metalearning”, but the jury is still out on where things will ultimately go.
- 有大量算法试图为强化学习获得更好的探索、或更好的内在动机,或者试图更安全地探索,等等。很多聪明人(包括我)都研究过这些问题,并提出了一些看起来能带来改进的算法,但随着时间推移,这些方法已经变成了一团杂乱而笨重的拼凑物,多少让人想起 CNN 出现之前的边缘检测器。我们很容易认为我们已经超越了那个阶段,因为这些方法使用的是端到端的深度 RL,但也许我们需要上升一个抽象层次,在许多个智能体的学习过程之上,去学习如何探索、如何产生动机、以及如何安全地探索。我们已经在这一类“元学习”上看到了早期的进展,但事情最终会走向何方,目前还没有定论。
- What about non-ML systems, like search or reasoning systems? Aren’t they doing something genuinely different, and wouldn’t you expect them to generalize better because they contain a symbolic recipe, while ML systems don’t? What I believe is going on here is a misunderstanding -- ML systems are perfectly capable of driving symbolic computation, as all the work in neural-net driven theorem proving shows, and they can also search over algorithms or programs based on how they fit the data (you can use neural nets to evolve code, though not well yet). In fact I think neural net systems exposed to a wide variety of environments with only high-level similarities (e.g. the laws of physics) will, if their shape and structure is organized in the right way, inevitably converge on parameters that implement simple programs (e.g. they will derive or represent the laws of physics) -- our brains did this despite being based on very fuzzy computation. Once that happens, neural nets will be searching over programs whereas non-ML methods represent fixed programs (or small families of programs, like tree search); at that point non-ML methods will just be doing a fixed, rigid subset of the computation done by ML methods[3]. So we can expect that ML systems will do much better than non-ML systems (even at symbolic tasks) while using more computation to do so, in line with BBOC. We can already see this with MCTS -- DeepMind has published some papers showing that learning to search can outperform fixed tree search algorithms, and it’s only a matter of time until this is incorporated into e.g. AlphaGo.
- 那么非 ML 系统呢,比如搜索或推理系统?它们难道不是在做着某种真正不同的事情吗?而且既然它们包含一个符号化的配方,而 ML 系统没有,你难道不会预期它们泛化得更好吗?我认为这里发生的事情是一种误解——ML 系统完全有能力驱动符号计算,所有由神经网络驱动的定理证明工作都表明了这一点;它们也可以根据算法或程序与数据的契合程度来搜索算法或程序(你可以用神经网络来演化代码,尽管目前还做不好)。事实上我认为,如果神经网络系统的形状和结构以正确的方式组织起来,那么当它接触到各种各样仅在高层面上相似的环境(例如物理定律)时,就必然会收敛到实现了简单程序的参数上(例如,它们会推导出或表示出物理定律)——我们的大脑尽管基于非常模糊的计算,也做到了这一点。一旦这成为现实,神经网络将在程序空间中进行搜索,而非 ML 方法表示的是固定的程序(或很小的程序族,比如树搜索);到那时,非 ML 方法就只是在做 ML 方法所做计算的一个固定而僵硬的子集[3:1]。因此我们可以预期,ML 系统会做得比非 ML 系统好得多(即使在符号任务上也是如此),同时为此使用更多的计算,这与 BBOC 一致。我们已经能在 MCTS 上看到这一点——DeepMind 发表了一些论文,表明学习如何搜索可以胜过固定的树搜索算法,而这被纳入例如 AlphaGo 只是时间问题。
- Consider human evolution. How did nature manage to discover and develop brains capable of uncovering the laws of physics, proving Fermat’s last theorem, and composing symphonies? You can trace step by step how it happened, and give explanations for each step, yet the process as a whole is still surprising in its power. BBOC offers a simple high-level explanation -- enough compute, enough minimally structured brain memory, a broad enough distribution of environment experience and challenges, and a driving objective that involved playing agents off against each other were sufficient to make the result inevitable eventually.
- 想想人类的演化。大自然究竟是如何发现并发展出能够揭示物理定律、证明费马大定理、创作交响乐的大脑的?你可以一步步追溯它是如何发生的,并为每一步给出解释,然而整个过程就其力量而言仍然令人惊讶。BBOC 提供了一个简单的高层解释——足够的算力、足够多只预设最低限度结构的大脑记忆、足够宽广的环境经验与挑战的分布,以及一个让智能体相互对抗的驱动性目标,这些就足以让这一结果最终成为必然。
Hopefully the above gives some flavor of why I see BBOC as so likely, despite seeming strange at first glance and not being obvious from looking at small-scale benchmarks.
希望以上内容能让人大致体会到我为何认为 BBOC 如此可能,尽管它乍看之下显得古怪,也不容易从观察小规模基准测试中看出来。
Note that none of this implies that scaling up today’s exact algorithms will lead to AGI. It does suggest that any algorithmic changes will likely be simple things that exploit symmetries or sparsity (good candidates might be memory-based models, or neural nets that grow in size over the course of training), or innovations in the training process, and that scale will be a huge part of the story. It also suggests that “alternatives to deep learning” are at very high risk of either not working or just recapitulating deep learning, and that algorithmic innovations said to introduce some desirable property should be treated with skepticism.
请注意,这一切都不意味着把今天的算法原样扩大规模就会通向 AGI。它确实表明,任何算法上的改动都很可能是利用对称性或稀疏性的简单东西(不错的候选者可能是基于记忆的模型,或者在训练过程中不断增大的神经网络),或者是训练过程中的创新,而规模将是整个故事中的极大部分。它还表明,“深度学习的替代方案”极有可能要么根本行不通,要么只是在重演深度学习;而对于那些号称引入了某种理想特性的算法创新,也应当持怀疑态度。
A striking feature of BBOC is that a research direction may look perfectly reasonable for a long time (visual edge detectors, model-based RL, hierarchical RL, exploration) and have many of the smartest people working on it, before it becomes clear that it’s simply rearranging compute or trying to be too prescriptive about how to solve a problem. Over time, however, you develop an intuition for when an approach is “too handcrafted”, similar to a programmer’s intuition about when an abstraction is insufficiently general or modular. Model-based, hierarchical, and exploration are still controversial -- maybe we’ll still get something out of them -- but they are starting to look increasingly suspicious, increasingly like the edge detectors of 2011.
BBOC 的一个显著特点是,某个研究方向可能很长时间里都看起来完全合理(视觉边缘检测器、基于模型的强化学习、分层强化学习、探索),而且有许多最聪明的人在研究它,直到后来才看清楚,它不过是在重新安排算力,或者是在如何解决问题上规定得过于细致。然而随着时间推移,你会逐渐培养出一种直觉,判断某种方法何时“手工设计的痕迹过重”,就像程序员对某个抽象是否足够通用或模块化所拥有的直觉一样。基于模型的强化学习、分层强化学习和探索仍然有争议——也许我们还能从它们身上得到些什么——但它们正开始显得越来越可疑,越来越像 2011 年的边缘检测器。
A final note is that sometimes it’s hard or impossible to take the BBOC-recommended approach to a problem right away. Sometimes you need to start with something lower-level on a narrower problem (e.g. start with ordinary deep RL on one video game before you try meta-RL on many video games) to get your bearings and understand how to design the higher-level algorithm. So narrower projects where you design lower-level architectures can be useful as stepping stones, something I’ll come back to when I discuss my research strategy on AI safety (section 6).
最后一点是,有时候很难、甚至不可能立刻就采用 BBOC 所推荐的方法来处理某个问题。有时候你需要先从更窄的问题上、更低层次的东西入手(例如先在一个电子游戏上做普通的深度强化学习,然后再尝试在许多电子游戏上做元强化学习),以便找准方向、理解该如何设计更高层次的算法。因此,那些设计较低层次架构的较窄项目可以作为有用的踏脚石,这一点我会在讨论我在 AI 安全方面的研究策略时(第 6 节)再回来谈。
译者注|Translator's Notes
1. 关于超链接。 原文中下列技术名词带有超链接样式(下划线):Adam、BatchNormalization、ResNets、RL from human feedback、now evidence、can be shown to be equivalent、Bayesian neural nets、performance after fast adaptation、neural-net driven theorem proving shows。该 PDF 为打印/截图式导出,未保留可解析的链接地址,因此译文中一并按普通文本处理,未添加链接。
2. 关于原文笔误。 英文原文(转录部分)中存在的重复词、多余标点等印刷瑕疵一律照原样保留,例如:stand out as as clearly superior、equivalent, to a Bayesian approximation、whereas as the number increases as it becomes easier、new update rule (in fact ... good) but。中文译文按作者显然想表达的意思翻译,不做“纠正后更漂亮”的改写。
3. 关于脚注。 原文共 3 条脚注(第 1 条在第 5 页底部,第 2 条在第 7 页底部,第 3 条在第 9 页底部),正文中的上标标记为 ¹、²、³,本译文中统一改用 Markdown 脚注标记 [^1]、[^2]、[^3],三条脚注均照译。
4. 关于分段。 原文正文在页与页之间连排,有 6 处段落被页边界切断(例如第 4 页末尾的 “By contrast, experience with”与第 5 页开头的 “huge projects ...”属于同一段)。本译文已按语义把跨页断开的段落合并为同一段,因此段落划分与 PDF 的物理分页不完全一一对应。
5. 术语处理。 缩写与专有名词(BBOC、AGI、RL、ML、CNN、LSTM、RNN、GRU、GAN、MCTS、HRAD、MIRI、OpenAI、Google、DeepMind、Ilya Sutskever、Adam、Atari、ImageNet 等)保留英文原形;关键术语首次出现时在中文译文中给出对应表达,全文术语保持统一(如 compute → 算力,agent → 智能体,environment → 环境,training signal → 训练信号,blob → 一团算力)。
My first job in AI was working on what turned out to be what I now believe was the single largest use of compute in one training process up to that time, Deep Speech 2, which used about 0.3 petaflop-days per trained model in 2015.
我在 AI 领域的第一份工作所参与的项目,后来成了我如今所认为的、截至当时单次训练过程中规模最大的一次算力使用,那就是 Deep Speech 2;2015 年时,每训练一个模型大约要消耗 0.3 petaflop-day 的算力。 ↩︎ ↩︎
I believe this comes from the experience with pre-ML AI systems, where picking the right algorithm meant you could generalize perfectly within a single algorithmic task, basically because this kind of “AI” was just programming. I discuss this a bit in the previous section.
我认为这来自机器学习之前的 AI 系统所带来的经验:在那些系统里,选对了算法就意味着你能在单个算法任务内完美泛化,这基本上是因为这类“AI”不过就是编程而已。我在上一节里对此略有讨论。 ↩︎ ↩︎
This is also discussed in the section above: early AI work focused on fixed programs or small families of programs that solve specific narrow problems; calling this work “AI” obscures BBOC and erroneously leads us to think there’s something different about symbolic computation.
这一点在上一节里也有讨论:早期的 AI 工作聚焦于解决特定窄问题的固定程序或规模很小的程序族;把这类工作称为“AI”会遮蔽 BBOC,并错误地让我们以为符号计算有什么不同之处。 ↩︎ ↩︎

浙公网安备 33010602011771号