翻译[11]-基于神经网络的噪音抑制
RNNoise:基于神经网络的噪音抑制
- 原文地址: [https://jmvalin.ca/demo/rnnoise/]
- 原作者: Jean-Marc Valin(Xiph.Org / Mozilla,Opus 与 Speex 作者之一),2017 年 9 月 27 日
- 许可: (C) Copyright 2017 Mozilla and Xiph.Org
- 存档时间: 2026-10-02(本地 SingleFile HTML 经 markitdown 转换整理)
译者注:本文是安卓扩音器工程 AudioSuitZulu(Kotlin,见
[https://github.com/qsbye/audio-suit-zulu])消噪模块的设计参考。工程
.trae/specs/silencer-anc/spec.md中明确借鉴了本文的"混合路线"思想——
确定性信号处理(基频检测、回声剥离、时延标定)与小型神经网络各司其职,
网络只承担它擅长的部分:该工程把波形形状拟合交给端上训练的小型 MLP,
而宽带/非周期噪声(人声、音乐、白噪、突发声)则被划为 RNNoise 类方案的
目标场景。原文页面含可交互的频谱图、音频对比(无处理 / RNNoise /
Speexdsp,0–20 dB 六种信噪比、嘈杂人声/车内/街道三类噪声)与麦克风实时
试听,译文无法内嵌这些音频,建议直接打开原文听效果。文中插图均为
base64 占位,译稿以图注文字说明,图片请看原页面。
RNNoise:学习噪声抑制(Learning Noise Suppression)
原文顶部是一张语谱图:鼠标悬停可在降噪前/降噪后的音频语谱图之间切换。
RNNoise 登场
本演示介绍 RNNoise 项目,展示如何将深度学习应用于噪声抑制。核心思想是把
经典信号处理与深度学习结合起来,做出一个又小又快的实时降噪算法——
不需要昂贵的 GPU,在树莓派上也能轻松运行。与传统降噪系统相比,它更简单
(更容易调参)、听感也更好(这坑我踩过!)。
噪声抑制
噪声抑制是语音处理中一个相当古老的课题,至少可以追溯到上世纪 70 年代。
顾名思义,它的目标是:拿到一个带噪信号,尽可能多地去掉噪声,同时让我们
关心的语音失真最小。
图:传统噪声抑制算法的概念示意图(见原文页面)。
这是传统降噪算法的概念视图。语音活动检测(Voice Activity Detection,VAD)
模块判断信号里什么时候是人声、什么时候只有噪声;噪声谱估计模块据此摸清
噪声的频谱特征(每个频率上有多少能量);知道了噪声长什么样,就可以把它从
输入音频中"减去"(做起来可不像听起来那么简单)。
光看上面这张图,降噪似乎很简单:不就三个概念上很简单的任务吗?对——也不对!
任何一个电子工程专业的本科生都能写出一个"在某些时候……勉强……能用"的降噪
算法。难的是让它在任何时间、对任何类型的噪声都表现良好。这需要对算法中
每一个旋钮都极其小心地调参,为各种怪异信号准备大量特殊处理,还要做海量
测试。总会有某种奇怪的信号跳出来制造麻烦,逼你继续调参,而调参时极容易
按下葫芦浮起瓢——修一个坏三个。这活儿一半是科学,一半是艺术。我以前在
speexdsp 库里做过的降噪器就是如此:勉强能用,但远
谈不上好。
深度学习与循环神经网络
深度学习是一个老想法的新版本:人工神经网络。神经网络从上世纪 60 年代就
存在了,近年来的新变化是:
- 我们现在知道如何把网络做到两个隐藏层以上;
- 我们知道如何让循环网络记住很久以前的模式;
- 我们有了真正能训练它们的算力。
循环神经网络(Recurrent Neural Network,RNN)在这里非常关键,因为它能对
时间序列建模,而不是把输入帧和输出帧各自独立看待。这一点对降噪尤其重要——
我们需要时间才能对噪声做出好的估计。很长一段时间里,RNN 的能力受到严重
限制:它们无法长时间保存信息,而且沿时间反向传播时的梯度下降非常低效
(梯度消失问题)。门控单元的发明同时解决了这两个问题,例如长短期记忆
网络(Long Short-Term Memory,LSTM)、门控循环单元(Gated Recurrent
Unit,GRU)以及它们的众多变体。
RNNoise 选用 GRU,因为在这个任务上它的表现略好于 LSTM,且所需资源更少
(CPU 和权重占用的内存都更少)。与简单循环单元相比,GRU 多了两个
门。重置门(reset gate)控制计算新状态时是否使用旧状态(记忆);
更新门(update gate)控制状态在多大程度上随新输入改变。正是这个更新门
(关闭时)让 GRU 能够(并且很容易地)长时间保存信息,这也是 GRU(和
LSTM)远强于简单循环单元的原因。
图:简单循环单元与 GRU 的对比(见原文页面)。差别在于 GRU 的 r、z
两个门,它们让学习更长时间跨度的模式成为可能。两者都是软开关(取值在
0 到 1 之间),由整层上一时刻的状态和输入经 sigmoid 激活函数算出。当
更新门 z 切到左边时,状态可以在很长一段时间内保持不变——直到某个条件
让 z 切到右边。
混合方案
深度学习的成功让一种做法变得流行:把整个问题一股脑儿丢给深度神经网络。
这类方法称为端到端(end-to-end)——从上到下全是神经元。端到端方法已经
被应用于语音识别和
语音合成。
一方面,这些端到端系统证明了深度神经网络可以多么强大;另一方面,它们有时
既非最优,又浪费资源。例如有些降噪方案使用每层数千个神经元、数千万个权重
的网络来做降噪——代价不仅是运行网络的计算量,还有模型本身的体积:你的
库变成一千行代码外加上几十 MB(甚至更多)的神经元权重。
因此我们在这里走了一条不同的路:所有反正都需要的基础信号处理全部保留
(不让神经网络去模仿它们),而让神经网络去学习信号处理旁边那些需要无穷
无尽微调的刁钻部分。与某些已有的深度学习降噪工作的另一点不同是:我们面向
的是实时通信而非语音识别,所以我们前瞻(look ahead)不起几毫秒——本例
中是 10 ms。
定义问题
为了避免输出数量(进而神经元数量)过于庞大,我们决定不直接处理采样点,也
不直接处理频谱,而是考虑遵循巴克刻度(Bark scale)的频率频带——这是一个
与人耳感知方式相匹配的频率刻度。我们总共使用 22 个频带,否则就得处理 480
个(复数)谱值。
图:Opus 频带布局与真实巴克刻度的对比(见原文页面)。RNNoise 使用与
Opus 相同的基础布局;由于频带相互重叠,Opus 频带的边界恰好成为重叠后
RNNoise 频带的中心。频率越高频带越宽,因为人耳在高频处的频率分辨力
越差;低频处频带较窄,但又没有巴克刻度给出的那么窄——否则我们就没有
足够的数据来做出好的估计。
当然,仅凭 22 个频带的能量无法重建音频。但我们可以为每个频带计算一个施加
于信号的增益。你可以把它想象成一个 22 段均衡器,快速地调节每一段的电平:
压低噪声、放行语音。
按频带计算增益有几个好处。第一,模型简单得多,要算的频带少。第二,它从
根本上杜绝了所谓音乐噪声(musical noise)伪影——即相邻频率都被压下去、
唯独单个频率成分漏网的现象;这类伪影在降噪中很常见,也非常恼人。有了
足够宽的频带,我们要么整段放行,要么整段切掉。第三个好处与模型的优化方式
有关:增益始终被限制在 0 到 1 之间,直接用 sigmoid 激活函数(输出同样在
0 到 1 之间)来计算增益,就能保证我们永远不会干出什么特别蠢的事,比如
凭空添加上原本不存在的噪声。
技术细节(点击展开)
输出端我们也可以选择修正线性(rectified linear)激活函数,用以表示从 0 到
无穷大的 dB 衰减量。为了在训练中更好地优化增益,损失函数采用施加于增益
α 次方的均方误差(MSE)。到目前为止,我们发现 α=0.5 在主观听感上效果最好。
α→0 等价于最小化对数谱距离,但那会有问题——最优增益可能非常接近零。
使用频带带来的较低分辨率,主要缺点是我们没有足够细的分辨率去压制基音谐波
之间的噪声。幸运的是这并不那么重要,而且还有一个简单的技巧可以处理
(见下文的基音滤波部分)。
既然输出基于 22 个频带,输入保留更高的频率分辨率就没什么意义,所以我们用
同样的 22 个频向来给神经网络喂频谱信息。由于音频的动态范围巨大,计算
能量的对数远比直接喂能量要好;既然都做了,再用 DCT(离散余弦变换)给特征
去相关也没有坏处。最终得到的数据是基于巴克刻度的倒谱(cepstrum),它与
语音识别中极为常用的梅尔频率倒谱系数(MFCC)关系密切。
除倒谱系数外,我们还加入了:
- 前 6 个系数在帧间的一阶和二阶导数;
- 基音周期(基频的倒数);
- 6 个频带的基音增益(清浊音强度);
- 一个特殊的非平稳性度量值,对检测语音很有用(超出本演示范围,不展开)。
这样神经网络一共拿到 42 个输入特征。
深度架构
我们使用的深度架构借鉴了传统降噪的思路,主要工作由 3 个 GRU 层完成。下图
展示了计算各频带增益所用的各层,以及这套架构如何对应到传统降噪的各个步骤。
当然,和神经网络领域常见的情况一样,我们无法真正证明网络确实按我们的设想
使用了这些层,但这种拓扑比我们试过的其他拓扑效果更好,因此有理由认为它的
行为与设计意图一致。
图:本项目所用神经网络的拓扑(见原文页面)。每个方框代表一层神经元,
括号内是单元数量。Dense(全连接)层是不带循环的全连接层。网络的一组
输出是施加在不同频率上的一组增益;另一组输出是语音活动概率——它不用于
降噪,但算是网络的一个有用副产品。
关键全在数据
深度神经网络有时也会相当蠢。它们对自己了解的东西很在行,但对于偏离已知
分布太远的输入,可能错得相当壮观。更糟的是,它们是非常懒惰的学生:只要
训练流程里存在任何能让它们逃避学习困难内容的漏洞,它们就一定会钻。这就是
训练数据质量至关重要的原因。
技术细节(点击展开)
流传很广的一个故事:很久以前,一些军方研究人员试图训练一个神经网络来识别
伪装在树林里的坦克。他们拍了有坦克和没坦克的树林照片,然后训练网络识别
哪些照片里有坦克。网络成功得超乎预期!只有一个问题:有坦克的照片是阴天
拍的,没坦克的照片是晴天拍的——网络真正学会的是区分阴天和晴天。如今
研究人员已经意识到这个问题,会避免这种明显的错误,但更隐蔽的版本仍然可能
出现(本人过去就中过招)。
在降噪这个问题上,我们没法直接采集可用于监督学习的输入/输出数据对,因为
我们几乎不可能同时拿到干净语音和它对应的带噪语音。我们只能用分开录制
的干净语音和噪声人工合成训练数据。棘手之处在于要凑出种类足够丰富的噪声
加进语音,还必须确保覆盖各种各样的录音条件。例如,一个只用全频带音频
(0–20 kHz)训练的早期版本,在音频被 8 kHz 低通滤波后就会失败。
技术细节(点击展开)
与语音识别中的常见做法不同,我们选择不对特征做倒谱均值归一化,并保留了
代表能量的第一个倒谱系数。因此我们必须保证数据涵盖所有真实电平的音频。
我们还对音频施加随机滤波器,使系统对各种麦克风频响具有鲁棒性(这一点通常
由倒谱均值归一化来处理)。
基音滤波
由于频带的频率分辨率太粗,无法滤掉基音谐波之间的噪声,这部分就用基础信号
处理来完成——这是混合方案的又一部分。当你对同一个变量进行多次测量时,提高
精度(降低噪声)最简单的办法就是求平均。显然,直接对相邻音频采样求平均
不是我们想要的,因为那等价于低通滤波。但当信号具有周期性时(例如浊音
语音),我们可以对相隔一个基音周期的采样求平均。这构成了一个梳状滤波器
(comb filter):放行基音谐波,同时衰减谐波之间的频率成分——噪声恰好就藏在
那里。为避免信号失真,梳状滤波按频带独立施加,滤波强度同时取决于基音相关
度和神经网络算出的频带增益。
技术细节(点击展开)
我们目前用 FIR 滤波器做基音滤波,但也可以用 IIR 滤波器(已在 TODO 清单
上):IIR 能带来更大的噪声衰减量,风险是强度过大时失真也更严重。
从 Python 到 C
网络的全部设计与训练都在 Python 中用很棒的 Keras 深度
学习库完成。由于 Python 通常不是实时系统的首选语言,运行时代码必须用 C
实现。幸运的是,运行神经网络远比训练神经网络简单,我们只需实现前馈层和
GRU 层。为了让权重占用更小的体积,训练时我们把权重的幅度约束在 ±0.5
以内,这样就能方便地用 8 位数值存储权重。最终模型只有 85 kB(若用 32 位
浮点数存储权重则需要 340 kB)。
C 代码以 BSD 许可开放在 GitLab
(GitHub 镜像)。撰写本演示时代码尚未
做优化,但在 x86 CPU 上已经比实时快约 60 倍,在树莓派 3 上也比实时快约
7 倍。做好向量化(SSE/AVX)后,预计还能再快约 4 倍。
来听听样例吧!
道理说得挺好,但它实际听起来如何?这里是 RNNoise 实际工作的几个例子,
去除三种不同类型的噪声。所用噪声和干净语音均未出现在训练集中。
原文页面此处是一个交互式播放器,可选项包括:
- 降噪算法:不做处理 / RNNoise / Speexdsp;
- 噪声电平(SNR):0 dB / 5 dB / 10 dB / 15 dB / 20 dB / 干净;
- 噪声类型:嘈杂人声(babble)/ 车内噪声 / 街道噪声;
- 切换样例时可以选择从头播放,或从当前位置继续播放。
该对比用于评估 RNNoise 相对于不做处理、相对于 Speexdsp 降噪器的效果。虽然
提供的信噪比最低到 0 dB,但我们瞄准的大多数应用(例如 WebRTC 通话)的
信噪比其实更接近 20 dB 而非 0 dB。
那么到底该听什么?说来奇怪,你不应该期待可懂度(intelligibility)的
提升。人类在噪声中理解语音的能力实在太强,而一个增强算法——尤其是不允许
前瞻待降噪语音的算法——只能破坏信息。那我们为什么还要做这件事?为了音质。
增强后的语音听着不那么烦,也更不容易让听者疲劳。
实际上,确实有少数几种情况它能提升可懂度。第一种是视频会议中多个讲话者
被混音到一起的场景:降噪可以防止所有未发言者的噪声被混进当前发言者的声音,
音质和可懂度都得到改善。第二种是语音经过低码率编解码器的情况:这类编解码器
对带噪语音的劣化通常比对干净语音更严重,先去掉噪声能让编解码器工作得更好。
用你自己的声音试试!
不满足于上面的样例?你可以直接用麦克风录音,(近)实时地给你的音频降噪。
点击原文页面上的按钮后,RNNoise 会在浏览器里用 JavaScript 执行降噪。算法
本身是实时运行的,但我们特意延迟了几秒输出,好让你更容易听出降噪效果。
请务必戴上耳机,否则会听到反馈啸叫。 在"No suppression"和"RNNoise"
之间切换即可对比效果;如果你的输入噪声不够,还可以点"white noise"按钮
人工加入白噪声。
把你的噪声捐献给科学
如果你觉得这项工作有用,有个简单的办法能让它变得更好:只需花你一分钟。
点击链接,让我们录下你所在环境一分钟的噪声,这些噪声可用于改进神经网络的
训练。附带的好处是:网络将来就认得你那里的噪声,当你在视频会议(例如
WebRTC)中用到它时,效果可能更好。我们对任何可能进行语音通信的环境中的
噪声都感兴趣:办公室、车里、街上,或任何你可能使用手机/电脑的地方。
感谢所有捐献噪声的人。数据现已免费开放下载
(6.4 GB),详见压缩包内的 README。
后续方向
想了解 RNNoise 更多技术细节,请看这篇论文
(当时尚未投稿;正式版为 J.-M. Valin, A Hybrid DSP/Deep Learning
Approach to Real-Time Full-Band Speech Enhancement,MMSP 2018,
PDF,arXiv:
1709.08243)。代码仍在活跃开发中
(API 尚未冻结),但已经可以用于实际应用。它目前面向 VoIP/视频会议应用,
稍加调整后大概还能用于许多其他任务。一个明显的方向是自动语音识别(ASR):
虽然可以先把带噪语音降噪再送给 ASR,但这种做法并不最优——它丢弃了关于
处理过程固有不确定性的有用信息;如果 ASR 不仅知道最可能的干净语音,还
知道这个估计有多可信,那会有用得多。RNNoise 另一个可能的"再就业"方向是
为电声乐器做一个聪明得多的噪声门(noise gate):只需要好的训练数据和
几处代码改动,就能把树莓派变成一个相当不错的吉他噪声门。有人接招吗?想必
还有许多我们尚未想到的潜在应用。
想对本演示发表评论,请看作者的 Dreamwidth 博文。
——Jean-Marc Valin(jmvalin@jmvalin.ca),2017 年 9 月 27 日
补充资源
- 代码:RNNoise Git 仓库
(GitHub 镜像) - J.-M. Valin,A Hybrid DSP/Deep Learning Approach to Real-Time Full-Band Speech Enhancement,
IEEE MMSP 2018(arXiv: 1709.08243) - 作者的 Dreamwidth 博文(评论区)
致谢
特别感谢 Michael Bebenita、Thomas Daede 和 Yury Delendik 协助制作本演示,
感谢 Reuben Morais 在 Keras 方面的帮助。
Jean-Marc 的文档工作由 Mozilla Emerging Technologies 赞助。
(C) Copyright 2017 Mozilla and Xiph.Org
RNNoise: Learning Noise Suppression(原文存档)
以下为原文存档,由本地保存的网页 HTML 经 markitdown 转换整理;图片在
存档中为data:image/...;base64...占位引用,此处保留占位标记,图片请看
https://jmvalin.ca/demo/rnnoise/。
RNNoise: Learning Noise Suppression
The image above shows the spectrogram of the audio before and after (when moving the mouse over) noise suppression.
Here's RNNoise
This demo presents the RNNoise project, showing how deep learning can be applied to noise suppression.
The main idea is to combine classic signal processing with
deep learning to create a real-time noise suppression algorithm that's small and fast. No expensive GPUs required
— it runs easily on a Raspberry Pi. The result is much simpler (easier to tune) and sounds better than traditional noise
suppression systems (been there!).
Noise Suppression
Noise suppression is a pretty old topic in speech processing,
dating back to at least the 70s.
As the name implies, the idea is to take a noisy signal and remove as much noise as possible
while causing minimum distortion to the speech of interest.
This is a conceptual view of a conventional noise suppression algorithm.
A voice activity detection (VAD) module detects when the signal contains voice and
when it's just noise. This is used by a noise spectral estimation module to figure out
the spectral characteristics of the noise (how much power at each frequency). Then, knowing
how the noise looks like, it can be "subtracted" (not as simple as it sounds) from the input
audio.
From looking at the figure above, noise suppression looks simple enough: just three conceptually
simple tasks and we're done, right? Right — and wrong! Any undergrad EE student can write
a noise suppression algorithm that works... kinda... sometimes. The hard part is to make it work
well, all the time, for all kinds of noise. That requires very careful tuning of every knob in the
algorithm, many special cases for strange signals and lots of testing. There's always some weird
signal that will cause problems and require more tuning and it's very easy to break more things
than you fix. It's 50% science, 50% art. I've been there before with the noise suppressor
in the speexdsp library. It kinda works, but it's not great.
Deep Learning and Recurrent Neural Networks
Deep learning is the new version of an old idea: artificial neural networks. Although those have
been around since the 60s, what's new in recent years is that:
- We now know how to make them deeper than two hidden layers
- We know how to make recurrent networks remember patterns long in the past
- We have the computational resources to actually train them
Recurrent neural networks (RNN) are very important here because they make it possible to model time
sequences instead of just considering input and output frames independently. This is
especially important for noise suppression because we need time to get a good estimate
of the noise. For a long time, RNNs were heavily limited in their ability because
they could not hold information for a long period of time and because the gradient descent
process involved when back-propagating through time was very inefficient (the vanishing gradient
problem). Both problems were solved by the invention of gated units, such as the
Long Short-Term Memory (LSTM), the Gated Recurrent Unit (GRU), and their many variants.
RNNoise uses the Gated Recurrent Unit (GRU) because it performs slightly better than LSTM on
this task and requires fewer resources (both CPU and memory for weights). Compared to simple
recurrent units, GRUs have two extra gates. The reset gate controls whether the
state (memory) is used in computing the new state, whereas the update gate controls
how much the state will change based on the new input. This update gate (when off) makes it
possible (and easy) for the GRU to remember information for a long period of time and is the
reason GRUs (and LSTMs) perform much better than simple recurrent units.
Comparing a simple recurrent unit with a GRU. The difference lies in the GRU's r and z
gates, which make it possible to learn longer-term patterns. Both are soft switches (value
between 0 and 1) computed based on the previous state of the whole layer and the inputs, with a sigmoid
activation function. When the update gate z is on the left, then the state can remain
constant over a long period of time — until a condition causes z to switch to the right.
A Hybrid Approach
Thanks to the successes of deep learning, it is now popular to throw deep neural networks
at an entire problem. These approaches are called end-to-end — it's neurons all
the way down. End-to-end approaches have been applied to
speech recognition and to
speech synthesis.
On the one hand, these end-to-end systems have proven just how powerful deep
neural networks can be. On the other hand, these systems can sometimes be both suboptimal,
and wasteful in terms of resources. For example, some approaches to noise suppression use
layers with thousands of
neurons — and tens of millions of weights — to perform noise suppression. The drawback
is not only the computational cost of running the network, but also the size of the model
itself because your library is now a thousand lines of code along with tens of megabytes (if not more)
worth of neuron weights.
That's why we went with a different approach here: keep all the basic signal processing
that's needed anyway (not have a neural network attempt to emulate it), but let the neural
network learn all the tricky parts that require endless tweaking next to the signal processing.
Another thing that's different from some existing work on noise suppression with deep learning
is that we're targeting real-time communication rather than speech recognition, so we can't afford
to look ahead more than a few milliseconds (in this case 10 ms).
Defining the problem
To avoid having a very large number of outputs — and thus a large number of neurons —
we decided against working directly with samples or with a spectrum. Instead, we consider frequency
bands that follow the Bark scale, a frequency scale that matches how we perceive sounds. We use
a total of 22 bands, instead of the 480 (complex) spectral values we would otherwise have to consider.
Layout of the Opus bands vs the actual Bark scale. For RNNoise, we use the same base layout
as Opus. Since we overlap the bands, the boundaries between the Opus bands become the center
of the overlapped RNNoise bands. The bands are wider at higher frequency because the ear has
poorer frequency resolution there. At low frequencies, the bands are narrower, but not as narrow
as the Bark scale would give because then we would not have enough data to make good estimates.
Of course, we cannot reconstruct audio from just the energy in 22 bands. What we can do though, is
compute a gain to apply to the signal for each of these bands. You can think about it as using a 22-band
equalizer and rapidly changing the level of each band so as to attenuate the noise, but let the
signal through.
There are several advantages to operating with per-band gains. First, it makes for a much
simpler model since there are fewer bands to compute. Second, it makes it impossible to create
so-called musical noise artifacts, where only a single tone gets though while its neighbours
are attenuated. These artifacts are common in noise suppression and quite annoying. With bands that
are wide enough, we either let a whole band through, or we cut it all. The third advantage comes from
how we optimize the model. Since the gains are always bounded between 0 and 1, simply using a
sigmoid activation function (whose output is also between 0 and 1) to compute them ensures that
we can never do something really stupid, like adding noise that wasn't there in the first place.
▶ Show nerdy details
▼ Hide nerdy details
For the output, we could also have chosen a rectified linear activation function to represent
an attenuation in dB between 0 and infinity. To better optimize the gain during
training, the loss function is the mean squared error (MSE) applied to the gain raised to the
power α. So far, we have found that α=0.5 produces the best results perceptually.
Using α→0 would be equivalent to minimizing the log spectral distance, and
is problematic because the optimal gain can be very close to zero.
The main drawback of the lower resolution we get from using bands is that we do not
have a fine enough resolution to suppress the noise between pitch harmonics. Fortunately,
it's not so important and there is even an easy trick to do it (see the pitch filtering part below).
Since the output we're computing is based on 22 bands, it makes little sense to have
more frequency resolution on the input, so we use the same 22 bands to feed spectral information
to the neural network. Because audio has a huge dynamic range, it's much better to
compute the log of the energy rather than to feed the energy directly. And while we're at it, it never
hurts to decorrelate the features using a DCT. The resulting data is a cepstrum
based on the Bark scale, which is closely related to the Mel-Frequency Cepstral Coefficients (MFCC)
that are very commonly used in speech recognition.
In addition to our cepstral coefficients, we also include:
- The first and second derivatives of the first 6 coefficients across frames
- The pitch period (1/frequency of the fundamental)
- The pitch gain (voicing strength) in 6 bands
- A special non-stationarity value that's useful for detecting speech (but beyond the scope of this demo)
That makes a total of 42 input features to the neural network.
Deep architecture
The deep architecture we use is inspired from the traditional approach to noise suppression.
Most of the work is done by 3 GRU layers. The figure below shows the layers we use to compute the
band gains and how the architecture maps to the traditional steps in noise suppression. Of course,
as is often the case with neural networks we have no actual proof that the network is using its
layers as we intend, but the fact that the topology works better than others we tried makes it
reasonable to think it is behaving as we designed it.
Topology of the neural network used in this project. Each box represents a layer of
neurons, with the number of units indicated in parentheses. Dense layers are
fully-connected, non-recurrent layers.
One of the outputs of the network is a set of gains to apply at different frequencies. The
other output is a voice activity probability, which is not used for noise suppression
but is a useful by-product of the network.
It's all about the data
Even deep neural networks can be pretty dumb sometimes. They're very good at what they know
about, but they can make pretty spectacular mistakes on inputs that are too far from what they
know about. Even worse, they're really lazy students. If they can use any sort of loophole in the training
process to avoid learning something hard, then they will. That is why the quality of the training
data is critical.
▶ Show nerdy details
▼ Hide nerdy details
A widely-circulated story is that a long time ago, some army researchers were trying
to train a neural network to recognize tanks camouflaged in trees. They took pictures of
trees with and without tanks, then trained a neural network to identify the ones that had
a tank. The network succeeded beyond expectations! There was just one problem. Since the
photos with tanks had been taken on cloudy days while the photos without tanks had been
taken on sunny days, what the network really learned was how to tell a cloudy day from a
sunny day. While researchers are now aware of the issue and avoid such obvious mistakes,
more subtle versions of it can still occur (and have occurred to yours truly in the past).
In the case of noise suppression, we can't just collect input/output data that can
be used for supervised learning since we can rarely get both the clean speech and the noisy
speech at the same time. Instead, we have to artificially create that data from separate
recordings of clean speech and noise. The tricky part is getting a wide variety of noise data
to add to the speech. We also have to make sure to cover all kinds of recording conditions. For
example, an early version trained only on full-band audio (0-20 kHz) would fail when the
audio was low-pass filtered at 8 kHz.
▶ Show nerdy details
▼ Hide nerdy details
Unlike what is common for speech recognition, we choose not to apply cepstral mean
normalization to our features and we retain the first cepstral coefficient that represents
the energy. Because of that, we have to ensure that the data includes audio at all
realistic levels. We also apply random filters to the audio to make the system
robust to a variety of microphone frequency responses (which is normally handled by cepstral
mean normalization).
Pitch filtering
Since the frequency resolution of our bands is too coarse to filter noise between
pitch harmonics, we do it using basic signal processing. This is another part of the
hybrid approach. When one has multiple measurements of the same variable, the easiest
way to improve the accuracy (reduce noise) is simply to compute the average.
Obviously, just computing the average of adjacent audio samples isn't what we want since
it would result in low-pass filtering. However, when the signal is periodic (such as voiced
speech), then we can compute the average of samples offset by the pitch period. This results
in a comb filter that lets pitch harmonics through, while attenuating the frequencies
between them — where the noise lies. To avoid distorting the signal, the
comb filter is applied independently for each band and its filter strength depends both on the pitch correlation
and on the band gain computed by the neural network.
▶ Show nerdy details
▼ Hide nerdy details
We currently use an FIR filter for the pitch filtering, but it is also possible (and on the TODO list)
to use an IIR filter, which would result in greater noise attenuation at the risk of higher
distortion if the strength is too aggressive.
From Python to C
All the design and training of the neural network is done in Python using the awesome
Keras deep learning library. Since Python is usually
not the language of choice for real-time systems, we have to implement the run-time code
in C. Fortunately, running a neural network is by far easier than training one, so
all we had to do was implement feed-forward and GRU layers. To make it easier to fit the
weights in a reasonable footprint, we constrain the magnitude of the weights to +/- 0.5 during
training, which makes it easy to store them using 8-bit values. The resulting model fits
in just 85 kB (instead of the 340 kB required to store the weights as 32-bit floats).
The C code is available under a BSD license. Although as of writing this demo, the code
is not yet optimized, it already runs about 60x faster than real-time on an x86 CPU. It even runs about 7x faster than
real-time on a Raspberry Pi 3.
With good vectorization (SSE/AVX), it should be possible to make it about 4x faster than it currently is.
Show Me the Samples!
OK, that's nice and all, but how does it actually sound? Here's some examples of RNNoise
in action, removing three different types of noise. Neither the noise nor the clean speech were
used during the training.
Your browser does not support the audio tag.
Suppression algorithm
- No suppression
- RNNoise
- Speexdsp
Noise level (SNR)
- 0 dB
- 5 dB
- 10 dB
- 15 dB
- 20 dB
- Clean
Noise type
- Babble noise
- Car noise
- Street noise
Select where to start playing when selecting a new sample
Keep playing
Set current position as restart point
Player will continue when changing sample.
Evaluating the effect of RNNoise compared to no suppression and to the Speexdsp noise suppressor. Although the
SNRs provided go as low as 0 dB, most applications we are targeting (e.g. WebRTC calls) tend to have SNRs
closer to 20 dB than to 0 dB.
So what should you listen for anyway? As strange as it may sound, you should not be expecting
an increase in intelligibility. Humans are so good at understanding speech in noise that an enhancement
algorithm — especially one that isn't allowed to look ahead of the speech it's denoising —
can only destroy information. So why are we doing this in the first place? For quality. The enhanced speech is
much less annoying to listen to and likely causes less listener fatigue.
Actually, there are still a few cases where it can actually help
intelligibility. The first is videoconferencing, when multiple speakers are being mixed together. For that
application, noise suppression prevents the noise from all the inactive speakers from being mixed in with the active
speaker, improving both quality and intelligibility. A second case is when the speech goes through a
low bitrate codec. Those tend to degrade noisy speech more than clean speech, so removing the noise
allows the codec to do a better job.
Try it on your voice!
Not happy with the samples above? You can actually record from your microphone and have your audio denoised in (near)
real-time. If you click on the button below, RNNoise will perform noise suppression in Javascript from your browser.
The algorithm runs in real-time but we've purposely delayed it by a few seconds to make it easier to hear the denoised output.
Make sure to wear headphones otherwise you'll hear a feedback loop. To start the demo, select either "No suppression"
or "RNNoise". You can toggle between the two to see the effect of the suppression. If your input doesn't have enough noise,
you can artificially add some by clicking the "white noise" button.
Suppression algorithm
- Off
- No suppression
- RNNoise
Noise type
- None
- White noise
Donate Your Noise to Science
If you think this work is useful, there's an easy way to help make it even better! All it takes is a minute
of your time. Click on the link below, to let us record one minute of noise from where you are. This noise can
be used to improve the training of the neural network. As a side benefit,
it means that the network will know what kind of noise you have and might do a better job when you get to use it
for videoconferencing (e.g. in WebRTC). We're interested in noise from any environment where you might communicate using voice.
That can be your office, your car, on the street, or anywhere you might use your phone or computer.
Thanks to everyone who donated their noise. The data is now freely available for download (6.4 GB).
See the included README file for more details.
Where from here?
If you'd like to know more about the technical details of RNNoise, see this paper (not yet submitted).
The code is still under active development (with no frozen API), but is already usable in applications. It is currently
targeted at VoIP/videoconferencing applications, but with a few tweaks, it can probably be applied to many other tasks.
An obvious target is automatic speech recognition (ASR), and while we can just denoise the noisy speech and send the
output to the ASR, this is sub-optimal because it discards useful information about the inherent uncertainty of the process.
It's a lot more useful when the ASR knows not only the most likely clean speech, but also how much it can rely on that estimate.
Another possible "retargeting" for RNNoise is making a much smarter noise gate for electric musical instruments.
All it should take is good training data and a few changes to the code to turn a Raspberry Pi into a really good guitar noise gate.
Any takers? There are probably many other potential applications we haven't considered yet.
If you would like to comment on this demo, you can do so on here.
—Jean-Marc Valin
(jmvalin@jmvalin.ca) September 27, 2017
Additional Resources
- The code: RNNoise Git repository
(Github mirror) - J.-M. Valin, A Hybrid DSP/Deep Learning Approach to Real-Time Full-Band Speech Enhancement, International Workshop on Multimedia Signal Processing, 2018. (arXiv:1709.08243)
- Jean-Marc's blog post for comments
Acknowledgements
Special thanks to Michael Bebenita, Thomas Daede and Yury Delendik for their help putting together this demo. Thanks to Reuben Morais for
his help with Keras.
Jean-Marc's documentation work is sponsored by Mozilla Emerging Technologies.
(C) Copyright 2017 Mozilla and Xiph.Org

本演示介绍 RNNoise 项目,展示如何将深度学习应用于噪声抑制。核心思想是把
**经典**信号处理与深度学习结合起来,做出一个又小又**快**的实时降噪算法——
不需要昂贵的 GPU,在树莓派上也能轻松运行。与传统降噪系统相比,它更简单
(更容易调参)、听感也更好(这坑我踩过!)。
浙公网安备 33010602011771号