RTL8713 dsp-SDK、 ai-voice 和 tflite-micro 简介、AI Voice 开发指南
ameba-dsp is the development framework for Realtek AmebaLite HiFi5 DSP.
| Chip | master |
|---|---|
| RTL8726E | |
| RTL8713E |
git clone https://github.com/Ameba-AIoT/ameba-dsp
git clone --recursive https://github.com/Ameba-AIoT/ameba-dsp
The extension SDK supports two advanced modules:
Documentation for latest version: https://aiot.realmcu.com/en/latest/dsp/dsp_sdk_index.html
Note: Each SoC series has its own documentation, please find documentation with the specified chip.
See the ApplicationNote chapter Build Environment from above links for a detailed setup guide.
MCU SDK repository address: https://github.com/Ameba-AIoT/ameba-rtos
SoC 系统架构介绍
Realtek Ameba SoC 采用异构多核架构,由 MCU 核和 DSP 核组成,两者协同工作实现完整的系统功能。 传统单片机架构在处理计算密集型任务(如 AI 推理、复杂音频处理)时往往性能不足或者任务间相互干扰。异构多核架构通过分工协作解决了这一问题:
-
MCU 核:专注于系统控制和通信,负责 Wi-Fi/蓝牙协议栈、外设驱动等,保证系统稳定性和实时性;
-
DSP 核:专注于算法计算,负责 AI 算法、音频处理、传感器数据融合等计算密集型任务,发挥硬件加速器的性能优势。
DSP SDK 不能独立运行,必须配合 MCU SDK (FreeRTOS SDK) 使用。Ameba SDK 的整体系统架构如下:
DSP 对 MCU 的依赖关系:
|
依赖类型 |
说明 |
MCU 承担的角色 |
|---|---|---|
|
启动依赖 |
DSP 无自举能力,无法独立启动 |
MCU 在系统启动时将 DSP 固件加载到内存并触发启动 |
|
硬件资源共享 |
外设时钟速度较慢,不建议 DSP 直接访问所有硬件外设 |
MCU 提供驱动服务,DSP 通过 IPC 调用 |
|
内存管理 |
DSP 使用的内存区域需要预先配置 |
MCU 在 Boot 阶段完成内存配置和分配 |
|
时钟与电源 |
DSP 需要独立的时钟域管理 |
MCU 负责 DSP 时钟使能和电源管理 |
DSP SDK 专注于算法处理:
-
DSP 核心:Cadence HiFi 5 高性能音频/语音 DSP 处理器
-
算法库:AIVoice、TFLite Micro、Neural Network Library
-
驱动支持:iDMA、GDMA 等专用加速器驱动
-
开发工具:Xplorer IDE、调试插件、ISS 仿真器
MCU SDK 提供完整的系统功能:
-
无线连接:Wi-Fi 和蓝牙协议栈,提供网络通信能力
-
外设驱动:UART、I2C、SPI、GPIO、Timer 等,DSP 可以通过 IPC 间接调用
-
安全功能:Secure Boot、TrustZone、加密引擎等
-
DSP 管理:DSP 固件加载、启动、监控
备注
关于 MCU SDK 的详细内容,请参考: MCU SDK 使用指南 。
DSP SDK 组成结构
本节介绍 DSP SDK 的目录结构和各组件的功能定位。
DSP SDK 总体结构:
SDK
├── bsp 外设驱动 (IPC、GDMA、GPIO、Timer、OTP 等)
├── configurations Windows、Linux 平台的编译配置包
├── example GDMA、iDMA 等使用示例
├── lib FreeRTOS、IPC、HiFi 5 等库文件
└── project DSP 工程工作区
Project 工作区
project 目录是 Xplorer IDE 的工作区,包含主程序入口、链接脚本、工程配置文件等:
project
├── auto_build Linux 命令行编译脚本及中间文件
├── image 编译生成的固件和反汇编文件
├── img_utility 后处理脚本、LSP 修改脚本
├── project_dsp DSP 工程文件(主函数、MPU 配置等)
└── RTK_LSP 链接脚本(Linker Script)
Example 示例
example 目录包含了基本外设的使用方法。AIVoice、TFLM 等高级应用的示例位于各自独立仓库。 关于例程的编译方法和编译机制说明,请参考: 例程编译
|
例程名称 |
功能类型 |
简介 |
|---|---|---|
|
example_gdma |
DMA 操作 |
演示 GDMA 的单块和多块传输模式 |
|
example_idma |
DMA 操作 |
演示 iDMA 将数据从 PSRAM 搬运到 DTCM |
|
example_idma_nn |
算法加速 |
演示 iDMA ping-pong 策略加速神经网络计算 |
Library 库文件
lib 目录包含 DSP 开发所需的组件库和算法库,提供不同 ABI (Call0/Window) 和工具链的编译版本,包括 aivoice、ipc、freertos、HiFi 5、tflite_micro、xa_nnlib 等。 关于库文件的详细说明和使用方法,请参考: DSP 库
版本兼容性
DSP SDK 与 MCU SDK 存在版本依赖关系。 由于两个 SDK 由独立维护,且涉及异构多核架构的紧密协作,使用不兼容的版本组合可能会出问题。
组件 Ai-Voice
Overview
AIVoice is an offline AI solution developed by Realtek, including local algorithm modules like Audio Front End (Signal Processing), Keyword Spotting, Voice Activity Detection, Speech Recognition etc. It can be used to build smart voice related applications on Realtek Ameba SoCs.
Note that this repository is not recommended for standalone use. Please use it with SDK (ameba-rtos or ameba-dsp or ameba-linux ) based on the chosen SoC and OS.
Documentation: AIVoice Documentation
Supported SoCs
| Chip | OS | Processor | master | SDK link |
|---|---|---|---|---|
| RTL8730E | Linux | CA32 | ameba-linux | |
| RTL8730E | RTOS | CA32 | ameba-rtos | |
| RTL8713E/RTL8726E | RTOS | HiFi5 DSP | ameba-dsp | |
| RTL8721Dx | RTOS | KM4 | ameba-rtos | |
| RTL8721F | RTOS | KM4 | ameba-rtos |
SDK 选择取决于芯片型号与操作系统:
|
芯片 |
操作系统 |
SDK |
AIVoice 路径 |
|---|---|---|---|
|
RTL8721Dx |
FreeRTOS |
ameba-rtos |
{SDK}/component/aivoice |
|
RTL8721F |
FreeRTOS |
ameba-rtos |
{SDK}/component/aivoice |
|
RTL8730E |
FreeRTOS |
ameba-rtos |
{SDK}/component/aivoice |
|
RTL8713E/RTL8726E |
FreeRTOS |
ameba-dsp |
{SDK}/lib/aivoice |
|
RTL8730E |
Linux |
ameba-linux |
{SDK}/apps/aivoice |
Modules
AFE (Audio Front End)
AFE is audio signal processing module for enhancing speech signals. It can improve robustness of speech recognition system or improve signal quality of communication system.
In AIVoice, AFE includes submodules:
- AEC (Acoustic Echo Cancelling)
- BF(Beamforming)
- NS(Noise Suppression)
- AGC (Automatic gain control)
- SSL (Sound Source Localization)
Currently SDK provides libraries for five microphone arrays:
- 1mic
- 2mic_30mm
- 2mic_50mm
- 2mic_70mm
- 3mic_50mm
Other microphone arrays or performance optimizations can be provided through customized services.
Support two modes:
- Speech recognition
- Voice communication
KWS (Keyword Spotting)
KWS is the module to detect specific wakeup words from audio. It is usually the first step in a voice interaction system. The device will enter the state of waiting voice commands after detecting the keyword.
AIVoice provides two solutions:
- Fixed keyword
- User-defined keyword
VAD (Voice Activity Detection)
VAD is the module to detect the presence of human speech in audio.
In AIVoice, a neural network based VAD is provided and can be used in speech enhancement, ASR system etc.
ASR (Automatic Speech Recognition)
ASR is the module to recognize speech to text.
In AIVoice, ASR supports recognition of Chinese speech command words offline.
Flows
Some algorithm flows have been implemented to facilitate user development.
- full flow: AFE+KWS+ASR
- AFE+KWS
- AFE+KWS+VAD
Examples
AIVoice Offline: Full flow with pre-recorded audio
This example shows how to use AIVoice full flow with a pre-recorded 3 channel audio and will run only once after EVB reset. Audio functions such as recording and playback are not integrated.
Please refer to examples/full_flow_offline/README.md for details.
SpeechMind Realtime: Microphone audio stream input
SpeechMind is an intelligent voice assistant framework that integrates AIVoice with audio functions such as recording and playback. Its demo implementation varies by chip:
- RTL8726E and RTL8713E: Refer to examples/speechmind_demo/README.md for DSP part, and refer to SpeechMind for MCU part.
- RTL8730E: RTOS is supported. Refer to SpeechMind for details.
- Other chip: Coming soon...
流程
为了方便用户开发,部分算法流程已在 AIVoice 中实现:
-
AFE+KWS+ASR(full_flow):一个完整的本地算法流程,包括 AFE、KWS 和 ASR。AFE 和 KWS 始终开启,当 KWS 检测到唤醒词时,ASR 开启并支持持续识别,超时后 ASR 退出。
-
AFE+KWS:流程包括 AFE 和 KWS,始终开启。
-
AFE+KWS+VAD:流程包括 AFE、KWS 和 VAD。AFE 和 KWS 始终开启,当 KWS 检测到唤醒词时,VAD 开启并支持持续检测语音端点,超时后 VAD 退出。
如果需要其他模块组合,如 AFE+VAD 等,或是部分模块不使用 AIVoice,可通过调用单独的模块接口实现自定义流程。
组件 TFLite-Micro(TensorFlow Lite Micro for Ameba SoCs)
Overview
TensorFlow Lite for Microcontrollers is a port of TensorFlow Lite designed to run machine learning models on DSPs, microcontrollers and other devices with limited memory.
This repository is a version of the TensorFlow Lite Micro library for Realtek Ameba SoCs.
Note that this repository is not recommended for standalone use. Please use it with SDK (ameba-rtos or ameba-dsp) based on the chosen SoC.
Documentation: ameba-tflite-micro
Supported SoCs
| Chip | OS | Processor | master | SDK link |
|---|---|---|---|---|
| RTL8730E | RTOS | CA32 | ameba-rtos | |
| RTL8713E/RTL8726E | RTOS | HiFi5 DSP | ameba-dsp | |
| RTL8713E/RTL8726E | RTOS | KM4 | ameba-rtos | |
| RTL8721Dx/RTL8720E | RTOS | KM4 | ameba-rtos |
Getting Started
Build Tensorflow Lite Micro Library and Example
-
ameba-rtos
-
Step 1. Enable tflite_micro by menuconfig.py in gcc_project directory
-
Step 2. Build library with example
Run script ./build.py -a {example_name}
-
-
ameba-dsp
-
Step 1. Build library
Run script build/build_amebalite_dsp.sh
-
Step 2. Build example (support both Xplorer and command line)
Refer to examples/{example_name}/README.md for details.
-
Version Sync
This repository has been automatically generated from the master TensorFlow Lite for Microcontrollers repository at https://github.com/tensorflow/tflite-micro.
To sync to the latest tflite-micro and create a ameba SoCs compatible project, you can run the script:
sync/sync_from_tflite_micro.sh
License
This repository is provided under Apache 2.0 license, see LICENSE file for details.
TensorFlow library source code and third_party code contain their own licenses specified under respective repos.
Ai-Voice 接口
模块接口
|
接口 |
模块 |
|---|---|
|
aivoice_iface_afe_v1 |
AFE |
|
aivoice_iface_vad_v1 |
VAD |
|
aivoice_iface_kws_v1 |
KWS |
|
aivoice_iface_asr_v1 |
ASR |
所有接口均支持以下函数:
-
create()
-
destroy()
-
reset()
-
feed()
详情请参考 ${aivoice_lib_dir}/include/aivoice_interface.h。
流程接口
|
接口 |
流程 |
|---|---|
|
aivoice_iface_full_flow_v1 |
AFE+KWS+ASR |
|
aivoice_iface_afe_kws_v1 |
AFE+KWS |
|
aivoice_iface_afe_kws_vad_v1 |
AFE+KWS+VAD |
所有接口均支持以下函数:
-
create()
-
destroy()
-
reset()
-
feed()
详情请参考 ${aivoice_lib_dir}/include/aivoice_interface.h。
事件及回调信息
|
aivoice输出事件 |
事件触发时间 |
回调信息 |
|---|---|---|
|
AIVOICE_EVOUT_VAD |
当VAD检测到语音段开始或结束 |
包含VAD状态,偏移的结构体 |
|
AIVOICE_EVOUT_WAKEUP |
当KWS检测到唤醒词 |
包含ID,唤醒词,唤醒得分的JSON字符串。示例: {"id":2,"keyword":"ni-hao-xiao-qiang","score":0.9} |
|
AIVOICE_EVOUT_ASR_RESULT |
当ASR检测到命令词 |
包含FST类型,命令词,ID的JSON字符串。 示例: {"type":0,"commands":[{"rec":"打开空调","id":14}]} |
|
AIVOICE_EVOUT_AFE |
AFE收到输入的每一帧 |
包含AFE输出数据,通道数等的结构体 |
|
AIVOICE_EVOUT_ASR_REC_TIMEOUT |
ASR/VAD超时 |
NULL |
AFE 事件定义
struct aivoice_evout_afe {
int ch_num; /* 输出音频信号的通道数,默认值:1 */
short* data; /* 增强后的音频信号 */
char* out_others_json; /* 保留用于其他输出数据(例如标志位),key: value 形式 */
};
VAD 事件定义
struct aivoice_evout_vad {
int status; /* 0: VAD 从语音变为静音,表示语音段的结束点
1: VAD 从静音变为语音,表示语音段的起始点 */
unsigned int offset_ms; /* 相对于重置点的时间偏移量 */
};
通用配置
AIVoice 可配参数:
- timeout:
-
在 full flow 中,如果在此持续时间内未检测到命令词,则 ASR 退出。在 AFE+KWS+VAD 流程中,唤醒后 VAD 仅在此持续时间内工作,当 timeout=-1 时,关闭超时。
- memory_alloc_mode:
-
默认使用 SDK 默认堆。SRAM 模式使用 SDK 默认堆,同时还从 SRAM 分配空间用于内存关键数据。 SRAM 模式目前仅适用于 RTL8713E 和 RTL8726E DSP。
详情请参考 ${aivoice_lib_dir}/include/aivoice_sdk_config.h。
非通用配置参数请见算法模块章节。
示例
AIVoice 使用基础
本节解释调用 AIVoice 算法的标准步骤,在后续两个示例的代码中均有应用。建议先阅读本节以掌握基础使用步骤。
-
选择需要的 aivoice 流程或模块。
/* 步骤 1:
* 选择需要的 aivoice 流程.
* 参考 aivoice_interface.h 文件末尾查看支持的流程
*/
const struct rtk_aivoice_iface *aivoice = &aivoice_iface_full_flow_v1;
/* 步骤 2:
* 按需修改默认配置。
* 您可以修改 afe/vad/kws/...的 0 个或多个配置项
*/
struct aivoice_config config;
memset(&config, 0, sizeof(config));
/*
* 这里使用 afe_res_2mic50mm 作为示例。
* 可以根据实际使用的 afe 资源修改这些配置。
* 详情请参考 aivoce_afe_config.h;
*
* afe_config.mic_array 必须与您链接的 afe 资源匹配
*/
struct afe_config afe_param = AFE_CONFIG_ASR_DEFAULT_2MIC50MM; // 根据链接的 afe 资源修改此项
config.afe = &afe_param;
/*
* 仅在完全理解参数含义时修改这些设置。
* 如果不了解这些参数的含义,
* 建议使用默认配置
*/
struct vad_config vad_param = VAD_CONFIG_DEFAULT();
vad_param.left_margin = 300; // 可根据需要修改配置
config.vad = &vad_param; // 可设置为 NULL
struct kws_config kws_param = KWS_CONFIG_DEFAULT();
config.kws = &kws_param; // 可设置为 NULL
struct asr_config asr_param = ASR_CONFIG_DEFAULT();
config.asr = &asr_param; // 可设置为 NULL
struct aivoice_sdk_config aivoice_param = AIVOICE_SDK_CONFIG_DEFAULT();
aivoice_param.timeout = 10;
config.common = &aivoice_param; // 可设置为 NULL
-
使用
create()和指定配置来创建并初始化 aivoice 实例。
/* 步骤 3:
* 创建 aivoice 实例
*/
void *handle = aivoice->create(&config);
if (!handle) {
return;
}
-
注册回调函数。
/* 步骤 4:
* 注册一个回调函数。
* 在本示例中,您可能只会接收到部分 aivoice_out_event_type 事件类型,
* 具体取决于您使用的流程。
* */
rtk_aivoice_register_callback(handle, aivoice_callback_process, NULL);
回调函数可以按实际使用需求进行修改:
static int aivoice_callback_process(void *userdata,
enum aivoice_out_event_type event_type,
const void *msg, int len)
{
(void)userdata;
struct aivoice_evout_vad *vad_out;
struct aivoice_evout_afe *afe_out;
switch (event_type) {
case AIVOICE_EVOUT_VAD:
vad_out = (struct aivoice_evout_vad *)msg;
printf("[user] vad. status = %d, offset = %d\n", vad_out->status, vad_out->offset_ms);
break;
case AIVOICE_EVOUT_WAKEUP:
printf("[user] wakeup. %.*s\n", len, (char *)msg);
break;
case AIVOICE_EVOUT_ASR_RESULT:
printf("[user] asr. %.*s\n", len, (char *)msg);
break;
case AIVOICE_EVOUT_ASR_REC_TIMEOUT:
printf("[user] asr timeout\n");
break;
case AIVOICE_EVOUT_AFE:
afe_out = (struct aivoice_evout_afe *)msg;
// afe 每帧都会输出音频
// 本示例中,为了让日志清晰仅打印一次
static int afe_out_printed = false;
if (!afe_out_printed) {
afe_out_printed = true;
printf("[user] afe output %d channels raw audio, others: %s\n",
afe_out->ch_num, afe_out->out_others_json ? afe_out->out_others_json : "null");
}
// 按需处理 afe 输出的音频
break;
default:
break;
}
return 0;
}
-
使用
feed()给 aivoice 输入音频数据。
/* 在芯片上运行时,通常使用麦克风采集的实时音频流,
* 本示例中使用一条固定音频
* */
const char *audio = (const char *)get_test_wav();
int len = get_test_wav_len();
int audio_offset = 44;
int mics_num = 2;
int afe_frame_bytes = (mics_num + afe_param.ref_num) * afe_param.frame_size * sizeof(short);
while (audio_offset <= len - afe_frame_bytes) {
/* step 5:
* Feed the audio to the aivoice instance.
* */
aivoice->feed(handle,
(char *)audio + audio_offset,
afe_frame_bytes);
audio_offset += afe_frame_bytes;
}
-
(可选) 如果需要重置状态,使用
reset()。 -
如果不再需要 aivoice,使用
destroy()销毁实例。
/* 步骤 6:
* 销毁 aivoice 实例 */
aivoice->destroy(handle);
示例一:AIVoice 离线示例(使用预录音频)
该例子通过一条提前录制的三通道音频演示如何使用 AIVoice 的全流程,在开发板启动后仅运行一次。 未整合录音、播放等音频功能。
示例代码在 aivoice/examples/full_flow_offline 目录下。
参考资料:
1.
AIVoice 开发指南
https://aiot.realmcu.com/zh/latest/rtos/ai/aivoice/aivoice_overview/index.html
Realtek Ameba SoC DSP SDK 架构
https://aiot.realmcu.com/zh/latest/dsp/dsp_sdk_introduction/index.html?refresh=1790470389502
DSP 开发环境搭建
https://aiot.realmcu.com/zh/latest/dsp/dsp_environment/index.html


浙公网安备 33010602011771号