Research on AI-Enabled Real-Time Audio Joint Source-Channel Coding
-
摘要: 目前,第三代合作伙伴计划(3rd Generation Partnership Project, 3GPP)的业务与系统技术规范组第4工作组(TSG SA Working Group 4, SA4) 已经开始研究基于人工智能(Artificial Intelligence, AI)的语音编解码器。本文首次以3GPP地球静止轨道(Geostationary Earth Orbit, GEO)语音典型的物理层参数配置与目前SA4正在研究的Descript 音频编解码器(Descript Audio Codec, DAC)作为评估基线,保证了基线结果的公平性,并首次提出了总功率限制的DAC信源信道调制联合优化方案、恒模限制的DAC信源信道调制联合优化方案和DAC信源信道联合编解码方案。在语音质量可接受的前提下,所提方案相比于基线方案有2.5 dB~4 dB的覆盖增益。希望为GEO语音的后续研究提供有益思路。
-
关键词:
- GEO语音 /
- 信源信道联合编码 /
- 联合信源信道编码调制 /
- AI
Abstract:Currently, the 3rd Generation Partnership Project (3GPP) Technical Specification Group Service and System Aspects Working Group 4 (TSG SA Working Group 4, SA4) has begun researching Artificial Intelligence (AI)-based audio codecs. Based on the Descript Audio Codec (DAC) being studied by SA4, this paper presents a joint optimization scheme for DAC source-channel coding and modulation with total power constraints, a joint optimization scheme for DAC source-channel coding and modulation with constant modulus constraints, and a DAC joint source-channel coding scheme. Furthermore, simulation evaluation is conducted using the typical physical layer parameter configuration of 3GPP Geostationary Earth Orbit (GEO) voice scenario. Finally, preliminary research suggestions are made, with the hope of providing useful insights for the subsequent research on GEO voice. Objective In recent years, Joint Source-Channel Coding (JSCC) has gained widespread attention in academia and industry. In academia, the sources studied for JSCC include video, image, and audio. However, the industry has not focused on these sources as in academia, but rather on JSCC where channel state information (CSI) is considered as the source. The reason the industry has not considered JSCC for video and images is due to two main factors. First, for video and images, 3GPP generally does not conduct independent research, but instead directly reuses source compression standards developed by external organizations such as the Moving Picture Experts Group (MPEG) and the Joint Photographic Experts Group (JPEG). Second, it is constrained by the challenges of the source-channel coding architecture. For example, in joint source-channel coding at the application layer, the channel information obtained at the application layer is often outdated, which limits the potential gains of source-channel coding. However, the situation is quite different for audio. Audio coding and decoding are designed independently by 3GPP, and currently, 3GPP SA4 is researching AI-based speech coding and decoding. Moreover, in GEO speech scenarios, where satellites remain largely stationary relative to ground stations and the channel changes slowly, the physical layer CSI can be transmitted to the application layer, overcoming the issue of outdated channel information. This makes the research on JSCC for audio more promising within 3GPP, compared to JSCC for video or image sources. Methods This paper firstly uses the typical physical layer parameter configurations of 3GPP GEO voice and the DAC encoder, currently being studied by SA4, as an evaluation baseline, ensuring the fairness of the proposed scheme’s gain evaluation. Based on this, the paper proposes the integration of DAC with channel coding and modulation to improve audio quality in low-bitrate scenarios. First, the total power constrained (TPC) DAC-JSCCM scheme is considered, which offers the maximum potential gain due to the joint optimization of source coding, channel coding, and modulation. Then, considering the high PAPR (Peak-to-Average Power Ratio) impact on the power amplifier, the paper proposes a constant-envelope constrained (CEC) DAC-JSCCM scheme. Finally, since the modulation symbols included in RTP payload require significant changes to the existing protocol, a DAC-JSCC scheme that is more compatible with current protocols is proposed, where bit sequences are included in RTP payload. Results and Discussions In the evaluation of baseline scheme 1, the link-level simulation at the physical layer is first conducted to obtain the relationship between SNR and BLER for an input length of 64 bits using a 1/3 Turbo code. Then, assuming that the error rate of an audio frame is the same as the BLER, the transmission error rate of an RTP packet containing L audio frames is calculated. According to the existing processing logic, if one audio frame in an RTP packet is transmitted with an error, the entire RTP packet will be discarded. In real-time communication, VoIP, or speech enhancement scenarios, a PESQ score of 2.5 is generally considered the minimum usable threshold, and voice quality below this score is regarded as unsuitable for professional or commercial services. Therefore, in GEO voice calls, we choose PESQ = 2.5 as the reference point. When PESQ = 2.5, the TPC DAC-JSCCM and CEC DAC-JSCCM provide a coverage gain of approximately 4 dB compared to baseline scheme 1 ( Fig.4 ). The DAC-JSCC scheme offers a coverage gain of 2.5 dB compared to baseline scheme 1 (Fig.4 ). The coverage gains of DAC-JSCCM and DAC-JSCC come from the application layer performing joint source-channel coding, which avoids the cliff effect caused by using Turbo coding at the physical layer. When the Complementary Cumulative Distribution Function (CCDF) of Peak-to-Average Power Ratio (PAPR) is $ {10}^{-3} $, the TPC DAC-JSCCM is 4 dB higher than the QPSK modulation, with sharper signal peaks, higher requirements for power amplifier linearity, and is more prone to distortion. In contrast, the CEC DAC-JSCCM is very close to the QPSK modulated PAPR.Conclusions This paper proposes the Total Power Constrained DAC-JSCCM, Constant Modulus Constrained DAC-JSCCM, and DAC-JSCCM schemes, based on the DAC codec currently being researched by SA4. Simulations are conducted using the typical physical layer configuration of the existing 3GPP GEO voice scenario. The simulation results show that, when the PESQ is 2.5 (the minimum usable threshold), compared to DAC combined with existing physical layer transmission technologies, the proposed DAC-JSCCM provides a coverage gain of 4 dB, while the proposed DAC-JSCC scheme provides a coverage gain of 2.5 dB. SA4 is currently researching more advanced GEO voice codecs beyond DAC, and we will extend our research by combining better GEO voice codecs in the future. -
表 1 仿真参数设置
参数 基线
方案1基线
方案2总功率限制
DAC-
JSCCM恒模
限制
DAC-
JSCCMDAC-
JSCC应用层输出
(N=64)N比特 3N
比特3N /2
复数3N /2
复数3N比特 物理层调制方式 QPSK QPSK 无 无 QPSK 物理层信道编码 1/3
Turbo1/3
无无 无 无 训练信噪比
(dB)无 无 [–5,10) [–5,10) [–5,10) DAC下采样输出维度 1024 1024 192 96 192 RVQ码本数 8 24 无 无 无 码本大小 256 256 无 无 无 表 2 信噪比与误块率表(64比特长度,1/3 Turbo编码)
信噪比(dB) 误块率 –6 1 –5 1 –4 1 –3 0.925 –2 0.665 –1 0.236 0 0.0464 1 0.0039 2 0.0004 3 0 4 0 5 0 6 0 7 0 -
[1] 周通, 姜大洁, 谭俊杰, 等. 语义通信的用例、挑战与标准化影响浅析[J]. 移动通信, 2025, 49(7): 2–11. doi: 10.3969/j.issn.1006-1010.20250522-0001.ZHOU Tong, JIANG Dajie, TAN Junjie, et al. A survey of semantic communication: Use cases, open challenges, and standardization trends[J]. Mobile Communications, 2025, 49(7): 2–11. doi: 10.3969/j.issn.1006-1010.20250522-0001. [2] OPPO. 基于信源信道联合编码的CSI反馈[R]. IMT-2030-Semantic_2024004, 2024. (查阅网上资料, 未找到本条文献信息, 请确认). [3] OPPO. 信道与无线环境语义通信_基于联合信源信道编码的CSI反馈[R]. IMT-2030-Semantic_2024036, 2024. (查阅网上资料, 未找到本条文献信息, 请确认). [4] 小米. 面向CSI反馈的信道估计与联合信源信道编码的联合设计[R]. IMT-2030-Semantic_2024056, 2024. (查阅网上资料, 未找到本条文献信息, 请确认). [5] YANG Ke, WANG Sixian, DAI Jincheng, et al. SwinJSCC: Taming swin transformer for deep joint source-channel coding[J]. IEEE Transactions on Cognitive Communications and Networking, 2025, 11(1): 90–104. doi: 10.1109/TCCN.2024.3424842. [6] GUO Jiangyuan, CHEN Wei, SUN Yuxuan, et al. VideoQA-SC: Adaptive semantic communication for video question answering[J]. IEEE Journal on Selected Areas in Communications, 2025, 43(7): 2462–2477. doi: 10.1109/JSAC.2025.3559160. [7] TUNG T Y and GÜNDÜZ D. DeepWiVe: Deep-learning-aided wireless video transmission[J]. IEEE Journal on Selected Areas in Communications, 2022, 40(9): 2570–2583. doi: 10.1109/JSAC.2022.3191354. [8] BOURTSOULATZE E, KURKA D B, and GÜNDÜZ D. Deep joint source-channel coding for wireless image transmission[J]. IEEE Transactions on Cognitive Communications and Networking, 2019, 5(3): 567–579. doi: 10.1109/TCCN.2019.2919300. [9] YILMAZ S F, NIU Xueyan, BAI Bo, et al. High perceptual quality wireless image delivery with denoising diffusion models[C]. IEEE Conference on Computer Communications Workshops, Vancouver, Canada, 2024: 1–5. doi: 10.1109/INFOCOMWKSHPS61880.2024.10620904. [10] DAI Jincheng, WANG Sixian, TAN Kailin, et al. Nonlinear transform source-channel coding for semantic communications[J]. IEEE Journal on Selected Areas in Communications, 2022, 40(8): 2300–2316. doi: 10.1109/JSAC.2022.3180802. [11] WENG Zhenzi, QIN Zhijin, and LI G Y. Robust semantic communications for speech transmission[C]. 2025 IEEE International Conference on Acoustics, Speech and Signal Processing, Hyderabad, India, 2025: 1–5. doi: 10.1109/ICASSP49660.2025.10889160. (查阅网上资料,未能确认本条文献修改是否正确,请确认). [12] HAN Tianxiao, YANG Qianqian, SHI Zhiguo, et al. Semantic-preserved communication system for highly efficient speech transmission[J]. IEEE Journal on Selected Areas in Communications, 2023, 41(1): 245–259. doi: 10.1109/JSAC.2022.3221952. [13] WENG Zhenzi, QIN Zhijin, TAO Xiaoming, et al. Deep learning enabled semantic communications with speech recognition and synthesis[J]. IEEE Transactions on Wireless Communications, 2023, 22(9): 6227–6240. doi: 10.1109/TWC.2023.3240969. [14] CHEN Xiaojiao, WANG Jing, XU Liang, et al. A perceptually motivated approach for low-complexity speech semantic communication[J]. IEEE Internet of Things Journal, 2024, 11(12): 22054–22065. doi: 10.1109/JIOT.2024.3378779. [15] 3GPP. 3GPP TS 23.501 System architecture for the 5G System (5GS)[S]. 2025. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认). [16] 3GPP. 3GPP TS 38.323 Packet data convergence protocol (PDCP) specification[S]. 2025. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认). [17] IMT-2030(6G)推进组. 面向6G的语义通信技术研究报告[M]. 北京: IMT-2030(6G)推进组, 2025. (查阅网上资料, 未找到对应的英文翻译, 请补充). [18] Descript. Descript-audio-codec[EB/OL]. https://github.com/descriptinc/descript-audio-codec, 2026. [19] 3GPP. 3GPP TR 26.940 Study on ultra low bit rate speech codecs[S]. 2026. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认). [20] 3GPP. 3GPP TS 36. 212 Multiplexing and channel coding[S]. 2025. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认). [21] ETSI. ETSI TS 136 213 V19.1. 0 (2025-10) Physical layer procedures[S]. 2025. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认). [22] OpenDataLab. VCTK[EB/OL]. https://opendatalab.com/OpenDataLab/VCTK, 2026. [23] Common voice[EB/OL]. https://commonvoice.mozilla.org/zh-CN, 2026.(查阅网上资料,请补充作者信息). [24] Librispeech[EB/OL]. https://huggingface.co/datasets/openslr/librispeech_asr, 2026. (查阅网上资料,请补充作者信息). [25] ITU. ITU-T P. 863-2018 Perceptual objective listening quality prediction[S]. 2018. (查阅网上资料, 未找到出版信息, 请补充). -
下载: