An Adaptive Kalman Speech Enhancement Method Driven by Burst Noise Suppression and Dual-time-scale Perception
-
摘要: 针对传统基于自回归(AR)模型的卡尔曼滤波语音增强算法在非平稳噪声环境下噪声模型难以自适应更新、突发噪声易导致模型失配以及协方差参数调节缺乏环境感知机制等问题,该文提出一种用于突发噪声抑制与双时间尺度感知驱动的自适应卡尔曼语音增强方法。首先,引入基于频谱能量熵(EER)构建的环境偏离比,用于刻画当前帧与背景环境统计特性的偏离程度。在此基础上,建立短时与长时双时间尺度EER跟踪机制,分别用于表示瞬时变化与稳定环境统计,并利用两者之间差异构造自适应偏离强度因子,从而实现对语音过程噪声协方差与观测噪声协方差的联合调节。其次,结合突发噪声的核心特性,提出一种基于能量突变阈值和频谱平坦度结合的突发噪声帧判断策略,用于有效区分时域能量突变且带有弱有色或白噪声特性的突发噪声,另外,提出一种基于线性预测残差方差比的语音可建模性因子判决准则,用于进一步区分语音结构与突发噪声成分,避免中等能量强度突发噪声误判导致语音模型污染,从而提升噪声建模的鲁棒性与可靠性。该文算法实验采用NOIZEUS数据库上的音频样本,结果表明,该文所用方法在短时客观可懂度(STOI)、感知语音质量评价指标(PESQ)和分段信噪比(SegSNR)等客观指标上优于传统的AR卡尔曼滤波,以及基于$ {J}_{1} $灵敏度的增广卡尔曼滤波,尤其在非平稳噪声和突发噪声条件下表现出更强的自适应能力与增强稳定性。Abstract:
Objective Traditional Auto-Regressive (AR) Kalman speech enhancement algorithms face three critical challenges under non-stationary noise: limited adaptability of noise model updates, model mismatch caused by burst noise, and a lack of environment-aware covariance adjustment. These limitations degrade enhancement performance and restrict their use in practical speech communication. An improved adaptive Kalman speech enhancement method is therefore proposed to improve speech quality and intelligibility under complex noise conditions. Methods First, an environmental deviation measure based on the Energy Entropy Ratio (EER) is constructed to quantify the statistical deviation between the current frame and the background environment. A dual-time-scale EER tracking mechanism is then established to capture instantaneous variations and stable background statistics. Their difference is used to generate an adaptive deviation intensity factor for the joint adjustment of the process-noise and observation-noise covariances. Second, a two-stage burst noise discrimination scheme is developed based on the characteristics of burst noise. Frame energy changes and the Spectral Flatness Measure (SFM) are jointly used for preliminary burst noise detection. A Speech Modeling Metric (SMM) based on the linear prediction residual variance ratio is further used to distinguish speech from weak-colored burst noise and prevent contamination of the speech AR model. Results and Discussions Experiments are conducted on the NOIZEUS database. The proposed method outperforms the conventional AR-based Kalman filter and the sensitivity-based Augmented Kalman Filter (AKF) in terms of Short-Time Objective Intelligibility (STOI), Perceptual Evaluation of Speech Quality (PESQ), and Segmental Signal-to-Noise Ratio (SegSNR). It shows improved adaptability and stability under non-stationary and burst noise conditions. Dual-time-scale environmental perception and burst noise discrimination facilitate rapid adaptation to environmental changes and reduce speech distortion. Conclusions The proposed burst noise suppression and dual-time-scale adaptive Kalman method addresses the limitations of conventional AR-based Kalman speech enhancement under complex noise conditions. EER-based tracking enables environment-aware covariance adjustment, whereas the two-stage discrimination scheme improves burst noise detection and reduces contamination of the speech AR model. The experimental results demonstrate the robust performance of the proposed method and support its application to speech enhancement under non-stationary noise. -
表 1 :不同SNR下STOI, PESQ, SegSNR评分
SNR(dB) STOI PESQ SegSNR SS KF AKF 本文方法 SS KF AKF 本文方法 SS KF AKF 本文方法 0 0.64 0.64 0.705 0.70 1.89 1.95 1.82 1.87 –3.72 –4.77 –2.60 –1.96 5 0.76 0.76 0.791 0.80 2.10 2.10 2.02 2.05 –0.94 –2.02 –0.33 –0.33 10 0.89 0.881 0.87 0.89 2.31 2.35 2.35 2.42 2.65 1.42 2.01 2.83 15 0.95 0.951 0.92 0.94 2.52 2.69 2.50 2.67 5.69 5.04 6.05 5.92 注:SS:基础谱减;KF:传统卡尔曼滤波;AKF:基于$ {J}_{1} $的增广卡尔曼滤波. 表 2 不同SNR下STOI, PESQ, SegSNR评分
噪声类型 AKF 本文方法 Babble 0.66 0.66 Car 0.82 0.84 Restaurant 0.79 0.79 Station 0.79 0.78 表 3 消融实验指标结果
方法 噪声AR条件更新 噪声污染抑制 双尺度EER STOI PESQ SegSNR 完整方法 √ √ √ 0.81 2.16 –1.31 噪声AR更新 × √ √ 0.82 2.06 –2.18 噪声污染抑制 √ × √ 0.73 1.92 –2.48 双尺度 √ √ × 0.81 2.05 –1.90 未降噪前评分 0.80 2.03 –2.47 -
[1] 张殿熙, 乔兆亮. 语音增强技术及应用[C]. 天津市电子工业协会2025年年会论文集, 天津, 2025: 29–31. doi: 10.26914/c.cnkihy.2025.026058.ZHANG Dianxi and QIAO Zhaoliang. Speech enhancement technology and applications[C]. 2025 Annual Conference of Tianjin Electronic Industry Association, Tianjin, China, 2025: 29–31. doi: 10.26914/c.cnkihy.2025.026058. [2] 曹丽静. 语音增强技术研究综述[J]. 河北省科学院学报, 2020, 37(2): 30–36. doi: 10.16191/j.cnki.hbkx.2020.02.006.CAO Lijing. Overview of speech enhancement algorithms[J]. Journal of the Hebei Academy of Sciences, 2020, 37(2): 30–36. doi: 10.16191/j.cnki.hbkx.2020.02.006. [3] 杜扶遥, 姜囡, 刘浠辰. 涉案语音的降噪处理分析研究[J]. 广东公安科技, 2024, 32(4): 30–35.DU Fuyao, JIANG Nan, and LIU Xichen. Research on noise reduction processing of involved speech[J]. Guangdong Public Security Science and Technology, 2024, 32(4): 30–35. [4] 王涛, 鲁怀伟, 刘宝成. 基于AR模型的Kalman语音增强算法[J]. 青岛大学学报(自然科学版), 2018, 31(2): 48–53. doi: 10.3969/j.issn.1006-1037.2018.05.09.WANG Tao, LU Huaiwei, and LIU Baocheng. Kalman speech enhancement algorithm based on AR model[J]. Journal of Qingdao University (Natural Science Edition), 2018, 31(2): 48–53. doi: 10.3969/j.issn.1006-1037.2018.05.09. [5] JODWAL M, KUMAR S, COLNEY L, et al. Performance analysis of speech enhancement techniques[C]. 2024 First International Conference on Electronics, Communication and Signal Processing, New Delhi, India, 2024: 1–7. doi: 10.1109/ICECSP61809.2024.10698182. [6] 王华朋, 冯嘉琪. 基于深度学习的语音增强方法综述[J]. 科学技术与工程, 2025, 25(20): 8331–8346. doi: 10.12404/j.issn.1671-1815.2404954.WANG Huapeng and FENG Jiaqi. Review of speech enhancement methods based on deep learning[J]. Science Technology and Engineering, 2025, 25(20): 8331–8346. doi: 10.12404/j.issn.1671-1815.2404954. [7] NIAN Zhaoxu, TU Yanhui, DU Jun, et al. A progressive learning approach to adaptive noise and speech estimation for speech enhancement and noisy speech recognition[C]. 2021 IEEE International Conference on Acoustics, Speech and Signal Processing, Toronto, Canada, 2021: 6913–6917. doi: 10.1109/ICASSP39728.2021.9413395. [8] WEI Haimeng and ZHANG Xiaobo. Time-frequency conformer and Kalman filter-based speech enhancement method[C]. 2025 7th International Conference on Intelligent Control, Measurement and Signal Processing, Xian, China, 2025: 325–329. doi: 10.1109/ICMSP68723.2025.11407752. [9] GEORGE A E W, SO S, GHOSH R, et al. Robustness metric-based tuning of the augmented Kalman filter for the enhancement of speech corrupted with coloured noise[J]. Speech Communication, 2018, 105: 62–76. doi: 10.1016/j.specom.2018.10.002. [10] ROY S K and PALIWAL K K. Sensitivity metric-based tuning of the augmented Kalman filter for speech enhancement[C]. 2020 14th International Conference on Signal Processing and Communication Systems, Adelaide, Australia, 2020: 1–6. doi: 10.1109/ICSPCS50536.2020.9310005. [11] 文玉梅, 朱宇. 传感信号宽带噪声实时自适应抑制方法[J]. 电子与信息学报, 2025, 47(8): 2746–2756. doi: 10.11999/JEIT250018.WEN Yumei, ZHU Yu. Real-time Adaptive Suppression of Broadband Noise in General Sensing Signals[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2746–2756. doi: 10.11999/JEIT250018. [12] 邓洪高, 余润华, 纪元法, 等. 偏差未补偿自适应边缘化容积卡尔曼滤波跟踪方法[J]. 电子与信息学报, 2025, 47(1): 156–166. doi: 10.11999/JEIT240469.DENG Honggao, YU Runhua, JI Yuanfa, et al. An Adaptive Target Tracking Method Utilizing Marginalized Cubature Kalman Filter with Uncompensated Biases[J]. Journal of Electronics & Information Technology, 2025, 47(1): 156–166. doi: 10.11999/JEIT240469. [13] XU Xiaodong, FLYNN R, and RUSSELL M. Speech intelligibility and quality: A comparative study of speech enhancement algorithms[C]. 2017 28th Irish Signals and Systems Conference, Killarney, Ireland, 2017: 1–6. doi: 10.1109/ISSC.2017.7983599. doi: 10.1109/ISSC.2017.7983599. [14] VASEGHI S V. Linear prediction models[M]. VASEGHI S V. Advanced Digital Signal Processing and Noise Reduction. 2nd ed. Chichester: John Wiley & Sons, 2001: 227–262. doi: 10.1002/0470841621.ch8. [15] KOO B, GIBSON J D, and GRAY S D. Filtering of colored noise for speech enhancement and coding[C]. International Conference on Acoustics, Speech, and Signal Processing, Glasgow, UK, 1989: 349–352. doi: 10.1109/ICASSP.1989.266437. [16] 王文益, 伊雪. 基于改进语音存在概率的自适应噪声跟踪算法[J]. 信号处理, 2020, 36(1): 32–41. doi: 10.16798/j.issn.1003-0530.2020.01.005.WANG Wenyi and YI Xue. An adaptive noise tracking algorithm using improved speech presence probability[J]. Journal of Signal Processing, 2020, 36(1): 32–41. doi: 10.16798/j.issn.1003-0530.2020.01.005. [17] DUBNOV S. Generalization of spectral flatness measure for non-Gaussian linear processes[J]. IEEE Signal Processing Letters, 2004, 11(8): 698–701. doi: 10.1109/LSP.2004.831663. [18] 兰朝凤, 蒋朋威, 陈欢, 等. 基于双路径递归网络与Conv-TasNet的多头注意力机制视听语音分离[J]. 电子与信息学报, 2024, 46(3): 1005–1012. doi: 10.11999/JEIT230260.LAN Chaofeng, JIANG Pengwei, CHEN Huan, et al. Multi-head attention time domain audiovisual speech separation based on dual-path recurrent network and Conv-TasNet[J]. Journal of Electronics & Information Technology, 2024, 46(3): 1005–1012. doi: 10.11999/JEIT230260. -
下载: