UPM Institutional Repository

Improved speech enhancement using parallel MVDR beamforming and Coherent-to-Diffuse Power Ratio post-filtering


Citation

Natarajan, Sureshkumar and Hassan, Mohd Khair and Raja Ahmad, Raja Kamil and Ahmad, Faisul Arif and Azrad, Syaril and Macleans, June Francis and Abdalla, Hussna Elnoor M. and Saparkhojayev, Nurbek and Al-Haddad, Syed Abdul Rahman (2026) Improved speech enhancement using parallel MVDR beamforming and Coherent-to-Diffuse Power Ratio post-filtering. Pertanika Journal of Science and Technology, 34 (2). pp. 951-973. ISSN 0128-7680; eISSN: 2231-8526

Abstract

Speech communication involves the exchange of information between two or more individuals; however, background noise and reverberation often degrade speech clarity and intelligibility. The impact of these degradations depends on factors such as the number, intensity, and spatial characteristics of noise sources, as well as reflections from surrounding surfaces. These effects pose significant challenges for applications including teleconferencing, hearing-aid devices, and human-machine voice interfaces. This study aims to address the combined effects of background noise and reverberation by proposing a speech enhancement framework that integrates Minimum Variance Distortionless Response (MVDR) beamforming with a Coherent-to-Diffuse Power Ratio (CDR) based post-filter. The methodology relies on parallel processing, where MVDR beamforming performs spatial noise suppression, while CDR values are estimated from microphone-domain signals and used to compute post-filter gains for suppressing residual diffuse noise and reverberation in the beamformer output. The proposed algorithm was evaluated over an input signal-to-noise ratio range from 0 dB to 40 dB and compared against the Integrated Sidelobe Cancellation Linear Prediction (ISCLP) baseline using four objective metrics: Perceptual Evaluation of Speech Quality (PESQ), Extended Short-Time Objective Intelligibility (ESTOI), Cepstral Distance (CD), and Weighted Spectral Slope Distance (WSS). The results demonstrate consistent improvements in PESQ and reductions in CD and WSS across all tested conditions, while gains in ESTOI remain modest. These findings indicate improved spectral fidelity and speech naturalness, highlighting the practical relevance of the proposed framework for real-time, low-complexity speech enhancement in noise and reverberation-prone communication systems.


Download File

[img] Text
126593.pdf - Published Version
Available under License Creative Commons Attribution Non-commercial No Derivatives.

Download (1MB)

Additional Metadata

Item Type: Article
Subject: Computer Science (all)
Subject: Chemical Engineering (all)
Subject: Environmental Science (all)
Divisions: Faculty of Engineering
DOI Number: https://doi.org/10.47836/pjst.34.2.15
Publisher: Universiti Putra Malaysia Press
Keywords: Acoustics; CDR; Dereverberation; MVDR beamformer; Noise suppression; Speech enhancement
Sustainable Development Goals (SDGs): SDG 9: Industry, Innovation and Infrastructure, SDG 11: Sustainable Cities and Communities, SDG 3: Good Health and Well-being
Depositing User: Ms. Siti Radziah Mohamed@mahmod
Date Deposited: 30 Jun 2026 07:22
Last Modified: 30 Jun 2026 07:22
Altmetrics: https://www.altmetric.com/details.php?domain=psasir.upm.edu.my&doi=10.47836/pjst.34.2.15
URI: http://psasir.upm.edu.my/id/eprint/126593
Statistic Details: View Download Statistic

Actions (login required)

View Item View Item