Original Research (Published On: 03-Oct-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.65348Parveen Lehana and Arfana Chowdhary
Adv. Artif. Intell. Mach. Learn., - (-):-
1. Parveen Lehana: Department of Electronics University of Jammu India
2. Arfana Chowdhary: DSP Laboratory, Department of Electronics, University of Jammu Jammu-180006, India
DOI: 10.54364/AAIML.2026.65348
Article History: Received on: 06-Feb-26, Accepted on: 26-Sep-26, Published on: 03-Oct-26
Corresponding Author: Parveen Lehana
Email: pklehana@gmail.com
Citation: Arfana Chowdhary and Parveen Kumar Lehana. Performance Evaluation of DenseNet-121 and Bi-LSTM with Attention Mechanism based Vocal Signal Separation. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.65348
Abstract
Audio source separation is
the process of separating a mixed audio signal into its individual sources.
This task remains difficult because of significant overlap in time and
frequency. The presence of noise and fluctuating signal characteristics further
deteriorate the task. Traditional signal processing methods rely on strong
statistical assumptions, which often limit their effectiveness in real-world
acoustic signal processing. To overcome some of these challenges, this paper
presents a supervised attention-based deep learning approach for separating
single-channel audio signals. This approach
directly predicts the individual source magnitude spectrograms without
employing traditional time-frequency masking methods. The approach is based on
fixed-length segments of audio signals, which are represented at a constant
rate and mapped to magnitude spectrograms using short-time Fourier transform
(STFT). A structured preprocessing employed in the algorithm ensure normalisation,
cropping, and temporal alignment, maintains constant input data size during
training and testing phases. The separation approach employs a DenseNet-121
network that uses hierarchical spectral feature extraction. Thereby enabling
the effective extraction of complex harmonic and formant characteristics. To
predict time-frequency regions containing salient perceptual significance, circular
spatial attention mechanism has been employed for enabling the network to
concentrate on significant acoustic elements. The temporal characteristics and
contextual parameters from spectrum frames are captured using a bidirectional
long short-term memory (Bi-LSTM) network and multi-head self-attention. The combined feature representations are
processed through fully connected layers of the DenseNet-121 network to estimate
the magnitude spectrograms of the constituent sources. The setup allows for end-to-end learning of nonlinear source mappings.
The network is trained using scale-invariant signal-to-distortion ratio
(SI-SDR) as loss function and the trained network is evaluated using objective
metrics, such as short-time objective intelligibility (STOI) and perceptual
evaluation of speech quality (PESQ). The experimental results
demonstrate satisfactory performance of separation, intelligibility, and
perceptual quality across different mixing ratios. The approach establishes a low
complexity, scalable, and interpretable solution to realistic problems in
source separation.
Statistics
Article Views: 3
PDF Downloads: 0
