ISSN :2582-9793

An Intelligent Speech Analytics Framework for Depression Detection: Comparative Assessment of Machine Learning Architectures

Original Research (Published On: 10-Oct-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.65351

Deepak Kumar and Monika Khatkar

Adv. Artif. Intell. Mach. Learn., - (-):-

1. Deepak Kumar: K R Mangalam University- Gurgaon

2. Monika Khatkar: Assistant Professor- School Of Engineering and TechnologyK R Mangalam University- Gurgaon

Download PDF Here

DOI: 10.54364/AAIML.2026.65351

Article History: Received on: 15-Jun-26, Accepted on: 03-Oct-26, Published on: 10-Oct-26

Corresponding Author: Deepak Kumar

Email: javadevdeepak@gmail.com

Citation: Deepak Kumar and Monika Khatkar. An Intelligent Speech Analytics Framework for Depression Detection: Comparative Assessment of Machine Learning Architectures. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.65351


Abstract


ABSTRACT:

Depression affects emotional, cognitive, and behavioral functioning and it is a pretty ubiquitous mental health (MH) disorder. It is often described as subjective time consuming, and kind of expert-driven, so clinical interviews plus self reported questionnaires are still the traditional methods, for diagnosing.. Automated systems can detect depressed symptoms utilising voice biomarkers with the aid of AI and speech signal processing (SSP). Emotional and psychological states can greatly affect the variation in pitch, energy, pace, and spectrum characteristics of speech, rendering it a potential tool to assess depression. This research is designed to develop an independent depression detection system (DDS) using speech and voice (S&V) input and evaluate ML and DL algorithms on depression and non-DD. The audio files should be used to get the data, labels should be extracted from the audio files, features should be taken from preprocessed data, the model (Mod.) should be trained with the extracted features, and the performance of the Mod. should be evaluated. Librosa extracts Mel Frequency Cepstral Coefficients (MFCCs), Chroma features, and Mel Spectrogram (MS) features from speech and is able to gather them in a hybrid feature vector. We split the data straight into training (80%) and test (20%) sets in a normal random fashion. The accuracy, precision, recall for the MLP Mod. were observed as 98% , 99% and 99% respectively, while the ROC-AUC score came out to 0.96 in the trials. Overall the results propose that S&V signals can be leveraged for depression identification and to strengthen automated MH assessments. Also, screening for depression plus clinical decision support with speech AI appears to be efficient, non-invasive alongside scalable.


Statistics

Article Views: 31
PDF Downloads: 4