A Model for Synthetic Speech Detection Based on Emotional Variations Using Artificial Intelligence

ТАТУ хабарлари · 2025-yil

Annotatsiya

Today, the automatic recognition of emotions in human speech is gaining increased importance due to the growing role of speech interfaces in humancomputer interaction applications.In this context, the present study analyzes several methods for detecting synthetic speech and proposes a classification approach based on emotion recognition.Five primary human emotions -anger, boredom, happiness, neutral state, and sadness -are considered, and speech that does not correspond to any of these emotional categories is investigated as potentially synthetic.To achieve more accurate results, statistical methods were employed to combine various feature streams.For emotion detection in speech, a feature vector was constructed by combining the following acoustic features: 16 Linear Predictive Coding (LPC) coefficients, 12 Linear Predictive Cepstral Coefficients (LPCC), 16 Log Frequency Power Coefficients (LFPC), 16 Perceptual Linear Prediction (PLP) coefficients, 20 Mel-Frequency Cepstral Coefficients (MFCC), and jitter.The proposed method is based on three classification techniques: Linear Discriminant Analysis (LDA), k-Nearest Neighbors (K-NN), and Hidden Markov Models (HMM).The experimental results demonstrate that the selected features are stable and effective in recognizing emotions and yield high performance based on two evaluation metrics.

Maqola ma’lumotlari
JurnalТАТУ хабарлари
Nashr sanasi2025-10-22
DOI10.61663/253tuitmct2

ТАТУ хабарлари jurnalidan boshqa maqolalar

ТАТУ хабарлари — barcha maqolalar