ANALYZING AUDIO FEATURES TO DISTINGUISH HUMAN VOICE FROM AI

Muxriddin Abduganiev, Mamura Uzakova, Aygul Burxanova

Innores · 2026-yil

Annotatsiya

Telling a real human voice apart from AI-generated speech isbecoming incredibly important for things like security, forensics, and everyday tech.In this study, a more detailed analysis was performed by studying three acousticfeatures: Mel-Frequency Cepstral Coefficients (MFCC), Zero Crossing Rate (ZCR),and Spectral Centroid. The audio was preprocessed in Python using the librosalibrary, with the silent parts of the audio removed. The results were quite clear: thespectral centroid of real human speech is higher, and the variation of the MFCCs ismuch greater. Interestingly the ZCR values did not differ greatly between the two. Inconclusion, the results presented here demonstrate that simple audio parameters canbe effectively used for automatic detection of synthetic voices

Maqola ma’lumotlari
MualliflarMuxriddin Abduganiev, Mamura Uzakova, Aygul Burxanova
JurnalInnores
Nashr sanasi2026-05-30
Jild2
Son5
Betlar257-265
TilIngliz

Kalit so‘zlar

speech synthesis, acoustic features, MFCC, zero crossing rate, spectral centroid, librosa, human-computer interaction, voice forensics

Ilmiy soha

Innores jurnalidan boshqa maqolalar

Innores — barcha maqolalar