Performance evaluation of speaker identification algorithms using speech signal features

Шукуров, К.Э., Хасанов, У.К.

Рақамли технологияларнинг назарий ва амалий масалалари · 2026-yil

Annotatsiya

This article analyzes the effectiveness of using different models in speaker recognition processes and selects the best one for the system. In terms of accuracy and speed performance of the system, the classical MFCC + cosine similarity and modern x-vector, ECAPA-TDNN + PLDA architectures are compared. Based on the data set generated from different speakers, the accuracy, f1-score, EER, latency, and GPU load indicators of the models are evaluated. According to the experimental results, the ECAPA-TDNN model outperforms the other models with an accuracy of 95.7%. Since the speaker recognition stage is also important for speaker separation systems, accuracy indicators are of high relevance. The ECAPA-TDNN + PLDA model offers good solutions in terms of using computational resources, working with large data sets, and analyzing their data.

Maqola ma’lumotlari
MualliflarШукуров, К.Э., Хасанов, У.К.
JurnalРақамли технологияларнинг назарий ва амалий масалалари
Nashr sanasi2026-02-26
Jild9
Son1
Betlar80-89
TilRus
DOI10.62132/ijdt.v9i1.325

Kalit so‘zlar

идентификация говорящего, ECAPA-TDNN, x-вектор, MFCC, логарифмическое сходство, PLDA, косинусное сходство, глубокое обучение, вектор признаков, речевая биометрия, AM-softmax, speaker identification, ECAPA-TDNN, x-vector, MFCC, log-Mel, PLDA, cosine similarity, deep learning, feature vector, speech biometrics, AM-softmax

Ilmiy soha

Рақамли технологияларнинг назарий ва амалий масалалари jurnalidan boshqa maqolalar

Рақамли технологияларнинг назарий ва амалий масалалари — barcha maqolalar