While large-scale pre-trained models have significantly advanced multilingual Automatic Speech Recognition (ASR), many low-resource languages remain under-served due to the scarcity of high-quality annotated speech corpora. This paper introduces the Karakalpak Speech Corpus (KSC), the first publicly available benchmark dataset for Karakalpak, a Turkic language spoken by over two million people primarily in Karakalpakstan. The corpus encompasses 50 hours of predominantly read speech. The data was collected from 25 native speakers with a balanced gender distribution. To establish a performance benchmark, we fine-tuned the Wav2Vec 2.0 architecture, specifically evaluating the efficacy of transfer learning from multilingual pre-trained models
| Mualliflar | Kudaybergenov, Tangirbergen, Uteuliev, Niyetbay, Khudaybergenov, Kabul, Kudaybergenov, Jabbar |
|---|---|
| Jurnal | Al-Farg'oniy avlodlari |
| Nashr sanasi | 2026-03-18 |
| Son | 1 |
| Betlar | 262-269 |
| Til | Rus |
speech dataset, speech recognition, speech-to-text, Transfer Learning, Wav2Vec 2.0 model, Machine Learning, Deep Learning
This article describes the research conducted by training the neural network model proposed by NVIDIA in speech recognition with a speech dataset of the Uzbek language. Currently, various models of neural networks are…
This article studies the effectiveness of modern deep learning models for the task of deep contextual analysis of text data in social networks. As part of the study, the ability of RNN, LSTM and DistilBERT models based…
In this work, methods and algorithms for improving the quality of a given image are studied. Methods for reducing noise, improving contrast, increasing sharpness, and clearly displaying image elements during image…
This article studies the technical architecture and mathematical model of an adaptive digital environment based on the integration of AutoCAD + Python + AI in teaching engineering geometry and computer graphics. Within…
This paper proposes a speech recognition system for automatically recognizing separately pronounced Uzbek words. The system uses MFCC for acoustic feature extraction and a Gaussian Hidden Markov Model for word modeling…
This article describes a solution for systematizing the approach to personal data protection (PDP) using the method of de-identification in the context of regulatory pressure and growing cyber threats. It proposes…
This article analyzes two methods of adversarial attacks created against machine — training — based pest program detection systems-JSMA (Jacobian Saliency Map Attack) and Carlini & Wagner (C&W) - in a practical…
Currently, alongside the rapid development of information and communication technologies, not only the number of malware programs but also their functionality is increasing. This necessitates the use of intelligent…
This article presents a study of a dictionary-based pre-filter for safe educational content generation in a school environment. An Android application based on the MVVM architecture was developed as an experimental…
Abstract: This article proposes algorithms and software integration models for effective data analysis in the DNA (deoxyribonucleic acid) identification system based on STR (Short Tandem Repeat) markers. In the course…