EFFICIENCY AND EXISTING PROBLEMS OF ARTIFICIAL INTELLIGENCE MODELS IN THE DIGITALIZATION OF LOW-RESOURCE LANGUAGES

Aimbetova , Gulara, Sarsenbaeva, Hu’rlixa, Djumabaev , Alpamis

Techscience.uz - техника фанлари долзарб масалалри · 2026-yil

Annotatsiya

This article analyzes current problems in the development of artificial intelligence and natural language processing (NLP) models for low-resource languages, in particular, the Karakalpak language. The main goal of the research is to develop a methodology for data refinement and normalization to improve model accuracy under conditions of data scarcity. In the methods section, special normalization algorithms and Byte Pair Encoding (BPE) tokenization were used for specific letters of the Karakalpak language. The results showed that the proposed preprocessing methods made it possible to significantly reduce the model error (Loss) and increase the translation and search accuracy (BLEU score) from 12.4 to 24.8. The research findings serve as a theoretical and practical basis for creating digital dictionaries and intelligent systems for low-resource languages.  

Maqola ma’lumotlari
MualliflarAimbetova , Gulara, Sarsenbaeva, Hu’rlixa, Djumabaev , Alpamis
JurnalTechscience.uz - техника фанлари долзарб масалалри
Nashr sanasi2026-05-14
Jild4
Son5
Betlar43-48
TilO‘zbek
DOI10.47390/ts-v4i5y2026n07

Kalit so‘zlar

NLP, low-resource languages, Karakalpak language, preprocessing, normalization, artificial intelligence, machine translation, NLP, kam resursli tillar, qoraqalpoq tili, preprocessing, normallashtirish, sun’iy intellekt, mashinali tarjima.

Ilmiy soha

Techscience.uz - техника фанлари долзарб масалалри jurnalidan boshqa maqolalar

Techscience.uz - техника фанлари долзарб масалалри — barcha maqolalar