This research is dedicated to the development of ModernUzBERT, an advanced embedding model for the Uzbek language based on contemporary Transformer architecture with an extended context window of 8192 tokens. In low-resource and agglutinative languages such as Uzbek, the limited capacity for processing comprehensive context significantly hampers model performance. To address this challenge, a large-scale text corpus in the Uzbek Latin script comprising over 125 million tokens was curated, and an optimized vocabulary of 52,000 sub-word units was developed. The model training process was executed in two primary stages: initially, the model was pre-trained from scratch using the Masked Language Modeling (MLM) objective; subsequently, it was adapted for semantic search via Supervised Fine-Tuning (SFT) using a dataset of 30,000 question-answer pairs. Through the implementation of Flash Attention 2 and Unpadding technologies, GPU resource utilization efficiency was substantially enhanced. The 8192-token context window enables the analysis of large-scale documents while preserving semantic integrity. Upon conclusion of the study, the developed model and all its components were released for open-access to the scientific community on the Hugging Face platform.
| Mualliflar | Хужаяров, И.Ш., Очилов, М.М., Холматов, О.А., Жуманов, В.И. |
|---|---|
| Jurnal | Рақамли технологияларнинг назарий ва амалий масалалари |
| Nashr sanasi | 2026-05-15 |
| Jild | 9 |
| Son | 2 |
| Betlar | 45-53 |
| Til | Rus |
| DOI | 10.62132/ijdt.v9i2.375 |
DOI: 10.62132/ijdt.v9i2.375 · Maqolaning asl sahifasi
NLP, узбекский язык, ModernBERT, семантические эмбеддинги, информационный поиск, длинный контекст, NLP, uzbek language, ModernBERT, semantic embedding, information retrieval, long context, flash attention
This article examines methods of analyzing water composition and the possibilities of using modern technologies. The study compared traditional physical, chemical, and biological methods of water quality assessment, as…
In this paper, using the deformation compatibility condition, similar to the well-known displacement equation, differential deformation equations are written. These, in combination with the equilibrium equation and…
This article analyzes modern scientific research conducted in the field of Smart Environment infrastructure and examines existing approaches to the development of decision support models. In recent years, due to the…
Modern methods of data and space image processing - filtering, segmentation, classification, and restoration - require precise setup of a multitude of continuous parameters that determine the behavior of algorithms and…
Substitution boxes, commonly known as S-boxes, are one of the most important nonlinear components of modern symmetric-key block encryption algorithms. Their main task is to introduce confusion and nonlinear…
Semi-supervised Density Peak (SDenPeak) algorithm is known to be efficient and simple in tasks clustering. It improves clustering performance by adding pair-wise constraints, must-link and cannot-link constraints, that…
This paper proposes a new feature extraction method for speaker recognition, called Delta-Gated Spike Encoding (DGSE). The proposed approach combines log-Mel spectrograms, temporal delta, adaptive thresholding, an…
This review paper systematically examines modern architectures, algorithms, and approaches aimed at reducing latency and dynamically managing resources in real-time big data stream processing systems. The…
This study is devoted to modeling and optimizing information exchange processes in a virtual platform for local product trade. The research object is the information flow within the digital platform, while the subject…
This article examines monitoring seasonal changes in reservoir area based on Sentinel-2 data using the Google Earth Engine platform and the NDWI water index based on a time series for 2024, 2025, and the first three…
Рақамли технологияларнинг назарий ва амалий масалалари — barcha maqolalar