Methodology for parametrically efficient adaptation of large language models to agglutinative languages ​​with a deficit of training data: the experience of the Uzbek language

Jumakulov, Bakhodir, Гаипназаров , Рустам, Хусайдинова , Дилобар

Al-Farg'oniy avlodlari · 2026-yil

Annotatsiya

The article proposes a theoretical and methodological framework for the parameter-efficient adaptation of pre-trained large language models (specifically the LLaMA-3 and Mistral-7B classes) to the Uzbek language under conditions of data scarcity. The study examines the morphological constraints of the agglutinative structure of the Uzbek language in the context of BPE tokenization, a methodology for constructing a minimally sufficient corpus, the principles for selecting LoRA/QLoRA hyperparameters, and a generation quality evaluation system that accounts for morphological specificities. The practical significance of this approach lies in its application within Uzbekistan's e-government systems and educational technologies.

Maqola ma’lumotlari
MualliflarJumakulov, Bakhodir, Гаипназаров , Рустам, Хусайдинова , Дилобар
JurnalAl-Farg'oniy avlodlari
Nashr sanasi2026-04-28
Son2
Betlar46-54
TilRus

Kalit so‘zlar

large language models, parameter-efficient fine-tuning, LoRA, QLoRA, Uzbek language, agglutinative morphology, low-resource NLP, LLaMA, Mistral., большие языковые модели, параметрически эффективная настройка, LoRA, QLoRA, узбекский язык, агглютинативная морфология, низкоресурсный NLP, LLaMA, Mistral., katta til modellar, parametr jihatdan samarali sozlash, LoRA, QLoRA, o‘zbek tili, agglutinativ morfologiya, past resursli NLP, LLaMA, Mistral.

Ilmiy soha

Al-Farg'oniy avlodlari jurnalidan boshqa maqolalar

Al-Farg'oniy avlodlari — barcha maqolalar