An Integrated Analysis of Multilingual Texts Spanning Dual Alphabets

Адилова, Ф.Т., Давронов, Р.Р., Сафаров, Р.А.

Рақамли технологияларнинг назарий ва амалий масалалари · 2023-yil

Annotatsiya

Language recognition in natural language processing (NLP) aims to determine the specific language of a text or document. As the number of languages increases, this task becomes more complex. This study introduces a detailed model for detecting languages from text, with an emphasis on the Latin-Cyrillic script of the Uzbek language. Noting the research gap in this domain, we unveil a precise Uzbek Latin-Cyrillic script recognition model leveraging an apt transformer architecture. The model was tested on our self-compiled Uzbek language corpus, which also offers a robust benchmark for subsequent Uzbek language identification studies. Our approach encompasses 21 languages, including Uzbek, across both Latin and Cyrillic alphabets. Our findings highlight that the XLM-RoBERTa transformer-driven language detection model significantly outperforms its predecessors in terms of accuracy and efficiency.

Maqola ma’lumotlari
MualliflarАдилова, Ф.Т., Давронов, Р.Р., Сафаров, Р.А.
JurnalРақамли технологияларнинг назарий ва амалий масалалари
Nashr sanasi2023-10-02
Jild5
Son3
Betlar47-56
TilIngliz

Kalit so‘zlar

NLP, Multilingual Language Models, Cloud Natural Language API, Open AI, ChatGPT, model compression, transformer, NLP, Многоязычные языковые модели, Облачный API естественного языка, Открытый ИИ, ChatGPT, сжатие модели, преобразователь

Ilmiy soha

Рақамли технологияларнинг назарий ва амалий масалалари jurnalidan boshqa maqolalar

Рақамли технологияларнинг назарий ва амалий масалалари — barcha maqolalar