MODELING THE DEVELOPMENT OF LEMMATIZATION ALGORITHMS FOR THE UZBEK LANGUAGE

ТАТУ хабарлари · 2025-yil

Annotatsiya

Lemmatization is a vital task in Natural Language Processing (NLP), essential for various applications.However, the Uzbek language, categorized as a lowresource language in NLP, lacks dedicated lemmatization systems.This deficiency hinders the development of advanced NLP tools for Uzbek texts.To address this, we present a structured approach using IDEF0 and IDEF1X models to design an Uzbek lemmatization system.The IDEF0 model details the system's functionality, outlining the lemmatization process, while the IDEF1X model illustrates the relationships between data components.Furthermore, we developed a rule-based system that leverages Uzbek morphology to accurately determine word lemmas.This research contributes a well-defined architectural framework and a practical rule-based solution for lemmatization, aiding in the advancement of NLP for Uzbek and similar low-resource, morphologically rich languages.

Maqola ma’lumotlari
JurnalТАТУ хабарлари
Nashr sanasi2025-04-11
DOI10.61663/251tuitmct5

ТАТУ хабарлари jurnalidan boshqa maqolalar

ТАТУ хабарлари — barcha maqolalar