METADATA AND FORMATTING ISSUES IN THE UZBEK-ENGLISH PARALLEL CORPUS AND EXISTING NLP TOOLS FOR THE UZBEK LANGUAGE

Elov, Botir, Amirkulov, Ma’rufjon, Suyunova, Malika

Techscience.uz - техника фанлари долзарб масалалри · 2025-yil

Annotatsiya

This article explores the issues of metadata formatting, the use of TEI and CoNLL-U standards, and the analysis of existing Natural Language Processing (NLP) tools for the Uzbek language in the process of creating an Uzbek-English parallel corpus. The paper discusses each stage of corpus development, including text alignment, syntactic and morphological annotation, and structural encoding. Furthermore, it evaluates the performance of Uzbek morphological analyzers, lemmatizers, and POS taggers, emphasizing their practical significance in constructing high-quality bilingual corpora. The results of the study provide a methodological basis for accurately encoding, automatically analyzing, and integrating parallel corpora into linguistic search and processing systems.

Maqola ma’lumotlari
MualliflarElov, Botir, Amirkulov, Ma’rufjon, Suyunova, Malika
JurnalTechscience.uz - техника фанлари долзарб масалалри
Nashr sanasi2025-10-23
Jild3
Son9
Betlar14-21
TilO‘zbek
DOI10.47390/ts-v3i9y2025no3

Kalit so‘zlar

parallel corpus, metadata, TEI, CoNLL-U, Uzbek language, NLP, morphological analysis, tagging, lemmatization, linguistic resources., parallel korpus, metama’lumot, TEI, CoNLL-U, o‘zbek tili, NLP, morfologik tahlil, teglash, lemmatizatsiya, lingvistik resurs

Ilmiy soha

Techscience.uz - техника фанлари долзарб масалалри jurnalidan boshqa maqolalar

Techscience.uz - техника фанлари долзарб масалалри — barcha maqolalar