MATHEMATICAL MODEL OF AUTOMATIC STOP WORDS DETECTION OF TEXTS IN THE UZBEK LANGUAGE

Madatov, Kh.A., Мадатов, Х.А.

Ҳисоблаш ва амалий математика муаммолари · 2024-yil

Annotatsiya

Filtering stop words is an important task when processing text queries for information retrieval in large data sets. Existing mathematical models of this problem are not suitable for all families of natural languages. For example, they do not cover families of languages to which the Uzbek language may be classified. In this work, an attempt is made to construct a new mathematical model of this problem for Turkic languages, which include Uzbek. This model concerns the so-called agglutinative languages, in which the task of automatically recognizing unimportant words in Uzbek language texts is much more difficult, since stop words are “masked” in the text. This paper proposes a model of a mathematical structure that corresponds to the type of language being studied and allows filtering words that are not essential for information retrieval. This model allows you to compress texts to work with various methods for identifying stop words.

Maqola ma’lumotlari
MualliflarMadatov, Kh.A., Мадатов, Х.А.
JurnalҲисоблаш ва амалий математика муаммолари
Nashr sanasi2024-05-22
Son2
Betlar99-105
TilRus

Kalit so‘zlar

TF-IDF, stop words, tuple, dictionary, unique words, TF-IDF, несущественное слово, кортеж, словарь, уникальные слова

Ilmiy soha

Ҳисоблаш ва амалий математика муаммолари jurnalidan boshqa maqolalar

Ҳисоблаш ва амалий математика муаммолари — barcha maqolalar