Automated Linguistic Profiling of Disputed Texts in the Uzbek Language: A Model Based on Hybrid Vectorization and Support Vector Machine (SVM)

Raupov, Mekhroj

Lingvospektr · 2026-yil

Annotatsiya

The article proposes a model for the automated linguistic profiling of disputed texts in the Uzbek language. The aim is to develop a methodology that, for forensic-linguistic examination, determines in a quantitative, reproducible and interpretable manner the probabilistic socio-demographic characteristics of an author, namely gender, age and region, as well as the legal classification of a text into insult, defamation or neutral. The methodology integrates Biber’s register analysis, Lakoff’s language-and-gender theory and Nini’s theory of linguistic individuality, adapting them to the agglutinative nature of Uzbek. Features are vectorized using a hybrid TF-IDF and FastText method, while a separate Support Vector Machine classifier is applied to each profiling task. The results are demonstrated through worked examples of TF-IDF weighting, character n-gram extraction and a confusion matrix. The proposed model operates interpretably and accurately under conditions of mixed Latin-Cyrillic writing and morphological richness. Thus, the study offers a codeable, interpretable and ethically constrained model for Uzbek forensic linguistics.

Maqola ma’lumotlari
MualliflarRaupov, Mekhroj
JurnalLingvospektr
Nashr sanasi2026-07-03
Jild6
Son1
Betlar54-61
TilIngliz

Kalit so‘zlar

linguistic profiling, disputed text, forensic linguistics, support vector machine, hybrid vectorization, character n-grams, Uzbek language, idiolect

Ilmiy soha

Lingvospektr jurnalidan boshqa maqolalar

Lingvospektr — barcha maqolalar