FUNDAMENTAL METHODS FOR IDENTIFYING AI-GENERATED DATA IN ACADEMIC AND ONLINE PUBLICATIONS USING MACHINE LEARNING AND NLP MODELS

Abdullaev , Munis, Kungratov, Ilmurod

Innovation science and technologiy · 2026-yil

Annotatsiya

This article provides a comprehensive scientific analysis of the fundamental and practical methodologies for identifying textsgenerated by Artificial Intelligence (AI) within academic and digital publishing ecosystems. The rapid maturation of generative languagearchitectures is fundamentally transforming traditional copyright paradigms and the principles of academic integrity. The primaryobjective of this research is to develop, test, and evaluate innovative methodologies for distinguishing synthetic data using NaturalLanguage Processing (NLP) and Machine Learning (ML) classification models. The study comparatively evaluates the effectivenessof stylometric feature extraction, zero-shot probability distribution analysis, and transformer-based deep learning classifiers. Empiricalresults confirm that traditional plagiarism systems based on exact lexical matching have completely lost their functional viability.Concurrently, the hybrid-ensemble architecture proposed in this study demonstrated high resilience against complex adversarialevasion attacks. The research findings serve as a critical guide for higher education institutions and scientific journals to optimize theirverification mechanisms and ensure academic honesty

Maqola ma’lumotlari
MualliflarAbdullaev , Munis, Kungratov, Ilmurod
JurnalInnovation science and technologiy
Nashr sanasi2026-06-01
Jild2
Son6
TilIngliz
DOI10.5281/zenodo.20780989

Kalit so‘zlar

Synthetic, text, generative, architecture, natural, language, processing, transformer, academic, integrity, stylometry, verification, classifier, probability, algorithm

Ilmiy soha

Innovation science and technologiy jurnalidan boshqa maqolalar

Innovation science and technologiy — barcha maqolalar