Development of a system for real-time translation and voice synthesis of english audiovisual content into uzbekistan based on artificial intelligence

Ахмедов, Муроджон, Омаджон, Уришев

Al-Farg'oniy avlodlari · 2026-yil

Annotatsiya

This paper presents an AI-based pipeline that translates English YouTube videos into Uzbek and produces synchronized dubbed speech in near real time. The system integrates audio acquisition, preprocessing (16 kHz mono PCM), automatic speech recognition with OpenAI Whisper, neural machine translation, and text-to-speech synthesis. A task-queue architecture with status and log monitoring improves robustness for long videos, while fallback translation and TTS services maintain continuity under failures. Experiments on diverse video genres demonstrate practical latency and intelligible Uzbek output, making the approach suitable for localization, education, and media applications.

Maqola ma’lumotlari
MualliflarАхмедов, Муроджон, Омаджон, Уришев
JurnalAl-Farg'oniy avlodlari
Nashr sanasi2026-02-28
Son1
Betlar74-79
TilRus

Kalit so‘zlar

искусственный интеллект, распознавание речи, машинный перевод, синтез речи, перевод в реальном времени, мультимедийный контент, artificial intelligence, speech recognition, machine translation, speech synthesis, , real-time translation, multimedia content

Ilmiy soha

Al-Farg'oniy avlodlari jurnalidan boshqa maqolalar

Al-Farg'oniy avlodlari — barcha maqolalar