TYPOLOGY OF LINGUISTIC DISTORTIONS OCCURRING IN OCR-GENERATED TEXTS

Auteurs

  • Mohinur Alikulova 1 Автор

Résumé

 

This study investigates the typology and quantitative impact of linguistic distortions caused by Optical Character Recognition (OCR) in digitized Uzbek texts. Analyzing an experimental corpus of 250,000 words across different historical writing systems (Latin, Cyrillic, and Arabic scripts), the research evaluates errors based on character (CER), word (WER), and morpheme (MER) error rates. The findings reveal that graphical errors constitute the dominant distortion category, accounting for 58–65% of all OCR bugs, driven primarily by visual homomorphy and diacritical omissions. Since graphical substitutions trigger chain-reaction distortions at the lexical and morphological levels, this paper underscores the critical need for developing automated multi-system post-correction frameworks to preserve research validity in Uzbek digital philology. 

Téléchargements

Publiée

2026-08-20

Numéro

Rubrique

Articles

Comment citer

Alikulova, M. (2026). TYPOLOGY OF LINGUISTIC DISTORTIONS OCCURRING IN OCR-GENERATED TEXTS. International Conference on Social Sciences & Humanities, 2(8), 58-61. https://uniconflix.com/index.php/ICSH/article/view/5734