METHODOLOGICAL GAPS IN PREPARING DIGITIZED TEXTS FOR SCHOLARLY ANALYSIS
Аннотация
This study investigates the methodological and algorithmic gaps in preparing digitized Uzbek texts for systematic corpus analysis. Due to an over-reliance on exact character matching, national information retrieval systems suffer a 25–40% “silent loss “ of data caused by OCR degradation. Furthermore, the absence of fuzzy matching, lemmatization, and semantic search sharply reduces linguistic query efficiency. To transition these resources from simple electronic collections to high-tech scholarly corpora, this paper underscores the urgent need for multi-layered linguistic annotations aligned with TEI P5 and LAF standards. Ultimately, we propose a strategic roadmap for adopting standardized national TEI schemas and integrating semantic search layers to align Uzbek computational linguistics with modern global NLP standards.