TOKENS AT THE BOUNDARY OF THE TAGSET IN A MANUALLY ANNOTATED UZBEK MICROCORPUS
Abstract
Annotation projects normally report what a tagset covers. This paper reports what it does not. In a manually annotated 12,048-token microcorpus of written Uzbek, 575 tokens carry a transposition flag; each was audited against five functional directions of transposition established for Uzbek. Five hundred and twelve tokens (89.0 per cent) matched one direction; 63 (11.0 per cent) were kept as an explicit residual instead of being assigned to the nearest available label. Setting the formal pre-classification, built from tags and morphological features, against the contextual audit shows that formal cues predict function unevenly: the dominant direction lost 6.5 per cent of its members, while the two peripheral directions lost 31.5 and 51.6 per cent. The residual is not confined to one style (31 / 16 / 14 across fiction, academic and journalistic samples) and draws on four source categories. The size and internal composition of a residual class is therefore a measurable property of an annotation scheme, and reporting it openly protects the counts that follow rather than weakening them.