UzNER-5Style: A 75,000-Sentence Multi-Style Uzbek Named Entity Recognition Dataset

Published: 26 August 2026| Version 1 | DOI: 10.17632/frw4m5j77d.1
Contributor:
Javohirjon Nazarov

Description

UzNER-5Style is a large-scale Uzbek named entity recognition dataset created across five functional styles: official, journalistic, scientific, literary, and conversational. The dataset contains 75,000 sentences, with 15,000 sentences for each style. The corpus is annotated using the BIOES tagging scheme and includes 30 named entity types. These include person names, organizations, geopolitical entities, dates, monetary values, quantities, durations, percentages, locations, events, laws, documents, positions, phone numbers, e-mail addresses, URLs, times, titles, works of art, academic degrees, and several other entity categories. The final dataset contains approximately 750,000 token-level records. Each record includes a token identifier, source information, functional style, sentence identifier, token, BIOES tag, domain information, and validation status. The dataset was synthetically generated and then subjected to several quality-control stages. These included duplicate checking, BIOES sequence validation, entity-boundary correction, semantic error correction, and text-level cleaning. A subset of 5,000 sentences, consisting of 1,000 sentences from each functional style, was independently reviewed by two annotators. Disagreements were subsequently resolved to obtain final adjudicated labels for the validation subset. This dataset is intended to support research on Uzbek named entity recognition, low-resource natural language processing, cross-style generalization, domain-aware NER, and robustness across different functional styles.

Files

Steps to reproduce

Prepare sentence templates for five Uzbek functional styles: official, journalistic, scientific, literary, and conversational. Insert named entities from the 30 predefined entity categories into the sentence templates. Tokenize each generated sentence. Assign BIOES labels to the named entity spans. Check BIOES sequences and remove invalid or inconsistent annotations. Review duplicate sentences, incorrect entity boundaries, false-positive entities, and unnatural sentence patterns. Select 1,000 sentences from each style, giving a total validation subset of 5,000 sentences. Have two annotators independently review the selected sentences. Compare the annotations using token-level agreement, Cohen’s kappa, entity-type agreement, and span-level precision, recall, and F1. Resolve annotation disagreements and use the adjudicated labels in the final dataset.

Institutions

Categories

Computer Science, Natural Language Processing, Database

Licence