UzLex Dataset: Core Uzbek Lemmas, Sentiment Lexicon, and PER/ORG/LOC/ABBR Gazetteers with Quality Control

Published: 16 February 2026| Version 1 | DOI: 10.17632/yrjbpkzxbh.1
Contributor:
Bobur Saidov

Description

UzLex is a curated multi-table Uzbek lexical dataset comprising (i) a core lemma lexicon with POS labels and quality-control flags, (ii) a sentiment lexicon with polarity labels and uncertainty flags, and (iii) gazetteers for NER (PER/ORG/LOC/ABBR). The release includes audit artifacts, quality summaries, and high-precision subsets for strict downstream use

Files

Steps to reproduce

Download the files uzlex_core.csv, uzlex_sentiment.csv, and uzlex_gazetteers.csv (or the full release archive). Verify file integrity using SHA256SUMS.csv (optional but recommended). Load tables in Python/R (CSV) or any spreadsheet tool. UTF-8 encoding is used. For high-precision usage, load the subset files: uzlex_core_high_precision.csv (review_bucket=ok; confidence≥0.75) uzlex_sentiment_high_precision.csv (uncertain=0) uzlex_gazetteers_high_precision.csv (PER/ORG/LOC/ABBR only) Optional tables are provided: uzlex_numeric_tokens.csv (numeric items separated from the core lexicon) uzlex_roles.csv (role/title entries separated from gazetteers) Column definitions and QA notes are available in data_dictionary.csv, stats.csv, audit_report.xlsx, and QUALITY_ASSURANCE_SECTION.md.

Institutions

Categories

Computer Science, Artificial Intelligence, Information Retrieval, Computational Linguistics, Data Science, Natural Language Processing

Licence