UzWordNet-GoldCore: A Gold-Core Uzbek Lexical-Semantic WordNet Dataset

Published: 18 June 2026| Version 1 | DOI: 10.17632/v7m7v3jdn5.1
Contributors:
, Ruhillo Alaev

Description

UzWordNet-GoldCore is an open Uzbek lexical-semantic dataset organized in a WordNet-style structure. The dataset contains 11,178 lexical records and 14 fields, covering Uzbek words, synonym sets, antonyms, definitions, part-of-speech labels, style labels, domain labels, ontology categories, example sentences, and semantic relations such as hypernyms, hyponyms, holonyms, and meronyms. The dataset includes 6,506 unique words and 2,706 unique definitions. It is provided as a UTF-8 tab-separated values file and is intended for Uzbek natural language processing research and development, including lexical-semantic analysis, synonym and antonym modeling, semantic relation extraction, dictionary-based applications, and evaluation of Uzbek language models. The data contains normalized Uzbek Latin text. Whitespace, apostrophe variants, and character-level inconsistencies were standardized to support reliable computational processing. Optional relation fields are left empty when a relation is not available.

Files

Categories

Natural Language Processing, Machine Learning

Licence