Dataset of Khorezm dialect words of Uzbek language

Published: 14 May 2025| Version 1 | DOI: 10.17632/txrk7jm6x3.1
Contributor:
Davlatyor Mengliev

Description

As part of the study, a dataset was formed, which was used by a rule-oriented algorithm to standardize dialect forms into formal equivalents. In particular, the dataset contains 2249 dialect words: 1) The words in this dataset were compiled thanks to the joint work of expert linguists who are well versed not only in the Uzbek (formal) language, but also in the dialect forms of this language. 2) The sources of words in the dataset were not only expert linguists, but also native speakers of the dialect language, who are also co-authors of this article. The dataset was formed manually, no automation processes were carried out except for cases of transliteration of Cyrillic into Latin.

Files

Categories

Natural Language Processing, Uzbekistan, Database

Licence