Annotated Marathi Word Sense Disambiguation Dataset

Published: 6 September 2026| Version 1 | DOI: 10.17632/499w76zjn2.1
Contributors:
,
, Rasika Ransing,

Description

This dataset is a large-scale sense-annotated resource for Word Sense Disambiguation (WSD) in Marathi, a morphologically rich and relatively low-resource Indic language. It was developed to address the limited availability of reusable, sense-annotated Marathi corpora for supervised WSD research. The final dataset contains 50,242 binary sentence–gloss pair records, covering 16 ambiguous Marathi words and 50 distinct senses across different grammatical categories, including nouns, verbs, adjectives, adverbs, and postpositions. The target words and their candidate senses were derived from the Marathi WordNet. Naturally occurring Marathi sentences were collected from established Marathi news portals and blogs, including Loksatta, Esakal, Lokmat, ABP Live Marathi, BBC Marathi, TV9 Marathi, News18 Marathi, Maharashtra Times, Mitraho, Marathi Spandan, and Maayboli. The collected text was cleaned, sentence-segmented, filtered based on sentence length, and duplicate sentences were removed to obtain a suitable corpus for WSD. Since Marathi exhibits substantial inflectional variation, the Stanford Stanza Marathi lemmatizer was used to identify target words in their different morphological forms. A prefix-based matching strategy was additionally applied when lemmatization failed for particular cases. For each word-sense pair, seed sentences were generated using the corresponding Marathi WordNet gloss and manually reviewed for clarity, grammatical correctness, and sense consistency. These verified examples were subsequently used as reference sentences for constructing sense representations. Candidate sentences were assigned to senses using MuRIL token-level embeddings and cosine similarity. A similarity threshold of 0.65 was used for the primary assignment, while a relaxed threshold of 0.60 was used for controlled padding of underrepresented senses. The dataset follows a binary sentence–gloss pair formulation. For every sentence containing an ambiguous target word, the sentence is paired with each candidate sense gloss. The correct sentence–sense pairing receives label 1, whereas incorrect pairings receive label 0. This structure makes the dataset suitable for cross-encoder WSD models and also allows conversion into a conventional multi-class classification format. Each record contains seven fields: sentence, target_word, lemma, pos, sense_id, gloss, and label. The dataset can be used for Marathi WSD benchmarking, supervised learning, transformer-based cross-encoder research, lexical-semantic analysis, contextual representation learning, and NLP research for low-resource Indic languages.

Files

Institutions

Categories

Computer Science, Artificial Intelligence, Computational Linguistics, Natural Language Processing, Indian Language

Licence