Pahari-NER

Published: 23 February 2026| Version 1 | DOI: 10.17632/vh748dtkdj.1
Contributor:
Nadia Mushtaq Gardazi

Description

This dataset represents the first annotated Named Entity Recognition (NER) dataset developed for the Pahari language, created to support research in low-resource natural language processing and linguistic analysis. The corpus contains manually annotated text in which each token is labeled with its corresponding named entity category, such as person, location, organization, and other relevant entities, following standard sequence labeling conventions. The dataset has been prepared in a structured, machine-readable format suitable for training and evaluating computational models including traditional machine learning and deep learning approaches. It is intended to facilitate the development of NER systems, promote reproducible research, and contribute to the advancement of language technology resources for the Pahari language.

Files

Steps to reproduce

Not applicable

Institutions

Categories

Linguistics, Data Science, Natural Language Processing, Machine Learning, Deep Learning

Licence