NERSkill.Id

Name: NERSkill.Id
Creator: Meilany Nonsi Tentua
Published: 2024-04-02T15:42:58.508Z
Keywords: Natural Language Processing, Indonesian Language, Recognition, Text Mining

Tentua, Meilany Nonsi; Suprapto, Suprapto; Afiahayati, Afiahayati

doi:10.17632/5s8r9ndfvc.3

NERSkill.Id

Published: 2 April 2024| Version 3 | DOI: 10.17632/5s8r9ndfvc.3

Contributors:

Meilany Nonsi Tentua, Suprapto Suprapto, Afiahayati Afiahayati

Description

NERSkill.Id stands out as the initial annotated corpus designed specifically for NER datasets emphasizing skill entities in the Indonesian language. This marks a valuable addition to the existing resources for Natural Language Processing (NLP) in Indonesian. Despite its relatively compact size, NERSkill.Id holds considerable promise for refining language models. Moreover, its integration with larger pre-existing corpora can enhance the training of more extensive and versatile mixed Indonesian models tailored for diverse NLP tasks. The dataset categorizes named entities into three distinct classes: hard skill, soft skill, and technology. It consists of 418.868 tokens. Subsequently, these tokens are marked using the BIO format. The annotation table is presented in ConLL2003 format, consisting of three columns: Sentence#, word, and tag columns. *We already have paper at https://www.sciencedirect.com/science/article/pii/S235234092400163X (please cite)

Files

Steps to reproduce

The data used to create the NERSkill.Id were scraped from the Indeed , Jobstreet , loker.id and Job.Id. We attach the sourch code which we use to scraped the data. NERSkill.Id was annotated manually by eight anatators. We used label B-HSkill; I-HSkill; B-SSkill; I-SSkill; B-Tech; I-Tech; and O for annotation.

Institutions

Universitas PGRI Yogyakarta
Universitas Gadjah Mada

NERSkill.Id

Description

Files

Steps to reproduce

Institutions

Categories

Related Links

Licence