Relabeled Second HAREM's Golden Collection
Description
This dataset is a relabeling of the Second HAREM's Golden Collection, created as a benchmark for a Named Entity Recognition and Relation Extraction contest by the Linguateca organization around 2008. The original dataset consists of 129 documents, comprising thousands of mentions and the annotated relations between them. In particular, the "identity" (or "ident") relation encodes the information that two mentions refer to the same real world entity. However, since this relations were only annotated within each document, inter-document entity resolution was infeasible. For this reason, we inspected the annotated clusters of mentions and merged together those left unmerged, originating a new dataset, suitable for inter-document entity resolution. We restricted ourselves to only 3 out of the 10 original categories of mentions from the original dataset, namely: Organization, Location, and Person.
Files
Steps to reproduce
Further details can be found on the paper Named entity resolution: benchmarking and entity matching strategies (2026), by Rafael Omiya and Rubén Interian.
Institutions
- Universidade Estadual de Campinas (UNICAMP)São Paulo, Campinas