Relabeled Second HAREM's Golden Collection

Published: 30 June 2026| Version 1 | DOI: 10.17632/wh5sz652jk.1
Contributors:
Rafael Omiya,

Description

This dataset is a relabeling of the Second HAREM's Golden Collection, created as a benchmark for a Named Entity Recognition and Relation Extraction contest by the Linguateca organization around 2008. The original dataset consists of 129 documents, comprising thousands of mentions and the annotated relations between them. In particular, the "identity" (or "ident") relation encodes the information that two mentions refer to the same real world entity. However, since this relations were only annotated within each document, inter-document entity resolution was infeasible. For this reason, we inspected the annotated clusters of mentions and merged together those left unmerged, originating a new dataset, suitable for inter-document entity resolution. We restricted ourselves to only 3 out of the 10 original categories of mentions from the original dataset, namely: Organization, Location, and Person.

Files

Steps to reproduce

Further details can be found on the paper Named entity resolution: benchmarking and entity matching strategies (2026), by Rafael Omiya and Rubén Interian.

Institutions

Categories

Natural Language Processing, Benchmarking

Licence