Machine-Translated Texts with Human Annotations and Automatic Metric Evaluations

Published: 11 July 2025| Version 1 | DOI: 10.17632/zx3vt8r26k.1
Contributors:
,
,
, Lubomir Benko,

Description

This dataset consists of English journalistic texts translated into Slovak using both statistical and neural machine translation systems. Each translated segment was evaluated by human annotators, and errors were recorded in binary format across five error categories. Additionally, the dataset includes the scores of 68 different automatic evaluation metrics, commonly used to assess machine translation quality. The data is divided into training and testing subsets, allowing for the development of models to predict error categories based on the automatic metric scores.

Files

Institutions

  • Univerzita Konstantina Filozofa v Nitre

Categories

Natural Language Processing, Machine Translation

Licence