The cross-lingual transfer evaluation of sentiment score

Published: 2 August 2022| Version 1 | DOI: 10.17632/mgs3xjtfgv.1
Contributors:
, Ľubomír Benko

Description

The dataset represents the processed movie subtitle data adjusted for sentiment analysis, which was implemented using IBM Watson Natural Language Understanding (IBM NLU). The source data contains Slovak and English subtitles from 10 movies, which are matched into pairs. Each of the subtitles is matched with a machine translation generated using Google Translate and DeepL. In the next matrix, the results of the sentiment analysis from IBM NLU service for each segment are processed. The third file contains the results of validating the accuracy and error rates of the machine translations from the BLEU and TER metrics.

Files

Institutions

Univerzita Konstantina Filozofa v Nitre

Categories

Natural Language Processing, Machine Translation, Sentiment Analysis

Licence