Coded English-French-Japanese Health Translation Dataset and Reliability Data

Published: 29 September 2026| Version 1 | DOI: 10.17632/2bz2vdtcbs.1
Contributor:
Hirsh Garhwal

Description

This dataset contains the materials used in the study “Semantic Drift, Metaphor Vulnerability, and Emotional Meaning in Health Translation: A Corpus-Based Comparison of English-French and English-Japanese Subtitles.” Supplementary Data 1 contains the coded results for all 217 health-related English-French-Japanese segments analyzed in the study. Supplementary Data 2 contains the 25% validation subset used for independent reliability assessment. Supplementary Code 1 contains the Python code used to identify and extract the 217 study segments from the TED2020 parallel corpus.

Files

Steps to reproduce

Health-related English-French-Japanese segments were identified from the TED2020 parallel corpus using the provided Python script. The resulting 217 aligned segments were manually coded for semantic shift, emotional intensity, metaphor handling, and register. A 25% subset was independently re-coded for reliability assessment. The uploaded files contain the full coded dataset, the validation subset, and the extraction script.

Institutions

Categories

Computational Linguistics, Data Analysis, Corpus-Based Translation Studies

Licence