Coded English-French-Japanese Health Translation Dataset and Reliability Data
Description
This dataset contains the materials used in the study “Semantic Drift, Metaphor Vulnerability, and Emotional Meaning in Health Translation: A Corpus-Based Comparison of English-French and English-Japanese Subtitles.” Supplementary Data 1 contains the coded results for all 217 health-related English-French-Japanese segments analyzed in the study. Supplementary Data 2 contains the 25% validation subset used for independent reliability assessment. Supplementary Code 1 contains the Python code used to identify and extract the 217 study segments from the TED2020 parallel corpus.
Files
Steps to reproduce
Health-related English-French-Japanese segments were identified from the TED2020 parallel corpus using the provided Python script. The resulting 217 aligned segments were manually coded for semantic shift, emotional intensity, metaphor handling, and register. A 25% subset was independently re-coded for reliability assessment. The uploaded files contain the full coded dataset, the validation subset, and the extraction script.
Institutions
- University of WashingtonWashington, Seattle