Dataset of a Large-Scale English–Arabic Parallel Subtitle Corpus from English-Language Films (2000–2025)

Published: 1 July 2026| Version 2 | DOI: 10.17632/wn5kkc3ns4.2
Contributors:
, Ahmad S Haider

Description

This dataset is a large-scale English–Arabic parallel subtitle corpus compiled from 160 English-language films released between 2000 and 2025. The dataset comprises two complementary files. The first file contains approximately 296,190 aligned English–Arabic subtitle pairs (approximately 2.72 million words), where each record consists of a single English subtitle segment and its corresponding Arabic translation. The second file provides film-level metadata describing each source film used in the corpus. The corpus file includes the following fields: English subtitle segment Arabic subtitle translation Source film title The metadata file contains descriptive information for each film, including: Film title Release year Director(s) Genre British Board of Film Classification (BBFC) rating Runtime (hh:mm) IMDb rating Country of production USA box office revenue Cumulative worldwide box office revenue Number of awards and nominations Production house Source platform (e.g., iTunes, Netflix, Amazon Prime Video, DVD, OSN) The corpus was compiled by collecting publicly available subtitle files, extracting English and Arabic subtitle text, aligning corresponding subtitle segments, and organizing the data into a structured Excel spreadsheet. Basic preprocessing included removing subtitle numbering and timing information, normalizing formatting, verifying subtitle alignment, minimizing duplicate records, and preserving the original subtitle content where appropriate. Each subtitle pair is linked to its source film through the metadata file, enabling linguistic, translational, and film-specific analyses across different genres, production contexts, and release periods. This resource is valuable for: Researchers in audiovisual translation (AVT) and translation studies Corpus linguistics and bilingual corpus research Machine translation and natural language processing (NLP) Terminology extraction and bilingual sentence alignment Arabic–English language technologies Translation and interpreting education Evaluation and benchmarking of language models and subtitle translation systems The dataset supports both qualitative and quantitative investigations of subtitle translation, including translation strategies, lexical and phraseological variation, discourse and pragmatic analysis, cultural adaptation, subtitle segmentation, and corpus-based linguistic research. Furthermore, the accompanying film metadata facilitates analyses examining the influence of genre, production characteristics, release period, and other film-related variables on subtitle translation practices. Together, the two files provide a reusable benchmark resource for developing and evaluating computational and linguistic approaches to English–Arabic subtitle translation.

Files

Institutions

Categories

Natural Language Processing, Corpus Linguistics, Corpus-Based Translation Studies

Funders

  • The authors extend their appreciation to the Deanship of Research and Graduate Studies at King Khalid University for funding this work through Large Research Groups Program under grant number RGP2/433/47

Licence