YouTube-Based Nahwu and Shorof Learning Transcription Dataset
Description
The YouTube-Based Nahwu and Shorof Learning Transcription Dataset is a collection of textual transcriptions obtained from educational YouTube videos that focus on teaching Nahwu and Shorof, which are fundamental components of Arabic grammar. Nahwu generally deals with Arabic sentence structure and grammatical rules, while Shorof focuses on word formation and morphological patterns. This dataset was compiled by collecting relevant instructional videos from YouTube and converting their spoken explanations into structured text through transcription. The purpose of this dataset is to support academic research, language learning analysis, and natural language processing (NLP) studies related to Arabic language education. The dataset may include information such as video titles, source links, and the corresponding transcription text. It can be used by researchers, educators, and students to analyze teaching methods, linguistic patterns, vocabulary usage, and instructional approaches used in online Nahwu and Shorof learning materials. Overall, this dataset provides a useful resource for studies in Arabic linguistics, language education, and computational text analysis based on online educational content.
Files
Steps to reproduce
Search for Nahwu and Shorof learning videos on YouTube using keywords such as “Nahwu lesson”, “Shorof lesson”, or “Arabic grammar tutorial”. Select relevant educational videos that explain Nahwu and Shorof materials. Collect the video information and access the spoken content from the selected videos. Convert the spoken explanations into written text through transcription. Organize the transcription together with the video title and source link into a structured dataset file. Review and clean the transcription text to ensure clarity and accuracy. Save the dataset in a structured format such as TXT, CSV, or Excel for further analysis or research.
Institutions
- Universitas Islam Negeri Maulana Malik IbrahimEast Java, Malang