Dataset Student and Course Performance for Heterogeneous Institutional Monitoring

Published: 3 September 2025| Version 1 | DOI: 10.17632/ct3ps2pjhz.1
Contributors:
,
,
,
,
,
,

Description

This dataset supports the study “Adaptive Ontology-Enabled Data Retrieval Model for Learning Analytics Integration Across Heterogeneous Educational Platforms.” It addresses the challenge of integrating heterogeneous educational data sources—including Learning Management Systems (LMS), Student Information Systems (SIS), and MOOC platforms—into a unified repository for learning analytics. The research hypothesizes that an ontology-driven, graph-based model can more effectively integrate diverse data and enable accurate, flexible retrieval compared to traditional relational database approaches. The dataset is organized into eight sheets: Coursegroup – enrolment records linking students, lecturers, faculties, and programmes. Prerequisite – course metadata, credit hours, and prerequisite structures. Courseperformance – grade distributions across cohorts, lecturers, and groups. Studentperformance – detailed individual assessment scores, totals, and grades. Attendance – monthly student attendance by course and semester. Classactivity – engagement indicators measured through discussion posts. MOOCstudentprofile – learner enrolment, completion dates, time spent, progress, and certification. MOOCstudentperformance – quiz, test, and exercise results across MOOC modules. These sheets capture performance and engagement across conventional courses, classroom activities, and MOOC environments. Notable Findings - The ontology-based retrieval model successfully harmonized heterogeneous data without schema conflicts. - Confusion matrix evaluation achieved >99% accuracy, >98% precision, 100% recall, and >99% F1-scores. - No false negatives and only a small number of false positives were recorded. The dataset supports diverse analytics tasks such as tracking outcomes, correlating attendance with performance, identifying at-risk students, and comparing MOOC and traditional learning. How the Data Can Be Interpreted Records represent validated student/course performance data, interpretable at three levels: - Course: outcomes, grade distributions, prerequisites. - Student: assessments, attendance, engagement. - MOOC: persistence, time investment, quiz/test outcomes. Ontology cross-linking enables flexible queries, e.g., “Which students with poor attendance underperformed?” or “How does MOOC completion relate to course grades?” This dataset highlights the value of ontology-based integration for scalable, accurate, and flexible learning analytics across heterogeneous educational platforms.

Files

Steps to reproduce

Methods and Protocols The following is a structured workflow on how the data was gathered and how to reproduce our research: 1. Requirement Analysis – Conducted qualitative interviews with lecturers and administrators to identify the key data needed for learning analytics. 2. Data Extraction – Exported datasets from LMS, SIS, and MOOC platforms into structured Excel sheets (8 categories covering enrolment, prerequisites, course outcomes, student performance, attendance, engagement, and MOOC activity). 3. Preprocessing – Cleaned and anonymized records, standardized grading codes and course identifiers, and normalized missing or inconsistent values. 4. Ontology Mapping – Developed the Student Performance and Course Performance (SPC) ontology to semantically align data across institutions. To prepare the data, users must align the provided SPC ontology with the XLSX dataset using a tool named OpenRefine. 5. Validation – Used TDDonto2 in Protégé to validate ontology structure and consistency with Description Logics (DL) axioms; domain experts cross-checked mappings. 6. Integration – Loaded datasets into a NoSQL graph-based repository (e.g., Neo4j/GraphDB) to support semantic querying. 7. Querying – Designed SPARQL queries reflecting institutional learning analytics needs (e.g., ongoing assessment tracking, attendance-performance correlation). 8. Evaluation – Applied confusion matrix metrics (TP, FP, FN, TN) to measure retrieval accuracy, achieving >99% accuracy, >98% precision, 100% recall, and >99% F1-score. In summary, other researchers can reproduce this process by: 1. Conducting requirement analysis through interviews to identify relevant learning analytics questions. 2. Exporting LMS, SIS, and MOOC data into structured tabular format. 3. Mapping datasets to the SPC ontology (which is openly available). 4. Using ontology validation tools (e.g., TDDonto2 in Protégé) to check consistency. 5. Loading the integrated data into a graph database (e.g., Neo4j, GraphDB) and executing SPARQL queries to replicate retrieval tests.

Institutions

Categories

Ontology, Big Data Analytics, Heterogeneous Database, Data Analytics, Knowledge Graph

Funders

Licence