Dataset: Machine Learning applied to Neuroimaging
Description
Research Context Machine learning and deep learning techniques have become increasingly important for analyzing complex neuroimaging and electrophysiological data in the study of Alzheimer’s disease and mild cognitive impairment. These approaches enable the identification of subtle structural, functional, and signal-based patterns that are difficult to detect using traditional analytical methods, supporting early diagnosis and biomarker discovery. What the Dataset Contains This dataset compiles structured information extracted from peer-reviewed studies that apply machine learning and deep learning methods to neuroimaging and electroencephalography data for Alzheimer’s disease, mild cognitive impairment, and cognitively normal populations. The dataset includes curated metadata describing imaging modalities, machine learning algorithms, sample characteristics, validation strategies, main findings, and reported biomarkers. The included studies span multiple data modalities, including structural magnetic resonance imaging, resting-state functional magnetic resonance imaging, positron emission tomography, and electroencephalography, enabling cross-modal comparison of machine learning performance and biomarker relevance. Notable Insights Machine learning models applied to neuroimaging data can reliably differentiate Alzheimer’s disease from mild cognitive impairment and cognitively normal controls. Functional connectivity alterations, particularly within default mode and frontoparietal networks, are recurrent biomarkers in functional magnetic resonance imaging–based studies. Structural magnetic resonance imaging studies consistently identify hippocampal and temporal lobe atrophy as key discriminative features. Electroencephalography-based deep learning approaches demonstrate promising performance for non-invasive classification across cognitive stages. Model performance varies substantially depending on feature extraction methods and validation strategies, highlighting the importance of methodological transparency. Data Interpretation and Intended Use The dataset is designed for meta-analytical research, methodological comparison of machine learning approaches, and educational purposes. Researchers can use this dataset to: Conduct systematic reviews and meta-analyses of machine learning performance in neuroimaging. Compare biomarkers across imaging modalities. Evaluate the impact of validation strategies on reported classification accuracy. Support reproducibility and secondary analyses in computational neuroscience research. Data Acquisition Summary All data were manually extracted from peer-reviewed publications retrieved from major scientific databases and curated into a standardized structure compatible with Mendeley Data. Each study is documented through structured summary files and a consolidated dataset table, ensuring traceability, consistency, and reproducibility.
Files
Steps to reproduce
Data Acquisition and Reproducibility Statment The dataset was constructed through a systematic extraction of structured information from peer-reviewed scientific publications focused on machine learning and deep learning applications in neuroimaging and electrophysiological data for Alzheimer’s disease and mild cognitive impairment. The included studies were identified through targeted literature searches in major scientific databases, primarily PubMed and ScienceDirect. Publications were selected based on predefined inclusion and exclusion criteria emphasizing the use of machine learning techniques applied to magnetic resonance imaging, functional magnetic resonance imaging, positron emission tomography, or electroencephalography data. Literature Identification and Screening A comprehensive set of articles was retrieved using keyword-based search strategies combining terms related to machine learning, neuroimaging modalities, Alzheimer’s disease, and mild cognitive impairment. Titles and abstracts were screened to identify relevant studies, followed by full-text assessment to confirm eligibility. Data Extraction Protocol For each eligible study, structured information was manually extracted, including study objectives, disease category, neuroimaging or electrophysiological modality, machine learning algorithms, sample size and characteristics, validation strategies, main findings, and reported biomarkers. Extraction followed a standardized template to ensure consistency across studies. Curation and Structuring All extracted data were curated manually to ensure terminological consistency and methodological clarity. Variables were standardized and harmonized across studies to facilitate cross-study comparison and meta-analytical use. Each study was assigned a unique identifier and documented through both structured summary documents and a consolidated dataset table. Tools and Software Used PubMed and ScienceDirect for literature retrieval Mendeley Reference Manager for citation management Microsoft Word for structured study summaries Microsoft Excel for dataset consolidation and variable standardization Mendeley Data for public dataset deposition The methodology described above is fully reproducible by any research team with access to the same scientific databases and standard data extraction tools. By applying the same search strategy, inclusion criteria, and extraction template, the dataset can be independently reconstructed or expanded with additional studies.
Institutions
- Universidad de Guadalajara Centro Universitario de Ciencias de la SaludJalisco, Guadalajara