An EEG Dataset for Five-Class Imagined Alphabet Decoding Using OpenBCI Recordings
Description
This dataset contains electroencephalography (EEG) recordings collected for five-class imagined alphabet decoding using an OpenBCI system. EEG data were recorded from 20 healthy adult participants during an eyes-closed alphabet-imagery task. The five target alphabet classes are A, K, N, P and Z. Each participant completed one recording session for each alphabet class and each session contained 25 repeated trials. The dataset includes raw OpenBCI recordings, externally generated event-log files, labelled training-segment files, metadata tables, preprocessed MATLAB epoch files, channel-quality audit files, generated figures and baseline deep-learning result files. The raw recordings were acquired with a 16-channel OpenBCI Cyton Serial Daisy configuration at 125 Hz. The raw channel montage includes Fp1, Fp2, C3, C4, P7, P8, O1, O2, F7, F8, F3, F4, T7, T8, P3 and P4. During preprocessing, channel-quality assessment identified T8 as problematic; therefore, the final exported preprocessed epoch set contains 15 retained EEG channels. The preprocessed release contains 100 alphabet sessions, 2,500 task trials, and 390,000 exported EEG windows. Each 10-second task segment was divided into 156 non-overlapping windows of 8 samples. At the 125 Hz sampling rate, each 8-sample window corresponds to 64 ms. The exported EEG tensor is stored as windows × time points × channels with a final shape of 390,000 × 8 × 15. The dataset is accompanied by metadata files describing participant-level anonymized information, channel configuration and task timing. It also includes code and outputs for baseline deep-learning experiments using leave-one-subject-out and calibration-based evaluation settings. The baseline workflow evaluates CNN-LSTM and LSTM-based architectures for imagined alphabet classification. Ethical clearance for the data collection reported in this dataset was obtained from the Committee for Ethical Clearance of Research Proposal, Faculty of Biological Sciences, University of Dhaka, Bangladesh, under the project titled "Imagined Alphabet Recognition from EEG Signals Using Neuro-Symbolic AI" (Approval Ref: 394/Biol.Scs./2026-2027 , dated 27 August,2026). This Mendeley Data release corresponds to the EEG acquisition, preprocessing and baseline deep-learning benchmarking component of that approved project; the participants, recording protocol and consent procedures are identical to those covered by the approval. This dataset is intended for research on EEG-based brain-computer interfaces, imagined speech or imagined character decoding, subject-independent EEG classification, calibration-based learning, low-cost OpenBCI acquisition, EEG preprocessing reproducibility and benchmark development for deep-learning methods.
Files
Steps to reproduce
1. Download the dataset repository, including the raw participant ZIP files, metadata files, preprocessing outputs, figures, and code. 2. For each participant, open the corresponding raw data archive. Each participant contains five alphabet-imagery sessions for A, K, N, P, and Z. Each session includes a raw OpenBCI text file, an event-log CSV file, and a training-segment CSV file. 3. Inspect the raw OpenBCI files. EEG was recorded using an OpenBCI Cyton Serial Daisy configuration with 16 EXG channels at 125 Hz. The raw recordings contain the sample index, 16 EXG channels, accelerometer, digital, analog, timestamp, marker, and formatted timestamp columns. 4. Use the event-log and training-segment CSV files to identify the timing of each trial. At the beginning of each session, an intentional blink was used as a synchronization marker. Task timing was taken from the external CSV logs generated by the Python auditory cueing script, not from the OpenBCI marker channel. 5. Extract the task intervals labelled for the target alphabet classes A, K, N, P, and Z. Each alphabet session contains 25 trials. Each trial contains a preparation period, a 10-second alphabet-imagery task period, and a rest period. 6. Preprocess the EEG data using the supplied MATLAB preprocessing scripts. The preprocessing workflow loads the raw OpenBCI files, extracts the 16 EXG signal channels, matches each recording with its event-log and training-segment files, performs channel-quality assessment, applies median centering and 50 Hz notch filtering, extracts labelled task segments, and exports the epoch data. 7. Exclude T8 from the final preprocessed epoch export because it was identified as a problematic channel during channel-quality assessment. The final preprocessed epoch set contains 15 retained EEG channels: Fp1, Fp2, C3, C4, P7, P8, O1, O2, F7, F8, F3, F4, T7, P3, and P4. 8. Divide each 10-second task segment into 156 non-overlapping windows of 8 samples. At 125 Hz, each 8-sample window corresponds to 64 ms, and 156 windows correspond to 9.984 seconds of task data per trial. 9. Use the exported preprocessed files, including epochs.mat, metadata.csv, channel_names.csv, bad_channel_audit.csv, ica_log.csv, fdf_session_index.csv, and fdf_epoch_manifest.csv, to reproduce the processed dataset. The exported EEG tensor is stored as windows × time points × channels, with a final shape of 390,000 × 8 × 15. 10. To reproduce the baseline deep-learning results, run the supplied Jupyter Notebook: C3_Picard_50HzNotch_FDF_LOSO.ipynb. The notebook loads the preprocessed epochs and metadata, prepares labels and participant-wise folds, performs split-leakage checks, and evaluates CNN-LSTM and LSTM-based architectures using pure LOSO, 1-trial-per-class calibration, and 5-trials-per-class calibration settings.
Institutions
- University of Frontier Technology, BangladeshDhaka Division, Gazipur
- University of DhakaDhaka Division, Dhaka