Annotated Facial Image Dataset for Nigerian Ethnicity Recognition with Baseline Evaluation

Published: 18 August 2026| Version 1 | DOI: 10.17632/xh9sy3v5m6.1
Contributors:
Mohammed Ali Bizi, Mohd Shahrizal Sunar, Farhan Mohamed, Shamsul Mohamad

Description

This dataset represents 3,000 annotated facial image samples developed to support fine-grained ethnicity recognition research using deep learning techniques. The dataset supports investigations into facial feature representation learning using CNN-based architectures, attention mechanisms, and metric learning approaches. The underlying dataset represents three major Nigerian ethnic groups: Hausa, Igbo, and Yoruba. The public release provides anonymized image identifiers and corresponding ethnicity annotations organized into training, validation, and test splits. The dataset is designed to facilitate research on discriminative facial representation learning, benchmarking, demographic analysis, and fairness-aware computer vision. Each facial sample is associated with an annotated ethnicity class label for supervised machine learning experiments. Metadata and annotation files are provided to support reproducible research and transparent evaluation. Due to the sensitive biometric nature of facial image data, the original facial images are not released under unrestricted public access. Qualified researchers may request controlled access to the facial images for legitimate academic research purposes. Access requests will be evaluated by the corresponding author and may require completion of a Data Use Agreement. Users must comply with ethical research practices and must not attempt to re-identify individuals or use the dataset for surveillance, discrimination, or other harmful applications.

Files

Steps to reproduce

1. The dataset was developed from facial image samples obtained under approved ethical procedures and organized into three annotated ethnicity categories: Hausa, Igbo, and Yoruba. 2. The collected samples were reviewed to ensure annotation consistency and removal of invalid or unsuitable samples before dataset preparation. 3. A stratified random splitting procedure was performed using a fixed random seed (42) to ensure reproducible dataset partitioning. 4. The dataset was divided into training (70%), validation (15%), and testing (15%) subsets while maintaining balanced representation across the three ethnicity categories. 5. During dataset preparation, samples were organized into structured training, validation, and testing partitions with class-specific organization. 6. Corresponding anonymized annotation files (train_labels_clean_final.csv, val_labels_clean_final.csv, test_labels_clean_final.csv, and labels_clean_final.csv) were generated containing anonymous image identifiers and ethnicity class labels. 7. Dataset organization and metadata generation procedures were implemented using Python libraries, including os, shutil, random, and csv. 8. The publicly released dataset package contains documentation files, anonymized metadata (dataset_metadata.json), annotation files, and ethical documentation to support reproducible and responsible academic research.

Categories

Artificial Intelligence, Computer Vision, Machine Learning

Licence