PUST Cafeteria Food Image Dataset: Real-World Bangladeshi Meal-Platter Images with Bounding-Box and Polygon Annotations

Published: 28 August 2026| Version 2 | DOI: 10.17632/fn6yhzjz83.2
Contributors:
,
,
,
,

Description

The PUST Cafeteria Food Image Dataset is a real-world image resource developed for food recognition, object detection, and instance segmentation of Bangladeshi cafeteria meals. The dataset contains 720 original meal-platter photographs collected between 1 August and 5 September 2025 at the Central Cafeteria of Pabna University of Science and Technology, Bangladesh. Images were captured using the rear cameras of Realme GT Master Edition and OPPO A92 smartphones under realistic cafeteria conditions with variations in illumination, viewpoint, background, food arrangement, overlapping objects, and partial occlusions. Each visually distinguishable food instance was manually annotated using the Roboflow platform. The dataset provides bounding-box annotations for object detection and polygon-based annotations for instance segmentation. Annotation preparation involved two annotators with cross-checking for class assignment, bounding-box placement, polygon delineation, and image–annotation consistency. The dataset contains 16 food classes. Images were auto-oriented using available metadata and resized to 640 × 640 pixels. The original images were divided into training, validation, and test subsets using a 70:15:15 split ratio (504, 108, and 108 images). Augmentation and light class balancing were applied only to the training subset, including selective augmentation of five underrepresented classes: chaa, matha, ruti, singara, and vegetable_roll. The final processed dataset contains 1,728 images with 6,297 annotated food instances, including 1,512 training images (5,493 instances), 108 validation images (395 instances), and 108 test images (409 instances). The dataset is provided in three annotation formats: YOLO_detect (bounding-box TXT), YOLO_seg (polygon-based TXT), and COCO (JSON annotation files). The archive includes raw images, processed datasets, metadata files (class_names.csv, dataset_summary.csv, image_source_mapping_final.csv, instances_before_augmentation.csv, instances_after_augmentation.csv, and device_information.txt), together with README.md and LICENSE.txt files. This dataset supports research in food recognition, object detection, instance segmentation, food-instance counting, transfer learning, domain adaptation, benchmarking, and intelligent food-service systems. Applications involving billing, portion estimation, calorie estimation, or dietary assessment require additional validated external information. The dataset was collected from a single university cafeteria using two smartphone models. Class distributions remain non-uniform despite selective augmentation and light class balancing. Augmented images improve variability but cannot replace independently collected meal scenes. Polygon annotations represent only visible food regions and do not reconstruct hidden areas caused by occlusion.

Files

Steps to reproduce

Step 1: Original image collection and annotation The workflow starts with the collection of 720 original meal-platter images stored in the 01_raw_original_images directory. Each visually distinguishable food instance is manually annotated using the Roboflow platform according to the 16-class taxonomy provided in class_names.csv. Bounding-box annotations are prepared for object detection, while polygon annotations are generated for instance segmentation to represent the visible regions of food items. Annotation consistency is verified through class-label checking, bounding-box placement, polygon delineation, and image–annotation correspondence. Step 2: Image preprocessing and dataset partitioning The collected images are auto-oriented using available metadata and resized to a standardized resolution of 640 × 640 pixels. The original image collection is divided at the image level into training, validation, and test subsets using a 70:15:15 ratio, resulting in 504, 108, and 108 images, respectively. The validation and test subsets are preserved without modification throughout the workflow to avoid data leakage. Step 3: Training augmentation and class refinement Data augmentation is applied exclusively to the training subset using geometric and photometric transformations, including flipping, random zoom/cropping, rotation, brightness and saturation adjustment, Gaussian blur, and salt-and-pepper noise. Selective augmentation-based light class balancing is performed for underrepresented classes (chaa, matha, ruti, singara, and vegetable_roll). Generated samples are reviewed to maintain annotation accuracy and image quality. Step 4: Final dataset preparation and annotation conversion After augmentation, refinement, and quality verification, the final processed dataset is generated with 1,728 images containing 6,297 annotated food instances across 16 classes. The dataset is organized into training, validation, and test subsets containing 1,512, 108, and 108 images, respectively. Annotations are exported into three parallel formats: YOLO object detection format with normalized bounding-box coordinates, YOLO instance segmentation format with normalized polygon vertex coordinates, and COCO JSON format containing image, category, and instance-level annotation records. Step 5: Dataset organization and quality verification The final archive is structured into 01_raw_original_images, 02_processed_dataset/coco, 02_processed_dataset/yolo_detect, 02_processed_dataset/yolo_seg, and 03_metadata. Supporting metadata files include class_names.csv, dataset_summary.csv, image_source_mapping_final.csv, instances_before_augmentation.csv, instances_after_augmentation.csv, and device_information.txt, together with README.md and LICENSE.txt. Final quality control verifies annotation consistency, image–annotation matching, partition integrity, directory organization, and taxonomy consistency across all released formats.

Institutions

Categories

Food Science, Artificial Intelligence, Computer Vision, Image Processing, Data Science, Machine Learning, Deep Learning

Licence