Dataset of Packaging Defects and Associated Cost of Poor Quality in a Wine Bottling Process
Description
This dataset contains 9,651 images acquired from a real wine bottling packaging process. The dataset is intended for applications in quality control, computer vision, and process improvement in industrial environments. The images are organized into three defect classes: missing cap, missing bottle, and wrong bottle. Each image reflects actual production conditions, where bottles are arranged in boxes with a capacity of 12 units, consistent with the real packaging configuration. To ensure variability and robustness in visual features, multiple cap colors are included, namely black, red, orange, blue, purple, green, and silver. This diversity supports the development of more generalizable defect detection and classification models. The dataset was divided into training (70%), validation (20%), and testing (10%) subsets, facilitating its direct use in supervised learning tasks. Additionally, the dataset includes a configuration file (data.yaml) that specifies the directory paths for each subset (training, validation, and testing) as well as the corresponding class labels. This file enables straightforward integration with deep learning frameworks, particularly those based on YOLO architectures. This dataset can support research in defect detection, automated inspection systems, and quality improvement methodologies such as Six Sigma and Industry 4.0 applications.
Files
Steps to reproduce
Data acquisition was conducted in wine-producing facilities during the packaging stage of the bottling process. Images were captured under real operating conditions, focusing on packaging lines where bottles are arranged into boxes. During the data collection phase, multiple production batches were monitored to ensure variability in operating conditions. The acquisition process specifically targeted the occurrence of packaging anomalies, including missing caps, missing bottles, and incorrect bottle types. Images were obtained directly from the production environment without controlled laboratory conditions, ensuring that factors such as lighting variability, bottle arrangement, and cap color differences were naturally represented. This approach enables the dataset to reflect real-world industrial scenarios and supports the development of robust defect detection models. The collected images were subsequently labeled according to the identified defect classes and organized into training, validation, and testing subsets.
Institutions
- Universidad Autónoma de Baja CaliforniaBaja California, Mexicali