KU-BD-Fish: High-Resolution White-Background and Segmented Images for Species Classification
Description
This dataset comprises 940 high‑resolution images of 26 freshwater fish species native to Bangladesh. In order to ensure all procedure are performed on a uniform standard and scientific basis, each specimen was photographed individually against white background with controlled lighting using an iPhone 14 Plus and Samsung S23 Plus. The dataset contains two versions of every image: the original unprocessed photograph with the white background and a segmented version where the background has been removed to isolate the fish. Acquisition equipment: iPhone 14 Plus and Samsung S23 Plus smartphones. Image formats: 1.Raw images: Original photographs with a consistent white background. 2.Segmented images: Background removed to highlight each fish specimen. Representation of species: 26 different species of freshwater fish from Bangladesh, each with several photos. File organisation: The images are organised into folders that are specific to each species, and a metadata file that goes with them provides a list of the species names and other pertinent information.For example the folder named Rui contains all the photos of Rui fish in both Raw_images and Segmented_Images folder. The following are some possible uses for these datasets: 1. Morphological analysis: It is useful for accurately measuring and comparing the body structures of fish. 2. Machine-learning research: It can be used to train and assess models for segmentation, object detection, and species classification. 3. Educational use: It offers excellent specimen photos for conservation education, aquaria, and ichthyology classes. 4. Taxonomic and biodiversity research: Its standardized, high-resolution specimen images can facilitate accurate species identification and reliable assessments of biodiversity, supporting ecological and conservation research.
Files
Steps to reproduce
We discovered that Bangladeshi fish species datasets were sparse and unorganised, making them unsuitable for machine learning classification. This was evident in the country's fish business. Therefore, the objective was to save a clear, well-structured dataset that could be used for machine learning, taxonomic analysis, and species classification. The dataset includes 26 freshwater fish species, which were selected due to their ecological and economical importance in Bangladesh. These fish were shot in high resolution on iPhone 14 Plus and Samsung S23 Plus. For consistency, all fish were photographed in a standardized procedure with white background to facilitate background removal and isolation of the fish. The background was removed after taking the raw pictures. The first procedure was carried out with the MacBook’s background remover tool, that served well for most images. But for finer granularity and accuracy, we used Python-based tools. For the more advanced processing such as contour detection, edge detection, and masking we used OpenCV and Pillow (PIL), which ensured that the segmentation was clean as there is no overlap between the fish with its background. Metadata and Model When all images were obtained, they were classified based on species and stored in separate folders together with metadata information that accompanies the picture for its use in successive work. We automated the complete process and documented in Python, Jupyter Notebooks, which allowed us to obtain consistent results across all images. The resulting dataset, consisting of raw and segmented images, was made available under an open license to promote collaboration and re-use. Such a structured and reproducible effort can be used as a building-block of future exploratory, classification and machine learning efforts related to Bangladesh fresh water fish species and can be easily reproduced by others thus allowing further work.
Institutions
- Khulna University