ViFoodLabel: A Vietnamese Food Product Label Dataset for Key Information Extraction
Description
ViFoodLabel is a dataset of Vietnamese food product label images annotated for structured key information extraction. It comprises 600 images self-photographed with a smartphone camera (iPhone 13) from physical packaged-food products in Ho Chi Minh City, with no stock, e-commerce, or third-party packaging images included. Each image is paired with a hand-annotated ground-truth record covering nine fields: product name, ingredients, additives, warnings, a name/value nutrition table, origin, net weight, manufacturing date, and expiry date. A 200-image double-annotation pass, scored with the same matching procedure used for model evaluation, gives a mean macro field F1 of 0.873 (lenient) and a nutrition pairing accuracy of 0.999. Images were screened for readability and target content, then annotated by hand following a written guideline covering ingredient-versus-additive classification, warning identification, and language preference on multilingual labels. Annotations are released as one JSON file per image, each field a single string, a list of strings, or, for nutrition, a list of name/value pairs; empty fields reflect content absent from the photographed frame rather than annotation gaps. Images are released downscaled from capture resolution to keep the archive a manageable size while preserving label legibility. The deposit is organized as follows: images/ # the 600 label photographs (JPEG), each with a zero-padded # four-digit identifier from 0001 to 0600. labels/ # the 600 ground-truth records (JSON), one per image, sharing # the image identifier.
Files
Institutions
- HUTECH UniversityHo Chi Minh, Ho Chi Minh City
- Ho Chi Minh City University of Foreign Languages and Information TechnologyHo Chi Minh, Ho Chi Minh City
- Ho Chi Minh City University of Industry and TradeHo Chi Minh, Ho Chi Minh City