ViFoodLabel: A Vietnamese Food Product Label Dataset for Key Information Extraction
Description
ViFoodLabel is a layout-aware dataset of Vietnamese food product label images annotated for key information extraction. It comprises 550 images self-photographed with smartphone cameras from physical packaged-food products in Vietnamese supermarkets and convenience stores, with no stock, e-commerce, or third-party packaging images included. Across the collection, the annotation contains 158,302 word tokens, 23,516 entity spans, and 4,014 nutrition name-to-value relations. Each image was screened for readability and target content, then annotated in Label Studio at the word-token level. Every token carries a tight bounding box, a verbatim transcription that preserves casing, diacritics, units, and on-label artifacts, and one BIO tag drawn from eleven entity types: product name, ingredient, additive, nutrition name, nutrition value, manufacturer, origin, net weight, manufacturing date, expiry date, and warning. A directed HAS_VALUE relation links each nutrition attribute to its corresponding value. Bounding-box coordinates are normalized to a 0 to 1000 range on both axes for compatibility with the LayoutLM family of models. From the word-level annotation, a field-grouped per-image record is assembled, grouping the labels into single-value fields (e.g. product name, origin, dates), list fields (ingredients, additives, warnings), and paired nutrition values. The deposit is organized as follows: docs/ #a data dictionary documenting every file and field, and the full annotation guideline in English and Vietnamese. label_studio/ # the anonymized Label Studio annotation export (one task per image). images/ # the 550 label images (JPEG), each with a zero-padded four-digit identifier. Identifiers are non-contiguous because images rejected during screening were removed, so the numbering does not run consecutively. processed/ #the derived per-image key-information records (one JSON per image), all per-image metadata in a single file, the frozen 80/10/10 train/development/test splits, and computed dataset statistics. The data support semantic entity recognition, relation extraction, and end-to-end information extraction from raw label images, and are reusable for optical character recognition, text normalization, and multilingual layout-aware document modeling. The data construction and processing scripts are available separately under an MIT license at https://github.com/hoadm-net/ViFoodLabel. This deposit is the companion dataset of the Data in Brief article of the same name.
Files
Institutions
- HUTECH UniversityHo Chi Minh, Ho Chi Minh City
- Ho Chi Minh City University of Foreign Languages and Information TechnologyHo Chi Minh, Ho Chi Minh City
- Ho Chi Minh City University of Industry and TradeHo Chi Minh, Ho Chi Minh City