Planetary Health: A Comprehensive View of Food in Brazil - PHFood Brazil

Published: 11 May 2026| Version 12 | DOI: 10.17632/mt4mj23j73.12
Contributors:
,

Description

This dataset is part of the study “Integrating national open databases for a comprehensive view on food, environment and health in Brazil”, an initiative that integrates national open databases to provide harmonized information on food production, consumption, environmental sustainability, and public health in Brazil from 1974 to 2024. The dataset was developed to investigate relationships between food systems, nutrient availability, agricultural production, environmental pressures, and dietary patterns over time. By aggregating and harmonizing data from multiple national sources, the platform enables the analysis of longitudinal trends associated with food production, nutrition, sustainability, and health. The database includes information on harvested areas, food production (tons), nutrient availability (e.g., energy, protein, fiber, vitamins, and minerals), food acquisition and consumption, water use and deficit, pesticide use, maximum residue levels (MRL), residue monitoring, toxicity classifications, and environmental indicators across Brazilian regions and states. Data were obtained from official Brazilian governmental platforms, including IBGE, ANVISA, FAOSTAT, and the Ministry of Agriculture. The datasets were integrated using Extract, Transform, Load (ETL) procedures to ensure compatibility, standardization, and interoperability across sources, years, and regions. Food items were harmonized and classified into standardized food groups, enabling integrated analyses of food systems and environmental sustainability. The database was designed according to FAIR principles (Findable, Accessible, Interoperable, and Reusable) to support transparency and reproducibility. This dataset can support studies in agrifood systems, nutrition, sustainability, public health, environmental exposure, and climate change. It may also contribute to future modeling approaches involving dietary scenarios, environmental footprints, and food system sustainability. This work was initially supported by the CAPES Thesis Award (2023) and is currently being expanded through postdoctoral research funded by FAPESP (2025).

Files

Steps to reproduce

Data Collection: Aggregate data from official Brazilian national databases, including agricultural statistics, environmental data (water usage and deficits), food consumption patterns, and pesticide residue monitoring. Sources include government platforms such as IBGE (Brazilian Institute of Geography and Statistics), ANVISA (Brazilian Health Regulatory Agency), and the Ministry of Agriculture. Data Preprocessing: Clean the raw data by standardizing formats (e.g., units, dates) and handling missing or incomplete data. Use ETL (Extract, Transform, Load) processes to harmonize data from various sources into a consistent format. Data Integration: Merge datasets by common keys such as food types, regions, and years. Ensure compatibility of columns across datasets. Link food production data with water usage, pesticide residue, and nutrient composition data. Data Transformation: Calculate aggregate values for nutrients (e.g., protein, fiber, vitamins) based on food production and consumption. Classify food items according to food groups, regions, and other relevant categories. Machine Learning Application: Apply state-of-the-art machine learning techniques to identify trends and correlations between food production, environmental sustainability, and health outcomes. These machine learning models can be adapted to suit specific research queries based on the dataset. FAIR Principles Adherence: Ensure that the dataset is structured in a way that adheres to FAIR (Findable, Accessible, Interoperable, Reusable) principles to facilitate reproducibility. Versioning and Update Plan: This database is designed as an extensible platform. Future updates may incorporate additional sources to further enhance its environmental and nutritional dimensions. Updates will follow a structured versioning process, aligned with the release schedules of the original data providers. Changes between versions will be tracked and summarized to support transparency, traceability, and reproducibility. Guidelines for Responsible Use: Users are advised to interpret the PHFood Brazil database with care, considering the structure and limitations of the integrated sources. The dataset is organized in a wide format, where each row corresponds to a specific combination of variables (e.g., food item, year, region). Only values aligned in the same row should be interpreted as part of the same data point. Blank cells indicate unavailable or inapplicable data and should not be treated as zeros. Users should also consider that the database integrates historical data from various sources, which may differ in scope, methodology, or periodicity. These characteristics are detailed in the accompanying codebook and documentation.

Institutions

Categories

Nutrition, Data Science, Pesticide, Environment and Health, Agricultural System, Climate Change

Licence