Dataset of source verification, thematic coding and co-occurrence structure for agrifood system transition lock-ins in Zambia
Description
This dataset contains the reproducibility materials developed for a structured integrative review of agrifood system transition constraints in Zambia. It comprises 63 verified analytical sources coded across 14 themes grouped within political, economic, environmental and social system dimensions. The deposited files include theme definitions, a source by theme binary coding matrix, the complete list of 91 pairwise theme co-occurrences, a 14 × 14 co-occurrence matrix, source data used for the corpus profile, the original thematic coding workbook, a data dictionary, and supporting reproducibility files. Each source is assigned a unique identifier that permits cross-checking of bibliographic information, verification status, and thematic assignments. The dataset is intended to support reproducibility and secondary analysis of the associated Agricultural Systems study, including recalculation of theme frequencies, examination of alternative co-occurrence thresholds, auditing of coding decisions, and comparative application in other agrifood-system transition studies. The associated research article is submitted to Agricultural Systems as manuscript AGSY-D-26-04018.
Files
Steps to reproduce
The dataset was compiled for a structured integrative systems review of agrifood-system transition constraints in Zambia. The analytical corpus comprises 63 verified sources published between 1964 and 31 July 2026. Each eligible source was assigned a stable Source ID and bibliographic and verification information was recorded. Sources were manually coded according to their documented analytical use in the review across 14 predefined themes grouped within political, economic, environmental and social system dimensions. Coding was binary: 1 indicates that a source was used as evidence for a particular theme and 0 indicates that it was not. Coding was based on analytical use rather than automated keyword extraction. To reproduce the thematic summaries, use 02_source_matrix.csv together with the theme definitions in 01_theme_definitions.csv. The source-by-theme matrix contains 63 rows and 14 binary theme columns. Theme prevalence is obtained by summing each binary theme column. Pairwise theme co-occurrence is reproduced by multiplying the transpose of the 63 × 14 binary source-by-theme matrix by the matrix itself (C = XᵀX). Diagonal values give the number of sources assigned to each theme, while off-diagonal values give the number of sources jointly coded to each pair of themes. With 14 themes, there are 91 unique unordered theme pairs. The complete results are supplied in 03_complete_edge_list.csv and 04_theme_cooccurrence_matrix.csv. The original visualization retained theme pairs with at least six co-coded sources. An independent reproduction workflow is provided in 08_reproduce_theme_summary.py. To reproduce the calculations, extract all files to the same directory, install Python 3 and pandas, and run the script. It reads the source matrix and theme definitions, recalculates theme totals and the complete co-occurrence matrix, generates the 91-pair edge list, and identifies pairs meeting the ≥6-source display threshold. The generated outputs can be compared with the deposited co-occurrence files for verification. 05_figure2_corpus_profile.csv contains the source data used for the corpus-profile summaries, while 06_original_thematic_coding_workbook.xlsx preserves the original coding workbook. 07_data_dictionary.csv provides field-level definitions. No participant-level, animal, patient or social-media data are included, and copyrighted full-text publications are not redistributed.
Institutions
- University of ZambiaLusaka Province, Lusaka