District-week notifiable infectious disease surveillance data for Khyber Pakhtunkhwa, Pakistan, 2023-2026
Description
A structured, long-format panel dataset of district-week notifiable infectious disease case counts for Khyber Pakhtunkhwa province, Pakistan, extracted from 164 weekly Integrated Disease Surveillance and Response (IDSR) Public Health Bulletins published by the National Institute of Health, Pakistan, between April 2023 and July 2026. The dataset comprises 59,824 district-week-condition observations spanning 44 canonical district- and sub-division-level reporting units and 16 notifiable conditions. Six conditions - acute diarrhoea (non-cholera), malaria, influenza-like illness, acute lower respiratory infection in children under five, bloody diarrhoea and typhoid - are reported in all 164 bulletins. Ten further conditions appear only in a subset of weeks, because from 2023 week 32 the source table reports the highest-ranking conditions by provincial case total rather than a fixed set, producing 106 distinct column orderings across the corpus. Absence of a rotating condition from a given week indicates non-inclusion in that week's ranked columns, not a zero count. The deposit documents this rotating-column structure explicitly and includes a time-scoped district crosswalk resolving administrative-boundary instability over the study period, including one reporting label whose underlying geography is redefined mid-series. Machine-readable flags record every known data-quality issue, including a duplicated source table, four rows corrected for a dropped-cell transcription error, and weeks in which district rows do not reconcile with the printed provincial total. The raw transcribed source text and the parser that produces the published files from it are both included, so the path from source to panel is fully reproducible. No comparable structured, validated dataset derived from Pakistan's national disease surveillance bulletins is currently available in the public domain.
Files
Steps to reproduce
Data were extracted from Table 4, the Khyber Pakhtunkhwa district-wise surveillance table, of the weekly IDSR Public Health Bulletins published as PDFs by the National Institute of Health, Pakistan (https://www.nih.org.pk/). 1. RETRIEVAL. All bulletins covering 2023 week 16 to 2026 week 28 that exist in the publisher's archive were used: 164 of 169 possible weeks. The five absent weeks (2023 w17-21) were confirmed missing from the archive itself, not unreachable. Bulletin filenames follow no consistent convention, so the publisher's index pages were harvested for actual hyperlinks rather than constructing URLs from a template. Every resolved URL is in bulletin_links_verified.csv. 2. TRANSCRIPTION. Table 4 was transcribed from each PDF using an AI-assisted document-reading tool, preserving column headings, row labels and cell values exactly as printed, including the literal "NR" where a value was not reported. Values were transcribed as printed, not inferred. The complete raw transcription is preserved in tables_raw.txt, in blocks headed "### YEAR WEEK". 3. PARSING. build.py converts tables_raw.txt into every CSV in this deposit. It normalises all observed column-heading variants onto 16 canonical condition names; maps 50 raw district and sub-division labels onto 44 canonical reporting units, time-scoped for the one label whose referent changes mid-series; corrects four rows affected by a dropped-cell transcription error; and sorts chronologically. A build-time assertion fails the build if two source rows within a week collapse onto the same canonical district and condition. To regenerate: set the RAW and OUT paths at the top of build.py, then run "python3 build.py". It needs only the Python standard library. Output is byte-identical to the CSV files deposited here. 4. VALIDATION. Four internal consistency checks were applied. Each week's printed Total row was tested for descending order, confirming a rank-ordered column design from 2023 week 32 and identifying the 11 earlier weeks using a fixed layout. Every data row's cell count was checked against its own week's heading count; no mismatches across 164 bulletins. A suspected duplicate week was verified by re-retrieving both source PDFs. Finally, district rows were summed against the printed provincial Total for every column in every week carrying one. This last check located a class of transcription error in which a blank cell in the source causes subsequent values to be assigned to the wrong condition. Four such rows were corrected, one confirmed directly against the source PDF. After correction, 137 of the 163 weeks with a printed Total reconcile exactly on every column; residual discrepancies are documented in data_quality_flags.csv and in the accompanying article.
Institutions
- University of SwatKhyber Pakhtunkhwa, Mingora