Dataset for “When Ambition Outruns Instruments: Provision-Level Policy Coherence for an Effective Circular Economy Transition in Peru”

Published: 27 August 2026| Version 1 | DOI: 10.17632/sfbwdzpfgt.1
Contributors:
karina chuquispuma,

Description

This dataset contains the supplementary data, codebook, JSON schema, extracted legal provisions, and pair-level scoring results for the manuscript "When Ambition Outruns Instruments: Provision-Level Policy Coherence for an Effective Circular Economy Transition in Peru". Contents: -Corpus Registry: Complete list of the 142 legal-regulatory instruments analyzed in the study. -Codebook & Schema: 18-field JSON schema and formal coding definitions used for LLM-assisted extraction. -Coded Dataset: 8,658 atomic provisions coded across the CEAP taxonomy nodes and canonical regulated objects. -Scoring & Adjudication Data: Pair-level evaluations and consensus scores across the four coherence axes (transposition, vertical fidelity, horizontal coherence, and collision).

Files

Steps to reproduce

1.Corpus Collection & Pre-processing: Retrieve the legal-regulatory instruments from public repositories (SPIJ and El Peruano) in PDF format. Convert machine-readable documents using layout-preserving extraction and transcribe scanned norms using layout-aware/multimodal extraction. 2.Provision-Level Extraction & Validation: Run the fixed, zero-shot, schema-constrained prompt via the API on each instrument to extract atomic provisions into the 18-field JSON schema. Validate all outputs against the formal schema to ensure structural completeness and referential integrity. 3.Normalization & Blocking: Map extracted regulated objects to the controlled vocabulary of 44 canonical regulated objects. Generate candidate comparison pairs by blocking provisions that share both the same CEAP taxonomy node and the same canonical regulated object. 4.Pair-Level Relational Scoring: Apply the corresponding scoring framework to each candidate pair. For the Transposition and Vertical axes, use the four-level fidelity rubric: faithful, partial, distorted, and gap. For the Horizontal and Collision axes, use Nilsson’s seven-point interaction scale from −3 to +3. Use multi-model LLM evaluation with temperature set to 0 and determine provisional verdicts through majority-vote consensus. 5.Expert Adjudication: Conduct human expert review to resolve residual model disagreements or scoring ties and validate the finalized pair-level verdicts.

Institutions

Categories

Social Sciences, Artificial Intelligence, Environmental Science, Sustainability, Public Policy, Circular Economy

Licence