TRACE-ESG: Traceable Reliability Assessment of Corporate Evidence for ESG Disclosure Mining
Description
This dataset supports the MethodsX article introducing TRACE-ESG, a formal protocol for auditable ESG evidence construction from corporate sustainability reports. The package provides cleaned corpus metadata, corpus-audit logs, duplicate-resolution records, extraction-quality logs, construct dictionaries, reproducibility scripts, baseline lexical candidate evidence outputs, candidate-count summaries, and method documentation. The TRACE-ESG corpus was assembled from corporate sustainability report records covering 2023–2025. The initial corpus contained 254 report records. After duplicate-hash screening and corpus-cleaning review, the cleaned MethodsX working corpus contains 247 report records. The cleaned manifest currently includes 246 unique document identifiers and 235 unique file hashes because source-review records are retained for audit transparency. A deterministic baseline lexical extraction produced 95,082 candidate evidence rows across five ESG constructs: Climate Emissions, Thai Regulation, Energy Transition, Circularity Waste, and Governance Risk. These rows are candidate evidence only. They are not validated disclosure evidence, disclosure-quality scores, ESG performance measures, or corporate-quality ratings. Human validation, semantic assessment, reliability scoring, and calibration are required before substantive interpretation. Original sustainability-report PDFs are not redistributed in this package unless redistribution rights are confirmed. The package instead provides manifests, metadata, hashes, candidate evidence, audit logs, and scripts to support reproducibility and independent inspection.
Files
Steps to reproduce
1. Download the TRACE-ESG V2 package and review `README_TRACE_ESG_V2.md`, `DATA_DICTIONARY_TRACE_ESG_V2.csv`, and `REPRODUCIBILITY_INSTRUCTIONS.md`. 2. Inspect the cleaned corpus metadata in `corpus/corpus_manifest_2023_2025_cleaned.csv`. This file defines the 247 report records used in the MethodsX working corpus and preserves audit fields such as reporting year, company code, file hash, parser status, extraction status, and source-review flags. 3. Review the corpus-governance files in the `corpus/` folder, especially `corpus_cleaning_audit_log.md`, `duplicate_hash_review_2023_2025.csv`, and `duplicate_hash_manual_resolution.csv`. These files document how the initial 254 report records were screened and reduced to the cleaned working corpus. 4. Recreate or inspect text-extraction quality using the scripts in `scripts/`. Run `01_build_corpus_manifest.py`, `02_extract_pdf_text.py`, `03_build_quality_log.py`, and `04_build_corpus_summary.py` if the original PDFs are available locally. Original report PDFs are not redistributed in this package unless rights are confirmed. 5. Review the construct dictionary in `config/construct_dictionary.csv`. This dictionary defines the five baseline ESG constructs used for lexical candidate extraction: Climate Emissions, Thai Regulation, Energy Transition, Circularity Waste, and Governance Risk. 6. Run `scripts/05_run_cleaned_candidate_extraction.py` to reproduce the baseline lexical candidate evidence output from the cleaned corpus manifest and extracted text files. 7. Compare the generated output with `outputs/candidate_evidence_2023_2025_cleaned.csv` and `outputs/construct_candidate_counts_2023_2025_cleaned.csv`. The expected baseline output is 95,082 candidate evidence rows across the five ESG constructs. 8. Use `outputs/candidate_extraction_run_summary.md` and `outputs/construct_candidate_counts_2023_2025_cleaned.md` to verify the corpus-level and construct-level candidate counts. 9. Interpret the output carefully. Candidate rows are baseline lexical candidate evidence only. They are not validated disclosure evidence, disclosure-quality scores, ESG performance measures, or corporate-quality ratings. 10. For validation or extension, apply the guidance in `config/validation_protocol.md` and `config/scoring_specification_TRACE_ESG.md`. Human validation, semantic assessment, reliability scoring, and calibration are required before candidate rows can be interpreted as validated disclosure evidence.
Institutions
- Chulalongkorn UniversityBangkok, Bangkok