A harmonized corpus of quantum error correction publications and patents (1995–2025)
Description
This dataset is a harmonized corpus of research and invention output in quantum error correction, the set of techniques by which quantum information is protected against noise, and the central obstacle standing between present-day quantum processors and fault-tolerant computation. It brings together the scholarly literature and the patent record of the field in a single schema and under a single open license, spanning 1995 to 2025. The corpus contains 16,274 records: 14,385 scholarly publications retrieved from OpenAlex and enriched against Crossref, and 1,889 United States patent publications retrieved from Google Patents Public Data. Each record carries a common set of fields, including identifier, source, type, title, abstract, year, venue or issuing office, authors or assignees, citation count, cited references, and subfield tags describing which aspects of the field a record concerns, such as codes, fault tolerance, decoders, logical qubits, surface and topological codes, bosonic codes and magic-state distillation. A separate flag identifies records concerned with machine-learning decoders, and a full variable dictionary accompanies the data. Every record is assigned to one of three retrieval tiers according to the strength of the topical match. The core tier comprises records whose titles carry an explicit quantum-error-correction term and is the subset recommended for most analyses; the extended tier comprises records in which such a term appears only in the abstract; and a third tier holds patents carrying the official classification, CPC G06N10/70, yet containing no such term anywhere. The third group is labelled rather than discarded, because it documents a finding of practical importance: two-thirds of patents bearing this classification carry it as a secondary code on broad quantum-computing inventions, so classification alone is not a sufficient filter. The quality of every tier has been measured rather than asserted. A stratified random sample of one hundred records was drawn from each tier and screened by two annotators working independently, with disagreements resolved by discussion against a fixed boundary rule. Precision is 86.0 per cent in the core tier (Wilson 95 per cent confidence interval 77.9 to 91.5), 37.0 per cent in the extended tier (28.2 to 46.8), and 2.0 per cent in the weak patent tier (0.6 to 7.0); the labelled samples are included so that these figures can be checked. Recall was not measured, as no exhaustive reference set exists for this window. Out-of-scope material consists chiefly of error mitigation and suppression, quantum key distribution, and algorithms designed to run on fault-tolerant machines rather than error correction itself. Drawn entirely from openly licensed sources, the corpus may be redistributed and reused without restriction, supporting bibliometric study of the field, analysis of academic research and industrial invention, and use as a filtered text collection.
Files
Steps to reproduce
All records were harvested computationally from two openly licensed sources, and the complete pipeline is provided as an executable Jupyter notebook in this deposit. No proprietary databases were used, and the procedure is deterministic: re-running the notebook over the same fixed snapshot window reproduces the corpus. The pipeline runs on Python 3.12 and requires a free OpenAlex API key and a Google account linked to BigQuery; it was executed on Kaggle Notebooks. Package versions are recorded in requirements_freeze.txt, and retrieval parameters, environment, and SHA256 checksums in provenance.json and checksums.sha256. The topical boundary is defined by twenty scope phrases, among them quantum error correction, stabilizer code, surface code, magic state distillation and quantum LDPC, each mapped to a subfield tag; the full mapping is in scope_definition.json. Every query is intersected with the OpenAlex topic "Quantum Computing Algorithms and Architecture" so that matches from unrelated disciplines are excluded. Quantum key distribution, error mitigation and classical channel coding are adjacent but out of scope. Tags overlap rather than being mutually exclusive. Publications were retrieved from OpenAlex over 1 January 1995 to 31 December 2025 in two passes, a title pass and a title-and-abstract pass, paginated with cursor paging through the polite pool with exponential back-off. Abstracts were reconstructed from inverted indices, records deduplicated by work identifier, and missing venue names enriched against Crossref. Patents were retrieved from Google Patents Public Data through BigQuery under CPC class G06N10/70; records are patent publications rather than unique inventions, and no family deduplication was applied. Both types were mapped onto one schema and deduplicated within record type, publications by DOI then normalized title, patents by publication number. Each record was assigned to one of three tiers according to whether a scope term appears in the title, in the abstract only, or nowhere. Every tier was validated. A stratified random sample of one hundred records was drawn from each tier under a fixed seed of 42. Two annotators screened every record independently, judging a record in scope only if primarily concerned with encoding into a quantum code, syndrome extraction, and correction, and out of scope if it concerns error mitigation, suppression, dynamical decoupling, control-pulse engineering, leakage removal, quantum key distribution, or algorithms for fault-tolerant machines. Disagreements were resolved by discussion and documented with written reasons in precision_audit_disagreements_resolved.csv. Precision was computed with Wilson 95 per cent confidence intervals and agreement with Cohen's kappa. A large language model produced provisional labels for the core-tier sample only; every record was subsequently labelled independently by both annotators, and the model's labels do not enter the reported figures.
Institutions
- North South UniversityDhaka Division, Dhaka