Screened, Merged and Deduplicated Pre-Harmonisation Bibliographic Dataset on Blockchain Technology in Accounting, 2006–2026

Published: 30 June 2026| Version 2 | DOI: 10.17632/k7f5xccmfc.2
Contributor:

Description

This dataset contains a screened, merged, record-level deduplicated, pre-harmonisation bibliographic corpus on blockchain technology and related distributed-ledger applications in accounting. Bibliographic records were retrieved from Scopus and the Web of Science Core Collection in June 2026. The searches initially returned 6,909 records from Scopus and 6,096 records from Web of Science. An initial title-and-abstract screening for topical relevance was applied to both databases before the records were uploaded to BiblioSpy. This stage excluded 4,994 Scopus records and 4,133 Web of Science records, leaving 1,915 Scopus records and 1,963 Web of Science records. A total of 3,878 records therefore entered the database-integration stage. BiblioSpy identified and removed six intra-database duplicate records—three within Scopus and three within Web of Science—and identified 936 publications represented in both databases. After record-level deduplication and integration, 2,936 unique publications remained, comprising 976 Scopus-only records, 1,024 Web of Science-only records and 936 records indexed in both databases. A second relevance screening was conducted after database integration and deduplication. The titles, abstracts and keywords of the 2,936 unique publications were examined against the study scope. Records were retained when they substantively addressed blockchain technology, distributed-ledger technology, smart contracts, cryptocurrencies, Bitcoin or triple-entry accounting in relation to accounting, auditing, assurance, financial reporting, corporate reporting, management accounting, cost accounting, taxation, bookkeeping, internal control, accounting information systems, continuous auditing or real-time accounting. A further 2,217 records were excluded because they were unrelated or insufficiently related to the study topic. The deposited dataset contains the resulting 719 unique journal articles published between 2006 and 2026. These comprise 235 Scopus-only records, 208 Web of Science-only records and 276 records indexed in both databases. The `Data Source` field records the database provenance of each publication. The dataset represents the corpus before comprehensive field-level metadata cleaning and harmonisation. Variations and inconsistencies may remain in author names, institutional affiliations, source titles, countries, author keywords, indexed keywords and cited references. The dataset has undergone database parsing, record-level integration, duplicate handling and relevance screening, but further metadata cleaning and harmonisation are required before reproducing the final bibliometric indicators and science-mapping outputs reported in the associated study.

Files

Steps to reproduce

1. Formulate a bibliographic search covering blockchain technology, distributed-ledger technology, smart contracts, cryptocurrencies, Bitcoin and triple-entry accounting in relation to accounting, auditing, assurance, financial reporting, taxation, internal control, bookkeeping and accounting information systems. 2. Run the documented search strategies in Scopus and the Web of Science Core Collection. Apply the eligibility criteria reported in the associated manuscript. 3. Export the available bibliographic records and cited-reference information from both databases. Retain database identifiers and provenance information. 4. Upload the Scopus and Web of Science exports to BiblioSpy version 1.0 and use the database-specific parsers to convert the records into a common bibliographic structure. 5. Identify and remove intra-database duplicate records. In the reported workflow, three duplicates were removed from Scopus and three from Web of Science. 6. Merge the two databases using BiblioSpy’s cross-database matching procedure. Matching uses available identifiers and bibliographic information, including DOI, title, authors, publication year and source title. The procedure identified 936 publications represented in both databases. 7. Retain one integrated record for each cross-database match. Where available, use the Scopus record as the base record and retain the more complete content for selected fields, including abstracts, keywords, affiliations and cited references. 8. Screen the titles, abstracts and keywords of the 2,936 unique records against the study scope. Include records that substantively investigate blockchain or related distributed-ledger technologies in an accounting-related context. 9. Exclude records in which blockchain terminology is incidental or where the principal subject is unrelated to accounting. The screening process excluded 2,217 records. 10. Export the resulting corpus of 719 unique records as a CSV file. Preserve the Data Source field to distinguish Scopus-only, Web of Science-only and shared records. 11. To reproduce the deposited Version 1 dataset, do not apply comprehensive field-level harmonisation after screening. Author names, institutional affiliations, source titles, keywords and cited references should remain in their pre-harmonisation form. 12. Before conducting final bibliometric or science-mapping analyses, undertake and document the necessary field-level cleaning, disambiguation and harmonisation procedures. Analytical results may differ depending on the cleaning rules and thresholds applied.

Institutions

Categories

Accounting, Information System, Research Method, Bibliometrics, Blockchain, Scientometrics

Licence