Screening data for a systematic mapping of Scopus-indexed 2026 literature on artificial intelligence in Chinese education (SAMYRAD 2026, contribution 302)

Published: 8 September 2026| Version 1 | DOI: 10.17632/3nx7ysk4my.1
Contributor:
Juan Sebastián Fernández-Prados

Description

Record-level screening data of a systematic mapping of the 2026 literature on artificial intelligence in Chinese education, companion to the paper "Digital transformation and systemic architecture of Chinese education: trajectory and the AI + Education plan toward 2030" (SAMYRAD 2026, Seville, 5-6 October 2026, IEEE Xplore proceedings). The corpus was retrieved from Scopus on 19 May 2026 with the search string TITLE-ABS-KEY (China AND education AND "artificial intelligence") AND PUBYEAR = 2026 (N = 363 records: 265 articles, 40 conference papers, 24 reviews, 18 book chapters, 16 other document types; 35 in press). Records were screened at title and abstract level against two inclusion criteria (setting in the Chinese education system; at least one pedagogical, attitudinal or competency outcome) and four hierarchically applied exclusion criteria (document type; AI not the object or outside an instructional setting; no primary outcome data; non-Chinese or unspecified setting). Screening was performed in two passes: a first pass by a large language model (Claude, Anthropic) applying the criteria, and an independent verification of every record by one author. Retained records were assigned a primary and, where relevant, a secondary theme against a codebook (acceptance models and cognitive friction; socio-territorial stratification; AI literacy frameworks; disciplinary specialization and local language models; and an open category). Files: search protocol (query, date, export settings, record composition); screening workbook with one row per record (identifiers, title, document type, publication stage, decisions and reasons of both passes, themes, borderline flags), a Summary sheet (counts, percent agreement and Cohen's kappa between passes) and a README sheet with the codebook; a CSV of record identifiers (Scopus ID, EID, DOI); Scopus advanced-search queries that reconstruct the corpus exactly; and the codebook as a separate file. The raw Scopus export (abstracts, author lists, keywords) is not redistributed, in line with Scopus terms of use; every record keeps its Scopus ID, EID and DOI, so full metadata can be retrieved from Scopus or the publisher. Screening decisions, protocol and codebook are released under CC BY 4.0. Bibliographic identifiers and titles are factual metadata of third-party works provided for identification only.

Files

Steps to reproduce

1. Corpus. In Scopus advanced search, run TITLE-ABS-KEY (China AND education AND "artificial intelligence") AND PUBYEAR = 2026. On 19 May 2026 this returned 363 records. Because PUBYEAR resolves to whole years, a later run returns a larger set; to reconstruct the exact corpus, run the four EID queries in 05_eid_reconstruction_queries.txt (100 EIDs each) and take the union, or match 04_corpus_record_identifiers_2026-05-19.csv against your export. 2. Export. Select all records and export in Scopus "Plain text" format with the fields Citation information, Bibliographical information, Abstract and keywords (the abstracts are needed for screening but are not part of this dataset). 3. Criteria. Apply the inclusion and exclusion criteria of 06_codebook_screening_criteria.md at title and abstract level. Exclusion criteria are hierarchical: record the first one met (EC1 document type; EC2 AI not the object or outside an instructional setting; EC3 no primary outcome data; EC4 non-Chinese or unspecified setting). Evaluations of large language models are included only when the instrument is an official Chinese educational or professional examination (rule BENCH). 4. Decisions. Compare your decisions with columns J to Q (first pass, large language model) and R to T (second pass, author verification) of 03_screening_sheet_scopus_2026.xlsx. Sheet Summary recomputes counts, percent agreement and Cohen's kappa from those columns; the counts reported in the paper are those of columns R and S. 5. Themes. Assign each retained record a primary theme (and an optional secondary theme) using the theme codes of the codebook. Theme counts in the paper come from column S. 6. Verification. Check the record composition against 02_search_protocol_scopus_2026-05-19.txt: 265 articles, 40 conference papers, 24 reviews, 18 book chapters, 16 other types; 35 records in press; 362 records with a DOI.

Institutions

Categories

Artificial Intelligence, Education, China

Licence