Data of Beyond “AI-Generated”: How Disclosure Granularity Shapes Human–AI Attribution and Quality-Sensitive Evaluation
Description
This dataset supports the study “Beyond ‘AI-Generated’: How Disclosure Granularity Shapes Human–AI Attribution and Quality-Sensitive Evaluation.” The project examines how different forms of disclosure about generative AI use shape viewers’ attribution of human and AI production roles and their evaluation of AI-generated visual works. The repository contains de-identified data and supporting materials from two complementary studies. Study 1 examines Human–AI Attribution. Study 1a provides developmental and supporting evidence used to refine the disclosure materials and experimental design. Study 1b provides the primary inferential evidence for H1 and H2, testing whether basic AI disclosure increases awareness of AI involvement and whether process-level contribution disclosure improves recognition of the documented allocation of human and AI production roles. The deposited materials include participant- and trial-level analytic data, stimulus documentation, scoring keys, balanced stimulus-list records, and supporting validation materials. Study 2 examines Quality-Sensitive Evaluation. Phase A contains independent blind ratings used to establish a consensus-based quality benchmark for 24 AI-generated visual stimuli. Phase B contains pilot data used to evaluate the disclosure materials. The main experiment compares coarse source disclosure, an information-matched disclosure condition, and granular human–AI contribution disclosure. The deposited data include participant- and trial-level evaluation records, condition-specific evidential-probe responses and scoring keys, stimulus-level quality benchmarks, disclosure-material information, stimulus metadata, and supporting robustness and sensitivity-analysis variables. The repository also includes reviewer-oriented documentation, variable guides, data indexes, quality-assurance summaries, claim-to-data crosswalks, and supporting materials intended to facilitate inspection and reproducibility of the reported analyses. Where applicable, the accompanying documentation distinguishes primary confirmatory analyses from developmental, secondary, robustness, and post hoc analyses. The Phase A quality scores should be interpreted as an independently prevalidated consensus-based benchmark rather than as an objective measure of artistic quality. Similarly, objective attribution and evidential-probe accuracy are scored against the documented study-specific production-role or condition-specific evidential keys and should not be interpreted as forensic reconstruction of unobserved production provenance. All participant-level analytic datasets have been de-identified before deposit and contain no direct personal identifiers. The shared materials are provided for scholarly inspection, verification of the reported analyses, and reuse subject to the terms of the repository licence and applicable ethical requirements.
Files
Steps to reproduce
Steps to reproduce 1.Download and extract all files in the repository. Begin with the README, data index, variable guide, and claim-to-data crosswalk, which describe the role of each dataset and identify the primary files supporting H1–H4. 2.For Study 1b, use the de-identified participant- and trial-level datasets together with the frozen scoring key. Reconstruct the AI-involvement awareness outcome for H1 and the documented human–AI production-role recognition outcome for H2. Study 1a materials are developmental and supporting evidence and are not used as the primary inferential test of H1 or H2. 3.For Study 2 Phase A, calculate the stimulus-level mean overall-quality rating for each of the 24 visual stimuli. Standardise the locked Phase A stimulus means using: $$ \text{Quality}_{z,j} = \frac{\bar{Q}^{\text{Phase A}}_{j}-4.947125}{0.611522}. $$ These scores provide the independent consensus-based quality benchmark used in the main Study 2 analysis. 4.Merge the resulting stimulus-level quality_z values with the Study 2 main trial-level dataset using the stimulus identifier. Do not recompute the quality benchmark using ratings from the main experiment. 5.Reproduce the primary Study 2 analysis by fitting the crossed linear mixed-effects model: overall_quality ~ quality_z * condition + category + prompt_level + stimulus_list + (1 + quality_z || response_id) + (1 + D2 + D3 || stimulus_id) D1 is the reference disclosure condition. The condition-specific quality slopes are therefore \(\beta_1\) for D1, \(\beta_1+\beta_4\) for D2, and \(\beta_1+\beta_5\) for D3. H3 tests the D3–D1 slope contrast, \(\beta_5\), whereas H4 tests the D3–D2 slope contrast, \(\beta_5-\beta_4\). Random-effect correlations are fixed to zero. 6.Apply the planned multiplicity adjustment to the H3 and H4 contrasts as documented in the analysis materials. The accompanying scripts and output files reproduce the reported condition-specific slopes, confidence intervals, and planned contrasts. 7.For the objective evidential-calibration analysis, score responses using the condition-specific evidential key supplied in the repository. A response of “Cannot determine from the information provided” can be correct in D1 or D2 when the disclosed information does not support a more specific inference. 8.Run the supplied robustness and sensitivity analyses after reproducing the primary models. These include alternative specifications and supporting diagnostics reported in the manuscript and supplementary materials. Developmental, secondary, robustness, and post hoc analyses are identified separately in the documentation. 9.Compare the regenerated outputs with the archived result tables and quality-assurance summaries. Software requirements, analysis scripts, variable definitions, scoring rules, and supporting documentation are included with the reproducibility materials.