TXD-22: A Large-Scale Benchmark Dataset for Multi-Class AI-Generated and Mixed-Source Text Identification

Published: 1 July 2026| Version 1 | DOI: 10.17632/prcjcggtjf.1
Contributors:
,

Description

TXD-22 is a large-scale, class-balanced benchmark dataset for multi-class AI-generated text detection, source attribution, and mixed-authorship analysis. It contains 66,000 question-answer samples evenly distributed across 22 classes, with exactly 3,000 samples per class. The 22 classes comprise text generated by 17 different AI systems (ChatGPT, Gemini, Claude, LLaMA, Mistral AI, Grok, Copilot, Perplexity, DeepSeek, DeepAI, Qwen, Gemma, Poe, Nova, Pi, Z.ai, and BLACKBOXAI), one class of purely human-written text, and four mixed-authorship classes combining human-written, AI-generated, and AI-refined content: (1) AI-generated + Human-written, (2) AI-generated + AI-refined, (3) Human-written + AI-refined, and (4) AI-generated + AI-refined + Human-written. The data were collected using a structured and reproducible framework. A hierarchical question set spanning 30 academic and professional domains, each divided into thematic sub-domains, was used to elicit semantically rich responses, ensuring topic diversity while preserving contextual consistency. AI responses were generated between January and March 2025 using a unified prompting strategy with a maximum response length of approximately 300 words. Human responses were contributed by around 350 participants from diverse academic and professional backgrounds, who wrote original answers in English without AI assistance; participation was voluntary, informed consent was obtained, and no personally identifiable information is retained. All samples passed a standardized preprocessing and quality-control pipeline (Unicode normalization, whitespace cleaning, punctuation standardization, duplicate removal, and response-length normalization). AI-generated outputs were not manually rewritten, preserving each source's original linguistic and stylistic characteristics. File and structure: The dataset is provided as a single CSV file with three columns: Question (the prompt), Answer (the response text), and Source (the class label indicating the generating AI system, human author, or mixed-authorship configuration). Unlike prior work that focuses mainly on binary human-versus-AI classification using small secondary datasets, TXD-22 targets a realistic and challenging multi-class scenario in which many AI systems and hybrid samples share highly similar linguistic patterns. As a baseline, ALBERT features combined with a Support Vector Machine (SVM) classifier achieved 52.01% accuracy on the 22-class task, illustrating the inherent difficulty of fine-grained source attribution in this setting.

Files

Categories

Natural Language Processing

Licence