Prediction-Level Benchmark Metadata and Model Scores for Multi-Provider AI-Text Detection in Turkish Scholarly Abstracts
Description
This dataset contains prediction-level experimental metadata and derived model scores from a frozen benchmark used to evaluate AI-text detection in Turkish scholarly abstracts. The release comprises 1,678 benchmark records grouped into 60 independent source clusters: 60 human records and metadata for 1,618 AI-assisted variants. The benchmark covers seven academic domains, four generative-AI providers (OpenAI, Gemini, Claude, and DeepSeek), six prompt families, four contribution levels, and 179 adversarial transformations. Each record includes provenance identifiers, source-cluster identifiers, binary labels, contribution metadata, provider and prompt metadata, domain and document-type variables, adversarial-operation metadata, text hashes, word counts, and three frozen model-score fields. The repository also provides raw and processed tabular releases, a data dictionary, source registry, quality-control reports, descriptive tables, authoritative results registries, deterministic validation code, statistical-analysis code, figure-generation code, software requirements, licenses, citation metadata, and SHA-256 checksums. The release is designed to support reproducible investigation of source leakage, cross-domain reliability, human false-positive rates, provider and prompt variation, calibration, threshold sensitivity, and prediction-level information quality. No full copyrighted scholarly abstracts and no full LLM-generated texts are redistributed. The repository contains experimental metadata, numerical prediction outputs, provenance identifiers, truncated content hashes, and reproducibility materials. Dataset accuracy is verified through deterministic Python checks and cryptographic checksums rather than by generative AI.
Files
Steps to reproduce
1. Download all repository files while preserving the directory structure. 2. Read README.md, docs/RIGHTS_AND_REUSE.md, and data/metadata/data_dictionary.csv before analysis. 3. Verify file integrity against SHA256SUMS.txt. 4. Create an isolated Python environment with Python 3.10 or later. 5. Install the required packages with: pip install -r requirements.txt 6. Run the deterministic release checks with: python analysis/validate_release.py 7. Confirm that all checks report PASS, including 1,678 records, 60 source clusters, 60 human records, 1,618 AI-assisted records, 179 adversarial records, unique item identifiers, consistent labels, positive word counts, and score values within [0,1]. 8. Use data/processed/dib_prediction_metadata_v1.csv as the primary analysis table and consult data/metadata/data_dictionary.csv for variable definitions. 9. Reproduce descriptive and inferential analyses using analysis/compute_statistics.py and the statistical-analysis plan in docs/IPM_Statistical_Analysis_Plan.md. 10. Regenerate figures using analysis/make_ipm_figures.py. Compare reproduced outputs with the registries and quality-control files supplied in the repository. 11. For any updated release, increment the dataset version, retain the previous DOI-linked version, regenerate SHA256SUMS.txt, and document changes in README.md.
Institutions
- Fırat UniversityElazığ, Elâzığ