LLM-Based Synthetic Dataset for Mental Health Text Analysis

Published: 3 July 2026| Version 1 | DOI: 10.17632/55ky4f4d9p.1
Contributors:
,

Description

This dataset contains a paired parallel corpus of authentic and synthetic mental health disclosures. The human corpus consists of 5,184 real Reddit posts collected from three mental health communities: Reddit r/depression, r/anxiety, and r/mentalhealth. Each post was rewritten under standardized conditions by multiple contemporary large language models, producing a synthetic corpus designed for comparative analysis of linguistic, emotional, and stylistic variation between human and AI-generated mental health discourse. The dataset supports research in computational linguistics, machine behavior, AI psychology, and authenticity in computer-mediated mental health communication.

Files

Steps to reproduce

Collect authentic mental health posts from Reddit, preprocess them, and generate synthetic rewrites using standardized prompts across multiple contemporary LLMs to enable comparative analysis between human and AI-generated discourse.

Institutions

Categories

Psychology, Artificial Intelligence, Mental Health, Large Language Model

Licence