LLM-Based Synthetic Dataset for Mental Health Text Analysis
Description
This dataset contains a paired parallel corpus of authentic and synthetic mental health disclosures. The human corpus consists of 5,184 real Reddit posts collected from three mental health communities: Reddit r/depression, r/anxiety, and r/mentalhealth. Each post was rewritten under standardized conditions by multiple contemporary large language models, producing a synthetic corpus designed for comparative analysis of linguistic, emotional, and stylistic variation between human and AI-generated mental health discourse. The dataset supports research in computational linguistics, machine behavior, AI psychology, and authenticity in computer-mediated mental health communication.
Files
Steps to reproduce
Collect authentic mental health posts from Reddit, preprocess them, and generate synthetic rewrites using standardized prompts across multiple contemporary LLMs to enable comparative analysis between human and AI-generated discourse.
Institutions
- Amrita Vishwa VidyapeethamTamil Nadu, Coimbatore