MentalDistress: A Multi-Class Social Media Text Dataset for Mental Distress Classification
Description
MentalDistress is a human preference dataset developed for evaluating personalized alignment in Natural Language Processing (NLP). This dataset is developed to support mental health text classification research in the English language. It consists of manually curated and annotated English text samples categorized into five psychological states. The dataset is designed to facilitate supervised learning approaches for detecting emotional distress and high-risk mental health indicators from textual data. The corpus contains a total of 9,935 annotated text samples, distributed across five classes representing different mental health conditions. šššš¦š¦ ššš¦š§š„šššØš§šš¢š” The dataset includes the following categories: ⢠Suicidal: 2,131 samples ⢠Depressed: 1,458 samples ⢠Anxious: 1,958 samples ⢠Frustrated: 1,846 samples ⢠Others: 2,542 samples The class distribution is relatively balanced, making the dataset suitable for multi-class classification experiments without severe class imbalance issues. ššš§š šš¢ššššš§šš¢š” šš”š šš”š”š¢š§šš§šš¢š” The text samples were collected from publicly available sources and manually reviewed. Each instance was carefully annotated according to predefined psychological category guidelines to ensure labeling consistency. Quality control measures were applied to maintain annotation reliability. ššš¬ šššš§šØš„šš¦ ā¢ Five-class mental health categorization ⢠Manually annotated dataset ⢠Fairly balanced class distribution ⢠Suitable for classical ML and deep learning models š£š¢š§šš”š§ššš šØš¦š ššš¦šš¦ ā¢ Mental health text classification ⢠Early detection of psychological distress ⢠NLP research ⢠Transformer-based model benchmarking ⢠AI-assisted mental health screening research šššš šš¢š„š šš§ The dataset is provided in CSV format with the following columns: ⢠original_row: Original row identifier ⢠text: English textual content ⢠label: Class name (Anxious, Depressed, Frustrated, Others, Suicidal) ⢠text_length: Length of the corresponding text sample ⢠locked_split: Predefined dataset split (train, validation, test)
Files
Institutions
- Leading UniversitySylhet Division, Sylhet