Identifying Smiling Depression vai Sentiment Score

Published: 16 June 2025| Version 1 | DOI: 10.17632/gjcghsrgm8.1
Contributors:
,
,
,
,
,
,
,

Description

The dataset comprises self-reported responses from undergraduate students across various colleges and universities in South India, primarily from the states of Telangana, Andhra Pradesh, and Tamil Nadu. Participants are aged between 18 and 22 years and are enrolled in technical (e.g., B.Tech) and traditional undergraduate programs (e.g., B.A., B.Sc., B.Com.). The data captures a diverse mix of students from urban, semi-urban, and rural backgrounds, providing a broad perspective on student experiences. Each entry includes basic demographic information such as age, gender, place of living, educational qualification, and the name of the institution. The core of the dataset consists of responses to a series of statements designed to probe behaviors and attitudes associated with "smiling depression"—a phenomenon where individuals conceal depressive symptoms behind a façade of happiness. Statements address themes such as hiding negative emotions, reluctance to discuss problems, the use of humour to mask distress, and the perceived need to appear happy or in control around others. For each statement, students indicate the frequency or degree to which it applies to them, using scales such as "Always," "Often," "Sometimes," "Rarely," and "Never," or levels of agreement from "Strongly Agree" to "Strongly Disagree." These qualitative responses are numerically encoded for analysis. Additionally, the dataset includes sentiment scoring outputs, where each response is assigned a sentiment value (e.g., -2 for strongly negative, 0 for neutral, +2 for strongly positive). An average sentiment score is calculated for each participant, offering a quantitative measure of their overall emotional expression related to smiling depression behaviors. This scoring allows for the identification of students who may be at risk, despite outward appearances of well-being. The data is anonymised to protect participant privacy, with personal identifiers removed. The comprehensive structure and sentiment scoring make the dataset suitable for sentiment analysis, enabling researchers to detect patterns and prevalence of smiling depression among undergraduate students and to explore correlations with demographic factors.

Files

Steps to reproduce

Survey Design and Preparation A structured questionnaire was developed to assess smiling depression among undergraduate students. The survey included demographic questions (age, gender, place of living, education qualification, college/university, and state) and a series of statements related to smiling depression behaviors. Respondents were asked to rate each statement on a Likert-type scale (e.g., Always, Often, Sometimes, Rarely, Never or Strongly Agree to Strongly Disagree). Data Collection The survey was distributed to undergraduate students across multiple colleges and universities in India, covering technical and traditional degree programs. Data was collected online, ensuring voluntary participation and anonymity. Responses were timestamped and included all required demographic and behavioural information. Data Cleaning and Preprocessing The raw responses were reviewed for completeness and consistency. Any entries with missing or invalid data were removed. Text responses were standardised, and categorical options were mapped to numerical values for analysis (e.g., Always = 2, Never = -2). Sentiment Scoring Each behavioural response was assigned a sentiment score based on its position on the Likert scale. For example, positive expressions (e.g., “Always” hiding feelings) were mapped to negative sentiment scores, reflecting depressive tendencies. The scoring rubric was applied uniformly to all responses. Calculation of Aggregate Scores For each participant, an average sentiment score was computed by aggregating the sentiment values across all behavioral indicators. This provided a quantitative measure of the extent to which a student exhibited smiling depression characteristics. Data Structuring The cleaned and scored data was organized into a tabular format, with each row representing a participant and columns for demographic details, individual sentiment scores for each statement, and the average sentiment score. Analysis and Visualization The processed dataset was analyzed to identify patterns and prevalence of smiling depression among different demographic groups. Visualizations and summary statistics were generated to support interpretation and reporting. Ethical Considerations Throughout the process, participant anonymity and data confidentiality were maintained. No personal identifiers were included in the final dataset

Institutions

  • Anurag Engineering College
  • Velagapudi Ramakrishna Siddhartha Engineering College
  • Vidya Jyothi Institute of Technology
  • Alagappa University Faculty of Arts
  • Vardhaman College of Engineering
  • VIT-AP Campus
  • Koneru Lakshmaiah Education Foundation

Categories

Demographics, Behavioral Effect, Sentiment Analysis

Licence