A Bilingual Synthetic Dataset for Elderly Emotional Wellbeing and Agentic AI Support in Low-Resource Contexts
Description
The dataset consists of 2,000 synthetic records, divided equally into two subsets: 1,000 English records and 1,000 Roman Urdu records. Each record represents a simulated elderly participant aged between 50 and 90 years. Both subsets share an identical schema comprising 13 attributes, enabling direct cross-lingual comparison and analysis. No missing values are present in the dataset. Summary statistics confirm identical distributions across both language versions, indicating controlled generation and minimizing unintended linguistic bias. Dataset Attributes Each record includes the following attributes: Scenario_ID: Unique identifier representing the contextual scenario Age: Participant age (50–90) Gender: Male or Female Emotional_Tone: Annotated dominant emotional state Cognitive_Difficulty: Self-reported difficulty level (1–5) Coping_Strategy: Emotional or cognitive regulation behavior Learning_Outcome: Outcome of AI interaction Support_Needed: Type of AI assistance required Technology_Interaction: Mode of technology use Reflection_Tag: High-level reflective keyword Depression_Level: Emotional difficulty indicator (0–10) Perceived_Usefulness: AI usefulness rating (0–10) AI_Utilization_Purpose: Primary intent of AI usage Participant_Response: Free-text response in English or Roman Urdu
Files
Steps to reproduce
The dataset was generated using a scenario-driven prompt engineering framework designed to simulate realistic elderly interactions with an Agentic AI system while maintaining strict control over demographic, emotional, and linguistic variables. Synthetic data generation was selected to eliminate privacy risks and ensure ethical compliance. 1 Scenario Design A finite set of contextual scenarios (S001–S200) was defined to represent common elderly experiences, including emotional reflection, technology usage challenges, learning difficulties, and health-related concerns. Each scenario was parameterized by age, gender, emotional context, and AI interaction intent to ensure balanced coverage. 2 Prompt Engineering Strategy Structured prompts were designed to generate: Natural-language participant responses Emotional tone annotations Cognitive difficulty indicators Coping strategies AI support requirements Perceived usefulness ratings Constraint-based prompting enforced predefined value ranges and categorical consistency. Records failing logical consistency checks were discarded and regenerated. 3 Cross-Lingual Generation English and Roman Urdu datasets were generated independently using identical scenario constraints. Direct translation was avoided to preserve linguistic naturalness. Post-generation statistical analysis confirmed parity across both language subsets.