Patientdata
Description
The dataset comprised 906 naturalistic accounts drawn from publicly available discussion threads posted by healthcare professionals on Reddit between 2015 and 2025. Posts were selected based on their relevance to patient-initiated mistreatment, including verbal aggression, threats, accusations, and boundary violations. The dataset spans a range of clinical roles (e.g., physicians, nurses, residents) and care settings, capturing diverse experiences of frontline healthcare work. Each account consisted of user-generated narrative reflections describing specific encounters, workplace conditions, and coping responses. As naturally occurring, unsolicited discussions, these data provide access to candid sensemaking processes that are less susceptible to retrospective bias or researcher prompting. The dataset therefore enables examination of mistreatment as it is experienced, interpreted, and managed in situ across repeated interactions and contexts.
Files
Steps to reproduce
Platform selection Data were collected from publicly accessible discussion threads on Reddit, a platform widely used by healthcare professionals to share workplace experiences anonymously. Subreddit identification Relevant subreddits were identified based on their focus on healthcare work (e.g., forums for physicians, nurses, and medical trainees). Selection was guided by activity level, relevance to clinical work, and presence of discussion threads on patient interactions. Keyword-based search Threads were retrieved using search terms related to patient-initiated mistreatment (e.g., “abuse,” “threat,” “aggressive patient,” “yelled at,” “violent patient,” “difficult patient”). Searches were conducted across posts dated between 2015 and 2025. Thread screening Retrieved threads were screened for relevance. Threads were included if they contained first-hand accounts of mistreatment in healthcare settings. Threads focused on unrelated topics (e.g., administrative discussions without interpersonal content) were excluded. Post-level inclusion criteria Individual posts within selected threads were included if they: Described a specific interaction involving patient-initiated mistreatment Reflected the perspective of a healthcare professional Contained sufficient narrative detail for interpretation Posts were excluded if they were: Duplicates or reposts Non-substantive (e.g., one-line reactions without context) Written from non-clinician perspectives Data extraction and cleaning Relevant posts were extracted and compiled into a structured dataset. Usernames and identifying information were removed to preserve anonymity. Non-relevant content (e.g., emojis, hyperlinks, formatting artifacts) was cleaned. Final dataset construction The final dataset consisted of 906 accounts drawn from 262 threads. Each account was treated as a unit of analysis representing a distinct narrative of mistreatment and response.
Institutions
- Indian Institute of Management BangaloreKarnataka, Bengaluru