PRISM Model: Six-Dimension Psychosocial Classification Dataset
Description
This dataset contains the processed, review-level psychosocial classification output produced by PRISM (Psychosocial Review Intelligence and Sensemaking Model), a theory-guided generative AI system developed to interpret employee-generated reviews at situational depth. The dataset supports the manuscript "From Employee Voice to Organisational Intelligence: A Theory-Guided Generative AI System for Psychosocial Interpretation of Employee Reviews," currently under review at the International Journal of Information Management. The dataset comprises 25,204 employee reviews from a global telecommunications company, collected from the Indeed platform between 2011 and 2019. Each row represents one employee review; the 92 columns encode binary indicators (1/0) for the psychosocial categories identified within that review across all six dimensions of the PRISM framework: identity (captured via employee status, U.S. state, and organisational level metadata), emotion (Plutchik's eight primary emotions at High/Medium/Low intensity), context (ten organisational context categories), relationships (twenty-two relationship types), power (French and Raven's five bases of social power at three intensity levels), and personality traits (the Big Five traits at three levels). These classifications were generated through theory-guided LLM prompt engineering, in which each psychosocial dimension was extracted using a dedicated prompt explicitly grounded in established psychosocial theory rather than open-ended or ad hoc classification. Important note on scope: this dataset contains only the derived psychosocial classifications, not the original review text. Raw review text is not included because Indeed's terms of service restrict redistribution of user-generated content; sharing only the derived, aggregated classification output avoids this restriction while still supporting reproducibility of the study's analytical findings. The dataset includes a README sheet describing its structure and provenance, a full data dictionary explaining every column and its theoretical grounding, the classification data itself, and a lookup table of abbreviations used in the metadata fields (e.g., U.S. state codes, employee status codes). This dataset is intended to support verification, reanalysis, and extension of the study's findings by other researchers working in employee voice analytics, organisational text mining, people analytics, and AI-enabled qualitative research methods. The accompanying analytical code used to generate these classifications is openly available at: https://github.com/CDAC-lab/employee-review-analysis.
Files
Institutions
- La Trobe UniversityVictoria, Melbourne