Snapshots of a Culture War: Dataset of X/Twitter and Reddit Posts from Conservative Christian and LGBTQIA+ Youth Issues Discourse Across Four Time Periods, with Sentiment Scores

Published: 17 August 2026| Version 1 | DOI: 10.17632/xv98t86rf6.1
Contributors:
,
,

Description

This dataset consists of 1159 entries scraped from Reddit and X/Twitter between February 10 and March 20, 2026, including at least 100 posts from each platform representing each of four different political/cultural/historical/technological moments in the online cultural struggle between the interests of LGBTQIA+ youth and those of conservative Christian ideologies and institutions. Posts were collected from before the Trump era, during the Trump Era but before the COVID-19 lockdowns, during the COVID Era, and after the COVID Era. Each post has been assigned three sentiment codes: One by a machine learning approach created for general language, one by a machine learning approach created specifically for X/Twitter, and one by a social scientist who researches LGBTQIA+ youths’ issues with conservative Christianity, who has previously published research that involved coding short-format text data. To keep work on this dataset within the realm of activity that is not human subjects research, the dataset does not include the body text of posts marked as deleted or removed in their respective databases. For the sake of technical completeness, the small percentage of posts found by our algorithms that were not really part of the discourse (e.g., product advertisements, “AI slop”) were left in, along with the sentiment categories our machine learning approaches assigned to them. This dataset may be useful to investigators of LGBTQIA+ youths’ real-lived experiences with online discourse involving conservative Christian ideologies and institutions, and to computer scientists interested in training sentiment analysis models to address politically charged issues.

Files

Steps to reproduce

This dataset is the product of multiple scraping tools custom-created to work within the boundaries of terms of use and technological affordances of Reddit and X/Twitter at the time of data collection. X/Twitter data were collected using the twikit library, Reddit data from 2021 and earlier were collected from the Pushshift Reddit archive, and X/Twitter data were collected via the native Reddit search API. Sentiment scores were assigned through a Robustly optimized Bidirectional Encoder Representations from Transformers approach (RoBERTa) for general language, the X/Twitter-specific cardiffnlp/twitter-roberta-base-sentiment-latest architecture, and manual coding by a human expert on the subject matter, who left notes in the dataset about cases that required judgment calls. Data were then stripped of positively identifying information like usernames to satisfy publication requirements.

Institutions

Categories

Social Sciences, Psychology, Religion, Christianity, Social Media, Natural Language Processing, LGBTQIA+, LGBTQ Health, Sentiment Analysis, Social Media Discourse Analysis

Licence