Phailin tweets

Published: 26 February 2026| Version 2 | DOI: 10.17632/w2hknypkjw.2
Contributors:
Swarnalakshmi Umamaheswaran, Anuradha Goswami

Description

The dataset comprises 58,261 tweets collected through the Twitter Academic API, representing one of the few systematic archives of social media discourse around major disasters in South Asia. What makes this dataset particularly valuable is its multi-layered structure. Beyond the raw tweets, we have developed a processed user-level dataset with derived metrics such as spread_score and focus_rate that researchers can use to analyze information diffusion patterns. Additionally, we include 1,500 manually annotated tweets with labels for user types and content categories, which we used to train a few-shot classification model achieving a recall score of 0.72. This annotated subset can serve as a benchmark for future machine learning applications in disaster communication research. The dataset addresses a significant gap in available resources for studying disaster response in the Global South context. Researchers in fields ranging from computational social science to public policy and disaster management will find this resource useful for examining stakeholder communication dynamics, temporal patterns in information flow, and the role of different actors—governments, NGOs, media, and citizens—during crisis events.

Files

Categories

Natural Language Processing, Computer Modeling in Social Science, Social Network Analysis, Disaster Planning, Social Media Analytics, Few-Shot Learning, Transformer Language Model

Licence