Encrypted Mobile Social Media Traffic Fingerprinting Dataset

Published: 4 August 2026| Version 2 | DOI: 10.17632/6v3zcxysh7.2
Contributors:
Bright Jiwueze,
,

Description

This dataset contains encrypted mobile network traffic collected from ten widely used Android social media applications: Facebook, Instagram, LinkedIn, Reddit, Snapchat, Telegram, TikTok, Twitter, WhatsApp, and YouTube. Traffic was captured over five independent collection days under controlled experimental conditions and processed using NFStream to generate bidirectional flow-level records. The dataset comprises 25,116 labeled network flows extracted from 50 packet capture (PCAP/PCAPNG) files, with each application represented by five independent capture sessions. The published data contain the complete set of NFStream-extracted flow features, including statistical flow characteristics, packet-size statistics, timing information, transport-layer attributes, and encrypted-session metadata. This repository provides the raw NFStream feature dataset used in our study. Feature selection, leakage-control preprocessing, temporal train/test partitioning, and multi-flow aggregation were performed during the experimental pipeline and are described in the accompanying manuscript. The dataset was developed to support reproducible research in encrypted traffic analysis, mobile application fingerprinting, digital forensics, network security, and explainable artificial intelligence. It accompanies the manuscript "Mobile Social Media Traffic Fingerprinting Using Encrypted Network Flows: A Benchmark Dataset and Comparative Study."

Files

Steps to reproduce

The dataset was collected using ten Android social media applications: Facebook, Instagram, LinkedIn, Reddit, Snapchat, Telegram, TikTok, Twitter, WhatsApp, and YouTube. Each application was exercised under a controlled and consistent interaction protocol across five independent capture days to reduce variability in the data collection procedure. Network traffic generated by each application was captured in PCAP/PCAPNG format using Wireshark. The resulting packet captures were processed using an automated Python pipeline based on NFStream, which converted raw packets into bidirectional network-flow records. The pipeline automatically identified the application and capture day from the filenames, extracted all available NFStream flow features, and exported the results as CSV files. The application-specific CSV files from the five capture days were then merged to produce the final benchmark dataset containing 25,116 labeled encrypted network flows. The published dataset contains the complete set of NFStream-extracted flow features. The feature selection, leakage-control preprocessing, temporal train/test partitioning (Days 1–4 for training and validation, Day 5 for testing), multi-flow aggregation, and explainable AI analyses (SHAP and LIME) used in the accompanying study were performed during the experimental pipeline and are described in detail in the associated manuscript. Software and tools: Wireshark and NFStream.

Institutions

Categories

Social Media, Machine Learning, Mobile Device, Forensic Analysis

Licence