An Obfuscated and LLM-Augmented Multi-Source Cross-Site Scripting (XSS) Payload Dataset for Web Application Security
Description
Description : This dataset is a curated benchmark of 27,314 labeled text entries designed to train and evaluate machine learning and deep learning detectors against Cross-Site Scripting (XSS) attacks under real-world obfuscation and adversarial conditions. To eliminate single-source bias, raw payload samples were aggregated from three independent repositories: PayloadBox XSS Payload List, Deep-XSS (DAS-Lab), and the Kaggle XSS Dataset. The aggregated samples were cleaned, deduplicated, and subjected to four targeted obfuscation techniques—Base64 encoding, URI encoding, HTML comment keyword splitting, and JavaScript transformations (hexadecimal encoding, variable renaming, and inline comments). To further enhance attack diversity, a CodeT5-base generative model (~220M parameters) was fine-tuned on prompt-target payload pairs to synthesize novel, syntactically valid adversarial XSS payloads and balanced benign samples. Dataset Summary: Total Records: 27,314 samples Malicious (Label 1): 22,066 standard, obfuscated, and LLM-generated XSS payloads Benign (Label 0): 5,248 clean web queries and benign scripts File Format: CSV / Tabular format Column Descriptions: Sentence: The input text string containing standard script tags, obfuscated/encoded XSS vectors, LLM-generated payloads, or benign parameter values. Label: Binary classification target (1 for malicious XSS payload, 0 for benign). Intended Use: This corpus is ideal for benchmarking deep learning architectures (such as character-level CNNs, LSTMs, and BiLSTMs), evaluating WAF filter resilience, testing NLP/LLM security models, and conducting adversarial machine learning experiments in web security.
Files
Institutions
- Rajshahi University of Engineering and TechnologyRajshahi Division, Rajshahi