ASA_Phishing_URL_Dataset
Description
A balanced, synthetically generated dataset of 5,000 phishing and legitimate URLs (2,500 each) for training and benchmarking phishing URL detection models. URLs were generated using an LLM to reflect real-world lexical and structural patterns, including typosquatted domains, IP-address hosts, suspicious TLDs, and common legitimate URL conventions. The dataset was deduplicated and validated prior to publication. Full details are provided in the accompanying README.md. It is suitable for: 1. Training or evaluating machine learning / deep learning classifiers (e.g., tree-based models, neural networks, transformers) 2. Lexical, structural, and character-level URL feature extraction research 3. Benchmarking generalization of models trained on other phishing datasets 4. Educational and academic use in cybersecurity and NLP-for-security research
Files
Institutions
- Daffodil International UniversityDhaka Division, Dhaka