PhishBD_2026

Published: 22 September 2026| Version 3 | DOI: 10.17632/8d6zsfwc7z.3
Contributors:
,
,

Description

This dataset provides the raw, pre-processing pool of phishing indicators that formed the starting material for the phishing class of PhishBD_2026 (available at [10.17632/8d6zsfwc7z.2]). It contains 115,827 records collected from four external threat intelligence feeds and a Bangladesh-based financial institution, prior to normalisation, deduplication, or feature engineering. Each record includes: raw_url: the indicator as originally received, in defanged form (e.g., hxxps:// and [.] in place of active protocol and dot characters, standard practice for sharing threat data without creating live, clickable links). Label: the class assigned by the source feed. Of the 115,827 records, 115,822 are labelled Phishing and 5 Legitimate; the latter are a small number of benign entries present in the original feed exports and were excluded during downstream cleaning. This raw pool is substantially larger than the 91,817 phishing records in the final published PhishBD_2026.csv; the difference reflects deduplication, URL normalisation, and quality filtering steps applied during dataset construction, described in the accompanying data article. For confidentiality reasons, source-level attribution (which specific feed or institution contributed each record) has been withheld, and record order has been randomised. This preserves the aggregate composition of the raw collection for reproducibility purposes without identifying the confidential threat-intelligence providers or the partner financial institution, consistent with the data-sharing agreements under which this material was obtained. Related identifiers: PhishBD_2026 (final, fully-featured dataset): [10.17632/8d6zsfwc7z.2]

Files

Steps to reproduce

Environment: Python 3.10, run in Google Colab with GPU acceleration enabled for the deep learning cells. Main libraries: pandas 2.2.3, tldextract 5.3.2, NumPy 2.1.3, scikit-learn 1.6.1, tqdm 4.67.3, openpyxl 3.1.5, imbalanced-learn 0.14.2, XGBoost 3.4.1, LightGBM 4.6.0. Phishing URL collection: Pulled from four threat intelligence feeds (CSV format, filtered to URL-type indicators only) plus an incident spreadsheet from a partner financial institution. Feed providers are not named, per a confidentiality agreement. URLs from the feeds arrive defanged, so hxxp was converted back to http, bracketed dots and colons were restored, and a scheme was added to any URL missing one. Each URL was parsed with urllib.parse and tldextract to extract domain, subdomain, path, query, and TLD. Legitimate baseline: Started with the top 10,000 Tranco domains, prefixed with http://, and run through the same parsing and feature pipeline as the phishing URLs. Feature engineering: Built 66 features across six passes: lexical and structural (34), entropy and randomness (6), token-level (5), Bangladesh-specific indicators (7), typosquatting and homograph detection (5), and obfuscation/redirection signals (9). Class balancing: The raw collection skewed heavily phishing. Legitimate URLs were added to reach a 70:30 ratio, using n_inject = floor(n_phish × 30/70) − n_legit,existing. Additions came from the full Tranco list, ISCX, Majestic Million filtered to .bd domains, and a hand-compiled list of 260 Bangladeshi sites across government, banking, MFS, telecom, education, media, e-commerce, healthcare, and city corporations, deduplicated against the existing set before merging. Redundancy check: Three features were dropped after auditing the original 66: one constant column, one exact duplicate, one at r ≈ 0.98 with an existing feature, leaving 63. Ten more were added afterward, bringing the total to 73, and the full set was re-audited at a 0.95 correlation threshold, which caught nine correlated pairs, including a perfect duplicate (has_hex_encoding and new_has_encoded_chars, r = 1.00) that had been missed in the first audit. A further ten pairs sit in the 0.90 to 0.95 band, for example is_bd_tld and new_tld_is_bd at r ≈ 0.93; none of these were removed either, since each pair is computed from a distinct definition and may diverge on specific subpopulations of URLs. Final checks: Deduplicated on the URL column, confirmed no missing values, checked outliers past 100x the 99th percentile (none found), dropped two internal-only tracking columns. Final file: 131,167 rows, 75 columns, no missing values, no duplicate URLs.

Institutions

Categories

Cybersecurity, Cyber Attack

Licence