PhishVN: A Time-Stamped Vietnamese URL Phishing Dataset with Impersonation-Scenario Labels and Confidence Tiers

Published: 18 August 2026| Version 4 | DOI: 10.17632/b97hxbxtpd.4
Contributor:

Description

PhishVN is an open, time-stamped phishing-website dataset localised to the Vietnamese context: a verified core of 18,997 URL records (2,587 phishing from the national NCSC "Tin Nhiem Mang" feed, gold/silver tiers; 16,410 legitimate) plus an explicitly-tagged bronze expansion of 34,119 community/feed-sourced phishing (ChongLuaDao, OpenPhish), for 53,116 records in total. Each record carries a 21-feature lexical schema aligned with the CompPhish dataset for cross-dataset study, an impersonation-scenario label (bank, government, tax, e-commerce, telecom, delivery, social, gaming) inferred from brand tokens in the URL, and a label-confidence tier. Legitimate URLs combine the certified trusted-organisation registry (easy negatives), a .vn-filtered Tranco slice, and a global Tranco sample (hard negatives), avoiding the "trusted-vs-malicious" and "non-.vn = phishing" shortcuts. Records split train/validation/test by a group-aware temporal rule keyed on the registrable domain (dated groups oldest-to-train, newest-to-test; undated by deterministic per-group hash), so campaign subdomains never span splits. Changes since v2 (DOI .2, 51,362 records): three ingestion fixes in the merge pipeline; a site-level label-conflict review excludes five ambiguous .vn sites; 21 reviewed trusted-org misattributions removed; the trusted-organisation registry re-scraped with enrichment (benign rows 20 to 2,026; audited 2026-08-04, no label defect found). Full changelog in docs/datasheet.md. Files: dataset_url.csv (full record table); vn_compphish.csv (CompPhish-schema features); splits/url_{train,val,test}.csv; abuse_type.csv (rule-derived abuse type per phishing URL); attacks/ (a fixed LLM-paraphrase evaluation set); docs/ (datasheet, column schema, source notes); a MANIFEST with SHA-256 checksums. A second archive, PhishVN_v3.0.0_results_bundle.zip, deposits the reference benchmark's raw evidence (per-seed metrics, per-row prediction files, fitted seed-0 estimators, pinned environment), so the article's reported intervals are verifiable without retraining. Notes: is_https is a collection artefact - exclude it from modelling (features are computed on the scheme-stripped URL). Personal data is redacted; no id-to-PII mapping. A gated tier (captured page HTML and screenshots) is available on request under a research-only data-use agreement. Described in a Data in Brief article (DIB-D-26-01800, under revision). Credit the upstream sources: NCSC Tin Nhiem Mang; ChongLuaDao; Tranco. Version 4 (v3.1.0) is a documentation-only update: every data file is byte-identical to version 3 (verifiable via MANIFEST.txt), and docs/datasheet.md now records the completed inter-annotator label audit that version 3 registered as in progress (four-way Cohen's kappa 0.609, 0.725 on the abuse-vs-legitimate distinction; positive-arm label noise 12.1%, 95% CI 7.1-20.0%, concentrated in the bronze stratum; benign arm 4.0%). Audit instruments and arbitration record: code repository.

Files

Categories

Cybersecurity, Machine Learning, Vietnam

Licence