E-Scooter Accidents in UK, 2020-2024, 25 features
Description
The file contains anonymised, record‑level information on 5,166 personal‑injury collisions involving e‑scooters that were reported by police forces in Great Britain between January 2020 and November 2024 and released by the UK Department for Transport via the STATS19 open data service. Each row represents one crash and its associated casualties, vehicles and site conditions. The dataset has been pre‑processed to keep 25 explanatory variables and a single binary response (Class).
Files
Steps to reproduce
The UK Department for Transport publishes on an annual basis data on all traffic accidents that occurred on public roads in the UK; accidents involving e-scooters have been included into the data from 2020 (Department for Transport, 2024). The dataset is organised into three interlinked main tables, vehicles involved, accident conditions, and casualties. The records offer a comprehensive and granular view of each accident. These attributes relate to details about the vehicle (vehicle type, brand, make, engine capacity, age), its manoeuvre (direction of travel, point of impact), rider and casualty demographics (age, gender, Index of Multiple Deprivation (IMD) decile), physical environment (light and weather conditions, road surface, presence of carriageway hazards), traffic control measures (speed limits, type of roads and junctions, dual vs single carriage way), and consequences of the accident (accident severity, number of people and vehicles involved, injuries sustained). Selecting accidents that involved e-scooters, we obtained a dataset with 5,166 accidents that took place between January 2020 and November 2024. An initial examination of the data showed that it contains missing values, mis-typed values, inconsistent encoding of categorical variables and others. To rectify these issues, data transformation steps were required prior to model development. Outliers, such as improbable ages of riders, were detected by visual inspection of variable distributions and replaced with missing value symbols. Observations, where more than 5% of the values were missing, were deleted. Remaining missing values of categorical and discrete variables were replaced with the mode and missing values for continuous variables were replaced with the mean of the corresponding variables. From the date and time column, three new categorical features were created encoding the hour, day of the week, and the month of the year. The categorical variables were converted to dummy features.
Institutions
- Aston University