Used Car Price Prediction Dataset

Published: 30 June 2026| Version 2 | DOI: 10.17632/8d38h82fyt.2
Contributor:
Md Saikot Hossain

Description

Dataset Description Title: Used Car Price Prediction Dataset Description This dataset has been developed to support research and development in the areas of machine learning, predictive analytics, and automotive market intelligence. The dataset contains information on used vehicles collected from publicly available online sources and has been carefully preprocessed to facilitate depreciation prediction and resale value analysis. The dataset includes 5997 vehicle-specific attributes such as title, brand, model, transmission type, fuel type, manufacturing year, energy capacity, kilometers driven, vehicle age, and second-hand selling price. These features capture both the technical characteristics and market-related factors that influence vehicle depreciation over time. Data preprocessing and feature engineering were primarily conducted using the R programming language. This process included data cleaning, missing value handling, duplicate removal, feature standardization, and derivation of additional variables such as vehicle age. The processed data were subsequently validated and prepared using Python to ensure consistency, data quality, and suitability for machine learning applications. Researchers may utilize this dataset for a variety of tasks, including vehicle depreciation prediction, resale price estimation, regression modeling, feature importance analysis, automotive market trend analysis, and benchmarking of machine learning algorithms. The dataset is suitable for educational, academic, and industrial research purposes. Dataset Features * title * brand * model * transmission * fuel_type * price * year_of_manufacture * energy_capacity * kilometers_run * car_age Potential Applications * Used Car Depreciation Prediction * Resale Value Estimation * Machine Learning Regression Tasks * Automotive Market Analytics * Predictive Modeling Research * Data Science Education and Benchmarking

Files

Steps to reproduce

#Steps to Reproduce 1. Raw data were collected from the designated data sources and imported into the R programming environment. 2. Data cleaning and preprocessing were performed in R, including handling missing values, removing duplicates, standardizing feature formats, and validating data consistency. 3. Feature engineering was conducted to derive additional attributes such as vehicle age and other relevant variables required for analysis. 4. The processed dataset was exported from R into CSV format. 5. Further data preparation and validation were performed using Python. This included data type verification, exploratory data analysis, feature inspection, and quality assurance checks. 6. The final dataset was generated and saved as a CSV file for use in machine learning, statistical analysis, and predictive modeling tasks. 7. Researchers can reproduce the dataset preparation workflow by executing the provided R scripts for data preprocessing and feature engineering, followed by the Python scripts for validation and final dataset generation.

Institutions

Categories

Machine Learning

Licence