Hybrid Bankruptcy Dataset for Indian Firms (2015–2019)
Description
This dataset accompanies the manuscript Forecasting Corporate Bankruptcy in India: A Hybrid Model Integrating Financial Ratios, Macroeconomic Indicators, and Machine Learning, submitted to Investment Management and Financial Innovations. It contains 100 firm-year observations covering 20 Indian firms (10 bankrupt and 10 healthy peers) over the period 2015–2019. The dataset integrates: Firm-level financial ratios: Debt-to-Equity Ratio, Current Ratio, Return on Assets (ROA), Net Profit Margin. Macroeconomic indicators: GDP Growth (%), Inflation (%), Interest Rate (%). Firm status: Bankrupt or Healthy, based on bankruptcy filings or sector-matched controls. The dataset is designed to enable replication of the hybrid bankruptcy prediction framework developed in the study, which combines logistic regression and random forest machine learning techniques to forecast bankruptcy risk. All values are structured and normalized for comparative analysis across firms and years. While the dataset is a curated/simulated reconstruction aligned with reported cases of corporate failure in India (2012–2021), it reflects the patterns, ratios, and macroeconomic conditions described in the manuscript.
Files
Steps to reproduce
The dataset was generated to replicate the hybrid bankruptcy prediction framework described in the manuscript Forecasting Corporate Bankruptcy in India: A Hybrid Model Integrating Financial Ratios, Macroeconomic Indicators, and Machine Learning. The following steps outline how the data can be reconstructed and used: Firm Selection Identify 10 large Indian firms that filed for bankruptcy between 2012 and 2021 (e.g., DHFL, Reliance Communications, Jet Airways, Kingfisher Airlines, Essar Steel, Bhushan Steel, Amtek Auto, Videocon, Jaypee Infratech, Lanco Infratech). For each, select one matched healthy peer from the same sector (e.g., HDFC Ltd. for DHFL, Bharti Airtel for RCom) based on similar revenue scale and market exposure. Final sample: 20 firms (10 bankrupt, 10 healthy). Time Frame Collect data for a five-year period prior to bankruptcy (2015–2019). For healthy peers, the same five years are used to ensure comparability. This produces 100 firm-year observations (20 firms × 5 years). Variable Construction Financial Ratios: Debt-to-Equity, Current Ratio, Return on Assets (ROA), Net Profit Margin. Macroeconomic Indicators: Annual GDP growth rate, inflation (CPI), central bank interest rate (sourced from IMF, RBI, CEIC). Firm Status: Binary label (“Bankrupt” or “Healthy”). Data Sources Financial statements: Audited annual reports, Bloomberg, CMIE Prowess. Macroeconomic data: IMF World Economic Outlook, Reserve Bank of India, CEIC Data. Sectoral/regulatory events: Archival reports, news databases, policy papers. Processing Ratios and macro indicators aligned into a panel structure (Firm × Year). Outliers smoothed by min–max normalization (0–1 scaling for scoring model). Missing values handled through cross-verification with multiple data sources. Use in Analysis Logistic regression and random forest models can be trained on this dataset to replicate reported classification performance (accuracy ~86%, AUC = 0.78). Bankruptcy risk scores are calculated using the weighted formula described in the paper: Risk Score = 0.35 × (1−Debt-to-Equity) + 0.30 × (1−Current Ratio) + 0.20 × (1−GDP Growth) + 0.15 × (1−Net Profit Margin). Replication Users can reproduce results by: Importing the dataset into R, Python, or Stata. Running supervised classification (logit, random forest). Comparing outputs to reported firm-by-firm case analyses and sector summaries. The dataset is structured in a simple CSV format (rows = firm-year observations, columns = variables). Researchers can extend this by adding qualitative indicators (auditor resignations, governance events) for enriched analysis.
Institutions
- Universita Ca' Foscari