Assessing the Developmental Impact of Private DFIs on SME Performance: Evidence from Matched Firm-Level Data in Kenya

Published: 15 December 2025| Version 1 | DOI: 10.17632/2wpc5skh2h.1
Contributor:
Richard Wanzala

Description

Research Hypothesis, Dataset Description, and Interpretation Research Hypothesis The dataset was constructed to test the hypothesis that: H1: SMEs receiving financing from Private Development Finance Institutions (PDFIs) experience higher employment growth, stronger revenue performance, and higher survival probability compared with similar firms that do not receive PDFI financing. This hypothesis is grounded in Financial Market Imperfection Theory, which argues that SMEs face credit rationing due to information asymmetries and limited collateral, and Development Finance Theory, which posits that mission-driven financiers can alleviate these constraints through catalytic capital and technical support. What the Data Show The dataset contains 657 firm-level observations, of which 268 are PDFI-financed firms and 389 are matched non-PDFI firms, reflecting realistic Kenyan SME characteristics. Key variables include: • Employment Growth (percentage change) • Revenue Growth (percentage change) • Survival Probability (0–1) • Matching weights (CEM_Weight and PSM_Weight) Preliminary analysis generated from this dataset indicates that: • PDFI-financed firms exhibit higher average employment growth (≈12.4%) vs. control firms (≈5.8%). • PDFI-backed firms show stronger revenue growth (≈18.7% vs. 9.5%). • PDFI-financed SMEs demonstrate a higher probability of survival (0.89 vs. 0.78). These trends support the hypothesis that catalytic financing improves SME performance. Interpretation and Usage The dataset is structured to enable replication of: • Quasi-experimental causal inference methods (Coarsened Exact Matching—CEM; Propensity Score Matching—PSM) • Weighted regressions for employment, revenue, and survival outcomes • Robustness checks across multiple matching estimators Users can interpret the variables as representing realistic SME performance metrics commonly found in African enterprise datasets. The matching weights allow researchers to directly apply causal inference models without reconstructing the matching process. The dataset is ideal for: • Teaching econometrics • Demonstrating causal inference methods • Simulating development finance impacts • Replicating tables in the associated manuscript (Tables 4–7) Because the dataset is hypothetical but empirically structured, users should not treat results as representative of actual Kenyan firms; instead, it should be used for methodological or training purposes.

Files

Steps to reproduce

Although the dataset is hypothetical, it was generated using a rigorous simulation framework designed to replicate structural patterns in real Kenyan SME datasets. The process involved the following steps: 3.1 Data Structure and Variable Distribution Design Variables were defined based on: • KNBS MSME Survey metadata • PDFI investment portfolio structures • Standard SME performance indicators in econometric studies • Typical ranges of employment, revenue, and firm age in African SME literature This ensured realism and empirical credibility. 3.2 Matching Framework Creation The dataset simulates quasi-experimental procedures using: 1. Coarsened Exact Matching (CEM) o Firms were grouped into bins for age, size, sector, revenue, and location. o Matched groups were assigned CEM_Weight values following Iacus, King & Porro (2012). 2. Propensity Score Matching (PSM) o A logistic model was simulated to estimate the likelihood of receiving PDFI financing. o The resulting PSM_Weight approximates inverse-probability weights. 3.3 Outcome Generation Outcome variables (employment, revenue, survival) were simulated using: • Linear models with treatment effects consistent with empirical evidence • Realistic noise terms to reflect firm-level heterogeneity • Sector-specific and location-specific variation 3.4 Software and Workflow Dataset creation and validation were performed using: • Stata 17 (matching weights, random draws, summary statistics) • Python (Pandas, NumPy) for simulation and exporting clean data • Microsoft Excel for final formatting Reproduction is straightforward: 1. Load the dataset (Excel, CSV, DTA, or SAV). 2. Apply CEM_Weight or PSM_Weight as analytical weights. 3. Run OLS or GLM regressions depending on the dependent variable. 4. Verify balance using standardized mean differences or L1 distance. 3.5 Reproducible Pipeline Summary 1. Define firm characteristics (size, age, sector, baseline revenue, location). 2. Simulate treatment assignment (PDFI financing) based on these covariates. 3. Generate matching weights (CEM, PSM). 4. Simulate outcomes using treatment effects + noise. 5. Export final tidy dataset. This protocol enables any researcher to regenerate the dataset using the logic and code provided.

Institutions

  • Stellenbosch University

Categories

Finance

Licence