U.S. State-level Renewable Energy Production (2020-2023) Econometric Geospatial Dataset
Description
This dataset contains raw and analyzed data. The analyzed dataset is derived from a cross-sectional panel of all 50 U.S. states, constructed by averaging annual values from 2020–2023 to smooth short-term volatility. The primary outcome /dependent variable is renewable energy production per capita (Renew_prod_pc, Btu per person), transformed using a log1p function (log_Renew_prod_pc) to address skewness and enable elasticity-based interpretation. Key predictors capture demographic, economic, political, policy, and energy-structure characteristics. State size and economic scale are controlled using logged population (log_Population) and logged gross state product per capita (log_Gdp_pc). Policy context is measured by the count of renewable energy incentives (Policy_Incentives_count) from DSIRE and partisan control of state government (Political_persuasion, binary). Regional and structural factors include Appalachian designation (Appa_st), coal production per capita, renewable energy consumption per capita, and the number of active coal and renewable power plants. Additional controls describe industrial composition (Gdp_oilgas_share, Gdp_mfg_share), labor market conditions (Unemp_rate), human capital (Pop_18_25_lt9_pct), and socioeconomic vulnerability (Poverty_rate). All variables are state-level and sourced from authoritative U.S. agencies (EIA, BEA, Census Bureau, BLS) or established policy databases. The analytical outputs include: (1) Bayesian regression parameter estimates (posterior means, SDs, 95% HDIs, R-hat, ESS, and credibility). (2) Model-level comparison statistics across alternative specifications (Full, Parsimonious, Minimal), including Bayesian R² and LOO-ELPD. (3) Spatial diagnostics, comprising bivariate Moran’s I statistics and Local Indicators of Spatial Association (LISA) cluster summaries and state-level classifications.
Files
Steps to reproduce
To reproduce the study, start by gathering state-level data from the listed federal agencies and policy databases for each year from 2020 through 2023. Once the raw data are assembled, average each variable across the four years so that every state is represented by a single, stable observation rather than year-to-year fluctuations. Renewable energy production per capita, population, and gross state product per capita are then log-transformed to reduce skewness and make the coefficients easier to interpret. All variables are merged into one dataset using the state as the unit of analysis. The analysis proceeds by estimating a series of Bayesian linear regression models with logged renewable energy production per capita as the outcome. The first model includes the full set of political, economic, demographic, energy, and policy controls. From there, simpler models are fit by removing less central predictors to check whether the main results hold in more parsimonious specifications. Weakly informative priors are used so the data drive the results while still stabilizing estimation. Models are estimated using Markov Chain Monte Carlo sampling with multiple chains, and diagnostics such as R-hat values, effective sample sizes, and divergent transitions are checked to ensure the chains have converged and are sampling efficiently. Substantive effects are evaluated using posterior means and 95 percent highest density intervals. A predictor is treated as credible when its interval does not cross zero. Overall model fit is assessed using Bayesian R² to gauge explained variance and leave-one-out expected log predictive density to compare out-of-sample predictive performance across models. After estimating the regression models, spatial patterns are examined to see whether renewable energy production clusters geographically. A U.S. state contiguity weights matrix is constructed, and bivariate Moran’s I statistics are calculated between renewable energy production and the spatially lagged values of key predictors using permutation-based significance tests. Finally, Local Indicators of Spatial Association are computed to identify where significant high-high, low-low, and outlier clusters occur, with results summarized both across all states and at the individual state level.
Institutions
- Boston UniversityMassachusetts, Boston