Startup GitHub Engineering Velocity Panel, Q2 2025 to Q2 2026

Published: 19 August 2026| Version 1 | DOI: 10.17632/crcnwdn2py.1
Contributor:
Data Nerd

Description

This dataset is a quarterly panel of public GitHub engineering-velocity signals for 55 venture-backed startups, observed across five quarters from Q2 2025 through Q2 2026. It contains 219 startup-period observations. The data is intended for research on public software-development activity, venture-capital deal sourcing, and reproducible exploratory analysis. The primary table records one observation per startup and quarter. It includes commit velocity, contributor activity, new repositories, signal type, stage, geography, and a public GitHub URL. Two supporting tables provide quarterly sector-level aggregates and the time-series distribution of the observed signal types. All source activity was collected from public GitHub REST API endpoints. The dataset does not contain private repository data, personal data, or funding outcomes. It records public engineering activity and is suitable for studying associations between software-development patterns and startup activity. The dataset is released under CC BY 4.0. Please cite the linked SSRN preprint and the dataset DOI when using or adapting this work.

Files

Steps to reproduce

1. Identify the 55 venture-backed startups and their public GitHub organizations or repositories. 2. Collect public GitHub REST API metadata for each startup at quarterly observation points from Q2 2025 through Q2 2026. 3. Compute 14-day commit velocity, contributor counts, contributor growth, and new repository counts. 4. Classify observed changes into framework migration, engineering hiring burst, infrastructure buildout, or deploy frequency spike. 5. Export the primary observation table and the two aggregate CSV tables. Limitations: public GitHub activity is an imperfect proxy for company-wide engineering work. Startups may use private repositories, have incomplete public ownership signals, or change their GitHub organization structure. The data should be used for association and exploratory analysis, not as proof of causation or a prediction of funding events.

Categories

Computer Science

Licence