Data & Code
Data and code releases from my research. All repositories are available on my GitHub profile.
Composition of 7800 VC-backed startup boards from first VC financing to exit (or 2017). From Ewens and Malenko (2026), 'Board Dynamics over the Startup Life Cycle', Journal of Finance.
Acquired intangible valuations from Ewens, Peters, and Wang (2024). When a public firm acquires a significant target, it must disclose assets and liabilities in the purchase price allocation.
Parameter estimates for intangible capital accumulation and estimated knowledge and organization capital stocks. From Ewens, Peters and Wang (2024), 'Measuring Intangible Capital with Market Prices', Management Science.
State-level law changes in the U.S. from 1995-2016 to study the impact of founder replacement on startup outcomes. From Ewens and Marx (2017), 'Founder Replacement and Startup Performance', RFS.
A mapping file between SDC's 'sdc_dealno' to 'gvkey'. Connects SDC's M&A database to Compustat using name and date matching with fuzzy string search.
Data on U.S. bank 'venture capital revenue' used in 'Venture Capital and Startup Agglomeration' (Chen and Ewens 2025) to assess the importance of banking institutions as limited partners in VC.
Public float data from firms' 10-K filings disclosing the market value of outstanding common equity held by non-affiliates. From Ewens, Xiao and Xu (2023), 'Regulatory Costs of Being Public', JFE.
Index file and code to process raw MD&A data from 10-K filings. Build your own panel database of public firm MD&A text. From Ewens, Peters and Wang (2024).
Form D filings data from a FOIA request with the SEC. Used in 'The Deregulation of the Private Equity Markets and the Decline in IPOs' by Ewens and Farre-Mensa (2020), RFS.
The complete dataset connecting new firm formation and patents, allowing researchers to explore startup innovation. Includes mapping and novelty measures.
Replication Policy
I am committed to research transparency. Replication packages for many of my published papers are available on the journal websites, and when possible, the main datasets in my GitHub repositories. If you encounter any issues running the code, please open an issue on Github or email me.