The Academic Data Project That Turned Into $375 Million

What could be better than creating data so valuable that an institution is happy to host and update it forever, like the Sean Lahman baseball database?

Creating data that sells for $375 million, like the Center for Research in Security Prices. University of Chicago professors assembled this series of finance datasets over decades, starting in 1960 with an effort to track every transaction of every publicly traded security. U Chicago sold CRSP to Morningstar last year for $375 million.

Why could they sell it for so much? It helps to be working in finance, where the willingness to pay is the highest. It also represents 65 years of work from what became a large team that included Nobelists like Eugene Fama. The data was valuable enough to become widely used by key institutions even though CRSP charged for it:

Today, $3 trillion in fund assets are linked to CRSP Market Indexes, including U.S. equity ETFs run by Vanguard, and more than 600 subscribers across 35 countries use CRSP Research Data Products.  

Did U Chicago sell CRSP at the right time? On the one hand, I wonder if this was a fire sale driven by federal grant cuts putting pressure on the U Chicago budget. On the other hand, assembling datasets like this is only going to get easier in the age of AI, so perhaps Chicago sold at the top.

For now though there is still an edge in having restricted datasets that AIs haven’t trained on and can’t access. When I ask myself what advantage my human research assistants have over AIs in 2026, the most obvious answer is that they can legally access restricted databases like CRSP or, in my current case, HeinOnline.

Nimble Individuals, Enduring Institutions, and The Sean Lahman Baseball Database

The Lahman Baseball Database offers player- and team-level stats all the way back to 1871 as freely downloadable files. It includes over 20,000 players and has been cited by 192 academic papers. That sounds like something that takes an enormous amount of effort to put together, but it seems to have been compiled by just one guy, journalist Sean Lahman.

This looks like yet another example of a lone individual outperforming the huge, well-funded institutions you might expect to compile such datasets- this time not the government but MLB, ESPN, et c.

But as we saw last week, lone individuals can’t keep it up forever. If you want your creation to last, you will eventually need an institution. In this case, Lahman recently passed his database on to the Society for American Baseball Research:

Sean Lahman has graciously agreed to donate the Lahman Baseball Database, an open source collection of historical baseball statistics, to SABR.

The Lahman Baseball Database — which Lahman created in 1996 and has made freely available online every year since then — contains complete batting and pitching statistics back to 1871, plus fielding statistics, standings, team stats, managerial records, postseason data, and more. While Lahman and others had previously released smaller datasets online, his database allowed researchers to perform complex queries across the entire history of the game for the first time. The Lahman Baseball Database has served as the foundation for many popular baseball research projects and simulation games, including Out of the Park Baseball and Baseball Mogul.

SABR plans to continue to update the database and make it available for free online every year at SABR.org/lahman-database

I can only hope more of us will compile datasets worth handing off to an institution that will keep updating them.