I was writing up something for my graduating seniors about how to keep learning economics after school, and realized I might as well share it with everyone. This may not be the best way to do things, it is simply what I do, and I think it works reasonably well.
Blogs by Economists: There are many good ones, but besides ours Marginal Revolution is the only one where I aim to read every post
Podcasts on the Economy: NPR’s The Indicator (short, makes abstract concepts concrete), Bloomberg’s Odd Lots (deeper dives on subjects that move financial markets)
Podcasts by Economists: Conversations with Tyler and Econtalk (note that both often cover topics well outside of economics). Macro Musings goes the other way and stays super focused on monetary policy.
Twitter/X: This is a double-edged sword, or perhaps even a ring of power that grants the wearer great abilities even as it corrupts them. The fastest way to get informed or misinformed and angry, depending on who you follow and how you process information. Following the people I do gives you a fighting chance, but even this no guarantee; even assuming you totally trust my judgement, sometimes I follow people because they are a great source on one issue, even though I think they are wrong on lots of other things. Still, by revealed preference, I spend more time reading here than other single source.
Finance/Investing: Making this its own category because it isn’t exactly economics. Matt Levine has a column that somehow makes finance consistently interesting and often funny; unlike the rest of Bloomberg, you can subscribe for free. He also now has a podcast. If you’d like to run money yourself some day, try Meb Faber’s podcast. If you’d like things that touch on finance and economics but with more of a grounding in real-world business, try the Invest Like the Best podcast or The Diff newsletter.
Economics Papers: You can get a weekly e-mail of the new papers in each field you like from NBER. But most econ papers these days are tough to read even for someone with an undergrad econ degree (often even for PhDs). The big exception is the Journal of Economic Perspectives, which puts in a big effort to make its papers actually readable.
Books: This would have to be its own post, as there are too many specific ones to recommend, and I don’t know that I have any general principle of how to choose.
This is a lot and it would be crazy to just read all the same things I do, but I hope you will look into the things you haven’t heard of, and perhaps find one or two you think are worth sticking with. Also happy to hear your suggestions of what I’m missing.
This was surprising to me, as I kind of expected CON laws to harm workers. Certificate of Need laws require many types of health care providers to obtain the permission of a state board before they are allowed to open or expand. This could lead to fewer health care facilities, and so less demand for health care workers, lowering wages and employment. It could also lead to less competition among health care employers, to similar effect.
On the other hand, less competition in the market for health services could raise profits, with room to share them in the form of higher wages. Or, CON being primarily targeted at capital expenditures like facilities and equipment could increase the demand for labor (to the extent that labor and capital are substitutes in health care). All these competing theories seem to cancel out to one big null when we look at the data.
We use 1979-2019 data from the Current Population Survey and a generalized triple-difference approach comparing CON-repealing to CON-maintaining states, and find a bunch of fairly precise zeroes. This holds for many different definitions of “health care worker”: those who work in the health industry, in health occupations, in hospitals, in health care outside hospitals, nurses, physicians, and more.
This is the first word on the topic, not the last; I wouldn’t be too surprised if someone down the road finds that CON does significantly affect health care workers. In this paper we pushed hard on the definition of “health care workers”, but not on “Certificate of Need” or “wages”. We simply classify states as “CON” or “non-CON” because that is what we have data for, but some states have much stricter programs than others, and some day someone will compile the data on this back to the 1970’s. The easier thread to pull on is “wages”. We use one good measure (the natural log of inflation-adjusted hourly real wages), but don’t do any robustness checks around it; considering “business income” could be especially important here. It is also possible that CON affects workers in other ways; we only checked wages and employment.
The full paper is here (ungated here) if you want to read more.
I got to be a guest of Vignesh Swaminathan who is based in Mumbai. It’s fun to have a deep conversation with someone on the other side of the world and share it with the whole internet (and the AI’s).
The first 10 minutes are about Tyler’s GOAT book. Vignesh asked me to name some influential economists who did not make Tyler’s list.
Around minute 12 we talk about the experimental economics methodology.
The middle (minute 15-42) is a discussion of the pipeline into tech and my Willingness to be Paid paper. He adds his perspective on tech jobs in India.
Around minute 42, Vignesh makes a switch over to the Barbie movie and then Oppenheimer. He observes that Oppenheimer is a “brand.” I speculate on careers in Barbieland. We recorded this before Christmas of ’23, right after everyone had seen these summer movies. Both movies ended up in the 2024 Oscars awards ceremony.
I predicted that people will eventually be able to create a custom movie from a verbal prompt, because of the AI content revolution. Here in Spring of ’24 that has already come true. Sora is shocking everyone and even caused Tyler Perry to halt a physical film studio expansion.
Around minute 55, we pivot to Hayek and competition, which leads to a postmortem on Google Plus (RIP).
Skimming back through this conversation has me thinking about tech work. The market for IT workers and programmers has evolved since I first started the project that became “Willingness to be Paid: Who Trains for Tech Jobs?”
I like pointing people all the way back to this report on jobs from 1958. Learn to Code has been good advice for a long time, for the people who can tolerate the work. That does not mean it will be true forever, but I would argue that it is still true today.
Silicon Valley as a career might have peaked around 2021. It’s not going away, but it might not be growing anymore in terms of the number of talented people who can be absorbed there. (Might I suggest Huntsville instead?)
The rise of artificial intelligence is affecting job seekers in tech who, accustomed to high paychecks and robust demand for their skills, are facing a new reality: Learn AI and don’t expect the same pay packages you were getting a few years ago.
Jobs in areas like telecommunications, corporate systems management and entry-level IT have declined in recent months, while roles in cybersecurity, AI and data science continue to rise, according to Janco’s data. The average total compensation for IT workers is about $100,000, making the position a target for continued cost-cutting.
One reason tech jobs are less attractive than some other professional paths is that the skillset changes. We mentioned this as a drawback in our policy paper. Computers are constantly changing. Vignesh and I discuss the issue of risk. I suggested that companies could pay less for talent if they were willing to offer packages that carry less risk of getting fired.
Nevertheless, tech still has decent job prospects. An unemployment rate of about 5% is about normal for work, even though tech had seen lower rates at the peak of demand. I do not know what programming as a career will look like in 10 years, but I’d say the same about screenwriting and live sports commentary. The LLMs are coming for everything or nothing or something in between.
I’ve been on tour (regionally) with our ChatGPT paper and getting opportunities to query different audiences about their LLM use. Last week I talked to a young man in our business school who is using ChatGPT to write SQL code at his job. I said in the podcast that I would still advise young people in Alabama to learn to code, even if they are not going to move to Silicon Valley. I think coding is more fun in the LLM-age or at least less miserable.
To be on Cowen’s short list is a compliment. Of all the thinkers and writers in recorded history, Adam Smith is one of only six writers that Cowen gives serious consideration to. Next, readers will ask, “Did our guy win?”
Tyler’s book will make no one happy because he does not take anyone’s side unequivocally. A huge fan of Adam Smith (and I know several) might have wanted a book about why Adam Smith is designated as the GOAT. I don’t want to ruin the book for anyone who hasn’t read it. What you will get is very interesting and thoughtful, so I hope you’ll read the manuscript* sometime, even if your guy doesn’t win.
*completely free – can get it on your Kindle somehow I heard
What We Are Learning about Paper Books – I did write the AdamSmithWorks post in collaboration with the GPT version of the book, as a first step, along with my own memory of having read the book. And then, secondly, I consulted the book manuscript. The GPT performed fairly well… considering that it’s a GPT. I suppose I thought that interrogating the GPT would save me time. However, I can now say authoritatively that Tyler’s actual writing is so much better than what you will get from the GPT. Among other things, the GPT is much more boring than Tyler’s actual manuscript.
Daniel Kahneman, the psychologist who won a Nobel prize in economics and wrote the best-selling book “Thinking Fast and Slow“, died yesterday at age 90. Others will summarize his biography and the substance of his work, but I wanted to highlight two aspects of his style that I think fueled his unusual success among both the public and economists.
Daniel Kahneman’s new book amazes me. Not so much due to the content, though I’m sure that will blow your mind if you haven’t previously heard about it through studying behavioral economics or psychology or reading Less Wrong. It is the writing style: Kahneman is able to convey his message succinctly while making it seem intuitive and fascinating. Some academics can write tolerably well, but Kahneman seems to be on a level with those who write popularly for a living- the style of a Jonah Lehrer or Malcolm Gladwell, but no one can accuse the Nobel-prize-winning Kahneman of lacking substance.
This made me wonder if it is simply an unfair coincidence that Kahneman is great at both writing and research, or causation is at work here. True, in more abstract and mathematical fields great researchers do not seem especially likely to be great writers (Feynman aside). But to design and carry out great psychology experiments may require understanding the subject intuitively and through introspection. This kind of understanding- an intuitive understanding of everyday decision-making- may be naturally easier to share than other kinds of scientific knowledge, which use processes (say, math) or examine territories (say, subatomic particles) which are unfamiliar to most people. Kahneman says that he developed the ideas for most of his papers by talking with Amos Tversky on long walks. I suspect that this strategy leads to both good idea generation and a good, conversational writing style.
But how did a psychologist get economists to not just take his work seriously, but award him the top prize in our field? One key step was learning to speak the language of our field, or coauthor with people who do. For instance, summarizing the results of an experiment as showing indifference curves crossing where rationally they should not:
Finally, something that helped Kahneman appeal to all parties was that he avoided the potential trap of being the arrogant behavioral economist. Most economists have a natural tendency toward arrogance, kept somewhat in check by our belief that most people are fundamentally rational. Behavioral economists who think most people are irrational can be the most arrogant if they think they are the only sane one, and should therefore tell everyone else how to behave. But Kahneman avoided this by seeming to honestly believe he is just as subject to behavioral biases as everyone else.
I’ve always told my health economics students that Medicaid is both better and worse than all other insurance in the US for its enrollees.
Better, because its cost sharing is dramatically lower than typical private or Medicare plans. For instance, the maximum deductible for a Medicaid plan is $2.65. Not $2650 like you might see in a typical private plan, but two dollars and sixty five cents; and that is the maximum, many states simply set the deductible and copays to zero. Medicaid premiums are also typically set to zero. Medicaid is primarily taxpayer-financed insurance for those with low incomes, so it makes sense that it doesn’t charge its enrollees much.
But Medicaid is the worst insurance for finding care, because many providers don’t accept it. One recent survey of physicians found that 74% accept Medicaid, compared to 88% accepting Medicare and 96% accepting private insurance. I always thought these low acceptance rates were due to the low prices that Medicaid pays to providers. These low reimbursement rates are indeed part of the problem, but a new paper in the Quarterly Journal of Economics, “A Denial a Day Keeps the Doctor Away”, shows that Medicaid is also just hard to work with:
24% of Medicaid claims have payment denied for at least one service on doctors’ initial claim submission. Denials are much less frequent for Medicare (6.7%) and commercial insurance (4.1%)
Identifying off of physician movers and practices that span state boundaries, we find that physicians respond to billing problems by refusing to accept Medicaid patients in states with more severe billing hurdles. These hurdles are quantitatively just as important as payment rates for explaining variation in physicians’ willingness to treat Medicaid patients.
Of course, Medicaid is probably doing this for a reason- trying to save money (they are also trying to prevent fraud, but I have no reason to expect fraud attempts are any more common in Medicaid than other insurance, so I don’t think this can explain the 4-6x higher denial rate). This certainly wouldn’t be the only case where states tried to save money on Medicaid by introducing crazy rules hassling providers. You can of course argue that the state should simply spend more to benefit patients and providers, or spend less to benefit taxpayers. But the honest way to spend less is to officially cut provider payment rates or patient eligibility, rather than refusing to pay providers as advertised. In addition to being less honest, these administrative hassles also appear to be less efficient as a way to save money, probably because they cost providers time and annoyance as well as money:
We find that decreasing prices by 10%, while simultaneously reducing the denial probability by 20%, could hold Medicaid acceptance constant while saving an average of 10 per visit.
Medicaid is a joint state-federal program with enormous differences across states, and administrative hassle is no exception. For administrative hassle of providers, the worst states include Texas, Illinois, Pennsylvania, Georgia, North Dakota, and Wyoming:
Source: Figure 5 of A Denial a Day Keeps the Doctor Away, which notes: “The left column shows the mean estimated costs of incomplete payments (CIP) by state and payer. The right column shows the mean CIP as a share of visit value by state and payer. “
My paper “Missouri’s Medicaid Contraction and Consumer Financial Outcomes” is now out at the American Journal of Health Economics. It is coauthored by Nate Blascak and Slava Mikhed, researchers at the Federal Reserve Bank of Philadelphia. They noticed that Missouri had done a cut in 2005 that removed about 100,000 people from Medicaid and reduced covered services for the remaining enrollees. Economists have mostly studied Medicaid expansions, which have been more common than cuts; those studying Medicaid cuts have focused on Tennessee’s 2005 dis-enrollments, so we were interested to see if things went differently in Missouri.
In short, we find that after Medicaid is cut, people do more out-of-pocket spending on health care, leading to increases in both credit card borrowing and debt in third-party collections. Our back-of-the-envelope calculations suggest that debt in collections increased by $494 per Medicaid-eligible Missourian, which is actually smaller than has been estimated for the Tennessee cut, and smaller than most estimates of the debt reduction following Medicaid expansions.
We bring some great data to bear on this; I used the restricted version of the Medical Expenditure Panel Survey to estimate what happened to health spending in Missouri compared to neighboring states, and my coauthors used Equifax data on credit outcomes that lets them compare even finer geographies:
The paper is a clear case of modern econometrics at work, in that it is almost painfully thorough. Counting the appendix, the version currently up at AJHE shows 130 pages with 29 tables and 11 figures (many of which are actually made up of 6 sub-figures each). We put a lot of thought into questioning the assumptions behind our difference-in-difference estimation, and into figuring out how best to bootstrap our standard errors given the small number of clusters. Sometimes this feels like overkill but hopefully it means the final results are really solid.
For those who want to read more and can’t access the journal version, an earlier ungated version is here.
Disclaimer: The results and conclusions in this paper are those of the authors and do not indicate concurrence by the Agency for Healthcare Research and Quality or the US Department of Health and Human Services. The views expressed in this paper are solely those of the authors and do not necessarily reflect the views of the Federal Reserve Bank of Philadelphia or the Federal Reserve System. Any errors or omissions are the responsibility of the authors.
When you look across countries, it appears that the first $1000 per person per year spent on health buys a lot; spending beyond that buys a little, and eventually nothing. The US spends the most in the world on health care, but doesn’t appear to get much for it. A classic story of diminishing returns:
This might tempt you to go full Robin Hanson and say the US should spend dramatically less on health care. But when you look at the same measures across US states, it seems like health care spending helps after all:
Source: My calculations from 2019 IHME Life Expectancy and 2019 KFF Health Spending Per Capita
Last week though, I showed how health spending across states looks a lot different if we measure it as a share of GDP instead of in dollars per capita. When measured this way, the correlation of health spending and life expectancy turns sharply negative:
Source: My calculations from 2019 IHME life expectancy, Gross State Product, and NHEA provider spending
Does this mean states should be drastically cutting health care spending? Not necessarily; as we saw before, states spending more dollars per person on health is associated with longer lives. States having a high share of health spending does seem to be bad, but this is more because it means the rest of their economy is too small, rather than health care being too big. Having a larger GDP per capita doesn’t just mean people are materially better off, it also predicts longer life expectancy:
Source: My calculations from 2019 IHME life expectancy and 2019 Gross State Product
As you can see, higher GDP per capita predicts longer lives even more strongly than higher health spending per capita. Here’s what happens when we put them into a horse race in the same regression:
The effect of health spending goes negative and insignificant, while GDP per capita remains positive and strongly significant. The coefficient looks small because it is measured in dollars, but what it means is that a $10,000 increase in GDP per capita in a state is associated with 1.13 years more life expectancy.
My guess is that the correlation of GDP and life expectancy across states is real but mostly not caused by GDP itself; rather, various 3rd factors cause both. I think the lack of effect of health spending across states is real, between diminishing returns to spending and the fact that health is mostly not about health care. Perhaps Robin Hanson is right after all to suggest cutting medicine in half.
The Differences-in-Differences literature has blown up in the past several years. “Differences-in-Differences” refers to a statistical method that can be used to identify causal relationships (DID hereafter). If you’re interested in using the new methods in Stata, or just interested in what the big deal is, then this post is for you.
First, there’s the basic regression model where we have variables for time, treatment, and a variable that is the product of both. It looks like this:
The idea is that that there is that we can estimate the effect of time passing separately from the effect of the treatment. That allows us to ‘take out’ the effect of time’s passage and focus only on the effect of some treatment. Below is a common way of representing what’s going on in matrix form where the estimated y, yhat, is in each cell.
Each quadrant includes the estimated value for people who exist in each category. For the moment, let’s assume a one-time wave of treatment intervention that is applied to a subsample. That means that there is no one who is treated in the initial period. If the treatment was assigned randomly, then β=0 and we can simply use the differences between the two groups at time=1. But even if β≠0, then that difference between the treated and untreated groups at time=1 includes both the estimated effect of the treatment intervention and the effect of having already been treated prior to the intervention. In order to find the effect of the intervention, we need to take the 2nd difference. δ is the effect of the intervention. That’s what we want to know. We have δ and can start enacting policy and prescribing behavioral changes.
Easy Peasy Lemon Squeezy. Except… What if the treatment timing is different and those different treatment cohorts have different treatment effects (heterogeneous effects)?* What if the treatment effects change over time the longer an individual is treated (dynamic effects)**? Further, what if the there are non-parallel pre-existing time trends between the treated and untreated groups (non-parallel trends)?*** Are there design changes that allow us to estimate effects even if there are different time trends?**** There’re more problems, but these are enough for more than one blog post.
For the moment, I’ll focus on just the problem of non-parallel time trends.
What if untreated and the to-be-treated had different pre-treatment trends? Then, using the above design, the estimated δ doesn’t just measure the effect of the treatment intervention, it also detects the effect of the different time trend. In other words, if the treated group outcomes were already on a non-parallel trajectory with the untreated group, then it’s possible that the estimated δ is not at all the causal effect of the treatment, and that it’s partially or entirely detecting the different pre-existing trajectory.
Below are 3 figures. The first two show the causal interpretation of δ in which β=0 and β≠0. The 3rd illustrates how our estimated value of δ fails to be causal if there are non-parallel time trends between the treated and untreated groups. For ease, I’ve made β=0 in the 3rd graph (though it need not be – the graph is just messier). Note that the trends are not parallel and that the true δ differs from the estimated delta. Also important is that the direction of the bias is unknown without knowing the time trend for the treated group. It’s possible for the estimated δ to be positive or negative or zero, regardless of the true delta. This makes knowing the time trends really important.
STATA Implementation
If you’re worried about the problems that I mention above the short answer is that you want to install csdid2. This is the updated version of csdid & drdid. These allow us to address the first 3 asterisked threats to research design that I noted above (and more!). You can install these by running the below code:
program fra syntax anything, [all replace force] local from "https://friosavila.github.io/stpackages" tokenize `anything' if "`1'`2'"=="" net from `from' else if !inlist("`1'","describe", "install", "get") { display as error "`1' invalid subcommand" } else { net `1' `2', `all' `replace' from(`from') } qui:net from http://www.stata.com/ end fra install fra, replace fra install csdid2 ssc install coefplot
Once you have the methods installed, let’s examine an example by using the below code for a data set. The particulars of what we’re measuring aren’t important. I just want to get you started with the an application of the method.
local mixtape https://raw.githubusercontent.com/Mixtape-Sessions use `mixtape'/Advanced-DID/main/Exercises/Data/ehec_data.dta, clear qui sum year, meanonly replace yexp2 = cond(mi(yexp2), r(max) + 1, yexp2)
The csdid2 command is nice. You can use it to create an event study where stfips is the individual identifier, year is the time variable, and yexp2 denotes the times of treatment (the treatment cohorts).
The above output shows us many things, but I’ll address only a few of them. It shows us how treated individuals differ from not-yet treated individuals relative to the time just before the initial treatment. In the above table, we can see that the pre-treatment average effect is not statistically different from zero. We fail to reject the hypothesis that the treatment group pre-treatment average was identical to the not-yet treated average at the same time period. Hurrah! That’s good evidence for a significant effect of our treatment intervention. But… Those 8 preceding periods are all negative. That’s a little concerning. We can test the joint significance of those periods:
estat event, revent(-8/-1)
Uh oh. That small p-value means that the level of the 8 pretreatment periods significantly deviate from zero. Further, if you squint just a little, the coefficients appear to have a positive slope such that the post-treatment values would have been positive even without the treatment if the trend had continued. So, what now?
Wouldn’t it be cool if we knew the alternative scenario in which the treated individuals had not been treated? That’s the standard against which we’d test the observed post-treatment effects. Alas, we can’t see what didn’t happen. BUT, asserting some premises makes the job easier. Let’s say that the pre-treatment trend, whatever it is, would have continued had the treatment not been applied. That’s where the honestdid stata package comes in. Here’s the installation code:
local github https://raw.githubusercontent.com net install honestdid, from(`github'/mcaceresb/stata-honestdid/main) replace honestdid _plugin_check
What does this package do? It does exactly what we need. It assumes that the pre-treatment trend of the prior 8 periods continues, and then tests whether one or more post-treatment coefficients deviate from that trend. Further, as a matter of robustness, the trend that acts as the standard for comparison is allowed to deviate from the pre-treatment trend by a multiple, M, of the maximum pretreatment deviations from trend. If that’s kind of wonky – just imagine a cone that continues from the pre-treatment trend that plots the null hypotheses. Larger M’s imply larger cones. Let’s test to see whether the time-zero effect significantly differs from zero.
What does the above table tell us? It gives us several values of M and the confidence interval for the difference between the coefficient and the trend at the 95% level of confidence. The first CI is the original time-0 coefficient. When M is zero, then the null assumes the same linear trend as during the pretreatment. Again, M is the ratio by which maximum deviations from the trend during the pretreatment are used as the null hypothesis during the post-treatment period. So, above, we can see that the initial treatment effect deviates from the linear pretreatment trend. However, if our standard is the maximum deviation from trend that existed prior to the treatment, then we find that the alpha is just barely greater than 0.05 (because the CI just barely includes zero).
That’s the process. Of course, robustness checks are necessary and there are plenty of margins for kicking the tires. One can vary the pre-treatment periods which determine the pre-trend, which post-treatment coefficient(s) to test, and the value of M that should be the standard for inference. The creators of the honestdid seem to like the standard of identifying the minimum M at which the coefficient fails to be significant. I suspect that further updates to the program will come along that spits that specific number out by default.
I’ve left a lot out of the DID discussion and why it’s such a big deal. But I wanted to share some of what I’ve learned recently with an easy-to-implement example. Do you have questions, comments, or suggestions? Please let me know in the comments below.
The above code and description is heavily based on the original author’s support documentation and my own Statalist post. You can read more at the above links and the below references.
*Sun, Liyang, and Sarah Abraham. 2021. “Estimating Dynamic Treatment Effects in Event Studies with Heterogeneous Treatment Effects.” Journal of Econometrics, Themed Issue: Treatment Effect 1, 225 (2): 175–99. https://doi.org/10.1016/j.jeconom.2020.09.006.
**Sant’Anna, Pedro H. C., and Jun Zhao. 2020. “Doubly Robust Difference-in-Differences Estimators.” Journal of Econometrics 219 (1): 101–22. https://doi.org/10.1016/j.jeconom.2020.06.003.
***Callaway, Brantly, and Pedro H. C. Santa Anna. 2021. “Difference-in-Differences with Multiple Time Periods.” Journal of Econometrics, Themed Issue: Treatment Effect 1, 225 (2): 200–230. https://doi.org/10.1016/j.jeconom.2020.12.001.
****Rambachan, Ashesh, and Jonathan Roth. 2023. “A More Credible Approach to Parallel Trends.” The Review of Economic Studies 90 (5): 2555–91. https://doi.org/10.1093/restud/rdad018.
2023 continues to be a dangerous year for eminent economists. We have once again lost a Nobel laureate who was influential even by the standard of Nobelists, Robert Solow:
I’m sure you will soon see many tributes that discuss his namesake Solow Model (MR already has one), or discuss him as a person. I never got to meet him (just saw him give a talk) and the Solow Model is well known, so I thought I’d take this occasion to discuss one of his lesser-known papers- “Sustainability: An Economists Perspective“. What follows comes from my 2009 reaction to his paper: