Who owns model risk?

Update to readers after doing the CPE (read Internal Audit of AI Agents and Risk for background )

It’s not every day that you get to present to 36 genuine internal auditors for an hour. Not the rowdiest crowd, true to form, but they did answer my poll questions. Many of them work at a regional bank here in Birmingham, AL. Here is the result of an (unscientific) poll from the session:

Opinion Poll: Who should own the risk appetite built into an AI model?

Claude has a better understanding of internal audit procedures than I do and actually helped me come up with the question. The interpretation of the graph is a collaboration between me and Claude, so I will not suppress the em dashes.

The plurality (first line) is the “textbook-correct” instinct — but it’s hollow without capability. Putting ownership on the business unit that deploys the model matches the Three Lines Model cleanly: the first line owns and manages the risks it takes. Good instinct, and worth affirming. The catch is the whole premise of your talk: the deployers usually can’t see the risk preference embedded in their model, let alone measure or set it. So “the business owns it” is right in principle but nominal in practice unless that owner is given the tools to actually recover and govern the appetite. That’s the gap between the poll’s ideal and the room’s reality.

The committee vote (31%) reflects real emerging practice, with a trap. Nearly a third reached for a dedicated AI-governance body — consistent with where NIST’s AI RMF and ISO 42001 point. But a committee can quietly dilute accountability: “everyone owns it” becomes “no one owns it” if it isn’t paired with a clearly accountable first line. Worth naming that risk out loud.

The most important result is the one that isn’t in any single bar: there’s no consensus. If you had asked this room “who owns credit risk?” you’d have gotten a tight, near-unanimous answer. The fact that ownership of a model’s risk appetite scatters across all four choices tells you this accountability is genuinely unsettled in their organizations.

The settled view: credit risk is owned by the first line — the business that originates the exposure. The lending or client-facing unit that decides to extend credit owns the risk of that decision. This is the textbook first-line ownership case, and it’s why credit risk is often the cleanest example used to teach the Three Lines Model.

So, congrats to me and Claude for coming up with a question that split opinion in a room of expert practitioners?

Internal Audit of AI Agents and Risk

Members of the Institute of Internal Auditors know that they can “Join the Birmingham IIA on September 23, 2026 for an insightful CPE webinar featuring Dr. Joy Buchanan… “

I am pleased to get a chance to translate Buchanan and Foster (2026) to an industry audience. I have been reading up on audit controls to prepare.

Some help from Claude with the following: Enterprise risk management rests on a simple discipline: an organization decides how much risk it is willing to take in pursuit of its objectives — its risk appetite — sets tolerances around that level, and then works to keep actual decisions inside those limits. Under COSO ERM, the appetite statement and its tolerances are the structure; the ongoing question is one of conformance. It’s part of the IIA’s AI Auditing Framework and the Three Lines Model: management sets and owns the appetite, risk and compliance build guardrails and monitor, and internal audit provides independent assurance that what the organization actually does matches what it said it would tolerate. A systematic gap between the two is a finding.

For human decision-makers, we’ve built machinery to check this — credit policies, delegated authorities, four-eyes review, documented rationale, an auditable paper trail. We know how to reconstruct whether a loan officer’s or portfolio manager’s judgment stayed inside the lines.

Here is where our paper connects. When an LLM makes a risk-and-return decision — approving credit, weighting a portfolio, ranking procurement options — it too has a risk appetite. But almost none of that oversight infrastructure exists for it. We argue that AI agents are already making decisions with economic consequences and an element of risk.

By showing that the softmax mechanism inside an LLM is McFadden’s random utility model, Buchanan and Foster (2026) establish that the model’s choices reveal a genuine utility function — a measurable risk preference. Our portfolio experiment then recovers the parameters: the slope of the indifference curve we report is the model’s risk appetite, quantified.

I know an Internal Auditor. Some of their old functions will probably get automated. But they have new work to do: auditing the AI agents! Our paper is a step toward both measuring and manipulating the risk appetites of AI agents.

Buchanan, J., & Foster, J. (2026). The innate economic preferences of language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.26288

Image by Grok. I don’t sell that mug but EWED does have merch at https://shop.spreadshirt.com/economist-writing-every-day/

The World Keeps Spinning Faster

So many things are happening that it’s hard to decide which to write about. So I’ll briefly cover several that would each deserve a full post on a normal week:

1. The Fed voted unanimously to raise rates to 4%, showing that the Warsh Fed still has some independence from the President.

2. The 10-year Treasury bond hit 5%. The CBO’s long-term budget forecasts assume 4% rates for the 10-year, and looked unpleasant even with that rosy assumption. The UK dumped Prime Minister Liz Truss over their 10-year hitting 4.5%.

3. Jacob Coxon resigned from Anthropic over AI safety concerns and set off a surprisingly large firestorm over it.

4. The Trump administration reacted by insisting we move forward with AI, with the Department of War blaming Effective Altruists for trying to slow it down:

I suspect this could lead many people to get into Effective Altruism by promoting it from ‘weird thing for nerds‘ to ‘something you can do to oppose Trump’. If that interests you, this is a good place to start.

5. EA hero Michael Kremer (whose work inspired many donations to deworming) was just appointed Chief Economist of the World Bank.

6. But AI progress keeps rolling. Quantum theorist Scott Aaronson declares the Singularity is here:

I recoil from the neverending shell game where you say “oh sure, of course AI can now [escape from its sandbox / solve Millennium Problems / whichever dramatic thing it most recently did], no one ever denied that [I did deny it], wake me up when AI does [thing AI hasn’t yet done but is going to do next year], that’s when I’ll reevaluate my whole worldview [no I won’t].” Where no matter how fast the rollercoaster accelerates, even after your whole familiar world has vanished behind you, you’re still inventing reasons why it doesn’t count.

My position on AI is merely the conservative, skeptical position of 2006, updated with intellectual honesty for the reality of late 2026. And that position, if you need me to spell it out, is as follows:

AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA

It seems to me that the Singularity has already started; it’s just wildly unevenly distributed.

I’m honestly surprised we have yet to see an upsurge of millenarianism. Y2K concerns came and went 2000 years after Jesus’ birth. But see how much people worry about AI now, and realize it will only get more powerful as we approach 2000 years after his death- 2033, or if you trust the revisionists, 2030.

What If AI Earnings Fall Far Short of Expectations?

Trillions of dollars are being plowed into building U.S. data centers to handle future AI compute demand. These are being paid for by rich people and organizations, anticipating juicy returns on their investments. Those juicy returns depend on consumers (individuals or businesses) being willing to pay enormous amounts for access to AI.  Commentators are so enamored with the glorious prospects of how AI will end poverty and maybe even death, that it is hard to find a clear statement of who exactly will pay how much for all this. Lots of folks are happy to pay $20/month for AI. $200/month? Not so much.

I predict that those earnings will fall far short of expectations. We observe that AI consumption is bifurcating into two main markets, a commodity tier and a premium tier (with, of course, sub-tiers within each of those broad categories).  Chinese models are readily available over the internet, and they have proven capable of performing well enough to handle most AI tasks. The Chinese models are priced far lower (on the order of 10X lower) than the big U.S. “frontier” models (ChatGPT from Open AI, and Claude suite from Anthropic). By far the most U.S. compute usage depends on these two labs. Western users are starting to use the Chinese models more and more. This keeps prices so low that the frontier labs are losing money on their AI sales. Sophisticated users automatically route their AI workload to the cheapest feasible provider.

Their will always be a subset of AI usage in the West that requires the highest level of performance, or freedom from Chinese government spying or manipulation, that will be directed to a premium tier. But if that premium tier ends up being only, say, 20% of AI usage in 2028, there is a real question as to whether the financial bases of the data centers being built now can be sustained. There are huge bear and bull arguments on both sides here, which are tough to balance. I got only equivocal “it depends” answers from AI on this.

If the data centers don’t make expected profits, what then? It all depends on how they were financed. Most of the build-out to date has been the big 4 hyperscalers, Google, Amazon, Microsoft, and Meta spending their free cash flow from their other business lines. If it turns out they simply flushed that money down the toilet, no big deal. Just a trillion-dollar whoopsie. The CEOs will still get their bonuses, don’t worry.

But now as more debt financing enters in, the stakes get higher. Analysis seems to show that the debt loads that the big 4 hyperscalers have incurred is manageable – -their base cash flows are so huge that they can manage their own debt. But in the past year we have seen the emergence of monstrous “Special Purpose Vehicles” (SPVs) with a mixture of equity, debt, and guarantees, to finance practically all the upcoming trillion dollars of data centers. This pushes the financing of the balance sheets of the hyperscalers.

If those newer data centers flop, their equity investors will take a hit, leaving their creditors in the hole and in control. One really needs to analyze exactly who those equity and debt holders are for the SPVs.  If the creditors decide to recoup some of their investment by selling the data center for say 60 cents on the dollar, the most likely buyers would be…Google and Amazon. There is a school of thought that this (let the SPVs fail, scoop up their assets at discount) has been their plan all along. World domination in AI compute!  Just like they have achieved world domination in online search and video and shopping. Maybe.

The resulting slowdown in data center investing would likely throw the U.S. economy into slowdown or recession, considering that it’s estimated that fully half of our recent GDP growth has been from circularly-financed AI buildout. Chipmakers’ (Nvidia, Micron, AMD, etc.) profits depend on continued acceleration in AI build-out. If that build-out stalls, or even slows down, chipmaker profits will crater. Whether this risk is already priced into their share prices is debated.

Boilerplate disclaimer: Nothing here should be considered advice to buy or sell any security.

Temperature vs thinking level for LLM products

If using ChatGPT or an API in 2023, you might have followed the advice “turn the temperature down” when you wanted a more serious answer. That is going out of date. The following is a collaboration between me and Grok to lay this out simply:  

What temperature was

Imagine the model is always choosing the next word from a long list of options, ranked from “very likely” to “weird but possible.” Temperature changes how adventurous that choice is.

  • Low temperature: almost always pick the obvious next word. Answers feel steady, repetitive, “corporate.” Good when there is basically one right output, like extracting a date from a contract.
  • High temperature: more willing to pick a less obvious word. Answers feel livelier but sometimes incorrect.

On older products you could often set this yourself.

What thinking level is

Newer models (Gemini 3, OpenAI’s reasoning models) do extra work before they talk to you. They draft a private scratchpad: break the problem into steps, check themselves, then write the answer you see. You don’t see the scratchpad even though you pay for it. That hidden draft is the “thinking.”

Thinking level (sometimes called reasoning effort) is not “how random should the next word be?” It is “how long is the model allowed to work on that draft?”

  • Low / minimal: answer quickly and cheaply. Fine for “what’s the status of ticket 1842?”
  • High: spend more time (and more tokens) on hard things — a gnarly spreadsheet, a multi-step plan, a tricky piece of code.
Continue reading →

Nicholas Polson Has Written Over 200 Academic Papers in 2026 (so far)

UPDATE: I stupidly didn’t realize my co-blogger Mike wrote about this too. My bad! I skimmed last week’s posts, but it didn’t click in my head. Be sure to read his thoughts.

ALSO: lots of the papers by Polson seem to have vanished from SSRN since I wrote this blog post… yesterday. Everything after August 14th has been taken down. That brings his count of papers in 2026 down to a mere 168 papers. Still essentially several lifetimes of output from a typical academic.

CODA: As Andrew Gelman documented in real time, all of the papers appear to have been removed from SSRN for now. No explanation as of yet, but you can still see many of the papers listed (for now) on Polson’s Google Scholar page.

For most academics, writing papers that may be eventually published in peer-reviewed journals is an important part of what we do. For some academics, it is the main thing they do (others have more emphasis on teaching courses at their university). Most academics always have a few projects they are working on, with perhaps a goal of finishing 2 or 3 a year, and thus having a regular pipeline of a few publications every few years. But some academics are much more prolific.

Take Daron Acemoglu for example. He has long been considered extremely prolific. So far in 2026, according to his Google Scholar page, he has had 15 papers that have either been published this year or come out as new working papers (that’s the bulk of them). In the past year (2025 and 2026), he has had publications in the Quarterly Journal of Economics, the American Economic Review, and the Journal of Economic Literature, among others (the AER paper was his Nobel lecture). For lower tier academics, that’s almost a lifetime of publications in 12 months or so. Acemoglu is extremely productive.

But I just discovered an economist that is, apparently, even more productive than Acemoglu, at least as measured by working papers. Nicholas Polson has, by my count using his SSRN page, already written 258 working papers in 2026 alone. He’s already written (or at least published to SSRN), six papers today, August 26, 2026. In the month of August 2026, he has written and posted to SSRN a total of 104 papers — and counting, since the month isn’t quite over.

These aren’t just short notes. Most of the papers are of normal academic length: 32 pages, 27 pages, 58 pages. The papers are both theoretical — involving complex math in some cases — or empirical, with regressions. Read any single paper, and it feels like just a normal academic paper, the kind of thing that an academic might work on for a few months. He even has a frequent co-author, which is common for economics papers (and helps to be more productive), a systems engineering professor named Vadim Sokolov, who is a co-author on a little over 100 of the papers this year.

What is going on here? Obviously the research productivity of Polson and his co-author Sokolov is aided by AI. Who isn’t using AI to increase their writing and research productivity these days? But I don’t think I have seen any academic, at least not in economics, that has really pushed it to the limit.

Presumably, many of these papers will get submitted to academic journals. I can imagine the editors of journals have a very hard job these days, as the number of papers submitted has likely increased significantly, while the time that referees have available has not increased much (of course, AI is likely making referees more productive too, though many journals ask you not to upload the paper into an AI program as a referee, since it is unpublished work when you are reviewing it).

I really don’t know where academic publishing goes from here. AI has made us all more productive in terms of output, but are we better at answering important questions in our science than pre-AI? Probably, though it is hard to know. Journals and the peer review process has traditionally been the filter to sort real contributions from gibberish. I don’t know how the peer review process continues in its current form given the massive increase in output (much of it good!) that we are seeing from academics. Dr. Polson is just a leading example of a growing challenge for academia.

Intellectual Squatting

So a professor at a major institution wrote 200 papers last year. Unlike other commenters I’ve observed so far, I think this is neither true research nor pure AI fraud. It is likely AI “slop” to varying degrees, but unlike a lot of slop there is probably real value within it. What I have not yet seen ascertained is whether any of it has been vetted, investigated, or curated by the author in a meaningful way. The real question is: what is the actual ambition here? I think the tell is the lack of submission to peer review.

I think this is a form of intellectual squatting. The nice version is it’s putting out a series of half-baked papers in the hopes of establishing a property right to the underlying ideas at an earlier stage of the research process than previously possible. The less generous interpretation is it’s dumping a series of haystacks on the plains and laying claim to the needles probabilistically within each. Imagine you are a person who has highly esoteric, potentially important ideas every day. Many of those ideas you suspect, based on some combination of experience and ego, are new in at least one dimension. You would like to get credit for that newness. For being first. What’s the problem?

The problem is that scholarship remains more perspiration than inspiration. Having a new idea is great, but it takes years to work through the nuance in sufficient detail that you can convince your peers of the coherence and originality of the contribution. During the minutes each day you are not working on this singular project you have the inspiration for other ideas, sometimes multiple within a single day. How frustrating is the proposition that someone else gets credit for the originality of contribution just because they had time to reveal it to the world while you were embroiled in your investigation of what is only one of your many score ideas!?

Ah, but meta-level inspiration has struck you! What if you took each one of those ideas, spent an hour curating a series of prompts around it, and then let Chat GPT (or another LLM) fabricate an entire research paper around it? It might not be good, correct, or even coherent, but it does somethine far more important. It establishes an intellectual property right to the claim of being first. Now, to be clear, you are fully aware of the deficiciency of your paper as an actual scholarly contribution, but if somone else writes a full paper you at least have something to point to and say “I was here first. Cite me. Hell, if I’m close enough you might even have to name it after me. Well, sure, us. But definitely include me. Glory shared via hypenhnation is better than no glory at all.”

Is it a contribution? That’s something that will vary on a case-by-case basis, but I expect far more misses than hits. The work isn’t there. It’s like plopping down a block of marble with a dramatic-ish sketch of a man on an adhered post-it note and claiming that Michaelangelo needs to share credit with you on any subsequent sculptures. It’s like asking people to cite that one cool tweet you did about how DNA is cool but maybe RNA could be useful in vaccines one day. Intellectual property rights trolling via AI blunderbuss.

BTW, I’m not 100% sure this isn’t an AI take on a modern Sokal hoax. An attempt to show how much AI slop is introducing a whole new version of Gresham’s Law to scholarship. But if we treat it as earnest, it’s proof that a very smart person can potentially disrupt the market for scholarship, patents, or any other intellectual property by laying claim to ideas in much the same way that the printing press undermined the market for plenary indulges. Flood the market, leave it to someone else to sort through the ecumenical consequences.

Boy Wonder Leopold Aschenbrenner Blows Up His $45 Billion Situational Awareness Hedge Fund

Leopold Aschenbrenner is a very bright guy. Born in Germany to physician parents, he skipped enough grades to graduate from high school at age 15, allowing him to enroll at Columbia University in 2017 at that same age. He went on to graduate from Columbia at age 19 as valedictorian with a degree in economics and mathematics-statistics. An econ prof at the time said his “record of scholarship exceeds that of any student in the department in the previous 20 years.” The Mercatus Center, a think tank at George Mason University, recognized his potential, awarding him an Emergent Ventures grant.

Leopold Aschenbrenner, Columbia Class of 2021 Valedictorian

After graduation, Aschenbrenner moved into the effective-altruism research and grantmaking world, making notable contributions in various ways. In 2023, he joined the newly created “Superalignment” team at OpenAI, that was charged with making conceptual and engineering progress on aligning systems “much smarter than humans” before such systems were built. The next year he was discharged by OpenAI; the company said it was because of a security leak, but his version (which I find more credible) is that he was ousted as retaliation for circulating an internal memo arguing that the company’s protections against model-weight theft and algorithmic exfiltration were inadequate for AGI-relevant work.

Two months after his departure from OpenAI, he self-published Situational Awareness: The Decade Ahead, which you can download here.    This monograph synthesized scaling-law extrapolations, geopolitical analysis, and AI-lab security commentary into a single forecast: that the largest AI labs, on current trends, will plausibly reach AGI around 2027 and that an intelligence explosion to superintelligence could occur in the subsequent few years. This work went viral on Wall Street, and before you can say “monetization”, he was leading a hedge fund named, appropriately enough, Situational Awareness. There he put into practice his convictions that the demand for compute would be voracious, and would be limited by physical constraints such as electricity and chip fabs.

Thus, in his fund he went long companies like SanDisk (memory fab), Bloom Energy (makes solid oxide fuel cells), and Nebius (builds whole data centers). Very long, in fact, with leverage reportedly as high as 400%. He tried to hedge this long book by shorting software companies which are viewed as vulnerable to disruption by AI. This strategy worked fabulously for a while. His assets under management (AUM) climbed to $45 billion, with returns in the first 6-7 months of 2026 approaching 400%. Not bad for a 25-year-old.

But then, genius failed (yet again)- -in a stunning reversal, the market rebelled in late July against big AI capex spends, dumping memory fabs and infrastructure, and bought into the maligned software (SaaS) sector. So, BOTH his long and short legs went against him, followed in due course by the dreaded margin calls. Aschenbrenner’s $45 billion shrank to a measly $10 billion in a matter of days, as he was forced to sell off his public equities at a discount to those friendly capitalist sharks at Citadel. Wise old heads wagged, saying, yup, this sort of Black Swan event always happens sooner or later, and you then get carried out on a stretcher if you run a highly leveraged bet that is not truly hedged.

Down, but not out – – The word on the Street is that folks with money to invest are already lining up to entrust more bazillions to our altruistic wizard. We have not heard the last of Leopold Aschenbrenner.

AI Innate Preferences Paper on Arxiv

Please check out my new paper, with Joshua Foster

The Innate Economic Preferences of Language Models (arXiv link)

Abstract: Language models increasingly settle real resource tradeoffs on behalf of principals yet their economic preferences remain unobserved. We demonstrate their generation rule is isomorphic to the random utility model of discrete choice. This allows internal logit scores to structurally identify preferences. Estimating risk attitudes across twelve models in a portfolio task reveals universal but heterogeneous risk aversion. Although models reject strictly dominated options, their elicited preferences fail invariance tests and violate the independence of irrelevant alternatives across varying experimental prompts. Finally, fine tuning establishes that a principal can explicitly engineer a target risk attitude.

I hope you will refer to the manuscript for details, but I will share one picture here. This is panel (a) of Figure 3: Empirical indifference curves for open-weight models mapped over the portfolio space.

In simple language, what the red/blue picture shows is that the Qwen language model is picking the portfolios that offer more money (in expectation, with a distaste for excessive risk). That’s basically what a rational actor should do. We find that the language models make fairly consistent choices and rarely violate the monotonicity requirement for a well-behaved utility function.

How we describe this figure in the paper: “Starting from a base bundle with expected return µ = 10 and risk σ = 30, we sweep over the dense grid of alternative portfolios from our experimental protocol and record the position-corrected logit gap between each grid portfolio and the base. The yellow dashed line overlays the indifference curve implied by the mean-variance structural estimates, and the heatmap colors encode the sign and magnitude of the logit difference, with blue regions preferred to the base and red regions dispreferred. Several patterns emerge from these plots. All six models produce upward-sloping indifference curves, confirming that higher risk must be compensated by higher expected return.”

We think this basic research on behavior is important, for alignment research and for business applications with delegating work to AI agents. The first question to ask, before testing whether we can impose our preferences on AI agents, is whether those agents have preferences at all in a consistent sense.

Suggested citation: Buchanan, J., & Foster, J. (2026). The innate economic preferences of language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.26288

The Academic Data Project That Turned Into $375 Million

What could be better than creating data so valuable that an institution is happy to host and update it forever, like the Sean Lahman baseball database?

Creating data that sells for $375 million, like the Center for Research in Security Prices. University of Chicago professors assembled this series of finance datasets over decades, starting in 1960 with an effort to track every transaction of every publicly traded security. U Chicago sold CRSP to Morningstar last year for $375 million.

Why could they sell it for so much? It helps to be working in finance, where the willingness to pay is the highest. It also represents 65 years of work from what became a large team that included Nobelists like Eugene Fama. The data was valuable enough to become widely used by key institutions even though CRSP charged for it:

Today, $3 trillion in fund assets are linked to CRSP Market Indexes, including U.S. equity ETFs run by Vanguard, and more than 600 subscribers across 35 countries use CRSP Research Data Products.  

Did U Chicago sell CRSP at the right time? On the one hand, I wonder if this was a fire sale driven by federal grant cuts putting pressure on the U Chicago budget. On the other hand, assembling datasets like this is only going to get easier in the age of AI, so perhaps Chicago sold at the top.

For now though there is still an edge in having restricted datasets that AIs haven’t trained on and can’t access. When I ask myself what advantage my human research assistants have over AIs in 2026, the most obvious answer is that they can legally access restricted databases like CRSP or, in my current case, HeinOnline.