Internal Audit of AI Agents and Risk

Members of the Institute of Internal Auditors know that they can “Join the Birmingham IIA on September 23, 2026 for an insightful CPE webinar featuring Dr. Joy Buchanan… “

I am pleased to get a chance to translate Buchanan and Foster (2026) to an industry audience. I have been reading up on audit controls to prepare.

Some help from Claude with the following: Enterprise risk management rests on a simple discipline: an organization decides how much risk it is willing to take in pursuit of its objectives — its risk appetite — sets tolerances around that level, and then works to keep actual decisions inside those limits. Under COSO ERM, the appetite statement and its tolerances are the structure; the ongoing question is one of conformance. It’s part of the IIA’s AI Auditing Framework and the Three Lines Model: management sets and owns the appetite, risk and compliance build guardrails and monitor, and internal audit provides independent assurance that what the organization actually does matches what it said it would tolerate. A systematic gap between the two is a finding.

For human decision-makers, we’ve built machinery to check this — credit policies, delegated authorities, four-eyes review, documented rationale, an auditable paper trail. We know how to reconstruct whether a loan officer’s or portfolio manager’s judgment stayed inside the lines.

Here is where our paper connects. When an LLM makes a risk-and-return decision — approving credit, weighting a portfolio, ranking procurement options — it too has a risk appetite. But almost none of that oversight infrastructure exists for it. We argue that AI agents are already making decisions with economic consequences and an element of risk.

By showing that the softmax mechanism inside an LLM is McFadden’s random utility model, Buchanan and Foster (2026) establish that the model’s choices reveal a genuine utility function — a measurable risk preference. Our portfolio experiment then recovers the parameters: the slope of the indifference curve we report is the model’s risk appetite, quantified.

I know an Internal Auditor. Some of their old functions will probably get automated. But they have new work to do: auditing the AI agents! Our paper is a step toward both measuring and manipulating the risk appetites of AI agents.

Buchanan, J., & Foster, J. (2026). The innate economic preferences of language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.26288

Image by Grok. I don’t sell that mug but EWED does have merch at https://shop.spreadshirt.com/economist-writing-every-day/

What If AI Earnings Fall Far Short of Expectations?

Trillions of dollars are being plowed into building U.S. data centers to handle future AI compute demand. These are being paid for by rich people and organizations, anticipating juicy returns on their investments. Those juicy returns depend on consumers (individuals or businesses) being willing to pay enormous amounts for access to AI.  Commentators are so enamored with the glorious prospects of how AI will end poverty and maybe even death, that it is hard to find a clear statement of who exactly will pay how much for all this. Lots of folks are happy to pay $20/month for AI. $200/month? Not so much.

I predict that those earnings will fall far short of expectations. We observe that AI consumption is bifurcating into two main markets, a commodity tier and a premium tier (with, of course, sub-tiers within each of those broad categories).  Chinese models are readily available over the internet, and they have proven capable of performing well enough to handle most AI tasks. The Chinese models are priced far lower (on the order of 10X lower) than the big U.S. “frontier” models (ChatGPT from Open AI, and Claude suite from Anthropic). By far the most U.S. compute usage depends on these two labs. Western users are starting to use the Chinese models more and more. This keeps prices so low that the frontier labs are losing money on their AI sales. Sophisticated users automatically route their AI workload to the cheapest feasible provider.

Their will always be a subset of AI usage in the West that requires the highest level of performance, or freedom from Chinese government spying or manipulation, that will be directed to a premium tier. But if that premium tier ends up being only, say, 20% of AI usage in 2028, there is a real question as to whether the financial bases of the data centers being built now can be sustained. There are huge bear and bull arguments on both sides here, which are tough to balance. I got only equivocal “it depends” answers from AI on this.

If the data centers don’t make expected profits, what then? It all depends on how they were financed. Most of the build-out to date has been the big 4 hyperscalers, Google, Amazon, Microsoft, and Meta spending their free cash flow from their other business lines. If it turns out they simply flushed that money down the toilet, no big deal. Just a trillion-dollar whoopsie. The CEOs will still get their bonuses, don’t worry.

But now as more debt financing enters in, the stakes get higher. Analysis seems to show that the debt loads that the big 4 hyperscalers have incurred is manageable – -their base cash flows are so huge that they can manage their own debt. But in the past year we have seen the emergence of monstrous “Special Purpose Vehicles” (SPVs) with a mixture of equity, debt, and guarantees, to finance practically all the upcoming trillion dollars of data centers. This pushes the financing of the balance sheets of the hyperscalers.

If those newer data centers flop, their equity investors will take a hit, leaving their creditors in the hole and in control. One really needs to analyze exactly who those equity and debt holders are for the SPVs.  If the creditors decide to recoup some of their investment by selling the data center for say 60 cents on the dollar, the most likely buyers would be…Google and Amazon. There is a school of thought that this (let the SPVs fail, scoop up their assets at discount) has been their plan all along. World domination in AI compute!  Just like they have achieved world domination in online search and video and shopping. Maybe.

The resulting slowdown in data center investing would likely throw the U.S. economy into slowdown or recession, considering that it’s estimated that fully half of our recent GDP growth has been from circularly-financed AI buildout. Chipmakers’ (Nvidia, Micron, AMD, etc.) profits depend on continued acceleration in AI build-out. If that build-out stalls, or even slows down, chipmaker profits will crater. Whether this risk is already priced into their share prices is debated.

Boilerplate disclaimer: Nothing here should be considered advice to buy or sell any security.

Temperature vs thinking level for LLM products

If using ChatGPT or an API in 2023, you might have followed the advice “turn the temperature down” when you wanted a more serious answer. That is going out of date. The following is a collaboration between me and Grok to lay this out simply:  

What temperature was

Imagine the model is always choosing the next word from a long list of options, ranked from “very likely” to “weird but possible.” Temperature changes how adventurous that choice is.

  • Low temperature: almost always pick the obvious next word. Answers feel steady, repetitive, “corporate.” Good when there is basically one right output, like extracting a date from a contract.
  • High temperature: more willing to pick a less obvious word. Answers feel livelier but sometimes incorrect.

On older products you could often set this yourself.

What thinking level is

Newer models (Gemini 3, OpenAI’s reasoning models) do extra work before they talk to you. They draft a private scratchpad: break the problem into steps, check themselves, then write the answer you see. You don’t see the scratchpad even though you pay for it. That hidden draft is the “thinking.”

Thinking level (sometimes called reasoning effort) is not “how random should the next word be?” It is “how long is the model allowed to work on that draft?”

  • Low / minimal: answer quickly and cheaply. Fine for “what’s the status of ticket 1842?”
  • High: spend more time (and more tokens) on hard things — a gnarly spreadsheet, a multi-step plan, a tricky piece of code.
Continue reading →

Boy Wonder Leopold Aschenbrenner Blows Up His $45 Billion Situational Awareness Hedge Fund

Leopold Aschenbrenner is a very bright guy. Born in Germany to physician parents, he skipped enough grades to graduate from high school at age 15, allowing him to enroll at Columbia University in 2017 at that same age. He went on to graduate from Columbia at age 19 as valedictorian with a degree in economics and mathematics-statistics. An econ prof at the time said his “record of scholarship exceeds that of any student in the department in the previous 20 years.” The Mercatus Center, a think tank at George Mason University, recognized his potential, awarding him an Emergent Ventures grant.

Leopold Aschenbrenner, Columbia Class of 2021 Valedictorian

After graduation, Aschenbrenner moved into the effective-altruism research and grantmaking world, making notable contributions in various ways. In 2023, he joined the newly created “Superalignment” team at OpenAI, that was charged with making conceptual and engineering progress on aligning systems “much smarter than humans” before such systems were built. The next year he was discharged by OpenAI; the company said it was because of a security leak, but his version (which I find more credible) is that he was ousted as retaliation for circulating an internal memo arguing that the company’s protections against model-weight theft and algorithmic exfiltration were inadequate for AGI-relevant work.

Two months after his departure from OpenAI, he self-published Situational Awareness: The Decade Ahead, which you can download here.    This monograph synthesized scaling-law extrapolations, geopolitical analysis, and AI-lab security commentary into a single forecast: that the largest AI labs, on current trends, will plausibly reach AGI around 2027 and that an intelligence explosion to superintelligence could occur in the subsequent few years. This work went viral on Wall Street, and before you can say “monetization”, he was leading a hedge fund named, appropriately enough, Situational Awareness. There he put into practice his convictions that the demand for compute would be voracious, and would be limited by physical constraints such as electricity and chip fabs.

Thus, in his fund he went long companies like SanDisk (memory fab), Bloom Energy (makes solid oxide fuel cells), and Nebius (builds whole data centers). Very long, in fact, with leverage reportedly as high as 400%. He tried to hedge this long book by shorting software companies which are viewed as vulnerable to disruption by AI. This strategy worked fabulously for a while. His assets under management (AUM) climbed to $45 billion, with returns in the first 6-7 months of 2026 approaching 400%. Not bad for a 25-year-old.

But then, genius failed (yet again)- -in a stunning reversal, the market rebelled in late July against big AI capex spends, dumping memory fabs and infrastructure, and bought into the maligned software (SaaS) sector. So, BOTH his long and short legs went against him, followed in due course by the dreaded margin calls. Aschenbrenner’s $45 billion shrank to a measly $10 billion in a matter of days, as he was forced to sell off his public equities at a discount to those friendly capitalist sharks at Citadel. Wise old heads wagged, saying, yup, this sort of Black Swan event always happens sooner or later, and you then get carried out on a stretcher if you run a highly leveraged bet that is not truly hedged.

Down, but not out – – The word on the Street is that folks with money to invest are already lining up to entrust more bazillions to our altruistic wizard. We have not heard the last of Leopold Aschenbrenner.

Announcing the Disability Records Project

Did you know that we have access to digital copies of the historical US census rolls? You can also find the digitized data at IPUMS. However, the data for people with disabilities is not great. It depends on the year, but those data have error rates on the order of 20% or higher.  We have the digital census rolls, the data just doesn’t match them.

So, I created a non-install windows computer application that lets people identify disabled people on those digital census rolls. Complemented with machine learning, my goal is to improve the accuracy of historical records about people with disabilities. Historical and quantitative research about disabled populations is relatively thin. We can do better. If you have students who would benefit from this research experience, then do please let me know! I can approve your institution’s email domain and we can get started.

The application is really straightforward with basically two user-facing features.

Continue reading →

What Is So Special About Object-Oriented Programming Languages Like C++ and Python?

I first learned computer programming about 1974, using FORTRAN running on an IBM 360 system that, yes, filled a whole room. And yes, my source code existed in the form of a stack of cards with holes punched in them, which got run through a physical card reader. FORTRAN and similar old-school languages were efficient (b/c computer resources were so constrained) and syntactically simple for solving well-specified problems.

C++ started to become popular in the 1980s, and Java in the 1990s. A big part of their appeal was that they were “object-oriented programming” (OOP) languages. I repeatedly asked my computer-programming professional friends back then to help me understand the difference between OOP and conventional Fortran type programs. They would get misty-eyed and rhapsodize about how their program components were modularized.  I guess I just failed to ask the right questions, because I never could understand why what they were talking about was so very much better or different than a good clean FORTRAN program, where most of the work was compartmentalized into well-defined functions and sub routines.

So I had a good talk with Claude about all this, and achieved enlightenment.. The differences seem to come down to a couple of key concepts:


(1) Data Compartmentalization

 In FORTRAN, you can modularize the data manipulation steps into subroutines, but the data tends to be more in common. Thus, for a very large programs, it is hard to keep some far-distant subroutine from accidentally altering your data. But with OOP, the data and the manipulation methods are “encapsulated” into one airtight thing, so no outside routine can mess with that data.

(2) More Robust Relations Among Chunks of Code

With OOP, there is also a feature called “inheritance”, where some new method can take advantage of an existing method, in a cleaner way than (in the FORTAN world) having a new subroutine call an existing subroutine, which would involve explicitly passing a bunch of parameters back-and-forth (which is very easy to mess up).

For doing fairly straightforward scientific calculations, even big ones, I think FORTRAN is still easier and more efficient. But for modern financial programs, involving millions of lines, written by huge teams of people that cannot all talk to one another, the win goes to OOP. Besides C++ and Java (still popular), in OOP we now have C# (standard for many Windows and gaming applications), and the crowd favorite, Python.


(That’s about it simply as I could put it, without getting long-winded and technical… If you want more details, you can always ask my buddy Claude)

How To Get Spare Car Fob Made Inexpensively (And Why It Is Wise to Do So)

When we recently traded in for a newer car, we were handed two fobs. These are push-button start cars, so you absolutely need a fob to drive one. An old-fashioned metal key will not work. Even the hidden metal key in your fob is only good for unlocking the door, not for starting the car. Since life has schooled me that things get lost, I asked how much it would cost to get a spare fob made at time of purchase, hoping I might catch a break.  Nope, the dealer cost would be $436. Ouch.

So, I later asked AI what to do. I was advised to have a non-dealer locksmith or hardware store do it instead. With this route, you might save even more money if you buy the blank fob yourself on Amazon for something like $30, and just have the locksmith program it and cut the hidden metal key.

A local ACE hardware store quoted me a price of $330, which is better than $436, but still not great. But I located a mobile locksmith, who drives a big van loaded with spare keys and fobs, and the equipment to cut key copies and to program fobs. He specializes in going and helping folks who are locked out of their cars, and he will drive to your house to make spare fobs and keys.   He would come to our house to make a new fob for $270. But it gets even better – – if I had two fobs made in one visit, the price would be only $220 apiece. It was an offer I could not refuse, so now I have spares for both cars.

If you don’t have a working fob for the locksmith to work from:  It turns out that for many foreign cars (Subaru, Toyota, and especially most European brands), it can be very difficult for a non-dealer locksmith to create a fob for you. Normally, if you find yourself locked out of your car some dark, snowy night, the mobile locksmith can roll up and create a new fob for you so you can get back in business. For brands like Ford, GM, Honda, and Nissan, he can accomplish this even if you can’t provide him with a working fob to copy. He will charge you twice as much, because it is a much harder job with no working template fob. It gets even harder for, say, Subaru. But for Mercedes-Benz or BMW, it may be impossible for most locksmiths to get a fob made on the spot without a working model to copy.  In a big metro area, there may be a few locksmiths who have the specialized equipment for these brands; you’d have to call around and ask the locksmith especially if they can make a new, say, Mercedes fob for an “all-keys-lost” situation. If not, you may have to get towed on a flat-bed ($$$) to a dealer, who can read your vehicle and program the fob ($$$). There may be a middle ground (which will take some days to play out) where you order a blank fob from a dealer along with some essential info for your locksmith to complete the programming, or where you order a programmed fob by mail.

All this argues for having a spare fob made ahead of time; maybe stash it somewhere that a friend or family member could bring it to you, if you were in distress not too far away.

Bonus fob tip: If your fob battery dies, you can still start your car by using the conventional metal key hidden inside your fob to open the car door, and then hold the fob very close to the car ignition button as you push the button, with your foot on the brake as usual.

Fable on Legibility

Claude Fable is Anthropic’s most capable publicly available “Mythos-class” model. It is optimized for long-running autonomous tasks and deep knowledge work. The roll out of this product has been dramatic. Little people like me have access to it for only two weeks, and I doubt I will be able to afford it thereafter. With my window of access, I posed it the following prompt:

“This article indicates that a much smarter model might not be possible because the universe is opaque. Since Tyler Cowen made this statement, AI has helped people make breakthroughs in math and biology. Write this again in 2026 using the latest state of technology. Is the smartest LLM today much smarter than GPT-4? And is the universe legible?” and I copied in my blog post Is the Universe Legible to Intelligence?

Fable replied in 3 pages of text which you can download as a Word doc here.

One line from Fable’s response: “Notice where the wins clustered: mathematics with checkable proofs, protein structures with experimental ground truth, contest problems with known answers. These are the maximally legible domains — places with a fixed target and a way to verify that you hit it.”

The writing is coherent and contains no obvious hallucinations. Is the answer true, and does it tell the whole truth?

Whenever you read something think about who wrote it, and keep in mind that every author/model has a bias and limitations. In my paper with Will Hickman, we found that just reminding people that a paragraph has an author (whether the author is human or AI) increased the demand for fact checking from readers. LLMs will become more persuasive and closer to (but never completely) correct. Keep reading all things with some skepticism whether they are written by scientists, politicians, or AI.

Regardless, there is definitely such a thing as making the known world more legible to AI today. Thus, people are talking about increasing funding for data availability and the possible demise of the “research paper.”

Research papers are more like stories than facts. The demand for stories is not going away, but I definitely cannot predict the future of the write-for-pay scientist.

AI, potentially, could go and get its own new data, instead of waiting for humans to archive it. Thus, the self-improving AI might take us beyond the current models… unless they run up against something that is not legible to intelligence…

Expressionism Is To Cameras As…. Caity Weaver?…. Is To Writing?

I don’t think it’s a coincidence that movements like expressionism, impressionism, and abstract art took off after the invention of the camera. Photorealistic paintings are impressive, but once they are duplicating what a camera does, they’re less interesting.

We’re due for similar movements in other fields to emerge as reactions to AI. Like writing in a way totally different from how an AI would write- ideally better than an AI would write, but even writing worse than an AI can be interesting if it is at least different.

It’s still early days for both AI and our reactions to it. But since the release of ChatGPT in 2022, what is the good new essay or book that you’re most confident was not written by AI, one that was written in an almost deliberately extra-human manner?

Continue reading →

Life Among the Geeks: Impressions from an Amateur Radio (Ham) Field Day

When visiting a large county park the other day, I noticed a sign for “Amateur Radio Field Day, Public Welcome.” I have always had great respect for amateur radio operators, for their technical prowess and for their service in emergencies when other forms of communication go down. So, I wandered over to see what was happening.

There were maybe half a dozen large screened in gazebo type tents. In each gazebo there were one or two men seated before some radio apparatus on a table, with a couple other guys sitting in more of a spectator role, chatting away. Someone had used a potato cannon to shoot a fishing line over the top of a tall tree, which was then used to pull up a rope, which was then used to pull up a long antenna wire. There was also a tall standalone antenna mast, held erect by three guy wires staked to the ground.

What were these guys doing? There were mainly doing what hams mainly do, which is reaching out across the continent and across the globe, to make contact via radio with other amateur radio operators. Some were doing it by voice communication, using single side band (SSB) technology for more efficient transmissions. Some were using a computer interface to turn their transmissions into a digital format, which could be transmitted and received more effectively. The guy on the other end would have matching software to turn the numbers back into words.

And there was one guy doing it the old-fashioned way, listening to “beep, beep, beedily beep” Morse code coming in, and transcribing that to letters and numbers. The Morse code beeps can cut through radio clutter more strongly than voice signals, allowing more distant contacts with relatively low power. Which is kind of a bragging point.

The special thing about so-called short waves used by hams is that they can reflect or refract from the underside of charged layers high up in the atmosphere. Your signal can bounce off those layers and come back to earth hundreds or thousands of miles away. Before the Internet, this was actually a big deal. The only other ways to communicate back then with distant people were snail mail, telegraph, or long-distance telephone, which was super expensive and grid dependent. I put together a short-wave video kit around 1970, and it was pretty amazing to hear a station broadcasting from Ecuador, not to mention Radio Moscow. Voice of America and BBC short wave broadcasts helped to keep the ideal of freedom alive in Soviet occupied Europe for many decades.

The guys in the park (I didn’t see any women operators, although I know they exist) were enthusiastic and friendly. When I mentioned I have a grandson who seems interested in technology, a man pressed on me an envelope filled with electronic components, and a small circuit board, encouraging me to help my grandson solder the parts into the board to make a short-wave receiver for Morse code. I thanked him sincerely, but said I might wait till my grandson was seven or eight before embarking on this project.

I am something of a techie myself, with a professional career in chemical engineering. But I found myself far out-geeked by these gentlemen. They would be certainly eager to help in the amount of emergency, but that’s actually a very rare event. It seems much of the thrill is simply the technology itself. To truly understand antenna theory, and even to design and build your own radio, are marks of distinguishment in that crowd. What I found most interesting from a psychological point of view was that when they do connect with someone 1000 or 10,000 miles away, they usually do not stop to chitchat about things like, how’s the weather where you are, how long have you been doing this, etc. No, nothing of normal human interest, it’s just: dutifully log the other party’s location and call sign, and move on to make the next contact.

My takeaway: as I’ve done maybe once a decade, I peered into my soul to assess whether I wanted to invest the time and money into pursuing this hobby. I could crank up my aging brain to pass the technical exams required to get an amateur radio operator’s license. But to set up a working short-wave radio and antenna, and get proficient at using it, would take a lot of effort. And in the Internet age, reaching out and touching someone on a different continent in real time just is not the rush that it used to be. As a responsible citizen, the emergency preparedness aspect is of interest to me. However, there is a newer, alternative class of radio called GMRS, which I think may be more realistic for most emergency situations. My encounter the other day with the hams did cause me to reengage with GMRS, and I may write a blog post about that in the future.