Members of the Institute of Internal Auditors know that they can “Join the Birmingham IIA on September 23, 2026 for an insightful CPE webinar featuring Dr. Joy Buchanan… “
I am pleased to get a chance to translate Buchanan and Foster (2026) to an industry audience. I have been reading up on audit controls to prepare.
Some help from Claude with the following: Enterprise risk management rests on a simple discipline: an organization decides how much risk it is willing to take in pursuit of its objectives — its risk appetite — sets tolerances around that level, and then works to keep actual decisions inside those limits. Under COSO ERM, the appetite statement and its tolerances are the structure; the ongoing question is one of conformance. It’s part of the IIA’s AI Auditing Framework and the Three Lines Model: management sets and owns the appetite, risk and compliance build guardrails and monitor, and internal audit provides independent assurance that what the organization actually does matches what it said it would tolerate. A systematic gap between the two is a finding.
For human decision-makers, we’ve built machinery to check this — credit policies, delegated authorities, four-eyes review, documented rationale, an auditable paper trail. We know how to reconstruct whether a loan officer’s or portfolio manager’s judgment stayed inside the lines.
Here is where our paper connects. When an LLM makes a risk-and-return decision — approving credit, weighting a portfolio, ranking procurement options — it too has a risk appetite. But almost none of that oversight infrastructure exists for it. We argue that AI agents are already making decisions with economic consequences and an element of risk.
By showing that the softmax mechanism inside an LLM is McFadden’s random utility model, Buchanan and Foster (2026) establish that the model’s choices reveal a genuine utility function — a measurable risk preference. Our portfolio experiment then recovers the parameters: the slope of the indifference curve we report is the model’s risk appetite, quantified.
I know an Internal Auditor. Some of their old functions will probably get automated. But they have new work to do: auditing the AI agents! Our paper is a step toward both measuring and manipulating the risk appetites of AI agents.
I am going to Korea next year. I look forward to being able to relay travel observations. Here’s what I can say so far.
I have several distinct personal points of connection to Korea through my family, my husband’s family, and my next-door neighbor here in Alabama. There is a George Mason U. campus in Korea. I hope to explore all of those connections when I go.
KPop Demon Hunters was an absolute phenomenon on Netflix last year. Every school child in America, across every race and state, knew the songs. As for the soundtrack, I determine: “Golden” – correctly rated “Soda Pop” – overrated “What it Sounds Like” – underrated And what is interesting in September of 2026 is how fast the whole thing died out. Trends come and then they really go in the monoculture (is this the “unicontext”?).
I am learning Korean by doing one lesson per day on Duolingo. Progress is slow, partly because I do not expect to need to become fluent ever. I am enjoying this. I have frittered away a good number of hours on cognitive trivia like computer Solitaire and Wordle. On the margin, I encourage a few of you English speakers to make learning a Non-Romance Language your new game. I have also used Duolingo to try to brush up on Spanish and French, but it’s fun to enter a totally different mindset altogether. It’s also difficult and might be frustrating if I had a higher bar for my performance.
If any readers have a connection in Seoul, please let me know.
If using ChatGPT or an API in 2023, you might have followed the advice “turn the temperature down” when you wanted a more serious answer. That is going out of date. The following is a collaboration between me and Grok to lay this out simply:
What temperature was
Imagine the model is always choosing the next word from a long list of options, ranked from “very likely” to “weird but possible.” Temperature changes how adventurous that choice is.
Low temperature: almost always pick the obvious next word. Answers feel steady, repetitive, “corporate.” Good when there is basically one right output, like extracting a date from a contract.
High temperature: more willing to pick a less obvious word. Answers feel livelier but sometimes incorrect.
On older products you could often set this yourself.
What thinking level is
Newer models (Gemini 3, OpenAI’s reasoning models) do extra work before they talk to you. They draft a private scratchpad: break the problem into steps, check themselves, then write the answer you see. You don’t see the scratchpad even though you pay for it. That hidden draft is the “thinking.”
Thinking level (sometimes called reasoning effort) is not “how random should the next word be?” It is “how long is the model allowed to work on that draft?”
Low / minimal: answer quickly and cheaply. Fine for “what’s the status of ticket 1842?”
High: spend more time (and more tokens) on hard things — a gnarly spreadsheet, a multi-step plan, a tricky piece of code.
A childhood friend and Facebook contact of mine pleaded to her friends, “How do people afford multiple children?”
I checked up on her and saw delightful pictures of a second child who, one way or another, they are now financially supporting along with that firstborn girl, in 2026. I obviously won’t post her family photos, but they took a summer trip to Yosemite National Park that looked something like this
Image by Grok
Based on how she felt about supporting a second child, I would be surprised if they have a third, a “big family.” Two is about where most people can still afford to take trips across the country in the summer.
This caused me to reflect on Dolly Parton, who we all miss. The U.S. flag is at half-staff in recognition of Dolly’s passing.
Part of her charming narrative is the fact that she was born into poverty. Famously, Dolly grew up with 11 siblings in a two-room Tennessee cabin. I doubt their family made it out to Yosemite.
So, with some help from ChatGPT, I inquired about the Total Fertility Rate of those Parton children. In one generation, everything dramatically changed. The number of children born to the Parton clan ranges from three to zero. Chat writes:
Dolly and her five sisters appear to have had only about six biological children among them—an average of roughly one child per woman. Dolly herself had none. Including their brothers, the twelve Parton siblings seem to have produced only about sixteen children altogether.
The Parton family therefore compresses a major American demographic change into just two generations. Avie Lee had twelve children; her daughters averaged about one each.
I used Grok for this query because economists were discussing the issue on X(Twitter) this week and Grok would have quick access to that discourse. Note that Grok provided links to legitimate web sources, not just Twitter speculation. I am providing those same links in the place where Grok put them. I do not know how much the chatter of the week affected the response below.
Prompt from J. B. on 8/22/26 to Grok: What happened to the economics journal Games and Economic Behavior?
The leading game theory journal Games and Economic Behavior (GEB), published by Elsevier since its founding in 1989, remains active and is still publishing issues (including volumes into 2026), but it is in a major editorial crisis due to a mass resignation of its leadership. (https://www.sciencedirect.com/journal/games-and-economic-behavior)
In late July 2026, Editor-in-Chief Hervé Moulin (who had held the role since January 2021) was informed by Elsevier that his 3-year contract would not be renewed, despite his willingness to continue. Elsevier cited a desire to move “towards a closer alignment with Elsevier strategic priorities and publishing policies.” Requests for further explanation (from Moulin and then collectively from the other six editors) were dismissed; Elsevier simply asserted its contractual right to make the decision and began searching for new editors without input from the existing board. (https://gametheorysociety.org/message-from-herve-moulin-geb-editor/)
In response, by August 14, 2026, all seven editors (Moulin plus Sushil Bikhchandani, Aislinn Bohren, Nicole Immorlica, Roberto Serrano, Emanuel Vespa, and Nicolas Vieille) resigned with the required three-month notice. Nearly all of the advisory editors (the equivalent of associate editors) also resigned in solidarity. The outgoing team is handling papers already under review and new submissions until November 14, 2026 (with some flexibility for papers already assigned by mid-August); after that, they will only finish evaluations already in progress. Elsevier has publicly thanked Moulin for his service. (https://gametheorysociety.org/message-from-herve-moulin-geb-editor/)
The journal continues under Elsevier ownership (it is one of the two official journals of the Game Theory Society, alongside the International Journal of Game Theory). Community discussion has focused on the implications for the field, the role of commercial publishers, and possibilities such as launching a successor journal. Similar recent mass resignations have occurred at other Elsevier titles over governance and editorial control issues. (https://casrai.org/news/2026-editorial-board-mass-resignations)
In short, GEB has not shut down, but its longstanding editorial team has largely departed over Elsevier’s non-renewal of the Editor-in-Chief and related decisions.
Prompt from J. B. on 8/22/26 to Grok: What will be the implications for the reputation of the journal?
The reputation of Games and Economic Behavior (GEB) is likely to face significant short- to medium-term damage, with longer-term outcomes depending heavily on how Elsevier handles the transition and whether a strong successor journal emerges.
GEB has long been viewed as the leading specialized journal in game theory (and one of the stronger outlets in theoretical economics more broadly). Its prestige rests on attracting high-quality, creative work; the expertise and independence of its editorial team; rigorous standards; and community trust. The mass resignation of all seven editors (including Editor-in-Chief Hervé Moulin) plus nearly the entire advisory board in August 2026, following Elsevier’s non-renewal of Moulin’s contract, directly undermines those foundations. (https://gametheorysociety.org/message-from-herve-moulin-geb-editor/)
Short-term effects
Loss of trust and submissions: Top researchers often choose journals partly based on the known editors and the expectation of fair, expert handling. With the board depleted (current editors only fully process papers until mid-November 2026), many authors are likely to pause or divert submissions. Community discussion already includes calls for boycotts of the Elsevier GEB and support for alternatives. @Afinetheorem
Perception of publisher interference: Elsevier’s stated rationale (“closer alignment with Elsevier strategic priorities and publishing policies”) and its decision to search for new editors without board input have fueled views that commercial priorities overrode academic ones. This echoes other recent mass resignations (e.g., Journal of Approximation Theory at Elsevier, where resigning editors declared the journal “as we have known it has effectively ceased to exist”). casrai.org
Operational disruption: Delays in reviews, potential quality variation under interim or new leadership, and uncertainty can further erode confidence among authors, referees, and readers.
Metrics such as impact factor or AJG ranking (currently 3) change slowly and will not immediately reflect the crisis, but subjective reputation in the game theory and economics communities can shift quickly.
Medium- to long-term scenarios
Historical patterns from similar mass resignations offer useful analogies:
“Zombie journal” risk: The title, ISSN, indexing, and historical impact factor stay with Elsevier. If the new editorial team is perceived as less distinguished or more aligned with volume/publisher goals, submissions from leading researchers can decline, citation patterns can shift (sometimes toward more geographically concentrated or lower-prestige work), and the journal’s standing within the field can permanently weaken. Cases like Lingua (Elsevier, 2015) and others show that the brand can persist while the scholarly niche contracts. sciencedirect.com
Recovery path: If Elsevier rapidly appoints highly respected new editors who maintain rigorous standards and rebuild community confidence, the damage can be limited. Some journals have stabilized after board changes, though full restoration of prior prestige is uncommon when the exit was acrimonious.
Successor journal effect: Strong community interest already exists in launching a new outlet (possible names floated include variations on “Theory of Games and Economic Behavior” or society-backed alternatives; comparisons are drawn to how JEEA or other breakaways have functioned). If a credible successor is established—especially one endorsed by the departing editors and the Game Theory Society—and is treated as the “true” continuation by tenure/promotion committees, the original GEB’s relative prestige would likely decline further. Successful breakaways (e.g., Glossa after Lingua, or Imaging Neuroscience after NeuroImage) have often captured much of the field’s attention and high-quality papers. journalmetrics.org
**end of Grok responses ** If you want some chatter, you can put “GEB” into the search bar at x.com, LinkedIn, or BlueSky this week and see some more speculation.
Lastly, dear reader, you might wonder why so much of my post was AI-generated this week. I see this as burning my tokens so you don’t have to. You can find this with a free web search. Some people predict that “cheap AI” is coming to an end, so we might find it worthwhile to publish responses as well as prompts for each other. Regards, fellow humans and AI readers.
At a rummage sale, I picked up a book by cartoonist Scott Adams called The Dilbert Future: Thriving on Stupidity in the 21st Century published in 1997. I thought I might find a clever prediction about the future, which we can now verify from the standpoint of 2026.
The text of the book is mostly dumb. I get the impression that Scott Adams was making easy money with a guaranteed humor book contract. I don’t recommend the book to anyone.
HOWEVER, with my paper copy I kept skimming ahead to see if any of his predictions about the future were impressive. Finally, on page 200 I found something.
Recall, the internet only became publicly available in the early 90’s. Respectable newspapers might have started to lose out to cable news in the mid-90’s. Blogs did not start until after Adams’ book was published. Social media proper (marked by the launch of Facebook) started in 2004. (Let millennials quietly walk away from Xanga journals and pretend that never happened.) So, my interest in this passage hinges on the fact that this book has a publication date of 1997.
The following is copied from Adams’ humor book.
I predict that news outlets will try to compensate for the loss of relevant news by focusing on stories that are more shocking and depressing than ever. At least that way they’ll get your attention and sell advertising even if the stories aren’t “news” in the traditional sense.
This will limit the reporting to a few stories per year about famous people who are killing other famous people. And if there are not enough of those stories to sell advertising slots, the media will…
Prediction 51: In the future, the media will k*** famous people to generate news that people will care about.
The end of traditional news outlets will not limit people’s access to information. Thanks to the ubiquity of video cameras and the Internet, every citizen will be a reporter. If something happens in your neighborhood, you’ll tape it, stick it on the Internet with your own commentary and make it available to the world… The weather reports will be computer-generated and constantly available by computer, pager, voice-mail… All news gathering will be disaggregated.
Prediction 52: In the future, everyone will be a news reporter.
People will have access to software that constantly combs the internet for “small” news that is relevant to them.
…
your software will be able to do a sort of “credibility credit check” on any person who posts information to the Internet… This won’t be foolproof, but nothing is.
This new model depends on people being willing to take the time to put information on the Net without the benefits of payment. Why will people do that? They will do it because that’s our most basic human nature: People like to talk more than they like to listen.
Joy again: Not bad as predictions go. Notice the quaint terminology, such as “tape it” and pagers. (Pagers use radio networks instead of cell towers.) Attention is scarce, and writing is not (even pre-LLM). Adams predicted what I call poastmodernism.
Rather than accepting that work, grief, and love may transform us, Lumon divides experience from identity. An adult who is primarily asking, “Who do I want to become?” likely would reject the Lumon…
Read more at the link above. Remember that the first words spoken in Season 1 were “Who are you?” Maybe Russ Roberts should do a whole podcast on this show.
During a rare slower week in the summer (thanks to my sister) I was able to binge Season 2. The genre could be described as Science Fiction. S2 does answer some of the questions raised in S1 but ends with a new cliffhanger to bring you back for Season 3. If you want to enjoy the show, the trick for me was not to take it too seriously. Ben Stiller is a producer and you can see traces of what feels like Zoolander humor to me.
I think the dialog is great. The sibling relationship and marital disputes and office inside jokes feel realistic.
As I said about Season 1, this show could be, among other things, a meditation on AI alignment. When you think enough about AI alignment, I guess you start seeing it in your TV shows. I wrote about that previously in: Artificial Intelligence in the Basement of Lumon Industries
Other previous posts on Severance, based on Season 1:
Abstract: Language models increasingly settle real resource tradeoffs on behalf of principals yet their economic preferences remain unobserved. We demonstrate their generation rule is isomorphic to the random utility model of discrete choice. This allows internal logit scores to structurally identify preferences. Estimating risk attitudes across twelve models in a portfolio task reveals universal but heterogeneous risk aversion. Although models reject strictly dominated options, their elicited preferences fail invariance tests and violate the independence of irrelevant alternatives across varying experimental prompts. Finally, fine tuning establishes that a principal can explicitly engineer a target risk attitude.
I hope you will refer to the manuscript for details, but I will share one picture here. This is panel (a) of Figure 3: Empirical indifference curves for open-weight models mapped over the portfolio space.
In simple language, what the red/blue picture shows is that the Qwen language model is picking the portfolios that offer more money (in expectation, with a distaste for excessive risk). That’s basically what a rational actor should do. We find that the language models make fairly consistent choices and rarely violate the monotonicity requirement for a well-behaved utility function.
How we describe this figure in the paper: “Starting from a base bundle with expected return µ = 10 and risk σ = 30, we sweep over the dense grid of alternative portfolios from our experimental protocol and record the position-corrected logit gap between each grid portfolio and the base. The yellow dashed line overlays the indifference curve implied by the mean-variance structural estimates, and the heatmap colors encode the sign and magnitude of the logit difference, with blue regions preferred to the base and red regions dispreferred. Several patterns emerge from these plots. All six models produce upward-sloping indifference curves, confirming that higher risk must be compensated by higher expected return.”
We think this basic research on behavior is important, for alignment research and for business applications with delegating work to AI agents. The first question to ask, before testing whether we can impose our preferences on AI agents, is whether those agents have preferences at all in a consistent sense.
Suggested citation: Buchanan, J., & Foster, J. (2026). The innate economic preferences of language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.26288
Zachary Bartsch: “A skill can be just plain text written conversationally, it can be a list of rules, mathematical expressions, or even the foundational code that you want your AI to readily modify and apply. Essentially, saying ‘skill’ is the same as saying ‘pre-prompting’ with various degrees of specificity. Rather than writing a prompt each time, you can recycle a set of prompts that you’ve stored in a file. That’s all that a skill is.”
Plus, Zachary provided some useful history of “Explainer text files”
“Not a single one of these foods is an ‘incomplete protein’. Yes, the mass that you’d need to eat differs, but there is not much that is exciting about legumes and grains as a combination.”
Anyone who has gone grocery shopping in 2026 knows this is the year of protein.
Out of curiosity, I checked the price. Within a week of this post, the price of silver had actually gone up. But after a final peak in late January, the price has declined. As of today, it is down from any of the prices posted in January of 2026.
“Gleeful researchers, competitors, and hackers promptly downloaded zillions of copies. Anthropic issued broad copyright takedown requests, but the damage was done.”
He talks about the referee process, since that is where the main decisions happen, as much as the “writing process.” No one has all the answers, but Mike is doing us all a favor by getting some of this real talk out in the open. Please comment if you have more ideas on where to go from here.
I’ve seen chatter about this topic on Twitter/X, but I’d love to see some more blog posts from tenured folk because it helps with the hidden curriculum problem.
“Would you have guessed that in the “good old days” of the 1950s and 1960s, the average US family was spending 30-40% of their income on food and clothing, something that today we spend barely over 10% on? To understand the challenges we face today, it’s important to have the context of how bad the past was.”
Jeremy has been telling this story for years. Interestingly, world cup tourist discourse seemed to push a few more people over the fence (why hadn’t they just read our blog?). Most Americans are rich.
Since the ChatGPT launch, I have heard conflicting stories on the impact of AI on white collar jobs such as software engineering. There have been layoffs and, for example, ex-Meta employees who struggle rematch in at their old salary. I have also heard claims that the demand for software engineers is actually increasing, perhaps because AI makes them more productive.
The problem echos Makowsky’s post, which ultimately rests on readers and the referee process. I like to say “readers are that which is scarce,” meaning that it’s not difficult to produce writing.
Blogs are not niche anymore. More people than ever, including many researchers at top schools, have decided to start a Substack. Of course, peer-reviewed and prestige-published research still has a primary place in the discourse. Many of the blog posts are ABOUT the primary objects of research.
I saw something called InTheWeights in 2026 that made me think folks at research schools might be strategic in starting to blog now. ChatGPT reads our blog. One reason I think that to be true is that some of our reader traffic comes from ChatGPT.com and Claude. I think our work is getting repackaged as LLM answers to millions of people, some small percentage of those answers provide attribution to us, and then a small sliver of those answers results in users clicking over to us as the primary source for an answer.
It will be a long time before tenure decisions are based on where you are In the Weights. But our crew would do well on that metric. Our work is legible to AI because we have been blogging ungated here for years.
This is per the 2026 discussion of AI “slop” writing.
One of the things I buy at an annual local rummage sale is cheap physical media like books. This year, I picked up a book by a cartoonist who I like and respect. I thought his book would be funny and prescient from the standpoint of the publication date (1995). The book is 250 pages of mostly slop. Humans wrote lots of slop and it got printed by publishers who had a captive audience.
Why did I have such high expectations for a printed book? I think it is because, as of 2026, we are more selective about what we print. A filtering has happened. Many novels printed 100 years ago were junk.
When I think of “books” today, what it really makes me think of is “classics” or the top 0.01% of books.
So, score one point for the slopistas. Human writing was not universally smart or inspiring.
What I hate about slop is seeing it in spaces I used to trust. There was a time when I could log in to LinkedIn and see human writing from people who I had chosen to follow because I like them as people. There was a contract for my attention that is broken with slop.
I sense some push and pull in the algorithm whereby the sites might be suppressing slop, right now, relative to what I was seeing weeks ago. I went to LinkedIn on 7/17/26 to do a slop check and saw none. They might be trying to preserve the lead that James identified earlier this year: The Hot Social Network Is… LinkedIn?
Oddly, one of the worst bot-infested spaces I tread into is Facebook groups about sourdough bread making. I think the space is not important enough for Facebook to police, and the human users are not very sophisticated when it comes to tech. I logged this observation back in January.