AI Innate Preferences Paper on Arxiv

Please check out my new paper, with Joshua Foster

The Innate Economic Preferences of Language Models (arXiv link)

Abstract: Language models increasingly settle real resource tradeoffs on behalf of principals yet their economic preferences remain unobserved. We demonstrate their generation rule is isomorphic to the random utility model of discrete choice. This allows internal logit scores to structurally identify preferences. Estimating risk attitudes across twelve models in a portfolio task reveals universal but heterogeneous risk aversion. Although models reject strictly dominated options, their elicited preferences fail invariance tests and violate the independence of irrelevant alternatives across varying experimental prompts. Finally, fine tuning establishes that a principal can explicitly engineer a target risk attitude.

I hope you will refer to the manuscript for details, but I will share one picture here. This is panel (a) of Figure 3: Empirical indifference curves for open-weight models mapped over the portfolio space.

In simple language, what the red/blue picture shows is that the Qwen language model is picking the portfolios that offer more money (in expectation, with a distaste for excessive risk). That’s basically what a rational actor should do. We find that the language models make fairly consistent choices and rarely violate the monotonicity requirement for a well-behaved utility function.

How we describe this figure in the paper: “Starting from a base bundle with expected return µ = 10 and risk σ = 30, we sweep over the dense grid of alternative portfolios from our experimental protocol and record the position-corrected logit gap between each grid portfolio and the base. The yellow dashed line overlays the indifference curve implied by the mean-variance structural estimates, and the heatmap colors encode the sign and magnitude of the logit difference, with blue regions preferred to the base and red regions dispreferred. Several patterns emerge from these plots. All six models produce upward-sloping indifference curves, confirming that higher risk must be compensated by higher expected return.”

We think this basic research on behavior is important, for alignment research and for business applications with delegating work to AI agents. The first question to ask, before testing whether we can impose our preferences on AI agents, is whether those agents have preferences at all in a consistent sense.

Suggested citation: Buchanan, J., & Foster, J. (2026). The innate economic preferences of language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.26288

Leave a comment