Temperature vs thinking level for LLM products

If using ChatGPT or an API in 2023, you might have followed the advice “turn the temperature down” when you wanted a more serious answer. That is going out of date. The following is a collaboration between me and Grok to lay this out simply:  

What temperature was

Imagine the model is always choosing the next word from a long list of options, ranked from “very likely” to “weird but possible.” Temperature changes how adventurous that choice is.

  • Low temperature: almost always pick the obvious next word. Answers feel steady, repetitive, “corporate.” Good when there is basically one right output, like extracting a date from a contract.
  • High temperature: more willing to pick a less obvious word. Answers feel livelier but sometimes incorrect.

On older products you could often set this yourself.

What thinking level is

Newer models (Gemini 3, OpenAI’s reasoning models) do extra work before they talk to you. They draft a private scratchpad: break the problem into steps, check themselves, then write the answer you see. You don’t see the scratchpad even though you pay for it. That hidden draft is the “thinking.”

Thinking level (sometimes called reasoning effort) is not “how random should the next word be?” It is “how long is the model allowed to work on that draft?”

  • Low / minimal: answer quickly and cheaply. Fine for “what’s the status of ticket 1842?”
  • High: spend more time (and more tokens) on hard things — a gnarly spreadsheet, a multi-step plan, a tricky piece of code.
Continue reading

Nicholas Polson Has Written Over 200 Academic Papers in 2026 (so far)

UPDATE: I stupidly didn’t realize my co-blogger Mike wrote about this too. My bad! I skimmed last week’s posts, but it didn’t click in my head. Be sure to read his thoughts.

ALSO: lots of the papers by Polson seem to have vanished from SSRN since I wrote this blog post… yesterday. Everything after August 14th has been taken down. That brings his count of papers in 2026 down to a mere 168 papers. Still essentially several lifetimes of output from a typical academic.

CODA: As Andrew Gelman documented in real time, all of the papers appear to have been removed from SSRN for now. No explanation as of yet, but you can still see many of the papers listed (for now) on Polson’s Google Scholar page.

For most academics, writing papers that may be eventually published in peer-reviewed journals is an important part of what we do. For some academics, it is the main thing they do (others have more emphasis on teaching courses at their university). Most academics always have a few projects they are working on, with perhaps a goal of finishing 2 or 3 a year, and thus having a regular pipeline of a few publications every few years. But some academics are much more prolific.

Take Daron Acemoglu for example. He has long been considered extremely prolific. So far in 2026, according to his Google Scholar page, he has had 15 papers that have either been published this year or come out as new working papers (that’s the bulk of them). In the past year (2025 and 2026), he has had publications in the Quarterly Journal of Economics, the American Economic Review, and the Journal of Economic Literature, among others (the AER paper was his Nobel lecture). For lower tier academics, that’s almost a lifetime of publications in 12 months or so. Acemoglu is extremely productive.

But I just discovered an economist that is, apparently, even more productive than Acemoglu, at least as measured by working papers. Nicholas Polson has, by my count using his SSRN page, already written 258 working papers in 2026 alone. He’s already written (or at least published to SSRN), six papers today, August 26, 2026. In the month of August 2026, he has written and posted to SSRN a total of 104 papers — and counting, since the month isn’t quite over.

These aren’t just short notes. Most of the papers are of normal academic length: 32 pages, 27 pages, 58 pages. The papers are both theoretical — involving complex math in some cases — or empirical, with regressions. Read any single paper, and it feels like just a normal academic paper, the kind of thing that an academic might work on for a few months. He even has a frequent co-author, which is common for economics papers (and helps to be more productive), a systems engineering professor named Vadim Sokolov, who is a co-author on a little over 100 of the papers this year.

What is going on here? Obviously the research productivity of Polson and his co-author Sokolov is aided by AI. Who isn’t using AI to increase their writing and research productivity these days? But I don’t think I have seen any academic, at least not in economics, that has really pushed it to the limit.

Presumably, many of these papers will get submitted to academic journals. I can imagine the editors of journals have a very hard job these days, as the number of papers submitted has likely increased significantly, while the time that referees have available has not increased much (of course, AI is likely making referees more productive too, though many journals ask you not to upload the paper into an AI program as a referee, since it is unpublished work when you are reviewing it).

I really don’t know where academic publishing goes from here. AI has made us all more productive in terms of output, but are we better at answering important questions in our science than pre-AI? Probably, though it is hard to know. Journals and the peer review process has traditionally been the filter to sort real contributions from gibberish. I don’t know how the peer review process continues in its current form given the massive increase in output (much of it good!) that we are seeing from academics. Dr. Polson is just a leading example of a growing challenge for academia.

Intellectual Squatting

So a professor at a major institution wrote 200 papers last year. Unlike other commenters I’ve observed so far, I think this is neither true research nor pure AI fraud. It is likely AI “slop” to varying degrees, but unlike a lot of slop there is probably real value within it. What I have not yet seen ascertained is whether any of it has been vetted, investigated, or curated by the author in a meaningful way. The real question is: what is the actual ambition here? I think the tell is the lack of submission to peer review.

I think this is a form of intellectual squatting. The nice version is it’s putting out a series of half-baked papers in the hopes of establishing a property right to the underlying ideas at an earlier stage of the research process than previously possible. The less generous interpretation is it’s dumping a series of haystacks on the plains and laying claim to the needles probabilistically within each. Imagine you are a person who has highly esoteric, potentially important ideas every day. Many of those ideas you suspect, based on some combination of experience and ego, are new in at least one dimension. You would like to get credit for that newness. For being first. What’s the problem?

The problem is that scholarship remains more perspiration than inspiration. Having a new idea is great, but it takes years to work through the nuance in sufficient detail that you can convince your peers of the coherence and originality of the contribution. During the minutes each day you are not working on this singular project you have the inspiration for other ideas, sometimes multiple within a single day. How frustrating is the proposition that someone else gets credit for the originality of contribution just because they had time to reveal it to the world while you were embroiled in your investigation of what is only one of your many score ideas!?

Ah, but meta-level inspiration has struck you! What if you took each one of those ideas, spent an hour curating a series of prompts around it, and then let Chat GPT (or another LLM) fabricate an entire research paper around it? It might not be good, correct, or even coherent, but it does somethine far more important. It establishes an intellectual property right to the claim of being first. Now, to be clear, you are fully aware of the deficiciency of your paper as an actual scholarly contribution, but if somone else writes a full paper you at least have something to point to and say “I was here first. Cite me. Hell, if I’m close enough you might even have to name it after me. Well, sure, us. But definitely include me. Glory shared via hypenhnation is better than no glory at all.”

Is it a contribution? That’s something that will vary on a case-by-case basis, but I expect far more misses than hits. The work isn’t there. It’s like plopping down a block of marble with a dramatic-ish sketch of a man on an adhered post-it note and claiming that Michaelangelo needs to share credit with you on any subsequent sculptures. It’s like asking people to cite that one cool tweet you did about how DNA is cool but maybe RNA could be useful in vaccines one day. Intellectual property rights trolling via AI blunderbuss.

BTW, I’m not 100% sure this isn’t an AI take on a modern Sokal hoax. An attempt to show how much AI slop is introducing a whole new version of Gresham’s Law to scholarship. But if we treat it as earnest, it’s proof that a very smart person can potentially disrupt the market for scholarship, patents, or any other intellectual property by laying claim to ideas in much the same way that the printing press undermined the market for plenary indulges. Flood the market, leave it to someone else to sort through the ecumenical consequences.

Boy Wonder Leopold Aschenbrenner Blows Up His $45 Billion Situational Awareness Hedge Fund

Leopold Aschenbrenner is a very bright guy. Born in Germany to physician parents, he skipped enough grades to graduate from high school at age 15, allowing him to enroll at Columbia University in 2017 at that same age. He went on to graduate from Columbia at age 19 as valedictorian with a degree in economics and mathematics-statistics. An econ prof at the time said his “record of scholarship exceeds that of any student in the department in the previous 20 years.” The Mercatus Center, a think tank at George Mason University, recognized his potential, awarding him an Emergent Ventures grant.

Leopold Aschenbrenner, Columbia Class of 2021 Valedictorian

After graduation, Aschenbrenner moved into the effective-altruism research and grantmaking world, making notable contributions in various ways. In 2023, he joined the newly created “Superalignment” team at OpenAI, that was charged with making conceptual and engineering progress on aligning systems “much smarter than humans” before such systems were built. The next year he was discharged by OpenAI; the company said it was because of a security leak, but his version (which I find more credible) is that he was ousted as retaliation for circulating an internal memo arguing that the company’s protections against model-weight theft and algorithmic exfiltration were inadequate for AGI-relevant work.

Two months after his departure from OpenAI, he self-published Situational Awareness: The Decade Ahead, which you can download here.    This monograph synthesized scaling-law extrapolations, geopolitical analysis, and AI-lab security commentary into a single forecast: that the largest AI labs, on current trends, will plausibly reach AGI around 2027 and that an intelligence explosion to superintelligence could occur in the subsequent few years. This work went viral on Wall Street, and before you can say “monetization”, he was leading a hedge fund named, appropriately enough, Situational Awareness. There he put into practice his convictions that the demand for compute would be voracious, and would be limited by physical constraints such as electricity and chip fabs.

Thus, in his fund he went long companies like SanDisk (memory fab), Bloom Energy (makes solid oxide fuel cells), and Nebius (builds whole data centers). Very long, in fact, with leverage reportedly as high as 400%. He tried to hedge this long book by shorting software companies which are viewed as vulnerable to disruption by AI. This strategy worked fabulously for a while. His assets under management (AUM) climbed to $45 billion, with returns in the first 6-7 months of 2026 approaching 400%. Not bad for a 25-year-old.

But then, genius failed (yet again)- -in a stunning reversal, the market rebelled in late July against big AI capex spends, dumping memory fabs and infrastructure, and bought into the maligned software (SaaS) sector. So, BOTH his long and short legs went against him, followed in due course by the dreaded margin calls. Aschenbrenner’s $45 billion shrank to a measly $10 billion in a matter of days, as he was forced to sell off his public equities at a discount to those friendly capitalist sharks at Citadel. Wise old heads wagged, saying, yup, this sort of Black Swan event always happens sooner or later, and you then get carried out on a stretcher if you run a highly leveraged bet that is not truly hedged.

Down, but not out – – The word on the Street is that folks with money to invest are already lining up to entrust more bazillions to our altruistic wizard. We have not heard the last of Leopold Aschenbrenner.

AI Innate Preferences Paper on Arxiv

Please check out my new paper, with Joshua Foster

The Innate Economic Preferences of Language Models (arXiv link)

Abstract: Language models increasingly settle real resource tradeoffs on behalf of principals yet their economic preferences remain unobserved. We demonstrate their generation rule is isomorphic to the random utility model of discrete choice. This allows internal logit scores to structurally identify preferences. Estimating risk attitudes across twelve models in a portfolio task reveals universal but heterogeneous risk aversion. Although models reject strictly dominated options, their elicited preferences fail invariance tests and violate the independence of irrelevant alternatives across varying experimental prompts. Finally, fine tuning establishes that a principal can explicitly engineer a target risk attitude.

I hope you will refer to the manuscript for details, but I will share one picture here. This is panel (a) of Figure 3: Empirical indifference curves for open-weight models mapped over the portfolio space.

In simple language, what the red/blue picture shows is that the Qwen language model is picking the portfolios that offer more money (in expectation, with a distaste for excessive risk). That’s basically what a rational actor should do. We find that the language models make fairly consistent choices and rarely violate the monotonicity requirement for a well-behaved utility function.

How we describe this figure in the paper: “Starting from a base bundle with expected return µ = 10 and risk σ = 30, we sweep over the dense grid of alternative portfolios from our experimental protocol and record the position-corrected logit gap between each grid portfolio and the base. The yellow dashed line overlays the indifference curve implied by the mean-variance structural estimates, and the heatmap colors encode the sign and magnitude of the logit difference, with blue regions preferred to the base and red regions dispreferred. Several patterns emerge from these plots. All six models produce upward-sloping indifference curves, confirming that higher risk must be compensated by higher expected return.”

We think this basic research on behavior is important, for alignment research and for business applications with delegating work to AI agents. The first question to ask, before testing whether we can impose our preferences on AI agents, is whether those agents have preferences at all in a consistent sense.

Suggested citation: Buchanan, J., & Foster, J. (2026). The innate economic preferences of language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.26288

The Academic Data Project That Turned Into $375 Million

What could be better than creating data so valuable that an institution is happy to host and update it forever, like the Sean Lahman baseball database?

Creating data that sells for $375 million, like the Center for Research in Security Prices. University of Chicago professors assembled this series of finance datasets over decades, starting in 1960 with an effort to track every transaction of every publicly traded security. U Chicago sold CRSP to Morningstar last year for $375 million.

Why could they sell it for so much? It helps to be working in finance, where the willingness to pay is the highest. It also represents 65 years of work from what became a large team that included Nobelists like Eugene Fama. The data was valuable enough to become widely used by key institutions even though CRSP charged for it:

Today, $3 trillion in fund assets are linked to CRSP Market Indexes, including U.S. equity ETFs run by Vanguard, and more than 600 subscribers across 35 countries use CRSP Research Data Products.  

Did U Chicago sell CRSP at the right time? On the one hand, I wonder if this was a fire sale driven by federal grant cuts putting pressure on the U Chicago budget. On the other hand, assembling datasets like this is only going to get easier in the age of AI, so perhaps Chicago sold at the top.

For now though there is still an edge in having restricted datasets that AIs haven’t trained on and can’t access. When I ask myself what advantage my human research assistants have over AIs in 2026, the most obvious answer is that they can legally access restricted databases like CRSP or, in my current case, HeinOnline.

Fable on Legibility

Claude Fable is Anthropic’s most capable publicly available “Mythos-class” model. It is optimized for long-running autonomous tasks and deep knowledge work. The roll out of this product has been dramatic. Little people like me have access to it for only two weeks, and I doubt I will be able to afford it thereafter. With my window of access, I posed it the following prompt:

“This article indicates that a much smarter model might not be possible because the universe is opaque. Since Tyler Cowen made this statement, AI has helped people make breakthroughs in math and biology. Write this again in 2026 using the latest state of technology. Is the smartest LLM today much smarter than GPT-4? And is the universe legible?” and I copied in my blog post Is the Universe Legible to Intelligence?

Fable replied in 3 pages of text which you can download as a Word doc here.

One line from Fable’s response: “Notice where the wins clustered: mathematics with checkable proofs, protein structures with experimental ground truth, contest problems with known answers. These are the maximally legible domains — places with a fixed target and a way to verify that you hit it.”

The writing is coherent and contains no obvious hallucinations. Is the answer true, and does it tell the whole truth?

Whenever you read something think about who wrote it, and keep in mind that every author/model has a bias and limitations. In my paper with Will Hickman, we found that just reminding people that a paragraph has an author (whether the author is human or AI) increased the demand for fact checking from readers. LLMs will become more persuasive and closer to (but never completely) correct. Keep reading all things with some skepticism whether they are written by scientists, politicians, or AI.

Regardless, there is definitely such a thing as making the known world more legible to AI today. Thus, people are talking about increasing funding for data availability and the possible demise of the “research paper.”

Research papers are more like stories than facts. The demand for stories is not going away, but I definitely cannot predict the future of the write-for-pay scientist.

AI, potentially, could go and get its own new data, instead of waiting for humans to archive it. Thus, the self-improving AI might take us beyond the current models… unless they run up against something that is not legible to intelligence…

Expressionism Is To Cameras As…. Caity Weaver?…. Is To Writing?

I don’t think it’s a coincidence that movements like expressionism, impressionism, and abstract art took off after the invention of the camera. Photorealistic paintings are impressive, but once they are duplicating what a camera does, they’re less interesting.

We’re due for similar movements in other fields to emerge as reactions to AI. Like writing in a way totally different from how an AI would write- ideally better than an AI would write, but even writing worse than an AI can be interesting if it is at least different.

It’s still early days for both AI and our reactions to it. But since the release of ChatGPT in 2022, what is the good new essay or book that you’re most confident was not written by AI, one that was written in an almost deliberately extra-human manner?

Continue reading