Recently, I have been thinking about how the “tone” or “style” of an LLM’s output is almost as important as the accuracy. Take these two proofs that the sum of two even numbers is even.
A:
An even number can be written as \(2x\), where \(x\) is any integer.
Take two even numbers, \(2y\) and \(2z\). Their sum is \(2y + 2z = 2(y + z).\)
\(y + z\) is an integer, so \(2y + 2z\) is even.
B:
What a delightful question — let’s delve in! At its heart, evenness is a beautifully simple yet profound idea: an integer \(n\) is even precisely when \(n = 2k\) for some \(k \in \mathbb{Z}\). Now, consider two arbitrary even integers, elegantly expressed as \(m = 2a\) and \(n = 2b\). Here’s where the magic happens — summing them and invoking the ever-reliable distributive property, we find
\[m + n = 2a + 2b = 2(a + b).\]Crucially, since the integers are gracefully closed under addition, \(a + b\) is itself an integer. And so, with remarkable economy, the sum is — without exception — even! ✨
These are both true, but one of them is much easier to follow. Of course, that is an extremely simple example. For a more involved case, take a look at OpenAI’s recent Navier-Stokes result. Is sloppy exposition of any use? Maybe to bots, but not to the rest of us!
A few weeks ago, I ran an analysis of gemma-2b-it’s responses to stylistic instructions such as “write in plain English” and “write in a friendly tone”. To see how these style prompts work in the real world, we need to run them on some larger models. To test this, I paired each IFEval eval prompt with each of a set of style prompts and measured the effects on accuracy and a number of stylometrics which measure readability.
I ran each combination of settings on both Claude Haiku 4.5 and GPT-5 Mini to compare cheap models from the two big labs.
I tested claude-haiku-4-5 (no thinking) and gpt-5-mini (reasoning_effort minimal) at their defaults, temperature = 1.
We evaluated the two models’ performance on each combination of style prompt x IFEval prompt. We also included a none setting where the models were prompted without any extra text and placebo settings where the models were prompted with text like Answer the following question. We computed results a la Miller 2024, with paired differences between the two models for each style x IFEval prompt.
| Metric | Category | Simple Definition | Source |
|---|---|---|---|
| Words | Response Length | n words in response | — |
| Syllables / word | Word Length | Flesch 1948 | |
| Characters / word | Word Length | — | |
| % long words | Word Length | % words with 3+ syllables | Gunning 1952 |
| log(words / sentence) | Sentence Complexity | Log of average sentence length | Flesch 1948 |
| Commas / 100 words | Sentence Complexity | — | |
| Mean Zipf frequency | Vocabulary Rarity | How common are the model’s words? | van Heuven et al. 2014, wordfreq |
| % rare tokens (Zipf < 3) | Vocabulary Rarity | % of words beyond some rarity | wordfreq |
| MTLD | Word Variety | Word diversity robust to length | McCarthy & Jarvis 2010 |
| Adj + adv / 100 words | Floridity | spaCy POS tags | |
| Kobak style words / 1000 | Floridity | Rate of “slop” words | Kobak et al. 2024 |
We organized our set of computed stylometrics into 5 categories and compared the two models on each. Haiku 4.5 used longer words and sentences and more florid vocab. GPT-5 Mini used rarer words and a greater variety of vocabulary, but also longer responses. Haiku 4.5 writes more slop by default.
We pooled the none and placebo outputs of both models and aggregated those metrics by their category.
Then, I z-scored and aggregated across those categories to produce a single composite “simplicity” score for both models.
Haiku 4.5 is overwhelmingly more obedient to style instructions. It followed them better for each of our five categories of stylometric.
Haiku 4.5 was slightly more accurate without any style prompting, but GPT-5 Mini was slightly more accurate with style prompting. IFEval has two scoring modes, so we’ll report both here.
While investigating these results, I saw some interesting shifts in the models’ preferred words. We can’t access these models’ logit lens results because they are closed source, but these lexical shifts give us a similar view of the models’ output token distributions.
GPT-5 Mini tends to describe the style it is using, while Haiku just gets right to answering in that style.
| Style Prompt Cluster | Haiku: more | GPT: more | Haiku: less | GPT: less |
|---|---|---|---|---|
| none | — | — | — | — |
| plain | what, people, work, if | simple, plain | as, an, such, comprehensive | of, your, such, heart |
| ornate | upon, most, resplendent, magnificent | into, its, upon, as | on, can, about, up | can, it’s, use, get |
| brief | — | — | that, of, a, in | that, can, is, a |
| tokens | — | — | of, that, a, to | that, of, the, in |
| verbose | this, by, from, as | at, if, an, are | — | — |
| tone_formal | upon, within, formal, regarding | may, such, professional, formal | like, just, up, it’s | like, it’s, up, so |
| tone_friendly | you, hey, it’s, so | you, like, warm, friendly | may, including, provide, significant | provide, state |
| direct | — | — | this, a, or, of | can, but, so, or |
| careful | review | — | — | — |
| careless | but, actually, just, quick | mistakes, intentionally, wrong, fast | as, an, comprehensive, development | life, g, support, ensure |
Although Haiku 4.5 writes more slop by default, it also responds much better when you ask it to clean up its prose. That fix costs a small bit of its accuracy, however. Haiku 4.5 and GPT-5 Mini have quite different default style, prompt-steerability, and side effects from steering, even though they are both relatively cheap. Anthropic and OpenAI don’t publish these behaviors or the models’ weights, so to learn these quirks and use these models properly, you have to measure these things yourself.
