Matthew Ritch

About Projects Fun Blog
6 October 2026

Slop on Sale


Recently, I have been thinking about how the “tone” or “style” of an LLM’s output is almost as important as the accuracy. Take these two proofs that the sum of two even numbers is even.

A:

An even number can be written as \(2x\), where \(x\) is any integer.

Take two even numbers, \(2y\) and \(2z\). Their sum is \(2y + 2z = 2(y + z).\)

\(y + z\) is an integer, so \(2y + 2z\) is even.

B:

What a delightful question — let’s delve in! At its heart, evenness is a beautifully simple yet profound idea: an integer \(n\) is even precisely when \(n = 2k\) for some \(k \in \mathbb{Z}\). Now, consider two arbitrary even integers, elegantly expressed as \(m = 2a\) and \(n = 2b\). Here’s where the magic happens — summing them and invoking the ever-reliable distributive property, we find

\[m + n = 2a + 2b = 2(a + b).\]

Crucially, since the integers are gracefully closed under addition, \(a + b\) is itself an integer. And so, with remarkable economy, the sum is — without exception — even! ✨

These are both true, but one of them is much easier to follow. Of course, that is an extremely simple example. For a more involved case, take a look at OpenAI’s recent Navier-Stokes result. Is sloppy exposition of any use? Maybe to bots, but not to the rest of us!

A few weeks ago, I ran an analysis of gemma-2b-it’s responses to stylistic instructions such as “write in plain English” and “write in a friendly tone”. To see how these style prompts work in the real world, we need to run them on some larger models. To test this, I paired each IFEval eval prompt with each of a set of style prompts and measured the effects on accuracy and a number of stylometrics which measure readability.

I ran each combination of settings on both Claude Haiku 4.5 and GPT-5 Mini to compare cheap models from the two big labs.

Takeaways

  • Haiku 4.5 writes more slop by default.
  • Haiku 4.5 responds much more strongly to style prompts than GPT-5 Mini does.
  • Haiku 4.5 was slightly more accurate without any style prompting, but GPT-5 Mini was slightly more accurate with style prompting.
  • Adding style prompts made both Haiku 4.5’s and GPT-5 Mini’s less accurate.

Experiment

I tested claude-haiku-4-5 (no thinking) and gpt-5-mini (reasoning_effort minimal) at their defaults, temperature = 1.

We evaluated the two models’ performance on each combination of style prompt x IFEval prompt. We also included a none setting where the models were prompted without any extra text and placebo settings where the models were prompted with text like Answer the following question. We computed results a la Miller 2024, with paired differences between the two models for each style x IFEval prompt.

Metric Category Simple Definition Source
Words Response Length n words in response —
Syllables / word Word Length   Flesch 1948
Characters / word Word Length   —
% long words Word Length % words with 3+ syllables Gunning 1952
log(words / sentence) Sentence Complexity Log of average sentence length Flesch 1948
Commas / 100 words Sentence Complexity   —
Mean Zipf frequency Vocabulary Rarity How common are the model’s words? van Heuven et al. 2014, wordfreq
% rare tokens (Zipf < 3) Vocabulary Rarity % of words beyond some rarity wordfreq
MTLD Word Variety Word diversity robust to length McCarthy & Jarvis 2010
Adj + adv / 100 words Floridity   spaCy POS tags
Kobak style words / 1000 Floridity Rate of “slop” words Kobak et al. 2024

Does GPT-5 Mini or Haiku 4.5 write more “slop” out of the box?

We organized our set of computed stylometrics into 5 categories and compared the two models on each. Haiku 4.5 used longer words and sentences and more florid vocab. GPT-5 Mini used rarer words and a greater variety of vocabulary, but also longer responses. Haiku 4.5 writes more slop by default.

We pooled the none and placebo outputs of both models and aggregated those metrics by their category.

Then, I z-scored and aggregated across those categories to produce a single composite “simplicity” score for both models.

none placebo

Is GPT-5 Mini or Haiku 4.5 more obedient to style instructions?

Haiku 4.5 is overwhelmingly more obedient to style instructions. It followed them better for each of our five categories of stylometric.

GPT-5 Mini Haiku 4.5

Is GPT-5 Mini or Haiku 4.5 more accurate, with and without the style instructions?

Haiku 4.5 was slightly more accurate without any style prompting, but GPT-5 Mini was slightly more accurate with style prompting. IFEval has two scoring modes, so we’ll report both here.

GPT-5 Mini Haiku 4.5

Vocabulary Shifts

While investigating these results, I saw some interesting shifts in the models’ preferred words. We can’t access these models’ logit lens results because they are closed source, but these lexical shifts give us a similar view of the models’ output token distributions.

GPT-5 Mini tends to describe the style it is using, while Haiku just gets right to answering in that style.

Style Prompt Cluster Haiku: more GPT: more Haiku: less GPT: less
none — — — —
plain what, people, work, if simple, plain as, an, such, comprehensive of, your, such, heart
ornate upon, most, resplendent, magnificent into, its, upon, as on, can, about, up can, it’s, use, get
brief — — that, of, a, in that, can, is, a
tokens — — of, that, a, to that, of, the, in
verbose this, by, from, as at, if, an, are — —
tone_formal upon, within, formal, regarding may, such, professional, formal like, just, up, it’s like, it’s, up, so
tone_friendly you, hey, it’s, so you, like, warm, friendly may, including, provide, significant provide, state
direct — — this, a, or, of can, but, so, or
careful review — — —
careless but, actually, just, quick mistakes, intentionally, wrong, fast as, an, comprehensive, development life, g, support, ensure

Conclusion

Although Haiku 4.5 writes more slop by default, it also responds much better when you ask it to clean up its prose. That fix costs a small bit of its accuracy, however. Haiku 4.5 and GPT-5 Mini have quite different default style, prompt-steerability, and side effects from steering, even though they are both relatively cheap. Anthropic and OpenAI don’t publish these behaviors or the models’ weights, so to learn these quirks and use these models properly, you have to measure these things yourself.

A bowl of slop on sale