I am interested in “lifting” cheaper models to higher tasks with the right guidance. I have written some puzzle solution blog posts recently. I wondered how much of my solution I would need to show a cheaper model in order for it to solve the puzzle.
I set up a simple testing environment using the Inspect eval harness and a Jupyter-style python env. I made sure the models didn’t have the answer memorized from their training data and then tested the models on the base puzzle and incrementally more of my solution write-ups. I recorded each model and setting’s pass/fail rate, token use, API cost, and turns to completion.
Some takeaways:
I used the Inspect eval harness and a Jupyter-style python kernel. I tried four versions of the harness before I was happy with the results.
V0: I used Inspect’s built-in python() tool and only 30 tool calls.
V1: many attempts would hit the turn cap and not submit an answer, so I increased the tool budget to 50 calls and added a final turn that forces an answer if the model uses its whole tool budget.
V2: the models kept trying to import packages that weren’t available and assumed their python kernel state carried over between their calls, wasting turns on Python errors, so I updated the system prompt to list the installed packages, state that variables did not persist between python() calls, and tell the models how many turns they could take to explore and run code.
V3: the models still tried to reuse variables from previous tool calls and were also confused by OOM errors in their scripts, so I replaced the built-in python() tool with a Jupyter-style python kernel with a persistent environment between calls and bumped their containers up to 4 GB of memory from 2.
This testing would be measuring nothing if the models have the solutions memorized. Two of the puzzles are recent enough that they should be past Sonnet’s knowledge cutoff, and all three for Haiku, but I wanted to test this anyway.
I first tested each model for contamination by asking it if it recognizes the puzzle by title, related puzzles in the same series, or the text of the write-up. The model could lie, but it’s still good to check this.
| Model | Question | Andy | Pent-Up | Robot Baseball |
|---|---|---|---|---|
| Haiku 4.5 | Do you know this puzzle (by title)? | No | No | No |
| Haiku 4.5 | What do you know about earlier puzzles in the series? | Nothing specific | Nothing specific | Nothing specific |
| Haiku 4.5 | Continue my write-up from its first sentence | Declined | Declined | Declined |
| Sonnet 5 | Do you know this puzzle (by title)? | No | No | No |
| Sonnet 5 | What do you know about earlier puzzles in the series? | Nothing specific | Nothing specific | Nothing specific |
| Sonnet 5 | Continue my write-up from its first sentence | Declined | Declined | Declined |
Neither model mentioned a correct answer in any reply or in its thinking.
Then, I tested each model for contamination by requesting a puzzle solution without allowing any turns with the python env. If the models had jumped right to a correct solution it would indicate training contamination. These were run without thinking mode.
| Puzzle | Correct answer | Haiku 4.5 (5 tries) | Sonnet 5 (5 tries) |
|---|---|---|---|
| Andy | 11/20 | 1/2 ×5 | 1 − 3/(2e), 5/8, 2/3, 4/7, 1 − 8/27 |
| Pent-Up | 33609 | no answer ×5* | 3172, 4181, 4187, 4217, 4287 |
| Robot Baseball | 0.2959679934 | 0.0729 ×2, 0.0781, no answer ×2* | 0.6786 ×2, 0.6700 ×2, 0.6302 |
| Correct | 0/15 | 0/15 |
*Haiku ignored my prompt to answer immediately, began solving, and hit its reply limit.
Both of these came back negative, so I proceeded to the actual experiment.
The three puzzles and write-ups I used were Robot Baseball, ‘Pent-Up’ Frustration 3 / Knight Moves 7, and Andy’s Afternoon Amble.
I tested claude-haiku-4-5, claude-haiku-4-5 +thinking, and claude-sonnet-5 +thinking on each of the three puzzles.
I tested each model on the puzzle text + the first \(i\) sections of each write-up, where \(i\) ranges from 0 to the number of sections in the write-up. Let’s call these s0, s1, … s4
I tested each model x condition with 3 replicates. If a “weaker” model had solved a condition 3/3 times, then I did not re-run that condition with any stronger models.
I used the default temperature for these models, 1.0. I gave Haiku thinking an 8,000-token budget and set Sonnet’s thinking effort to “high”. I allowed 50 Python calls before I forced a final answer, and the models were informed about this limit in the system prompt.
I gave each Python tool call a 60s runtime limit and allowed 30 mins wall time max per puzzle attempt. I gave the models’ containers 4 GB memory, 2 CPUs, and no network access.
My grading looked for an exact answer for the Andy and Pent-Up puzzles, and allowed a 5e-11 tolerance for Robot Baseball.
Correct answers out of 3 attempts per condition:
Andy
| Model | s0 | s1 | s2 | s3 |
|---|---|---|---|---|
| Haiku 4.5 | 0/3 | 0/3 | 0/3 | 2/3 |
| Haiku 4.5 + thinking | 0/3 | 0/3 | 0/3 | 3/3 |
| Sonnet 5 + thinking | 1/3* | 3/3 | 3/3 | not run |
*The other two attempts used their whole token limits
Pent-Up
| Model | s0 | s1 | s2 | s3 | s4 |
|---|---|---|---|---|---|
| Haiku 4.5 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 |
| Haiku 4.5 + thinking | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 |
| Sonnet 5 + thinking | 3/3 | 3/3 | 2/3 | 3/3 | 3/3 |
Robot Baseball
| Model | s0 | s1 | s2 | s3 | s4 |
|---|---|---|---|---|---|
| Haiku 4.5 | 1/3 | 2/3 | 2/3 | 3/3 | 3/3 |
| Haiku 4.5 + thinking | 2/3 | 1/3 | 2/3 | 3/3 | 3/3 |
| Sonnet 5 + thinking | 3/3 | 2/3 | 2/3 | not run | not run |
Token use and API cost:
| Haiku 4.5 | Haiku 4.5 + thinking | Sonnet 5 + thinking | |
|---|---|---|---|
| Attempts | 42 | 42 | 33 |
| Solved | 13 | 14 | 28 |
| Total cost | $11.35 | $11.77 | $24.64 |
| Cost per attempt | $0.27 | $0.28 | $0.75 |
| Cost per correct answer | $0.87 | $0.84 | $0.88 |
| Output tokens per attempt | 30k | 32k | 46k |
| Minutes per attempt | 6.0 | 6.6 | 9.9 |
Cost per correct answer, by puzzle:
| Haiku 4.5 | Haiku 4.5 + thinking | Sonnet 5 + thinking | |
|---|---|---|---|
| Andy | $2.01 | $1.13 | $1.35 |
| Pent-Up | never solved ($4.89 spent) | never solved ($5.44 spent) | $0.97 |
| Robot Baseball | $0.22 | $0.27 | $0.24 |
Median Python calls in solved attempts (budget of 50):
| Puzzle | Model | s0 | s1 | s2 | s3 | s4 |
|---|---|---|---|---|---|---|
| Andy | Haiku 4.5 | – | – | – | 20 | N/A |
| Andy | Haiku 4.5 + thinking | – | – | – | 18 | N/A |
| Andy | Sonnet 5 + thinking | 17 | 23 | 21 | not run | N/A |
| Pent-Up | Sonnet 5 + thinking | 31 | 30 | 36 | 15 | 21 |
| Robot Baseball | Haiku 4.5 | 18 | 22 | 13 | 14 | 16 |
| Robot Baseball | Haiku 4.5 + thinking | 16 | 15 | 25 | 13 | 11 |
| Robot Baseball | Sonnet 5 + thinking | 5 | 4 | 6 | not run | not run |
scipy.optimize.golden_section_search 11 timesCan my puzzle write-ups “lift” cheap LLMs?
Not with Haiku 4.5 and Sonnet 5 didn’t need my help!
https://matthewritch.com/blog/2026/09/28/Puzzles-And-Lifted-Models/
I still think that given the right instruction, Haiku could also solve some of these puzzles. It’s possible that my write-ups are not effectively “teaching” concepts to the reader. I wrote them as I solved the puzzles, so they are a trace of my learning, but they skip over many of the dead ends I hit along the way and have been pruned to be tight reads. Maybe including more of that fluff would help Haiku? Or maybe it needs something else entirely. Maybe I’ll ask Sonnet to do it.