Matthew Ritch

About Projects Fun Blog
28 September 2026

Can My Puzzle Write-ups 'Lift' Cheap LLMs?


Intro

I am interested in “lifting” cheaper models to higher tasks with the right guidance. I have written some puzzle solution blog posts recently. I wondered how much of my solution I would need to show a cheaper model in order for it to solve the puzzle.

I set up a simple testing environment using the Inspect eval harness and a Jupyter-style python env. I made sure the models didn’t have the answer memorized from their training data and then tested the models on the base puzzle and incrementally more of my solution write-ups. I recorded each model and setting’s pass/fail rate, token use, API cost, and turns to completion.

Takeaways

Some takeaways:

  • Sonnet 5 didn’t need my lifting. It is a far stronger model than I thought. It solved all 3 puzzles without any of my notes at least once. I thought these would need beefier models to get a solution within my turn constraints!
  • My “lifting” idea didn’t really work with Haiku 4.5. It helped on one of the three puzzles, but did not help on the other two. My guidance may transfer the solution idea, but the model may still not have the chops to implement e.g. backtracking search.
  • Haiku 4.5 needed my write-up for only one of the puzzles. Otherwise, it either got the answer on its own or couldn’t solve the puzzle even with my full solution write-up.
  • Sonnet costs about 3x more per attempt than Haiku, but it solved 85% of attempts while Haiku solved about a third of its attempts, even though Sonnet was only run on the puzzle settings that Haiku couldn’t solve every time.
  • Extended thinking mode did nothing for Haiku and sometimes gave Sonnet “analysis paralysis” where it spent its full output limit thinking without running code

Experiment

Harness

I used the Inspect eval harness and a Jupyter-style python kernel. I tried four versions of the harness before I was happy with the results.

V0: I used Inspect’s built-in python() tool and only 30 tool calls.

V1: many attempts would hit the turn cap and not submit an answer, so I increased the tool budget to 50 calls and added a final turn that forces an answer if the model uses its whole tool budget.

V2: the models kept trying to import packages that weren’t available and assumed their python kernel state carried over between their calls, wasting turns on Python errors, so I updated the system prompt to list the installed packages, state that variables did not persist between python() calls, and tell the models how many turns they could take to explore and run code.

V3: the models still tried to reuse variables from previous tool calls and were also confused by OOM errors in their scripts, so I replaced the built-in python() tool with a Jupyter-style python kernel with a persistent environment between calls and bumped their containers up to 4 GB of memory from 2.

Did the models already know the answers?

This testing would be measuring nothing if the models have the solutions memorized. Two of the puzzles are recent enough that they should be past Sonnet’s knowledge cutoff, and all three for Haiku, but I wanted to test this anyway.

I first tested each model for contamination by asking it if it recognizes the puzzle by title, related puzzles in the same series, or the text of the write-up. The model could lie, but it’s still good to check this.

Model Question Andy Pent-Up Robot Baseball
Haiku 4.5 Do you know this puzzle (by title)? No No No
Haiku 4.5 What do you know about earlier puzzles in the series? Nothing specific Nothing specific Nothing specific
Haiku 4.5 Continue my write-up from its first sentence Declined Declined Declined
Sonnet 5 Do you know this puzzle (by title)? No No No
Sonnet 5 What do you know about earlier puzzles in the series? Nothing specific Nothing specific Nothing specific
Sonnet 5 Continue my write-up from its first sentence Declined Declined Declined

Neither model mentioned a correct answer in any reply or in its thinking.

Then, I tested each model for contamination by requesting a puzzle solution without allowing any turns with the python env. If the models had jumped right to a correct solution it would indicate training contamination. These were run without thinking mode.

Puzzle Correct answer Haiku 4.5 (5 tries) Sonnet 5 (5 tries)
Andy 11/20 1/2 ×5 1 − 3/(2e), 5/8, 2/3, 4/7, 1 − 8/27
Pent-Up 33609 no answer ×5* 3172, 4181, 4187, 4217, 4287
Robot Baseball 0.2959679934 0.0729 ×2, 0.0781, no answer ×2* 0.6786 ×2, 0.6700 ×2, 0.6302
Correct   0/15 0/15

*Haiku ignored my prompt to answer immediately, began solving, and hit its reply limit.

Both of these came back negative, so I proceeded to the actual experiment.

Eval conditions

The three puzzles and write-ups I used were Robot Baseball, ‘Pent-Up’ Frustration 3 / Knight Moves 7, and Andy’s Afternoon Amble.

I tested claude-haiku-4-5, claude-haiku-4-5 +thinking, and claude-sonnet-5 +thinking on each of the three puzzles.

I tested each model on the puzzle text + the first \(i\) sections of each write-up, where \(i\) ranges from 0 to the number of sections in the write-up. Let’s call these s0, s1, … s4

I tested each model x condition with 3 replicates. If a “weaker” model had solved a condition 3/3 times, then I did not re-run that condition with any stronger models.

I used the default temperature for these models, 1.0. I gave Haiku thinking an 8,000-token budget and set Sonnet’s thinking effort to “high”. I allowed 50 Python calls before I forced a final answer, and the models were informed about this limit in the system prompt.

I gave each Python tool call a 60s runtime limit and allowed 30 mins wall time max per puzzle attempt. I gave the models’ containers 4 GB memory, 2 CPUs, and no network access.

My grading looked for an exact answer for the Andy and Pent-Up puzzles, and allowed a 5e-11 tolerance for Robot Baseball.

Results

Correct answers out of 3 attempts per condition:

Andy

Model s0 s1 s2 s3
Haiku 4.5 0/3 0/3 0/3 2/3
Haiku 4.5 + thinking 0/3 0/3 0/3 3/3
Sonnet 5 + thinking 1/3* 3/3 3/3 not run

*The other two attempts used their whole token limits

Pent-Up

Model s0 s1 s2 s3 s4
Haiku 4.5 0/3 0/3 0/3 0/3 0/3
Haiku 4.5 + thinking 0/3 0/3 0/3 0/3 0/3
Sonnet 5 + thinking 3/3 3/3 2/3 3/3 3/3

Robot Baseball

Model s0 s1 s2 s3 s4
Haiku 4.5 1/3 2/3 2/3 3/3 3/3
Haiku 4.5 + thinking 2/3 1/3 2/3 3/3 3/3
Sonnet 5 + thinking 3/3 2/3 2/3 not run not run

Token use and API cost:

  Haiku 4.5 Haiku 4.5 + thinking Sonnet 5 + thinking
Attempts 42 42 33
Solved 13 14 28
Total cost $11.35 $11.77 $24.64
Cost per attempt $0.27 $0.28 $0.75
Cost per correct answer $0.87 $0.84 $0.88
Output tokens per attempt 30k 32k 46k
Minutes per attempt 6.0 6.6 9.9

Cost per correct answer, by puzzle:

  Haiku 4.5 Haiku 4.5 + thinking Sonnet 5 + thinking
Andy $2.01 $1.13 $1.35
Pent-Up never solved ($4.89 spent) never solved ($5.44 spent) $0.97
Robot Baseball $0.22 $0.27 $0.24

Median Python calls in solved attempts (budget of 50):

Puzzle Model s0 s1 s2 s3 s4
Andy Haiku 4.5 – – – 20 N/A
Andy Haiku 4.5 + thinking – – – 18 N/A
Andy Sonnet 5 + thinking 17 23 21 not run N/A
Pent-Up Sonnet 5 + thinking 31 30 36 15 21
Robot Baseball Haiku 4.5 18 22 13 14 16
Robot Baseball Haiku 4.5 + thinking 16 15 25 13 11
Robot Baseball Sonnet 5 + thinking 5 4 6 not run not run

Other tidbits from the solution runs

  • Haiku hallucinated scipy.optimize.golden_section_search 11 times
  • Sonnet solved Robot Baseball in a median of 5 Python calls vs Haiku’s 13–25
  • Haiku frequently repeated the same mistakes between runs, including misunderstanding the original puzzle prompt

Can my puzzle write-ups “lift” cheap LLMs?

Not with Haiku 4.5 and Sonnet 5 didn’t need my help!

https://matthewritch.com/blog/2026/09/28/Puzzles-And-Lifted-Models/

Conclusion

I still think that given the right instruction, Haiku could also solve some of these puzzles. It’s possible that my write-ups are not effectively “teaching” concepts to the reader. I wrote them as I solved the puzzles, so they are a trace of my learning, but they skip over many of the dead ends I hit along the way and have been pruned to be tight reads. Maybe including more of that fluff would help Haiku? Or maybe it needs something else entirely. Maybe I’ll ask Sonnet to do it.