Can LLMs one-shot the daily Poople word ladder? Models get the start word and the legal word list, and must output a complete ladder to POOP in a single response โ scored against the provable optimum (par).
Mean score (par รท steps; invalid = 0) vs mean cost per attempt. Ringed points are the Pareto frontier โ no cheaper model scores higher. Hover for detail.
A separate experiment, not the benchmark above. Jev is a classifier: it answers typed questions and cannot write a ladder. Here it plays one rung at a time through three small questions per turn: is each nearby string a word, which follow-up word looks most like POOP, and which move to take. Code only spells strings out and keeps the books; it never looks a word up to decide anything and never counts letters for the model. The Reference rows are scripts playing the same puzzles with the true word list, so Jev's number has something to stand against.
Strict scores par รท steps and 0 for a game with any rejected guess, the same rule as the main board. Assisted forgives rejected guesses. Right call is how often the move was on a shortest path when the choice mattered. Similarity is how often Jev's pick of the most POOP-like follow-up word was correct. Legality errors is the share of word-or-not calls Jev got wrong.
Every attempt for one puzzle day, with the ladders models actually produced.