7.1 attempts, said the model
Level 90 of Bus Tourist Craze had a difficulty rating of 7.1 expected attempts. It was not a guess. It came from simulated playthroughs with an exact solver underneath and a calibration curve on top.
A person played level 90 34 times and never finished it once.
- Incorrect claim: level 90: 7.1 expected attempts
- Correction: level 90: 34 attempts, no clear
Nothing in this project was measured carelessly. Every number was checked against another number. That was the problem: this is a project where the instrument was wrong more often than the thing it measured.
Two presenters who have never been to Canada
The episode is presented by Pip and Bo, who teach young children about money on a different channel. Neither of them built this game, and neither has been to Canada.
There is a human in this story. You won’t see them. They are the cursor: every human decision arrives as a typed line.
The brief was a document. It had everything a 100-level game needs, except the game.
What the game had: 100 levels and 60 tests
When the episode was made, the project had:
- a rules engine with no screen, no clock and no input, so the same code runs in a browser, a phone shell, and a server that re-checks your moves;
- a solver that proves a level can be finished, and a generator that builds the car park backwards so a gridlocked one cannot be made;
- 100 levels, one per leg from Vancouver airport to Parliament Hill, and 100 destination cards drawn by a second model;
- a route map, a win screen, a briefing and settings;
- 90 commits, 14 merged pull requests and 60 tests.
Every one of those levels carries a number saying how hard it is. That number was re-theorised three times, and the game a player meets changed far less than the machinery behind it.
One person had played it: the person who commissioned it.
Agreement between two models is not evidence
There were two instruments. solve() computes the clear rate exactly, with no sampling error, and only works on small levels. estimate() plays a level over and over and counts the clears, which works on anything and carries error.
They were checked against each other and disagreed badly, which is how a real bug surfaced: the lookahead player had been rejecting any move where a loss was merely possible, as if an opponent chose the follow-ups. In a solitaire puzzle the player makes every move. Fixed, the two agreed.
Corrected, they agree: exact 20.5% against sampled 23.3% +/- 4.1%.
That is a genuine result, and it proves nothing about people, because both instruments are the same imaginary player. That player has perfect recall of the 13-tourist queue on screen, never mis-taps, and never forgets what it was planning two moves ago.
Agreement between two models is not evidence. When two measurements agree, that is not the finish. Ask whether either of them has ever met the thing.
The loop: generate, gate, measure
The build process is not the story here. How a level gets its number is, and it runs in six stages.
Generate. The car park is built backwards. Every bus is placed only where it already has a clear way out, so taking them away in reverse order is always a legal way to empty the lot. An unfinishable level is impossible to make, and solvability never needs the solver.
Gate. Two greedy robot players try to cheat each level: one counts the wheel, the other matches bus sizes to the tourist groups it can see. If the shortcut clears a level more than a quarter of the time, the level is thrown away however hard it measures.
Measure. Small levels are solved exactly. Big ones cannot be: once the turntable held 30 seats, an 18-bus level explored 701,327 states and a 20-bus level blew past the 2,000,000-state cap. So exact solving stops at 12 buses, and above that the level is sampled and its figure carries a ~.
export const EXACT_BUS_LIMIT = 12;
Calibrate, then curate
Calibrate. The model’s figure is multiplied by a number that came from a person, because the model is not one. That multiplier is the only place in the whole pipeline where reality gets in, and it rested on three observations.
| tourists | measured | real | ratio | | --- | --- | --- | --- | | ~160 | 4.2 mean | 2.2 mean | ~0.5 | | ~264 | 4.9 | 15 | ~3.0 | | ~742 | 7.1 | 34+ | >=5 |
The correction is a function in calibration.ts, log-linear between the observed points and clamped outside them, so nothing predicts a multiplier for a level length nobody has played. The middle row was later rebuilt from 7 levels instead of 1, and the ratio there fell from 3.0x to a mean of 1.6x.
Curate. Candidates are screened cheaply, the most promising are verified at full precision, and the one closest to its target is kept. The target is in real attempts, not the model’s units. They are different scales, and choosing in the wrong one picks the wrong level.
Somebody plays it
The sixth stage is a person playing the level. It was added last, and it overturned something every time it ran.
The whole stack, ending in one person playing
| Tool | Its job |
|---|---|
| Claude Code (Opus) | wrote all of it |
| TypeScript (strict) | engine, solver, client |
src/engine |
no DOM, no clock, no I/O |
| exact solver | closed-form clear rate |
| sampled playouts | above 12 buses |
| two greedy adversaries | is there a shortcut |
| Vite · Vitest | dev server, 60 tests |
| Codex CLI | 100 destination cards |
| Raspberry Pi 5 | the tester build |
| git (+ worktrees) | one branch per change |
| one person, playing it | the only real instrument |
The last row is not a tool.
Lesson 1: the model can’t see what you left out
The cost. Level 90 was rated at 7.1 attempts and the person gave up at 34. So the correction was 5x, and it was written down as 5x.
Then 11 shorter levels were built and played, and on those the same model ran twice as pessimistic as the person, who beat it outright on several.
| level | measured | first clear | | --- | --- | --- | | 90 | 1.5 | 1 | | 91 | 2.0 | 1 | | 92 | 2.6 | 2 | | 93 | 3.1 | 8 | | 94 | 3.7 | 2 | | 95 | 4.3 | 1 | | 96 | 4.8 | 1 | | 97 | 5.3 | 1 | | 98 | 5.9 | 1 | | 99 | 6.5 | 5 | | 100 | 7.0 | 1 |
5x, then 0.5x. Same model, same player, opposite directions.
The evidence. The difference was length. The long level 90 had 742 tourists; the short ones had about 160. On the long one, the person got about two thirds of the way in and lost everything, and a failure there throws away 10 minutes of work, so you cannot recall which line you already tried. A short level you just play again. Length does not make the decisions harder. It makes learning harder, and the modelled player has nothing to learn with.
The second half of the lesson is the sharper one. The 5x was recorded as a calibration constant, and it was a point on a curve. Correcting a model with one observation is the same mistake one level up, and it made the next batch 4 times too easy.
The rule. Whatever your model leaves out does not come back as an error. It comes back as a number that is confidently wrong, in a direction you cannot predict.
Lesson 2: zero is a measurement
The cost. 11 levels shipped in one batch. The person gave up on two of them, at 34 and 42 attempts.
The evidence. Sampling each level directly, at three lookahead depths:
lvl human clear@0 clear@1 clear@2 90 cleared 1st 12.7% 9.3% 19.3% 91 GAVE UP @34 0.0% 0.0% 0.0% 99 - 0.0% 0.0% 0.0% 100 GAVE UP @42 0.0% 0.0% 0.0%
Three of the 11 never cleared once in 150 playouts at any skill level, and they carried the friendliest difficulty scores in the batch: 5.8, 9.3 and 9.9 attempts. When the sampler came back with no clears, the code fell back to a cheaper estimate instead of reporting the failure. So the levels nothing could finish got the kindest numbers.
Zero out of 150 is not missing data. It is a fact: the clear rate is below 1 in 150. That is what gets reported now, as a lower bound.
- Incorrect claim: ~9.9
- Correction: >200, flagged bounded
The fallback existed for screening, where being wrong is cheap. It stood in for verification, where being wrong is the whole cost.
The rule. Zero is a measurement, not a missing value.
Lesson 3: nothing found is not nothing there
The cost. It happened three times in one project. An attempt counter recorded play against level numbers while the levels were regenerated twice underneath it, pooling three different puzzles into one row. A set of structural checks sat under a branch that only ran on small levels, so they quietly skipped most of what they guarded. And the script that generates the destination artwork printed this:
Every leg already has a card. Nothing to generate.
The evidence. That one was two faults. npm run swallowed one of the script’s own flags, and macOS ships bash 3.2, which has no mapfile, so the list of missing cards came back empty. Every one of the three reported success, and not one could tell nothing wrong from nothing examined.
That difference is not automatic. A check that finds nothing is saying one of two things, and it only knows which if someone writes the difference in. The art runner now errors when its inventory cannot be read, instead of reporting that nothing is missing.
The rule. Nothing found is not nothing there.
Lesson 4: a finding that isn’t code is decoration
The cost. Early on, a rule was measured: past about two colours for every depot bay, the puzzle stops being playable. It was correct, and it was written down in the project’s own record with the numbers beside it. Then the chapter was audited.
levels above a 2.0 ratio: 89 of 100 1.0-2.0: 11 levels 2.33: 55 levels 2.67: 34 levels
The evidence. Nothing went wrong in between. A gate added later needed more colours to stop levels being cheatable, and nothing pushed back on the number of bays. Two locally sensible rules combined into an unplayable one, and the finding that would have caught it was sitting in a document.
It is a limit in the code now, MAX_COLOUR_BAY_RATIO, with a test that the schedule never breaches it. So are three other rules that had been prose first.
The rule. A document cannot reject a candidate. If a finding matters, it has to become a check, a test, or something that fails. Otherwise it is decoration.
Lesson 5: a constraint nobody measured is a guess in uniform
This one is about who was wrong.
The cost. The player asked why the queue’s entrance and its exit sat next to each other on the wheel, which made the board hard to read. The agent answered that moving them would change whether every level was still solvable, and reopen the difficulty work. That was accepted, and 2 rounds of work went into animating around a constraint nobody had tested.
The evidence. Asked again, the agent measured instead. The same dispatch sequences were replayed through both layouts across all 100 levels.
5,783 dispatches, zero behavioural differences in departure order, line consumption, wheel contents, or the win condition.
It was free, and it had been free the whole time. The claim sounded like the kind of thing that is obviously true, and the check was one script.
The measurement lied twice before it told the truth. The first run reported 1,164 mismatches that were only an empty seat sitting at a different index; the second reported 1,164 more that came from a test harness with no loss rule. Either would have confirmed the wrong answer.
The uncomfortable part is not that the agent was wrong. It is that a confident architectural claim from the agent was accepted by the human, and shaped the product, without anyone making the system answer.
The rule. A constraint nobody measured is a guess in uniform. And the measurement that agrees with you is the one to run again.
Three more: a stale checkout, a dead record, a silent edit
A whole session built on a stale checkout. Local main sat at 6d6fbee while origin/main was four commits ahead, carrying two merged pull requests. The project’s own record read as current and was not. A defect was diagnosed, a fix designed, a decision escalated to the human on a false premise, and the whole chapter regenerated over 90 minutes against targets that had already been replaced. Every step was sound. Local coherence is not evidence of a current premise: fetch before you read your notes, not before you push.
A personal best that could never move. The win screen nearly shipped a “best moves” record. A win means every bus dispatched, and one dispatch is one move, so all 100 levels have par.moves === buses.length. The best score is the bus count on your first clear, forever. One question caught it before anyone saw it: what would this look like on a second run?
A find-and-replace that matched nothing. A scripted edit replaced nothing, silently, and the feature was nearly reported as working. It was caught only because a probe printed travelling buses created: 0. It happened twice in this project. A silent no-op is worse than a crash, and it is the default behaviour of most scripted edits.
One player, and the testers still to come
When the episode was published, the correction between the model and a person rested on three observations, and the top one was taken before the zero-clear bug was found. That level’s “measured 7.1” may itself have been a fallback figure.
The whole chapter had been played by one person, who commissioned it and could see every level’s difficulty score. So the build went onto a Raspberry Pi with a note for testers asking four questions. The first:
Where did you get stuck or bored? Which leg number. This matters more than anything else.
There was a difficulty distribution, a calibration curve, and one player. The episode ends by promising to come back when the testers have said something, including the parts of it that turn out to be wrong. On its own evidence, some of them will be.
Since the episode, Bus Tourist Craze has been approved for the App Store. It is free, with no ads and no tracking, and testers’ reports go to its public support repo.
A making-of that is half corrections
The project’s own making-of record was 2,817 lines when the episode was made, and about half of it is corrections. It is the best part. The repo it lives in is private, so it is not linked here. The practice that produces a record like it is public: making-of is two files and a rule.