The measurement that said nothing
This is from the pull request that shipped the fix. Real script, real corpus, real numbers:
| | Items | |---|---| | unchanged, approved | 103 | | unchanged, rejected | 363 | | **newly rejected** | **1** | | newly approved | 0 |
And this is from the pull request that reverted it:
| | Items | |---|---| | newly rejected | **31** | | …correct | 3 — the rugby item, plus two great-white-shark items | | …**wrong** | **28** |
The fix merged at 07:09 UTC on 16 August and the revert merged at 07:25. It was live for 16 minutes.
Nothing here was written by a bad engineer, and nothing was skipped. This is how a careful measurement said nothing at all.
Two presenters who have never seen a hockey game
The episode is presented by Pip and Bo, who didn’t build this, have never seen a hockey game, and teach young children about money on a different channel.
There is a human in this story. You won’t see them. They are the cursor: every human decision arrives as a typed line.
What the feed had: 20 sources and 220 tests
When the episode was made, the Sharks News Aggregator had:
- 20 news sources, pulled every 10 minutes, with names extracted against a team roster that syncs itself daily;
- stories clustered across sources, tagged, and classified by event: trades, injuries, signings, call-ups;
- a website, an RSS feed, and a bot that posts to Bluesky;
- all of it on a Raspberry Pi, behind a tunnel;
- 182 commits, 88 merged pull requests and 220 tests.
Every one of those articles goes through one function that returns true or false, check_sharks_relevance(). About 1,900 articles reached it over the last month, and it is the only thing standing between the feed and the rest of the internet.
In August it let this through:
Longstaff agrees to new deal with Sharks - Yahoo Sport
It is a Sale Sharks rugby union story.
A snapshot is not a sample
One thing before the rest, because it is the part that transfers.
The change was measured before it merged, against real stored articles, with a script written for the job, app.scripts.measure_relevance_change. That is the correct thing to do, and more than most changes get.
It ran against a database dump saved in January. January is mid-season. The bug shipped in August.
Measuring a single season measured nothing.
Every file of old data you have lying around was collected at a time. If your change is sensitive to that time, you measured the time.
A snapshot is not a sample. Don’t just ask whether you measured; ask what your data was doing when you saved it.
Notice, diagnose, write
Notice. A human read the feed and saw a rugby story on it.
Diagnose. “Sharks” is not a hockey word. At least four professional clubs carry the name: Sale and Cell C/Natal in rugby union, Cronulla-Sutherland in rugby league, and Jacksonville in arena football. And the Google Alerts source that fed this article in queries the bare word.
Write. The fix split the keywords in two. san jose sharks, sj sharks, sjsharks and barracuda approve an article on their own. Bare sharks needs a second hockey signal: hockey-only vocabulary, another NHL club named in the title, a hockey beat in the URL, “san jose” in the title or URL, a player, coach or staff member from the roster, or a source flagged as hockey-only. On top of that, a veto on rugby, rugby league and cricket names and jargon, checked ahead of every approval path.
Whoever wrote it was thinking. The veto deliberately leaves out try, union, league, football and pitch, because all of them turn up in ordinary hockey coverage.
Measure it before it merges
Measure. Replay every article in the dump through the old rule and the new rule, and print every verdict that changed.
Getting to 1 took a second pass. The first version of the corroboration list rejected six real Sharks stories, among them:
Rangers at Sharks game 50 · Canucks Face The Sharks · Hamilton Blocked A Trade To Sharks
Two more signals — another NHL club in the title, a hockey beat in the URL — fixed five of the six, and all five became regression tests. The last one was genuine and still rejected, and it was pinned in the tests as a known loss rather than papered over.
Ship. 246 tests passing, including the verbatim rugby headline and the five it had broken. Merged three minutes after the pull request opened, and deployed to the Pi.
The pull request said, in its own verification section, that the change had not yet been run against production, because the snapshot was from January. It named the command that would. That command was the next stage.
Measure it again, on production
Measure again. The same replay, over the last 30 days of production: 1,902 articles, 31 newly rejected, and 3 of them correct — the rugby story and two articles about actual sharks. The other 28 were genuine Sharks news:
Sharks Hire Jeff Kealty as Assistant General Manager · Sharks Re-Sign Graf · Sharks GM Mike Grier completes substantial $12.75M move involving 23-year-old forward · Sharks Midsummer Roster Projection: Defense
The production server was rolled back and rebuilt, and the revert merged at 07:25.
+102 −790
The same change, subtracted. And the precision matters here, because the whole story is about precision: the filter did not run in production for 30 days. It ran for 16 minutes. The 30 days is a replay over stored articles — exactly what the pre-merge measurement was, over a corpus that had a different month in it.
The whole stack, and the file that is not a tool
| Tool | Its job |
|---|---|
| Claude Code (Opus) | wrote most of it |
| Python · FastAPI | the API |
| SQLAlchemy · Postgres 16 | stories, entities, clusters |
| Celery · Redis | ingest every 10 minutes |
| feedparser · NLTK | reads the feeds |
| OpenRouter (Gemma) | the second opinion |
| Next.js 14 · Tailwind | the site |
| Docker Compose | one file, two overlays |
| Raspberry Pi 5 | production |
| noBGP | the public URL |
| atproto | posts to Bluesky |
| GitHub Actions | lint, tests, build |
And one file that is not a tool:
- Incorrect claim: db_data_export.sql — the ground truth
Lesson 1: a snapshot is not a sample
The cost. 28 real stories a month.
The evidence. The corroboration list leaned hardest on one signal: another NHL club named in the headline. In season, that is in nearly every headline — Rangers at Sharks, Canucks face the Sharks, a trade blocked to the Sharks. In the offseason it is in almost none of them. July and August coverage is contracts, hirings and roster projections: no opponent, no city, no hockey vocabulary, on URLs with no hockey marker.
The measurement was not sloppy. It was seasonal, and nobody asked what season it was.
The rule. Whatever your test data is, it was captured under conditions. Find out what they were before you trust a number that came out of it. Any future relevance change here is now measured over both an in-season and an offseason window.
Lesson 2: if the signal isn’t in the data, no rule will find it
The cost. The idea that a better keyword list would do.
The evidence. Two headlines:
Sharks Re-Sign Graf
Longstaff agrees to new deal with Sharks
The first is real hockey news. The second is rugby. Club name, a contract, no hockey word, same publisher. There is nothing in either title or either URL that separates them, so no rule reading titles and URLs can separate them.
The diagnosis was right, and it still is: “Sharks” is not a hockey word. Being right about the problem does not hand you a fix.
The rule. Check the signal is in the data before you build a rule that needs it.
Lesson 3: print the diff, not the count
The cost. A review that approved a number.
The evidence. 31 is exactly as consistent with a great change as with a terrible one. It cannot say which. Reading the 31 headlines settles it completely.
Nobody read the first measurement, because the first measurement said 1. One item is a number you accept. 31 is a list you read. The difference between them was the month the file was saved in.
The project’s plan now says how the next attempt gets verified:
Verify with: a replay script over both a January and an August window, printing every changed verdict for reading by eye. A count cannot tell a rugby story from a Barracuda call-up — that is how the first attempt passed review.
The rule. Print the diff, not the count, and have a human read it.
Lesson 4: the system already knew
The cost. The uncomfortable one.
The evidence. Production runs the relevance check twice. The keyword rule decides. A language model is asked the same question, and its answer is written to a log table and not used. On the rugby article, the model said no — correctly. That answer was recorded, and the article went on the feed.
The log had been read once, nearly three weeks earlier, for a different problem: whether letting the model decide would stop articles about other teams getting in. It held a month of comparisons.
| keyword | LLM | count | |---|---|---| | approve | relevant | 1,070 | | approve | **not** relevant | **25** | | reject | **relevant** | **642** | | reject | not relevant | 881 |
Handing the model the decision would have removed 25 articles a month and added 642. And its rejections were right for the wrong reason. Its stated reason on those 25:
“Darnell Nurse is an Edmonton Oilers player and has no affiliation with the San Jose Sharks.”
The player had been traded to San Jose. The model was reasoning from stale training data.
The rule. A shadow mode is only worth what you read out of it, and what you read out of it is evidence, not a verdict.
Lesson 5: optimise for the revert
The cost. 16 minutes. This is the good news.
The evidence. Merged at 07:09, reverted at 07:25, and the revert is the same change subtracted. That was possible because the measurement script was written to hold the old rule frozen inside it:
The pre-RM-3 predicate is inlined and frozen; comparing against a moving target measures nothing.
So both versions could be replayed against anything, afterwards, on demand. The tooling that failed to catch the problem before the merge is the same tooling that caught it after.
The rule. You are going to ship wrong things. The number worth optimising is how long a wrong thing survives: one pull request, one rebuild, 16 minutes.
Three more: the query, failing open, the stale file
The bug started in the query, not the code. The news alert asked for the bare word “sharks”, and did exactly that. Adding negative terms, Sharks -rugby -NRL -cricket, is free. Narrowing it to "San Jose Sharks" OR "SJ Sharks" was simulated first, and it dropped five real articles to remove two bad ones. Filter at the source, and measure it.
Know which way you fail. The relevance check fails open: if the model errors, the keyword result stands. That is right for a news feed and wrong for a spam filter. Either way it is a decision, and it should be a sentence in the code, not whatever the error handling happened to do.
The file that lied was already on the list to be deleted.
| R3-A1 | P2 | Code quality | Remove `db_data_export.sql` from the repo root — a stale SQL dump is a data-leak risk and bloats clones. Move to private backup or `.gitignore` it. |
It survived because it was convenient. That is how a stale artifact becomes ground truth.
The hockey question is still not built
The agreed fix was not built when the episode was published, and it still is not. It asks the model one narrow question — is this article about ice hockey? — and only on the ambiguous articles, roughly 31 a month. The argument for it is that a model does not need to know who plays for whom to tell rugby from hockey. That is an argument, not a result.
The bigger problem is also still open. When it was measured, a third of the rumours page and more than a quarter of the trade page were about other teams — 32% and 27% — because a player’s name in a headline is enough to get in.
Read the revert
Everything here is public. The repo, the change that shipped it, the revert with the production numbers, and the open item, written up in full.
Read the revert. It is short, and it is better than most post-mortems.