The page that reduces to one space
An Agoda property page is 289,857 bytes of HTML. It renders entirely in the browser, so once the scripts and the styles are stripped and the visible text is taken, what is left is a single space.
html bytes 289857 reduced chars 1 reduced text ' ' $ occurrences []
That space was sent to a language model with the question “what does the room cost”. It answered 19.99 — for a night at the Royal Plaza in Hong Kong, where the real rate was 191.69.
It was run 8 times to be sure: 5 sending a literal single space, 3 sending the real page through the production code path.
--- user message is a literal single space, 5 trials
0 ('{"price": 19.99, "currency": "$"}', 155)
1 ('{"price": 19.99, "currency": "$"}', 155)
2 ('{"price": 19.99, "currency": "$"}', 155)
3 ('{"price": 19.99, "currency": "$"}', 155)
4 ('{"price": 19.99, "currency": "$"}', 155)
--- the real Agoda page, fetched fresh, through from_llm(), 3 trials
0 price= 19.99 llm_called= True
1 price= 19.99 llm_called= True
2 price= 19.99 llm_called= TrueThe same answer to the cent, every time. The error field stayed empty and found stayed true, so as far as the rest of the program could tell, that was a price read off a page.
None of it ever reached the database, and the reason it did not is the whole argument of this piece.
It is fixed now, merged and deployed the same afternoon it was found, and the diff is below. The fix is 4 characters. Getting there is the story.
Two presenters who have never paid for a hotel room
The subject is PriceReduced, a personal price monitor. The presenters, Pip and Bo, did not build it and have never paid for a hotel room; they teach young children about money on a different channel.
There is a human in this story and they do not appear. Every decision a person made arrives the way it actually arrived: as a typed line.
There is no write-up for this project. The whole history is 21 commits, a handful of docstrings, and a database that has been running since July — which turned out to be better evidence than a write-up would have been. Nobody smooths a commit message they wrote at midnight with the site name still in it.
What exists: 19 items, 1,897 checks, two and a half cents
Give it a product URL. It reads the price on a schedule and keeps the history until the item is deleted.
19 URLs: vacuum parts, a jacket, a trackpad, 4 Hong Kong hotel stays, 2 award-flight searches and one Google Flights itinerary. 1,897 price checks, 1,552 of which found a price. It runs in Docker on a Raspberry Pi on a shelf, behind a public HTTPS URL, and nothing has been touched on it since 10 July.
The part everybody asks about first is the model bill. It is two and a half cents, all time — an independent, provider-side figure rather than the project’s own arithmetic.
| | | |---|---| | Total, all time | $0.0252 | | This month (August) | $0.0096 | | This week | $0.0022 | | Today | $0.0001 | | Models | 1 — Gemini 2.5 Flash Lite |
There is a spending cap in the config set at 500 calls a month. The busiest month used 115. It has never once bound, so the cap is not what is keeping this cheap.
The shape of the spend is a better finding than its size. Daily spend across the 30-day window is a flat bar at roughly $0.00031 a day, one spike aside. A system whose model calls are meant to be exceptional does not draw a flat line. Hold on to that; it is lesson 4, and it is the expensive one.
Put the model where the ambiguity begins
Everywhere you can name what you are looking for, you have already written the rule.
A price in a JSON-LD offer block is named. A price in an Open Graph meta tag is named. A price in an element the page itself labelled price is nearly named. None of those need a model. They need about 40 lines each, they are auditable, and when they are wrong they are wrong somewhere you can point at.
And the corollary nobody puts in the design doc:
A fallback is not free just because it is rare.
A page with no price needs a door, not a bigger model
4 of the 19 URLs are Agoda hotel pages, and every single generic check on them failed. Not “found the wrong price” — found nothing. The page builds itself in the browser, so there is nothing there for any tier to read, and no model could have read it either.
The fix was not a better model. It was a different door. The page ships the URL of its own pricing API, dates already in the query string: load the page once for the session, pull that URL out of the markup, call it.
1,107 checks have run through that one regular expression, which makes it the single most productive line in the repository. It is also the only reason the 19.99 never got written down — the site handler runs before the generic cascade, so on those pages the model is never asked.
The rules were wrong twice, and both fixes were rules
Meanwhile the bottom tier — the one that reads elements a page has merely labelled price — was getting it wrong in two different ways in the same week. On an Epson paper page it read the paper; on an Amazon trackpad page it read a warranty.
epson.ca Size:8.5" x 11" ... Our Price:$19.99 -> read 8.5
amazon.ca first [class*="price"] is a $14.99 -> read 14.99
protection plan; the product is $159.00Neither of those is a model problem. The first fix is to count only a number attached to a currency symbol. The second is better: stop taking the first match, collect every price-shaped candidate on the page, and let the page vote. The real price repeats across the buy box, the sticky bar and the mobile view; the add-on appears once. That is a 5-line rule and it has not been wrong since.
Both were caught by a human noticing a wrong number on a dashboard, which does not scale past 19 items.
So a model went on top as a check
One model call to sanity-check any bottom-tier price, the first time it is seen or whenever it changes. Agree, and confidence goes up. Disagree, and the model wins.
The reasoning was written down and it was reasonable: both known failures were ones a model would have caught. Its entire production record is 2 events. It confirmed one price and overrode one price, and the override was the wrong call — a 5% clip coupon.
- Incorrect claim: the model read 151.05
- Correction: the page said 159.00
The rule as written handed the model the argument. The patch widened “agreement” to 10%, so coupon-sized differences count as agreeing.
And here is the part only a fresh check could show. That same page, 8 weeks later: the same two numbers, to the cent. The model has not got better at it. The tolerance just got wide enough to absorb it. That is a real fix, and it is not the same thing as a solved problem.
The last change is the quiet one. Every check now also scrapes the product’s barcode, part number and brand, whether or not that is where the price came from, so the same product tracked at two shops can be checked against itself. Hold that thought too.
The whole stack, and what each tier actually found
12 rows in the extraction cascade. One of them is a model, and it found 2.2% of the prices.
| Tier | Checks | Prices found | Share of prices | |---|---|---|---| | agoda-api (site handler) | 1,107 | 952 | 61.3% | | json-ld | 454 | 454 | 29.3% | | regex (price-flagged elements) | 112 | 112 | 7.2% | | llm | 34 | 34 | 2.2% | | none (nothing found) | 190 | 0 | — |
Deterministic tiers produced 1,518 of the 1,552 prices, or 97.8%. And 2 rows in that cascade have never won a single check: microdata and embedded JSON. They cost nothing and they have done nothing, which is its own small lesson about how cheap it is to over-build a cascade.
Lesson 1: a fallback has to be able to say it had nothing
The cost. A confident wrong price, indistinguishable from a real one at the call site.
The evidence. The model was sent one space and answered 19.99. The interesting part is not that it made something up. It is that the code asking the question could not tell: the result was identical, field for field, to a price genuinely read off a real page.
There is a guard against exactly this, and somebody thought of it.
text = reduce_page_text(html, settings.llm_max_input_chars)
if not text:
result.error = "no page text to send to llm"
return result # llm_called stays False - nothing was spentOne space is not falsy. It walked straight through.
- Incorrect claim: if not text:
- Correction: if not text.strip():
Measured against the running container immediately before and immediately after the rebuild, same item, same URL:
| | reduced text | llm_called | found | price | error | |---|---|---|---|---|---| | before | ' ' (len 1) | True | True | 19.99 | None | | after | '' (len 0) | False | False | None | no page text to send to llm |
Read the second column, not the fourth. The model did not get better; the model was never asked. And the third column is the one a caller can act on — before, a made-up 19.99 and a real reading were the same object, and after, a miss is reported as a miss.
The actual price still comes back, 191.69, from the site handler, exactly as it did before. The fix closed the hole behind the door. It is not what keeps the door shut.
The rule. The transferable version has nothing to do with models. Any time you hand work to something that can only answer and never decline, check what you handed it — because it will answer.
Lesson 2: the right number can still be the wrong number
The cost. Recorded hotel prices that disagreed with what a browser showed for the same hotel on the same night.
The evidence. The Agoda handler worked. The extraction was not wrong. It was reading the field the room grid displays by default, which in most markets is before taxes and fees.
| Item | Hotel / stay | displayPrice | inclusivePrice | Delta | |---|---|---|---|---| | 6 | Royal Plaza, 12–16 Oct | 181.69 | 205.31 | +13.00% | | 7 | Royal Plaza, 19 Oct | 200.81 | 226.92 | +13.00% | | 8 | New World Millennium | 210.92 | 238.34 | +13.00% |
13.00%, identical to two decimal places across 3 different hotels. That is a 10% service charge and a 3% hotel tax, compounding. The number was right and the meaning was wrong, and no amount of checking the number would have found it.
The rule. Record what a number is, not just what it says. Every price point now carries its basis, and a chart refuses to plot two different bases on the same line.
Lesson 3: the badge that has never said yes
The cost. A verified branch that is correct, tested, and unreachable.
The evidence. The same product can be grouped across shops — Amazon and Apple, say — and every member gets a badge saying how sure the system is that it really is the same thing. A barcode match, or brand plus part number, earns the tick.
The tick has never been awarded.
| Member | GTIN | MPN | Brand | Verdict | |---|---|---|---|---| | apple.com product page | — | — | Apple | probable | | amazon.ca listing | — | — | — | probable |
Apple’s own product page publishes a brand and no part number and no barcode. The Amazon listing publishes nothing at all. Same $159 trackpad, 2 pages, no comparable identifier between them — so it falls through to comparing the titles, and the strongest answer it can give is “probably”.
The rule. An identity check is only as good as the identifiers your sources publish, and the 2 biggest shops in this dataset publish none. Building the verifier was right; assuming it would fire was not.
Lesson 4: retry is a cost nobody models
The cost. 61% of everything this project has ever spent on a model, for 0 prices.
The evidence. 225 model calls, 62% of which came back with nothing usable. That sounds like hard pages. It is not.
| Item | URL shape | Calls | Successes | |---|---|---|---| | 14 | amex.point.me award search, economy | 58 | 0 | | 13 | amex.point.me award search, business | 57 | 0 | | 23 | amazon.ca trackpad (after it started 403ing) | 22 | 0 | | 26 | Google Flights itinerary | 33 | 31 | | others (9 items) | — | 55 | 54 |
137 calls to 3 URLs that have never once returned a price. 2 of them are award-flight searches that build themselves in the browser: the price is not in the page and never was. The model was handed an input that did not contain the answer, 115 times, and asked again the next morning.
That is the flat line on the dashboard. It is not a system with a rare fallback; it is a subscription to 3 URLs.
The rule. A cap counts calls. It does not notice that the same call keeps failing. The circuit breaker that would is 4 lines, it is not written, and the reason it is not written is that the cap made the budget feel like a solved problem.
Lesson 5: and the model is still worth having
The cost. Nothing. This is the part that stops the piece being an argument against models.
The evidence. Every tracked page was fetched once and run through the rules and the model side by side. 17 pages fetched; 9 of them had a deterministic answer to compare against.
| Model vs. the rules | Pages | |---|---| | Agreed exactly | 4 | | Returned nothing at all | 4 | | Returned a wrong number | 1 | | Beat the rules | 0 |
Zero. On every page where both had an answer, the model tied or lost, and 4 of its 5 misses were plain retail product pages with the price sitting in a JSON-LD block.
And then there is the one page where the rules have nothing: no offer block, no price meta tag, no stable element name, and a price that genuinely moved from 489 to 725 across the month. The model reads it. It read 718.00 on the day of the replay, against production’s 719.00 the day before.
The rule. One page in 19, where the ambiguity is real and the answer is actually in the input. That is not the fallback failing to be the product. That is the fallback being exactly the right size.
Three more, quickly
The default model on day one was a free-tier one, and it was replaced after a single call, because free models parse badly and disappear.
This is a single-user app and it still had a concurrency bug — a manual check and the scheduler could fetch the same item at the same moment. It is fixed with a database lease, which is what you would do for 100 users.
2 shops block it outright, and the dashboard says so. It says site blocks bots rather than showing an empty price and letting you assume. An honest status is a feature, and it is the exact opposite of the bug at the top of this piece, which was a dishonest status wearing a real one’s clothes.
What is still not known
The plan said: run the cross-check on for a fortnight, off for a fortnight, and count. It was not run, because it cannot answer anything. The cross-check only fires on a bottom-tier price that is new or changed, and that has happened twice. A fortnight arm would be a fortnight of waiting to observe nothing. The replay above took an afternoon and answered more.
That is worth saying out loud, because a test with no power looks exactly like a test in a planning document, and nobody checks before they run it.
The guard got fixed. The other thing did not, on purpose. A circuit breaker — something that notices a URL has failed 5 times running and stops asking — would have saved about 120 of those 137 calls. It is written up rather than built, because 4 of its decisions are judgement calls and not one of them is a patch: the threshold, the reset policy, the scope, and the visibility. The one that matters is the second.
A breaker that never closes is a silent permanent opt-out, which is worse than the waste, because at least the waste shows up on a bill.
And the genuine unknown. Those 2 flight-search pages fetch perfectly well and contain no price at all. Would a real browser get one? The static page reads 0 Flights, which is just what it says before the JavaScript runs — not evidence either way. Nobody has pointed a browser at it. The cheap answer might beat the clever one: 4 lines that stop asking cost nothing, while a headless browser is a whole dependency that fights a site for a living.
Next time this project comes up, it is whichever of those got built, and what the call log says afterwards.
Read the pull request
19 URLs. One of them needs a model.
Everything here is public: the repo, the fix, with the reasoning and a regression test, the circuit breaker, proposed and deliberately not built, and the tiered extractor itself.
The pull request is 4 characters and a regression test that fails loudly if a whitespace-only page ever reaches the model again. It is the shortest thing here and the one worth reading.