Three claims, all wrong
This project made three claims, out loud and in writing.
- Incorrect claim: The channel name is free.
- Incorrect claim: This line of CSS stops the numbers jittering.
- Incorrect claim: This command won't print the API key.
- Correction: It printed the API key.
The third one was a check. The agent needed a voice API key, the human had put it in a shell profile so the agent could use it without seeing it, and the agent wrote itself a line to confirm the variable was visible:
echo "sees key: ${ELEVENLABS_API_KEY:+YES}${ELEVENLABS_API_KEY:-NO}"In that shell, ${VAR:-NO} expands to the value when the variable is set, not to NO. The key went into the transcript. The human was told at once and rotated it.
Nought for three. Ten episodes were made anyway. This is how, and what it cost to find out.
Pip and Bo are what got made
The episode is presented by Pip and Bo, who are not the people who made this. They are what got made: the two characters of a channel that teaches young children about money. This is not that channel, and nobody here is seven.
There is a human in this story. You won’t see them. They are the cursor: every human decision arrives as a typed line.
The brief was one sentence, typed into an empty folder. That was all of it.
What came out
When the episode was made, the project had produced:
- 10 episodes, each cut twice: a 3-minute landscape version and a 45-second vertical one;
- 21 videos, counting the channel trailer, and 35 minutes of finished film;
- 2 characters at 8 expressions each, and 24 props, all drawn in code;
- 10 thumbnails, a banner, and 21 titles and descriptions;
- a brand kit of more than 900 lines, every section locked.
3 of the videos were live and 18 were scheduled.
And the honest part: nobody had watched any of it. There was a schedule and no data, so every decision in the episode was a guess with a reason attached. Finding out became the next episode.
Don’t prompt for output. Build the factory
One idea before the process, because it is the part that transfers.
- Incorrect claim: make a video
- Correction: build the thing that makes videos
Nobody asked an AI to make a video. The agent was asked to build the thing that makes videos, and then that ran 21 times. The tenth episode took a fraction of the effort the first one did, because every fix from the first was still in the machine.
Don’t prompt for output. Build the factory.
Decide, specify, draw, write
Decide. 10 topics, in an order: money exists, choices cost something, money comes from work, and then you can store it, grow it, borrow it and make it. Each episode assumes the one before it. That is a spine, not a playlist.
Specify. This is the stage people skip. Before any video existed there was a spec: colour, type, characters, motion, sound and file names, written down in plain text. Contrast ratios were computed rather than eyeballed, so even the paper colour of the background is a measurement rather than a matter of taste.
Draw. Pip and Bo are six shapes each, rounded rectangles and one circle, written by a Python script. No illustrator was involved, which means one colour change reaches every episode, every thumbnail and the logo on the next build.
Write the short version first. The 45-second cut is the tightest constraint and the way anyone finds the channel at all. Write the 3-minute version first and cut it down, and the short one tastes like leftovers.
Gaps declared, durations measured
Neither character has a throat. The voices are synthetic, one audio file per line — not per scene, per line, so a two-hander can be placed exactly.
This is the trick worth stealing. Each episode’s build declares the pause before each line, measures how long each generated recording actually turned out, and derives every start time, scene boundary, crossfade and animation cue from those measurements.
So re-cutting the voice re-times the whole video. That stopped being a nicety when the voice provider changed and every line came back longer: nothing needed re-timing by hand.
Not one timestamp is ever typed.
HTML in, MP4 out
The video is a web page. Python writes the HTML, and the animation is one paused timeline. A renderer drives a headless browser frame by frame and hands the frames to ffmpeg. Then a checker runs over layout, motion and contrast, and anything with an error does not ship.
The episode called that pipeline deterministic: same input, same file, every time. The project’s own record later proved otherwise, using this episode:
run A 102 / 104 lines in place b1 and p36 misplaced run B 104 / 104 lines in place same bytes in
What is deterministic is the build: it produces byte-identical HTML every run. The renderer underneath it is not, at least for audio placement with 104 clips on one track, which is why an audio check now gates every upload.
The step nobody plans for
Thumbnails, banner, titles and descriptions are all generated from the same tokens, and read out of the episodes themselves rather than retyped.
Then the files have to get somewhere a scheduler can see them, because a scheduler cannot read a file on your laptop. There was no step in any plan that turned a rendered file into something on the open web. A shared cloud folder became the bridge, and the part worth keeping is how it is verified: every public link is requested and its byte count compared with the local file, because a link that returns a success code with a web page instead of the video is the normal failure, not an exotic one.
Public folder, verified links, scheduled, live.
The whole stack, and no video editor
| Tool | Its job |
|---|---|
| Claude Code (Opus) | wrote all of it |
| Python | generates the compositions |
| HTML · CSS · GSAP | the composition itself |
| HyperFrames | renders HTML to MP4 |
| headless Chrome · ffmpeg | shoots and encodes it |
| ElevenLabs | two voices, one file per line |
| fonttools | .ttf → .woff2 |
| Google Drive | the public-link bridge |
| Metricool | schedules to YouTube |
| git (+ worktrees) | one branch per change |
Nothing on that list is a video editor. Nobody dragged a clip onto a timeline.
Lesson 1: ask what the check actually proves
The cost. A channel name.
The evidence. Whether a handle was free was checked by requesting its page and treating a 404 as available. On that basis one name was declared free and recommended. It looked rigorous: there were percentages, tables and a methodology note. The human tried to register it, and it was taken.
A 404 means the channel is not displaying. It does not mean the handle can be claimed. Handles held by hidden, terminated or reserved channels return exactly the same status as free ones, so the check was measuring a stand-in for the thing, not the thing.
The rule now is that a tool like that can rule a name out. It can never rule one in.
The rule. Ask what your check actually proves, not what it looks like it proves.
Lesson 2: correct in general, wrong in the specific
The cost. Every counting number in the series, very nearly.
The evidence. The spec told every animated number to use this, so digits would not change width mid-count:
font-variant-numeric: tabular-nums;
That is correct advice for most fonts. Neither of the channel’s fonts supports it. In Fredoka the digit 1 is 396 units wide and the 2 is 608, a 53% difference, and the property does nothing at all. Every counting number would have visibly jittered, and nothing would have raised an error. It was fixed with fixed-width digit slots instead.
The same spec also invented a rule about waiting for fonts to load before rendering. The framework embeds fonts when it compiles and forbids building the timeline asynchronously, so the rule would have broken the render outright. Confident, specific, technical, and wrong.
The rule. Correct in general, wrong right here: that is the class of error to go hunting for.
Lesson 3: model the expensive step
The cost. An hour, instead of a rebuild.
The evidence. Before any real episode, the smallest possible video was rendered: two characters, one prop, one number. It broke four things. The worst was the sizing system: everything was scaled off the frame’s width, which leaves a vertical frame two-thirds empty. The others were the invented font rule above, characters rendering solid black, and fonts in the wrong format. Finding the sizing bug then cost an hour. After 10 episodes it would have been a rebuild.
There is a sharper version. One episode had to be built before its voice existed, and generating voice costs money. So a model was fitted to the recordings that already existed:
pip: duration = 0.540 + 0.297 × words (n=69, worst error 1.69s) bo: duration = 0.582 + 0.197 × words (n=47, worst error 0.66s)
34 files of pure silence were cut to the predicted lengths, and the whole episode was built and checked against them. It predicted 2:53. The real voice came in at 2:55, and two rounds of layout fixes had been made before a single credit was spent.
The rule. Model the expensive step, and let the cheap version drive development.
Lesson 4: a passing check is not a watched video
The cost. Three faults that every check passed.
The evidence. The episode about credit cards has a clock, and its hands moved as one lump in a circle instead of turning. The rotation was measured from each hand’s own bounding box, which does not contain the clock’s centre. It rendered and passed every check: the animation was legal, nothing overflowed, the contrast was fine. No checker has an opinion about whether a rotation looks like a clock.
The same week, every vertical video turned out to open on a quarter of a second of empty paper. A rule written for the landscape cut, where a beat of settling reads as composure, had been applied to a format where the viewer decides in under a second.
And one tag went live as two, what is money and anyway?, because the platform splits that field on commas. The build now refuses a tag containing one.
None of the three came from tooling. All three came from a human pressing play.
The rule. A passing check is not a watched video.
Lesson 5: taste doesn’t delegate
The cost. A character’s proportions, and nearly the sound of the show.
The evidence. Every audio level in the project was set by measurement: the voice, the music bed and five sound effects, all normalised against one reference.
−19.1 LUFS, −1.9 dB peak
And it could not tell whether any of it sounded right for the show. That went to the human, who approved it.
Earlier, the same human looked at a character design and said Bo’s head was too big. It was: 49% of body height, a lollipop. The fix was not a smaller head but moving Bo’s difference from Pip off head size entirely. That glance was faster and more correct than any measurement in the file, because one variable had been optimised beautifully, in isolation.
An agent will hold a constraint across 10 scripts without being reminded, and compute every ratio you ask for. It cannot tell you the thing is good.
The rule. Taste doesn’t delegate.
Three more: drift, repetition and done
Documentation that nothing depends on will drift. The stylesheet that supposedly enforced every colour in the project was imported by nothing, and agreed with reality by luck. It is read at build time now, so a wrong value breaks the video instead of sitting there quietly.
Don’t factor out a repetition that is carrying meaning. A script note said one episode’s visual device was reusable and should become a component, which is what a good engineer writes and was exactly the wrong instinct. The later episodes already had ideas of their own: a clock, and 11 coins where you expected 10. In a series, some of the duplication is the argument.
Done means done in every output you claim to ship. 10 episodes were finished, rendered and documented while the landscape cut, the canonical one, had never once been checked. All 10 were broken in the same place: an end card built for a tall frame, running off the top of a wide one.
No audience, no data, a schedule
When the episode was published, none of this had been validated. Not one of its design arguments had met an audience. There was a schedule, a spec and no data.
The episode promised that the next one would be whatever the numbers said, including the parts of this one that turned out to be wrong. It was, and some of them did.
The write-up, shared as a document
Everything here comes from the project’s write-up, which is mostly things that went wrong. It is the best part, and it is shared as a document. The channel it describes is Pipedia, on YouTube.