Leave the House Out of It is the original project that I started building decades ago to keep my hands dirty on coding, even as my job turned more into powerpoint and management. Over the holidays in 2021 I expanded it to include a “robot” player, BOOK-E (a play on the WALL-E robot movie that was popular at the time). You can read that three-part saga starting here. It was a Jupyter-notebooks-and-SageMaker affair: I gathered data by hand, let AWS Autopilot pick a model for me, and manually ran predictions each week. While tt worked well enough to convince me the idea had legs, I didn’t have time to keep it maintained.
The problem with the old BOOK-E was that I was most of the system. The notebooks didn’t run themselves, the data went stale the moment I stopped babysitting it, and “deploying the model” meant me remembering to do things on a Thursday. So this summer I gave BOOK-E the full brain transplant, with an important twist: this time I had a pair programmer. I built the whole thing with Anthropic’s Claude Code, and the difference between this project and 2021 is roughly the difference between building a shed alone and building a house with a contractor who never sleeps (and occasionally ships a bug at 2am, but we’ll get to that).
Introducing BOOK-E 2.0 (Now With Three Tiers)

The new BOOK-E isn’t a notebook, it’s a small platform running on Google Cloud, and it’s built in three layers that each mind their own business:
Tier 1 — Data. A BigQuery warehouse fed by scheduled jobs that pull from three sources: MySportsFeeds for live games, box scores, player stats, injuries, and — the crown jewel — intraday betting line movement from multiple sportsbooks; the free nflverse archives for history back to 2016 (including closing lines); and LTHOI itself, for our league’s lines and everybody’s bets. About 2,700 games, 70,000 player stat lines, and 68,000 timestamped line movements so far.
Tier 2 — Prediction. A set of pluggable “engines,” each a different theory about how to beat the lines. The two that matter are boosted-tree models trained right inside BigQuery ML: one that only knows football fundamentals (Elo ratings, recent form, rest, travel, weather, injuries), and one that knows all of that plus how the Vegas lines have been moving — the latter turned out to be the smarter of the two, and it’s the one BOOK-E listens to. There’s also a baseline engine that just parrots the sportsbook consensus, which exists to answer the only question that matters: is any of my cleverness actually better than doing nothing?
Tier 3 — The Agent. BOOK-E himself. As I described in my update on revamping LTHOI recently, I added the capability to have an agent play for you in the league. I’m using this with BOOK-E and every couple of hours during betting windows, a little service wakes up, generates fresh predictions, compares them to LTHOI’s lines, and places real bets — no human in the loop (see image below). He bets whenever his edge is at least 2 points on a spread or 3 on a total, respects the league’s caps and freeze windows, always flips his bet if the model changes its mind, and writes every decision (including every game he passed on, and why) to an audit log so I can interrogate him later.
The whole thing is deployed with Terraform, scales to zero when nothing’s happening, and costs me between $1 and $4 a month. The 2021 version cost more than that in SageMaker charges per training run.

The Data Is Still the Hard Part
In 2021 I wrote that the real challenge of machine learning is collecting and organizing the data. Five years, better tools, and one AI pair programmer later: still true.
Every interesting bug in this project was a data bug, and most of them were because Claude is WAY WAY WAY too confident (this is why I still don’t recommend vibe coding to non-technical folks). My favorite: an early backtest announced that BOOK-E could pick over/unders at 73%. For about an hour I mentally planned my early retirement. Then I asked Claude to explain how he got that number and found the odds feed was quietly mixing quarter and half-game lines in with the full-game ones. Claude was confidently proclaiming that he could predict there would be more than the 1Q over/under scored in a full game. With the garbage filtered out, that 73% collapsed to exactly 50.0%. No early retirement.
The other design obsession this time was leakage. Every feature in the warehouse is computed strictly from what was knowable before kickoff — injuries as-of the snapshot before the game, rolling stats that exclude the game itself, lines frozen at their pre-kickoff values. It’s the boring plumbing that decides whether your backtest is science or astrology.
The Season Ahead
BOOK-E is live for the 2026 season in our league, The Originals, and for the first time he runs entirely without me. The data refreshes itself, the model retrains itself every Tuesday after the week’s games settle, and the bets place themselves while I’m at my kid’s soccer game. Every decision he makes — every bet, every game he passes on, and the edge he saw when he did it — lands in an audit log, so at the end of the season there will be no hiding from the numbers, for him or for me.
Can a robot built on public data beat a league full of humans picking games on their commute? That’s the experiment, and I’m deliberately not peeking at the early returns before there’s enough of a sample to mean anything. Five years ago this project took me a holiday break and a stack of notebooks; this summer it took a few evenings of talking to an AI about football. I’ll report back when the season’s told its story — assuming BOOK-E hasn’t asked for a raise.
