loading...
The use case fits in a sentence: you just finished dinner with friends, you want boba, and you want to know what's open near you right now. I started it for the San Gabriel Valley, the densest boba corridor in the US, as a tool for me and my friends.
The map worked on day one. What didn't exist anywhere was the data underneath it. Google has no boba category: shops I'd call boba are typed tea house, cafe, restaurant, bakery, or juice shop, while hookah bars and donut shops carry "Boba" in their names. Its nearby search stops at twenty results with no way to page, so the densest blocks truncate silently. And "open now" from a stored file has to survive stale hours, ranges that cross midnight, holidays, and a viewer whose clock could be in Tokyo. Each of those is a classification or freshness problem, and each one costs money to ask Google about. What counts as a boba shop, and is it open right now? That question was the project.
Google bills per call, so a public map on a metered API is a bill that scales with strangers' curiosity. I decided early that nothing about a shop is fetched from Google at runtime: a pipeline pre-fetches everything, the browser downloads one file, and it works out "open now" by itself.
That buys speed and a bill of zero. It costs freshness, and everything downstream is the answer to that. Three refresh loops run at three cadences, and only the monthly one can add or remove a shop: it opens a pull request and stops, so I review the diff first. Corrections land in a thin delta layer that every full refresh supersedes, so it stays a cache and never becomes a second source of truth.
Absence from a search never hides a shop. Removal takes an explicit closed status from a confirmation call, failures get flagged for me instead of acted on, and a wrongly-closed shop heals itself on the next pass.
The expensive way to answer this is to ask a model about every place the search returns. The cheap way is to notice that most places aren't ambiguous at all. Google's own "Bubble tea store" label settles 1,178 of the 1,720 shops. A list of known chain names settles 338 more. The model only gets what's left: 674 cached decisions, 269 include and 405 exclude, each keyed by place and stamped with the model and date that made it.
The stamp is what makes the model safe to use. When Google's review summaries became available as a signal, I deleted the 198 decisions that had been made without one and re-ran them: 12 flipped from exclude to include, none regressed. Confidence also carries past the gate. The popularity sort multiplies by a tier weight (1.0 for a Google-labelled shop, 0.7 for a model include), so the model's doubt demotes a shop in ranking instead of only deciding whether it exists.
Guiji Tea was invisible for four months, and the pipeline found it on every refresh and threw it away every time. Its ID sat in a derived blocklist, filed by a classifier rule I had since deleted, and that blocklist was evaluated above Google's live "Bubble tea store" label. I reordered the ladder so fresh evidence rescues a stale verdict and hand-curated lists stay absolute, because curation is evidence and a cached verdict isn't. Fixing it also exposed an alias bug that had 22 shops displaying under the wrong brand.
Coverage had a second-order problem. Searches in the densest parts of the corridor come back full, which means the results are ranked rather than complete: a new shop with a handful of reviews is systematically unseeable. That's a ranking problem wearing a coverage problem's clothes.
Three times I tested an obvious next feature against data first, and three times the data said no:
The scoring question was real, though: a hookah lounge with a boba menu isn't what "let's get boba" means. What worked was a learned model over Google's own prose, trained on chain locations as positives and the language model's 320 cached excludes as negatives, a labelled set that had been sitting unused. It reproduced the mental model I'd have written by hand (pearl milk positive, breakfast negative), and it was wrong in three ways I only caught when one shop's score looked off: brand names as vocabulary, nationality as a proxy for business format, and ambience words as signal. Boba Bear, a clean tea shop, scored 12 because Google's description read "Hookah bar in Hong Kong Plaza". Penalized for its address. I stripped each proxy, renamed the metric teaFocusScore so nobody could read it as quality, then reviewed every shop under 40 and cut all of them.
Each no is written down with the numbers that made the call, so it stays decided.
Computing this in the browser is deceptively deep: shops open past midnight, holidays override the week, and the viewer's clock might be in Tokyo. Everything normalizes to Pacific wall-clock and re-checks itself every minute, so an idle tab never shows "open" past closing. Overnight hours are stored as a close past 24:00 on the opening day, which is how 1:30 AM reads "Open until 2 AM" instead of "Closed today".
Some of this was wrong for months, and I didn't find it by looking. I ran a full audit of the application and it caught the "Open now" filter as a no-op: it had an ARIA state, a default, and a spec entry, and nothing ever read its value. It also found the removal logic running backwards, with 355 shops sitting one refresh away from being deleted for no reason, and holiday hours being stored as the permanent weekly schedule. Each fix came with a test, and the removal path now keys off per-shop confirmation calls, which saturation can't touch.
A brand-new shop with no listed hours can't render, because there's nothing truthful to show. A holiday is only covered if a refresh ran in the week before it. I document limits like these instead of papering over them.
People search from memory, not from the sign above the door. They run a brand together as one word, leave the accent off, or type the shorthand their friends use. Every strictness in matching is one more way to show someone nothing while the shop sits right there on the map.
So I made matching loose. Punctuation, spacing, and accents are ignored, and a table of brand aliases maps what people type onto what the listing says: tptea lands on TP Tea, 50嵐 on Wushiland. A mainland-raised friend typing Simplified characters got nothing for most Taiwanese brands, so every CJK name is expanded into both scripts in the search key only, while the display keeps each brand's own spelling.
Leniency has a cost, and the cost is a wrong match. So I measured collision risk across the whole alias list before any of it shipped unguarded: 52 of 243 aliases changed which results came back, and just two reached across more than one brand, one a name two chains genuinely share, the other two unrelated shops. Cheap to check, and it's the difference between believing the leniency is safe and knowing it.
Every milestone build still boots from its commit, so the progression can be shown rather than described.
One rule, stated plainly: every surface is a cream, every ink is a brown, and there are no neutral grays anywhere, not even in the shadows. Round is the motif, with one piece of grammar holding it together: a status chip and an action button never share both silhouette and fill.
The system arrived the long way round: I ran a visual audit that recorded only problems, then built a Figma system, then wrote a spec, then swept the code. Code is the canonical source of tokens and Figma mirrors it, because the app is the artifact that ships.
The popup is the component I argued down to the pixel, but it didn't start as a design problem. For a long stretch it just accumulated: an action here, a line of metadata there, each addition sensible on its own.
Only once the action row ran out of room did it earn a proper sitting: the same card drawn several ways at once, arguing the open questions side by side rather than one release at a time. Where the address goes. Whether the chip carries the closing time. Whether the actions share a row with the primary or get one of their own.
What settled it was structural rather than visual: the list row is the source of truth, and every other surface mirrors its anatomy.
Mobile started as a derivative of the desktop layout, and it showed. A wordmark header ate the top of the screen, a filter bar ran off the edge, a floating pill hid the list, and selecting a shop could put that shop on screen twice, once in a map popup and once in a detail sheet underneath it.
Four pieces became one bottom sheet with three heights: the map with search docked at the bottom, the list or a selected shop's card, and reading height. Selection always resolves to the same place whichever way you got there, so tapping a pin and tapping a list row end in the same view instead of two different ones.
The header went next, on both breakpoints. The wordmark shrank to a floating chip, then moved into an About dialog behind a hamburger, and the map controls dropped to the bottom to ride up with the sheet. Chrome kept losing to the map, which is the point: the map is the product.
Two smaller decisions run on the same logic. Hours adopted the pattern people already know from Google Maps, the week rotated to start today, with today's row never moving as the other six expand below it, after I judged the custom version as weird. And the sheet's resting height is measured against the tallest real shop card, so its edge never shifts as you move between shops.
This was AI-assisted at volume, with Claude Code as the primary implementer. Nearly all of the code is its; my part is directing it, specifying behavior, testing every change in localhost or dev, finding the wrong behavior and describing it precisely enough to get it fixed, and reading the code where I know to look. The interesting part is the operating system around that: an instruction file that reads like an engineering process doc (investigate before changing, state the plan and its blast radius, flag rather than fix, test inline, append to the log), a session log with an entry for every working session since day one, and persistent memory for the things a repo can't record, like decisions already made and premises already tested and failed.
Where it helped most was work with a wide surface and a narrow judgment: research agents fanned out to find Instagram accounts for the chains, and one checked all 1,037 outbound URLs and found 80 broken, including a chain domain hijacked to serve gambling spam. Where it struggled was the opposite shape. "Flag, don't fix" let a known data-shape bug sit for two months as "harmless until something reads it," and by the time the audit caught it, the classifier was the thing reading it. The docs drifted 14 places behind the product. And the iOS keyboard, which can't be emulated, ended as an accepted imperfection: I built a viewport simulation harness and an on-device debug overlay so field reports became data, then stopped, because the constraint set made every alternative worse.
The habit underneath all of it: claims get checked by measurement, and the measurement gets written down. Data fixes are written as patterns rather than lookups ("tea token plus meal-category food noun means exclude"), so the next shop matching the pattern gets the same treatment without a code change.
There's no end-to-end suite. UI verification lived in rigorous but disposable probe scripts, and the best of them deserve to be regression tests.
The site is live and still moving: every boba shop in California, open-closed status computed identically for a viewer anywhere, and a runtime bill that doesn't move when traffic does. The next ambition is coverage beyond the state.
What's missing is usage. There are no visitor numbers on this page, because I never instrumented the project that way. The honest outcome is the product itself.
What it demonstrates is the shape of the work: a product whose easy-looking surface sits on classification and freshness problems, with cost shaping both, and a process where premises get measured and the failures are documented next to the fixes.