loading...
A portfolio is a strange product: its content is the designer, and its craft is the argument. I wanted this one to demonstrate a way of working, and the way I work now is code-first. The legacy process I was trained in runs mockup, review, handoff, then a build that drifts from the mockup. With an AI pair in the editor, the working product is cheaper to change than the mockup was, so I made the decisions there instead. Type scale, spacing, motion, and color were all settled by looking at the rendered page and measuring it, and the design-system document was written into the repo next to the code it describes.
To be precise about who does what: nearly all of the code is written by Claude Code. My part is to direct it, specify the behavior, test every change in localhost or the dev deployment, find the incorrect behavior, and articulate it back precisely enough to get it fixed. I read and follow the code when I know where to look; I do not dig through it to diagnose a bug myself. Small fixes ship through the same loop. That is the fluency on offer here, and it is the fluency this site was built with.
The AI side of that is uneven, and the site is honest about it. Drafting, auditing, and blind review were all faster with a model in the loop. Sourcing was not: a model will write a confident claim with nothing behind it, so every pipeline on this site that lets a model write also has a gate that checks it.
Visitors should also be able to ask questions the way they would in a first conversation, about experience, process, or a specific project, and get grounded answers in my voice instead of navigating a page hierarchy.
This site has a design system, and it did not start in a design tool. The source of truth is a markdown file checked into the repo: tokens, type scale, spacing rules, motion tiers, and a running record of which decisions are settled and which are still open. The token layer is implemented in lockstep as CSS custom properties, the Tailwind theme, and a small JS constants file, commented as a set so drift is visible.
Figma exists in this workflow, and it mirrors the code. The design-system file carries variable collections with the same token names as the CSS, and its baseline pages hold captures of the live site. The intent going forward is to explore there when it is useful and pull decisions back as one-file token diffs.
The first scoping decision came early and shaped everything after it: the rendered page overrules the abstraction. A unification pass narrowed case studies to a 768px measure on line-length theory, and a look at the rendered pages reversed it the same day. Longer pages and double-wrapped bullets cost more than the theory saved. Every surface now runs one 1024px frame, and both the attempt and the reversal are recorded with their reasons. That rule is why the rest of the system was built by measurement rather than by mockup.
The rule I held myself to: wherever a decision could be measured, it ships with its measurement.
The per-study tile colors shipped as a fully validated palette, and then I removed them. The hues carried no information; they were decoration. The grid is monochrome now, red reserved for interaction. The residue is a validated color system with no consumer, kept in the design-system file with the reasoning for the reversal.
Motion went through the same discipline. One near-symmetric ease had been driving everything, which read as mechanical. It became a small set of tokens: expo-out for entrances, sharp-in for exits, springs for direct manipulation. Feedback got a 160ms floor because the old 100ms, used in 27 places, was below the threshold where a transition reads as a response.
Case studies run about ten viewports, and the original mount-everything animation played 90% of its entrance off-screen. Content now scroll-reveals, with a separate slower cascade reserved for the blocks actually on screen at load. Parallax was considered and rejected: the void stays still; the type moves.
The landing page is a centered composer. Type a question and it rides into the chat as a single continuous element: the bar travels to its docked position, widens, and the message sends from there. The entrance runs as a ladder of beats. The bar widens, the rule draws in, the controls fill, the conversation fades up, and the suggestion chips arrive last. Every other entry point into the chat gets a quieter version that resolves in place, because a nav link is not a composer and nothing should pretend to travel.
Details that took real iteration: the suggested prompts are a recruiter-ordered pool that backfills as chips are consumed. The chip rail collapses by animating its height with the contents sliding, so the conversation above never jumps. A returning visitor's stored thread gets a "reconnected" seam line instead of silently reappearing. The chat itself stays mounted across every route, so navigating away mid-answer does not lose the response.
The persona answers from my actual content. The system prompt, resume, and case-study summaries are injected directly into every request, with an optional vector-store search on top. The response is a strict JSON contract enforced with structured output, including a related_slugs field where the model cites which case study an answer draws on. Those citations are validated server-side against the real content directory, so a hallucinated slug can never become a dead link. This is the pattern for AI on this site: the model is allowed to propose, and the code checks the proposal.
The unglamorous edges got design attention too. The static prompt prefix is hashed once into a stable cache key so requests share the provider's prompt cache. Timeouts, length caps, and an abuse guard bound the public endpoint. Upstream errors never reach the visitor; failures answer in persona, with contact info. When someone asks for access to the password-gated studies, the chat hands over the password: the gate is a soft filter, and the password is injected from an environment variable.
A chatbot that speaks as me can fail in ways a UI cannot: wrong self-labels, leaked instructions, broken formatting, out-of-scope helpfulness. Any prompt edit, content regeneration, or model change can silently regress it.
So the persona has a regression harness. Twenty cases across four categories (recruiter questions, out-of-scope requests, prompt-injection attacks, and ambiguous input), each scored by ten programmatic checks covering the JSON contract, formatting rules, forbidden labels, and system-prompt leaks, plus LLM-judge checks for response pattern and staying in persona. Prompt changes ship with before-and-after runs. The persona rewrite and the link-citation contract both landed at 20/20.
The harness earns its keep on the boring failures. One formatting case started failing after a recent change. Sampling the same question repeatedly showed the old prompt failed it about as often (3 of 6 samples against 2 of 6), so the "regression" was pre-existing flakiness, now documented in the eval's readme instead of triggering a rollback.
The eval builds its prompt through a mirror of the production code path and prints the same cache-key hash, so drift between what is tested and what ships is visible immediately.
The site's content had the same failure mode as its prompt: claims drifting from ground truth. The fix had the same shape. Build a pipeline, and put the gate where the model is weakest.
Case studies are rewritten from their source design files through a process that collects provenance-tagged facts first and drafts only against those facts. Then two independent gates run: a fresh-context audit agent that classifies every claim as grounded or not, and a blind panel of three recruiter personas who score the old and new versions without knowing which is which. A rewrite only ships if it survives the audit and wins the panel. Six client studies went through it, and the studies were then reordered on the site by an agent-read quality rubric that I directed.
The chatbot reads generated summaries of the same case-study files the site renders, so one source of truth feeds both the persona and the pages. This rewrite is itself running through that pipeline.
The site is live on Vercel and does the two jobs it was designed for: the work is readable, and the conversation is real, with grounded answers, cited sources, and honest failure states.
The process left artifacts a hiring panel can inspect: a checked-in design-system document with settled-and-open decisions, an eval harness with its dataset and rubric, a content pipeline with its audit gates, and a commit history where visual decisions ship alongside their measurements.
As a personal project the outcomes stay qualitative by choice. There are no analytics yet, and the claims on this page trace to repo artifacts.