I make AI useful after the demo.
I imagine new ways AI can be useful, design the systems behind them, and build them into products.
My work spans the imagined user experience, the architecture that makes it possible, and the engineering that makes it dependable. I speak engineering, product, and research, and still work at the keyboard.
I help founders and engineering teams explore possibilities, prototype ideas, and work through difficult technical decisions.
What we can work on →My career and references →
Ideas I originated
Three examples from my recent AI product work. In each, the experience and the architecture developed together.
Guided experiences, built on a common system.
I wanted an AI conversation to lead into something a person could do: a guided practice with a purpose, a sequence, and a way to pick it up again.
I conceived the interaction model and designed a common platform for it. Definitions and content describe the experience; shared execution and state support it. That lets a new experience build on the same machinery.
The guided experiences and their platform shipped. The architecture also absorbed interactions that had previously required separate agent types.
Reflections worth returning to after a conversation.
I designed a way to turn a conversation into a useful reflection. The architectural idea was to separate a reusable pattern from its application to an individual.
The system can match an existing pattern or generate a new one, then shape a personal reflection from it. A second model checks the generated result against a rubric and returns feedback for another attempt. Nothing publishes when verification fails.
I designed and implemented the core generation and verification pipeline that went into production.
A weekly audio reflection drawn from your interactions.
I saw an opportunity to turn a person's conversations over time into an audio reflection: a way to return to the material in a different form.
I originated the concept and built the original prototype. The design brings several analyses together into a coherent narrative, then turns that narrative into a script and audio.
My team shipped a production implementation carrying forward that architecture. My contribution was the concept, prototype, and design.
The engineering that carries an idea through
I built evaluation harnesses and anonymization machinery, and moved our voice agent to a lower-cost model with the quality threshold fixed before the deciding run.
How I check the work Two examples from practice
2026-08I believed That the four claims I was about to put on this site had been verified. A pass had checked them a few days earlier and marked all four confirmed.
What was true All four were misstated. One was flatly false: it reported a perfect result against a corpus that had never been scored at that number, and called synthetic conversations production traffic.
How it surfaced I re-checked them against primary sources instead of against the document that had blessed them. The original pass had verified that the numbers existed, never what they were measured on.
What got built A verification rule: every number carries its denominator, or it does not ship. A figure is not verified when it is found — it is verified when you can name the population it was measured on.
The failure was one class repeated four times. Every number was real and most recomputed exactly from raw data. What failed was scope — population, denominator, attribution, and the strength of a guarantee. The worst instance welded two separately-scored corpora together: a perfect score belonged to a small constructed-truth set, and it had been attached to a much larger corpus that measured lower and was never production data to begin with. The source file said, in bold, "do not pool them." The original pass had read that file and cited it, while confirming the pooled version elsewhere in its own tables. The lesson generalizes past my own bio: a grep can confirm a number appears somewhere. It cannot tell you what the number counts.
2026-08I believed That a finished model comparison was clean and I could act on its conclusion. I liked the conclusion.
What was true One side had not been running the way I believed it was. I had been comparing a hobbled configuration against a healthy one and calling the difference a finding.
How it surfaced Not in the result. In the run logs, whose numbers did not look like what that configuration should have produced.
What got built A pre-flight check that verifies the configuration each side actually ran under before any comparison is trusted — and the habit of trying to break my own finished results before anyone else has to.
The corrected result held. It did not have to, and the study would have looked just as finished either way. The point is where the check ran: before the claim left the building, on a result I wanted to be true.
The range behind the work
Production machine learning since 2000. Most recently co-founder and Head of AI at ElizaChat. Before that I built and led ML engineering at Proofpoint, was chief architect at Domo's launch, lead architect at FamilySearch, and led groupware teams at WordPerfect.
Earlier work includes compilers, real-time control, aircraft simulators, and FlipDog, an ML-built job board acquired by Monster.com. The recurring contribution is seeing a useful possibility and finding the system that makes it practical.
Explore the career record and references →
Diagnose · Design · Build · Guide
A two-week question, a build, a standing relationship, or the conversation before the expensive decision. I size the engagement to the problem.
Cooks for a hundred people up a canyon. Builds giant-bubble machines. Sits with hospice patients. Serious about the work, not about myself.