Practical notes from building and evaluating AI products. Explore quality, cost, latency, workflow adoption, and the evidence behind product decisions.
Claude Code Is a Workflow, Not a Chat Box
What sustained agentic work taught me about exit criteria, evaluation, context, and review
Evaluating AI Products Beyond Accuracy
A practical measurement stack for quality, cost, latency, and product value
The Missing Layer in AI Adoption Is Workflow Design
Why useful AI spreads through pain discovery, ownership, contribution paths, and habit formation
Trustworthy Decision Systems Begin With Model Skepticism
How to validate signals, separate mechanisms from artifacts, and make uncertainty useful
Reducing Hallucinations by 60% Without Changing the Model
Retrieval optimization, prompt engineering, and A/B testing for enterprise LLMs
Measuring AI Impact When You Can't A/B Test
Using quasi-experimental methods to evaluate feature value at scale