Selected thinking

Writing on AI Evaluation, Workflows & Data Science

Practical notes from building and evaluating AI products. Explore quality, cost, latency, workflow adoption, and the evidence behind product decisions.

Claude Code Is a Workflow, Not a Chat Box

What sustained agentic work taught me about exit criteria, evaluation, context, and review

Read More →


Evaluating AI Products Beyond Accuracy

A practical measurement stack for quality, cost, latency, and product value

Read More →


The Missing Layer in AI Adoption Is Workflow Design

Why useful AI spreads through pain discovery, ownership, contribution paths, and habit formation

Read More →


Trustworthy Decision Systems Begin With Model Skepticism

How to validate signals, separate mechanisms from artifacts, and make uncertainty useful

Read More →


Reducing Hallucinations by 60% Without Changing the Model

Retrieval optimization, prompt engineering, and A/B testing for enterprise LLMs

Read More →


Measuring AI Impact When You Can't A/B Test

Using quasi-experimental methods to evaluate feature value at scale

Read More →


Some Toy Algorithms - Sentiment Classification

Implementing commonly used models from scratch

Read More →