Sample frameDay 139·AI Fluency · Testing

Good morning. Everyone on your team has a favourite model. Not one of them can prove it.

AI Fluency.

The Held-Out Ten

5 min read·Apply by lunch

The question

When you last rewrote a prompt or moved to a newer model, what did you compare the new version against? If it was the two or three outputs you happened to glance at afterwards, you haven't measured anything. You've changed your mind and called it a test.

The idea

Ron Kohavi ran the experimentation platform at Microsoft and, before that, data mining at Amazon; with Diane Tang of Google and Ya Xu of LinkedIn he wrote Trustworthy Online Controlled Experiments (2020), the standard text on testing changes at scale. Its most useful number is also its most deflating: of the ideas Microsoft put through controlled experiments, only about a third improved the metric they were built to improve. A third changed nothing. A third made things worse. These weren't careless ideas. They were the surviving proposals of experienced product people who had argued for them in rooms full of other experienced people. The lesson carries straight into AI work, where you change a prompt, a model or a tool almost weekly and judge the change by whether the last few outputs felt better. The trap is that you're both the person who chose the change and the person marking it, working from memory, on a sample of two. Without a fixed set of cases and a known-good answer for each, you can't tell an improvement from a good mood.

Members only · 4 principles + template + AI mentor

364 more frameworks are waiting.

The Held-Out Ten is Day 1 of 365. One framework every morning for a year — across strategy, sales, negotiation, leadership, and more.

$1 /day

Billed $365/yr · cancel any time

365 frameworks

12 topics · templates · AI mentor

Reserve your place — $365/yr →

One payment a year · the whole library · no auto-renewal.

More AI Fluency frameworks

All 20 ai fluency frameworks →