Sample frameDay 157·AI Fluency · Testing

Before the renewal date. Every demo you've ever watched was driven by someone who already knew the answer.

AI Fluency.

The Bake-Off

5 min read·Apply by lunch

The question

When you last chose a tool your team now runs on every day, what sat under the decision beyond a demo somebody else drove — and why haven't you run it against your own work since?

The idea

In 2022 Percy Liang and the team at Stanford's Center for Research on Foundation Models published Holistic Evaluation of Language Models, known as HELM, and its lasting contribution was less about which system won than about the shape of the question. Before it, models were reported on wildly different tasks, and some prominent ones shared not a single benchmark in common. HELM ran thirty models across forty-two scenarios and reported seven things about each: accuracy, calibration, robustness, fairness, bias, toxicity and efficiency. Efficiency sits in that list on purpose, because cost and speed are part of what a system is rather than a footnote to it. The argument transfers directly to your buying decision. No single number describes a system, and the number a vendor leads with was chosen by the vendor. What you need isn't a better leaderboard; it's the same battery of measures, run on scenarios that belong to you.

Members only · 4 principles + template + AI mentor

364 more frameworks are waiting.

The Bake-Off is Day 1 of 365. One framework every morning for a year — across strategy, sales, negotiation, leadership, and more.

$1 /day

Billed $365/yr · cancel any time

365 frameworks

12 topics · templates · AI mentor

Reserve your place — $365/yr →

One payment a year · the whole library · no auto-renewal.

More AI Fluency frameworks

All 20 ai fluency frameworks →