Talk it through with Aurelius
Library›Aurelius›The problem
Aurelius · Work & Leadership
Knowledge + Guidance

Why can I spot good output but not evaluate a model?

You can tell when output is good. You call that instinct, and you trust it. Then someone asks you to evaluate the model, and you freeze — as if evaluation were a different skill, owned by other people. It is not. Your eye already knows what good looks like. What you lack is the discipline to write it down and test it twice. The work is not mysterious. It is translation — turn a feeling into a question, a question into a test, a test into a habit you run every time, not only when you remember.

◆ How this problem reads on the two dials
GuidanceKnowledge
Coaching
More to learn
1:1 with AureliusWith others (a Pod)
Mostly you & the coach
A little with peers
There is a real technique to teach — turning judgment into testable criteria — but the harder part is building the habit of running it every time, which calls for coaching.
How the two dials adapt to you →
What’s really going on

You already judge well — you lack a method, not an eye. Name three things you check when output looks right. Turn each into a question you can answer with yes or no. Run it on ten examples, not one. That written test, repeated, is rigor. Nothing more is required of you.

🔒 What you’ll build togetherUnlock by starting
A moveWrite down the three things your eye checks — before you look at the next output.
A moveTurn each into a yes-or-no question a stranger could answer without asking you.
A moveRun your test on ten outputs, not one — your eye lies to you on a single case.
A moveScore blind: hide which model made which output before you judge it.
A moveKeep a written log of every time your gut was wrong. Read it before you trust your gut again.

What changes unlock by starting

  • You can name, in writing, the three or four things that make output good — not just feel it.
  • You catch your own bias toward whichever model you already expected to win.
  • You can hand your test to someone else and get the same verdict you got.
  • You stop calling evaluation a mystery and start calling it a checklist you built yourself.
One object, two jobs: a public answer to a real problem, and — the moment you start the chat — Aurelius’s live plan for your version of it.