Why can I spot good output but not evaluate a model?
You can tell when output is good. You call that instinct, and you trust it. Then someone asks you to evaluate the model, and you freeze — as if evaluation were a different skill, owned by other people. It is not. Your eye already knows what good looks like. What you lack is the discipline to write it down and test it twice. The work is not mysterious. It is translation — turn a feeling into a question, a question into a test, a test into a habit you run every time, not only when you remember.
You already judge well — you lack a method, not an eye. Name three things you check when output looks right. Turn each into a question you can answer with yes or no. Run it on ten examples, not one. That written test, repeated, is rigor. Nothing more is required of you.
What changes unlock by starting
- You can name, in writing, the three or four things that make output good — not just feel it.
- You catch your own bias toward whichever model you already expected to win.
- You can hand your test to someone else and get the same verdict you got.
- You stop calling evaluation a mystery and start calling it a checklist you built yourself.