Talk it through with Aurelius
Library›Aurelius›The problem
Aurelius · Work & Leadership
Knowledge + Guidance

Why can't our team agree if a model's output is good?

You and your team look at the output and agree it seems good. That is not evaluation. That is taste. Taste is cheap — it costs nothing, and it flatters everyone who shares it. The defeat you feel is not because evaluation is impossible. It is because nothing is written down. There is no fixed target, so every disagreement becomes a clash of impressions, and impressions cannot be refuted, only repeated louder. You cannot make a model better by liking it harder. But you can choose, today, before the next output appears, to write three plain sentences of what 'good' means for your case. That is within your power. Do it before you look, not after.

◆ How this problem reads on the two dials
GuidanceKnowledge
More coaching
Some to learn
1:1 with AureliusWith others (a Pod)
Some one-to-one
Practise with peers
The team likely knows evaluation concepts already; what they lack is the discipline to apply criteria before judgment, which is a coaching problem more than a knowledge gap.
How the two dials adapt to you →
What’s really going on

Liking an output is not evaluating it. Your team feels defeated because you argue over impressions, not criteria. Choose this instead: before anyone looks at output, write down what 'good' means in plain sentences. Judge against that — alone first, then together. The confusion ends where the criteria begin.

🔒 What you’ll build togetherUnlock by starting
A moveBefore anyone sees new output, write down three plain sentences of what 'good' means for this case.
A moveChoose five hard cases that matter, not the easy ones that always look fine.
A moveEach person judges alone first, against the written criteria. Compare only after.
A moveWhen you disagree, find the exact criterion in dispute — not the overall feeling.
A moveRevisit and sharpen the criteria after every round. A vague rule taught you nothing.
PractiseCriteria Before Output · a Pod of 4 · 30 min

What changes unlock by starting

  • Fewer arguments that are really just different people liking different styles
  • A written standard the team points to, instead of relitigating taste each time
  • Failures caught that 'looks good' always missed
  • Faster decisions, because disagreement now has a clear object
One object, two jobs: a public answer to a real problem, and — the moment you start the chat — Aurelius’s live plan for your version of it.