Why can't our team agree if a model's output is good?
You and your team look at the output and agree it seems good. That is not evaluation. That is taste. Taste is cheap — it costs nothing, and it flatters everyone who shares it. The defeat you feel is not because evaluation is impossible. It is because nothing is written down. There is no fixed target, so every disagreement becomes a clash of impressions, and impressions cannot be refuted, only repeated louder. You cannot make a model better by liking it harder. But you can choose, today, before the next output appears, to write three plain sentences of what 'good' means for your case. That is within your power. Do it before you look, not after.
Liking an output is not evaluating it. Your team feels defeated because you argue over impressions, not criteria. Choose this instead: before anyone looks at output, write down what 'good' means in plain sentences. Judge against that — alone first, then together. The confusion ends where the criteria begin.
What changes unlock by starting
- Fewer arguments that are really just different people liking different styles
- A written standard the team points to, instead of relitigating taste each time
- Failures caught that 'looks good' always missed
- Faster decisions, because disagreement now has a clear object