We can train the model. Why does the system fall apart?
You built something that works on your machine. Then it met the world — other people's code, other people's timelines, a pipeline that has to run at three in the morning without you watching it — and it did not survive contact. You call this a failure of skill. It is not. No one told you that ML engineering and ML systems engineering are separate crafts. One is about the model. The other is about everything the model depends on to stay useful after you stop touching it. Your pod has people who can do the first. Almost no one has clearly claimed the second. That is the actual hole. You cannot fix a job no one has named. So name it this week, in front of your pod, out loud. Then choose who owns it. Not the most talented person — the one willing to hold it.
Because training a model and building a system are two different jobs, not one. The first is solitary craft. The second is coordination — pipelines, monitoring, handoffs, people. Your pod is failing at a job no one named, not at the job you already know. Name it. Assign it. Then work.
What changes unlock by starting
- Your pod names ML engineering and ML systems as two separate jobs, and stops pretending one person or one skill covers both.
- Someone owns system health specifically, so it is not everyone's job and therefore no one's.
- You stop diagnosing systems failures as modeling failures, and waste less time retraining what was never broken.
- Fewer surprises reach production, because someone is watching the seams instead of only the model.