How do I design systems to be reliable instead of just fixing them?
Keeping things running and designing them to stay running are different arts. One is reacting well; the other is building so you have less to react to. It is no shortcoming to be strong at the first while still learning the second. Most people never even notice the difference. You have. Reliability is not added at the end; it is designed in from the start by asking what happens when each part fails, not if. The shift is to stop assuming things work and start assuming they will break, then building so the break is small and survivable. Take one system you keep alive and ask a single question: what is the one failure that would hurt most, and what would soften it. Redundancy, a fallback, an alert sooner. Fix that one thing. Reliability is built one weak point at a time, and you already know where the weak points are.