A pod dying and a system recovering are not the same thing. That gap is where most of the interesting failures actually live, and it's what pushed me to build fault-sentinelhttps://github.com/aashiruu/fault-sentinel, a lightweight, Kubernetes-native ...