In the early days of running Kubernetes on premise we would get a user who would complain about ‘intermittent kubectl errors’. Typically this will be
This story starts when I noticed that nodes where going into a ‘NotReady’ status in a cluster. One ‘NotReady’ node was drained and rebooted only
I found this site online and thought it was too good not to share. There are lots of interesting failures and postmortems within production Kubernetes
