Kubernetes Troubleshooting

Learn about Kubernetes troubleshooting in this Cloud Native Short Take from RX-M. The module covers tools in cluster and workload troubleshooting, as well as troubleshooting workflows and some of the different steps in the process of workflow problem resolution. In this video Managing Partner Randy Abernethy reviews a troubleshooting scenario from RX-M’s open source Bust-a-Kube repo.

Video Transcript

Welcome to another Cloud Native Short Take. We’re going to take a look at the Kubernetes Troubleshooting module. In this module we discuss tools in cluster and workload troubleshooting, we look at troubleshooting workflows and some of the different steps in the process of workflow problem resolution, and we also take a look at a lot of the different failure modes in the Kubernetes cluster components.

We are going to begin with our Bust-a-Kube repo on GitHub. The idea behind this repo is to give you a bunch of different troubleshooting scenarios. The Bust-a-Kube repo gives you a whole bunch of different Kubernetes specifications in yaml file format that have something wrong with them. When you apply them to the cluster, it creates a problem and you need to sort that out. In the lab of this module we spend some time using some of these Bust-a-Kube examples. 

Let’s take a look at, for example, the workload-1 folder; here you have a README that explains how to use that particular Bust-a-Kube problem and a problem.yaml. You don’t want to necessarily look at this yaml. Some of the other READMEs may ask you to read the yaml and then apply it but most of the time you want to just apply it because in the real world, typically something else has created that file, like a CICD pipeline or perhaps it is the result of a Helm template. Now all the sudden you’ve got a workload that’s running and doesn’t work and you’re trying to figure out what’s happening. The idea is to read the README, follow the instructions and then see if you can solve the problem. If you get stuck the solution.md markdown has the example solution. Of course there’s 50 million ways to solve problems in Kubernetes. Quite often it may not be the only solution and yours may be just as good or even better. Let’s give this a try. Drop this in and run it. It says okay pod debug-pod1 “created”.

Let’s find out; one of the first things you might do is just poke your head around and see what’s happening using something like kubectl get pod. I can see a dead giveaway right here. I have zero out of one containers running in that pod. That’s a problem! And I can see that I have an init status of errImagePull.  We couldn’t even get started because the image failed to pull. There could be lots of reasons but this is one of the first things that you’re probably going to run into when you start working with Kubernetes: you’re going to put a typo in the image name. 

So when kubernetes looks at your spec,  it can tell if you’re providing a key that’s not in the Kubernetes specification. It can throw that out as an error but when you’re supposed to supply an image? It could be anything; there’s no way for Kubernetes to really know if that image is valid until it actually tries to pull it. We got the error at that exact point, so this is tricky. If you execute a command and it says “created” you think everything’s good. That’s not the case because Kubernetes has a lot of asynchronous parts that operate behind the scenes. 

There are a couple of ways that we could try to ferret out the problem. Number one, we could use kubectl get pod debug-pod1 -o yaml to get the entire object. We really just want to dump everything about it. There’s a lot of benefit to looking at the full yaml dump of anything in Kubernetes. You’re going to get all sorts of information from the Kubernetes side about what’s happening. You’ll also be able to reflect back on the things that you submitted. We were particularly interested in the containers–the error message we received was related to initContainers. This is telling us that there’s something wrong in the initContainer area of the spec. If we take a look at the image we’re trying to pull, it’s alpinelatest and the colon is missing. 

We could fix it in lots of ways. We could edit it with kubectl edit or we could just pull the spec and modify it. I am going to leave it like it is though and instead bounce back to the solution. The solution says “why don’t you get pods?” Right out of the gate if you run a pod and it says created all that really means is that the spec looked good and was put in etcd. I didn’t demo kubectl describe but that’s an obvious choice as your next step because you’re going to be able to get not only the highlights of the yaml that I dumped in full but you’re also going to be able to get the events that have occurred. This is nice because it gives you just the events for this resource and that can be pretty useful. You can see the Error: ImagePullBackOff right away and it spells it right out in the events. So there’s no question anymore once you use the events. If you wanted to download that file with curl or wget and make the change so that we have the correct tag you could then reapply that and everything would be good.

That’s an example of troubleshooting. It’s the simplest first of the troubleshooting scenarios in the Bust-a-Kube repo. If you get a chance, you might want to play with the Bust-a-Kube and we take pull requests, so if you’ve got great examples of tricky problems and ways to solve them that are fast and easy, we would love to have your contribution. We’re always adding stuff to the Bust-a-Kube repo as well so it will be growing over time. 

Pod troubleshooting is just one of the things that you’ll learn in the Kubernetes Troubleshooting module from RX-M. With the custom courseware builder you can build your own custom class with troubleshooting, stateful workloads, network policy, and everything else you could possibly want in the Kubernetes lexicon.

That is our Cloud Native Short Take on Kubernetes troubleshooting!

Secret Link