CNST: Kubernetes Job Basics

RX-M instructor Christian Lacsina explores the fundamentals of Kubernetes Jobs, their uses, and how to create them. Kubernetes Jobs are a type of controller that allows users to define one-off tasks to be executed within Kubernetes pods. These tasks are designed to be short-lived, terminating once a specific goal is achieved.

Video Transcript

Hello everybody and welcome to another RX-M Cloud Native Short Take! This is Christian Lacsina from RX-M and today I’ll be short taking the basics of the Kubernetes Job controller; informing you on what it is, what it can be used for, and of course demonstrating how to create one. Jobs are a type of controller in Kubernetes that allow users to define one-off tasks to run within Kubernetes pods. These tasks are meant to be short-lived application runs which exit once some kind of goal is reached. So the pods themselves are meant to run and then exit, showing the “completed” status. This is in contrast to other Kubernetes controllers like Deployments which run their applications and always keep them running. Jobs will run pods up until they exit; jobs are great for a few things like: cleanup work such as clearing image caches or rotating logs on a host. They’re also great for processing or deleting messages from message queues and they’re also really good for doing on-demand calculations which is why you might see jobs often used for serverless implementations that are run on Kubernetes. 

Now that we know what jobs are, let’s take a look at a sample job. I have a sample job right here and I’ve broken it down into three different categories. We have our API identifiers of course; that’s going to define what kind of object we’re creating in Kubernetes and how we can find it in the API. 

Next we’re going to have the spec and this is where some of the job specific settings are going to be defined. The goal of the job is one of those things such as the number of completions which I have defined here. Next are the number of pods to run in parallel if you wanted that to happen. So I’ve got that defined as 2 and I don’t have this shown here, but also a retention period before automatic removal of the job, which is something I’ll be addressing when we demo it. 

Finally there’s the Pod specification. If you’re familiar with Kubernetes Deployments then you’ll know that this pod specification is defined under the template and everything that you might know about a Kubernetes pod is going to apply here. This is an embedded pod specification so every field aside from the metadata name is valid. There is one caveat to this pod spec though in that the restart policy must either be OnFailure or Never. So you don’t want the kubelet to try and recreate any containers that have exited because that is the yardstick that the job uses to consider itself finished or still in progress. 

Let’s go ahead take a bit of a demo and see how we can use jobs. My Kubernetes cluster is a three node cluster with mixed architectures: one of the worker nodes is a Raspberry Pi named ian-pi8-2. I also have the job yaml that I showed previously. So let’s go ahead and apply this to my cluster; we’ll watch it actually run its course, see the job in action, and what kind of behavior you should expect when you’re running Kubernetes jobs. We’ve got that going here at the top of the screen; I have my kubectl get jobs,pods command set with a watch and we can see that it is currently running through all of its completions. I set a goal of 20 completions and I also set the parallelism of having 2 pods run at the same time–which is what we see on the screen. So right now we’ve got 2 and let’s just watch this and see exactly how this job is going to continue working for us. While that’s going on though you might notice that the job pods that have finished are just sitting here! Is that going to be a problem? Actually that’s not that bad of an arrangement; when the pods themselves are completed you are able to do things like look up their logs or maybe perform a describe command on them and otherwise get a pulse on them. Perhaps even see how they had actually worked as well as confirm if their outputs are actually there. So for now we’re going to keep watching this particular job run its course.

Currently at 8 completions so we need about 10 more or a little more than 10 of course. A couple of minutes later we are now at 20 of 20 completions and we can also see that the pods themselves are, as I mentioned earlier, just sticking around. Now at this point I am able to go through and do something like kubectl logs on one of the pods. We’ll do the one at the top and we can see that it has indeed calculated pi to I believe 2000 places. So that’s going to be the log output. I’ve confirmed that and we can see here that there are no longer any additional pods being run at this point. In order to clean up the job pods all one needs to do, just like with Deployments, is just delete the parent job. Now there is a setting that will allow you to determine how long a job can remain before it gets automatically removed but by default there is no such limit and so it is going to be up to you to actually remove the job on your own. We’ll go ahead and do that right now by using kubectl delete with the yaml file. Once that’s been removed if we take a look we see that there are no more pods and also there are no more jobs.

Those are the basics of Kubernetes jobs. There are some additional features that jobs are going to be subject to, such as a schedule, through another controller called the CronJob. There are a myriad of other settings that you can use to help customize your job runs.

Kubernetes jobs are just one of the things that you will learn when attending a course from RX-M. You can peruse our existing course catalog or make your own course with the custom course builder including the Batch Jobs module and any number of other modules from our cloud native topics!

I’ve been Christian Lacsina with RX-M. This has been another RX-M Cloud Native Short Take! Thanks for watching!

Secret Link