Infrastructure deployment challenges
Manual deployment of a new multi-node Kubernetes cluster can potentially be a slow, error-prone, and time-consuming process. It requires a bit of planning, allocation of the required resources, and then deployment and configuration of all the related software components. Once the cluster is up and running, we have to take care of its lifecycle management. What about if we have more than one Kubernetes cluster–for example, separate development, staging, and production environments? How could we be sure the configurations and the needed changes are done consistently across all of them? How could we reduce the risk of human error and track all the changes over time?
There are a number of solutions taking different approaches in solving this problem which are applicable to various use cases. Solutions that describe the deployment steps and configuration changes using an easy-to-read language like yaml allow multiple team members to collaborate, review, and version control the infrastructure changes, reducing the risks of human error–effectively abstracting the infrastructure as yet another piece of code.
Automation to the rescue!
Tools like kind or minikube enable users to quickly spin up and destroy container- or virtual machine-based Kubernetes clusters. There are also solutions based on Ansible playbooks, Terraform, or even scripts using the Kubernetes reference installer tool: kubeadm.In this article, we are going to explore a very popular OpenSource tool called kOps (Kubernetes Operations) which aims to solve Kubernetes infrastructure management in the cloud. In the project’s maintainer’s own words:
We like to think of it as kubectl for clusters.
kops will not only help you create, destroy, upgrade, and maintain production-grade, highly available Kubernetes cluster, but it will also provision the necessary cloud infrastructure.
With the help of kOps, we can automate various Kubernetes cluster management tasks in cloud environments like AWS and GCE. The DigitialOcean, Hetzner, and OpenStack clouds have beta support at the time of writing and support for Azure is also being worked on, although currently in alpha.
It is easy to get started
To use kOps, we need to take care of a few prerequisites:
- Have kops, kubectl, and aws cli client tools installed
- Have an account with the required minimal set of permissions assigned in our cloud provider of choice
- Choose the right DNS solution for creating DNS records for the new Kubernetes cluster
- Create a storage location for storing the state of the Kubernetes clusters managed by kOps
In order to illustrate how easy it is to create a new Kuberentes cluster in AWS, we are going to create a small three node cluster consisting of one Control Plane node and two worker nodes:

Tooling installation
We will start by installing the kops cli client. We will need it for all Kubernetes cluster lifecycle management tasks. Currently, there is no graphical user interface for it, but that is not really a problem as the client (based on the Cobra library) is quite intuitive and well-documented on the project page. The kOps client is supported on all major platforms – Linux, MacOS, and Windows and its installation is pretty straightforward using either the native package manager or by downloading the binary from the project releases.
- MacOS using brew:
$ brew update && brew install kops
- Linux using a binary:
$ curl -Lo kops https://github.com/kubernetes/kops/releases/download/$(curl -s https://api.github.com/repos/kubernetes/kops/releases/latest | grep tag_name | cut -d '"' -f 4)/kops-linux-amd64
$ chmod +x kops
$ sudo mv kops /usr/local/bin/kops
- Windows: get the latest kops-windows-amd64 binary from the project GitHub releases page.
Instructions on how to install the Kubernetes CLI client: kubectl also depend on your OS. In case you don’t have it already, you can check this section of the Kubernetes Documentation portal for more details.
Regarding the AWS cli client installation steps, please check this guide. If you are using another cloud provider, please check the respective kOps getting started page for the specific client that you might need to install.
Manage Cloud Access
In order to deploy resources in a Cloud environment, we will need an account for one of the supported providers. For the purpose of this guide, we are going to use AWS. If you don’t have an AWS account, you could create one for free using the AWS Free Tier.
Note: As with all cloud providers, if you are not careful with your resource usage, you can quickly go beyond what the free tier offers, incurring charges!
If you are building the cluster just for testing purposes, make sure you delete all the created resources once done. We can’t take any responsibility for any cloud charges resulting from following the instructions here.
In order for kOps to be authenticated with the cloud provider, it needs some credentials. If you have an existing user account, you can use it or create a dedicated service account for kOps as long as either has the required permissions:
AmazonEC2FullAccess
AmazonRoute53FullAccess
AmazonS3FullAccess
IAMFullAccess
AmazonVPCFullAccess
AmazonSQSFullAccess
AmazonEventBridgeFullAccess
The list of the required permissions for the other cloud platforms can be found underneath the kOps Getting Started documentation section for the respective provider. The documentation explains really well how to create a new IAM user and an IAM group, assign the user to the group, and configure the AWS cli with the new user credentials in case you want to use a separate service account.
Creating storage location for the kOps state
With the help of kOps, we are managing Kubernetes clusters in the cloud; it also makes sense to store the blueprints for the creation of those clusters in the cloud provider’s object storage solution. In the case of AWS, it uses an S3 bucket; in the case of GCE, it uses a Google Cloud Storage bucket, and so on.
To create an S3 bucket for storing the state of our new Kubernetes cluster, we could either use AWS web ui or the aws cli client:
aws s3api create-bucket \
--bucket frsca-k8s-cluster-state-store \
--region us-east-1
The new bucket that we created has the name frsca-k8s-cluster-state-store and it will be stored in the us-east-1 region.
Choosing the right DNS solution
The next thing we need to take into account is DNS for the cluster components. kOps offers great flexibility supporting a number of DNS setups:
- Using AWS purchased/hosted domain or subdomain
- Using AWS Route53 dns service for (sub)domains purchased outside of AWS
- Using private DNS servers
Besides those three options, there is also a fourth one: Gossip DNS. It uses peer-to-peer network communication instead of relying on an external DNS service. In this regard, it offers some advantages:
- No need for an external DNS service
- We can deploy Kubernetes clusters in regions where Route53 is not supported
- No impact on the cluster as a result of external DNS issues
To use the Gossip-based DNS deployment we have to use a special domain name for our newly deployed cluster ending with: .k8s.local. To keep the example simple, let’s call our kubernetes cluster: frsca.k8s.local.
As part of our FRSCA blog series, we are going to explore how to install tools from the FRSCA architecture on top of Kubernetes to help us secure our software supply chain.
Kubernetes cluster creation with kOps
Before we begin, let’s store some of our configuration in environment variables to reduce some typing and avoid potential typos.
$ export NAME=frsca.k8s.local
$ export KOPS_STATE_STORE=s3://frsca-k8s-cluster-state-store
Automating the creation of a Kubernetes cluster is pretty cool, but how are we going to access our nodes for management and troubleshooting? kOps uses ssh key-based authentication. We can use ssh-keygen to create a new public/private key pair, or in case you want to reuse some of your existing keys you can simply provide the respective public key by adding the following flag: --ssh-public-key.
Most cloud providers have their cloud infrastructure spread between multiple data centers (zones) in a given geographical location (region). Spreading our cluster nodes between multiple zones in a region will improve resilience in the case of a single data center outage; we just need to tell kOps which zones we want it to use!
In AWS we can list all zones belonging to a single region by using the aws cli:
$ aws ec2 describe-availability-zones --region us-east-1 |grep ZoneName
"ZoneName": "us-east-1a",
"ZoneName": "us-east-1b",
"ZoneName": "us-east-1c",
"ZoneName": "us-east-1d",
"ZoneName": "us-east-1e",
"ZoneName": "us-east-1f",
For this article, we are not going to disperse the cluster into multiple zones, so we are going to use just the zone: us-east-1a.
Next, lets create the Kubernetes cluster configuration using the command: kops create cluster:
$ kops create cluster --name=${NAME} \
--cloud=aws \
--zones="us-east-1a" \
--control-plane-zones="us-east-1a" \
--node-count=2 \
--control-plane-count=1 \
--node-size="t3a.medium" \
--control-plane-size="t3a.medium" \
--topology="public" \
--state ${KOPS_STATE_STORE}
I1120 00:13:55.511911 69067 new_cluster.go:1373] Cloud Provider ID: "aws"
I1120 00:13:57.612949 69067 subnets.go:224] Assigned CIDR 172.20.0.0/16 to subnet us-east-1a
Previewing changes that will be made:
I1120 00:14:26.016364 69067 executor.go:111] Tasks: 0 done / 109 total; 44 can run
W1120 00:14:26.620804 69067 vfs_keystorereader.go:143] CA private key was not found
I1120 00:14:27.175363 69067 executor.go:111] Tasks: 44 done / 109 total; 23 can run
...
Will create resources:
AutoscalingGroup/control-plane-us-east-1a.masters.frsca.k8s.local
Granularity 1Minute
InstanceProtection false
LaunchTemplate name:control-plane-us-east-1a.masters.frsca.k8s.local
LoadBalancers []
MaxInstanceLifetime 0
MaxSize 1
Metrics [GroupDesiredCapacity, GroupInServiceInstances, GroupMaxSize, GroupMinSize, GroupPendingInstances, GroupStandbyInstances, GroupTerminatingInstances, GroupTotalInstances]
MinSize 1
Subnets [name:us-east-1a.frsca.k8s.local]
SuspendProcesses []
...
Must specify --yes to apply changes
Cluster configuration has been created.
Suggestions:
* list clusters with: kops get cluster
* edit this cluster with: kops edit cluster frsca.k8s.local
* edit your node instance group: kops edit ig --name=frsca.k8s.local nodes-us-east-1a
* edit your control-plane instance group: kops edit ig --name=frsca.k8s.local control-plane-us-east-1a
Finally configure your cluster with: kops update cluster --name frsca.k8s.local --yes --admin
$
Note: By using the option: --topology="public", we’ve made our Kubernetes cluster nodes directly accessible from the internet. We should never do that for production environments! A list of recommendations for production setups is available here.
The configuration for our Kubernetes cluster has been created and stored within S3. Let’s take a look:
$ aws s3 ls --recursive --human-readable ${KOPS_STATE_STORE}
2023-11-20 00:48:20 0 Bytes frsca.k8s.local/clusteraddons/default
2023-11-20 00:48:19 1.2 KiB frsca.k8s.local/config
2023-11-20 00:48:19 324 Bytes frsca.k8s.local/instancegroup/control-plane-us-east-1a
2023-11-20 00:48:19 314 Bytes frsca.k8s.local/instancegroup/nodes-us-east-1a
$ kops get cluster
NAME CLOUD ZONES
frsca.k8s.local aws us-east-1a
$
We’ve created a base cluster configuration only; we have not created any cloud resources besides the configuration files in S3.
In order to edit the current cluster configuration we can use the following command:
$ kops edit cluster --name ${NAME}
It will open the yaml spec of the cluster configuration in our default text editor where we can review the generic cluster configuration:
apiVersion: kops.k8s.io/v1alpha2
kind: Cluster
metadata:
creationTimestamp: "2023-11-19T22:48:17Z"
name: frsca.k8s.local
spec:
api:
loadBalancer:
class: Network
type: Public
authorization:
rbac: {}
channel: stable
cloudProvider: aws
configBase: s3://frsca-k8s-cluster-state-store/frsca.k8s.local
etcdClusters:
- cpuRequest: 200m
etcdMembers:
- encryptedVolume: true
instanceGroup: control-plane-us-east-1a
name: a
manager:
backupRetentionDays: 90
memoryRequest: 100Mi
name: main
- cpuRequest: 100m
etcdMembers:
- encryptedVolume: true
instanceGroup: control-plane-us-east-1a
name: a
manager:
backupRetentionDays: 90
memoryRequest: 100Mi
name: events
iam:
allowContainerRegistry: true
legacy: false
kubeProxy:
enabled: false
kubelet:
anonymousAuth: false
kubernetesApiAccess:
- 0.0.0.0/0
- ::/0
kubernetesVersion: 1.28.3
networkCIDR: 172.20.0.0/16
networking:
cilium:
enableNodePort: true
nonMasqueradeCIDR: 100.64.0.0/10
sshAccess:
- 0.0.0.0/0
- ::/0
subnets:
- cidr: 172.20.0.0/16
name: us-east-1a
type: Public
zone: us-east-1a
topology:
dns:
type: Private
The configuration of the individual Control Plane and worker nodes are stored as part of InstanceGroup resources:
$ kops get instancegroup --name frsca.k8s.local
NAME ROLE MACHINETYPE MIN MAX ZONES
control-plane-us-east-1a ControlPlane t3a.medium 1 1 us-east-1a
nodes-us-east-1a Node t3a.medium 2 2 us-east-1a
$
From the output we can see the number and type of machines that will be created as part of each InstanceGroup. If we want to further modify them we can edit the InstanceGroup resources:
$ kops edit instancegroup nodes-us-east-1a --name frsca.k8s.local
apiVersion: kops.k8s.io/v1alpha2
kind: InstanceGroup
metadata:
creationTimestamp: "2023-11-19T22:48:18Z"
labels:
kops.k8s.io/cluster: frsca.k8s.local
name: nodes-us-east-1a
spec:
image: 099720109477/ubuntu/images/hvm-ssd/ubuntu-jammy-22.04-amd64-server-20230919
machineType: t3a.medium
maxSize: 2
minSize: 2
role: Node
subnets:
- us-east-1a
Having prepared and reviewed our new cluster configuration it is now time to actually create it!
Note: All instances created by kops will be built within ASGs (Auto Scaling Groups), which means each instance will be automatically monitored and rebuilt by AWS if it suffers a failure.
$ kops update cluster --name ${NAME} --yes --admin
I1120 11:55:31.500733 14866 executor.go:111] Tasks: 0 done / 109 total; 44 can run
W1120 11:55:32.127338 14866 vfs_keystorereader.go:143] CA private key was not found
I1120 11:55:32.198180 14866 keypair.go:226] Issuing new certificate: "etcd-manager-ca-main"
I1120 11:55:32.217093 14866 keypair.go:226] Issuing new certificate: "apiserver-aggregator-ca"
I1120 11:55:32.225256 14866 keypair.go:226] Issuing new certificate: "etcd-peers-ca-main"
I1120 11:55:32.231690 14866 keypair.go:226] Issuing new certificate: "etcd-manager-ca-events"
I1120 11:55:32.231705 14866 keypair.go:226] Issuing new certificate: "etcd-peers-ca-events"
I1120 11:55:32.246636 14866 keypair.go:226] Issuing new certificate: "etcd-clients-ca"
W1120 11:55:32.488991 14866 vfs_keystorereader.go:143] CA private key was not found
I1120 11:55:32.609868 14866 keypair.go:226] Issuing new certificate: "kubernetes-ca"
I1120 11:55:32.625286 14866 keypair.go:226] Issuing new certificate: "service-account"
I1120 11:55:34.770537 14866 executor.go:111] Tasks: 44 done / 109 total; 23 can run
I1120 11:55:36.941247 14866 executor.go:111] Tasks: 67 done / 109 total; 28 can run
I1120 11:55:38.803307 14866 executor.go:111] Tasks: 95 done / 109 total; 2 can run
I1120 11:55:39.322761 14866 executor.go:111] Tasks: 97 done / 109 total; 4 can run
I1120 11:55:41.186863 14866 executor.go:111] Tasks: 101 done / 109 total; 2 can run
I1120 11:55:43.886601 14866 executor.go:111] Tasks: 103 done / 109 total; 6 can run
I1120 11:55:44.614939 14866 executor.go:111] Tasks: 109 done / 109 total; 0 can run
I1120 11:55:44.620371 14866 update_cluster.go:322] Exporting kubeconfig for cluster
kOps has set your kubectl context to frsca.k8s.local
Cluster is starting. It should be ready in a few minutes.
Suggestions:
* validate cluster: kops validate cluster --wait 10m
* list nodes: kubectl get nodes --show-labels
* ssh to a control-plane node: ssh -i ~/.ssh/id_rsa ubuntu@
* the ubuntu user is specific to Ubuntu. If not using Ubuntu please use the appropriate user based on your OS.
* read about installing addons at: https://kops.sigs.k8s.io/addons.
$
Note: the creation of the ec2 instances can take a few minutes.
The Kubernetes cluster nodes are now created and being initialized:

kOps automatically updates our kubectl configuration file ~/.kube/config adding a new context for the newly created kubernetes cluster:
$ kubectl config get-contexts ${NAME}
CURRENT NAME CLUSTER AUTHINFO NAMESPACE
* frsca.k8s.local frsca.k8s.local frsca.k8s.local
$
We can validate the cluster readiness using the kops validate command:
$ kops validate cluster --wait 10m
Using cluster from kubectl context: frsca.k8s.local
Validating cluster frsca.k8s.local
INSTANCE GROUPS
NAME ROLE MACHINETYPE MIN MAX SUBNETS
control-plane-us-east-1a ControlPlane t3a.medium 1 1 us-east-1a
nodes-us-east-1a Node t3a.medium 2 2 us-east-1a
NODE STATUS
NAME ROLE READY
i-02c5243cea5f23694 node True
i-04784cb25b92c067b node True
i-0f270bd256adce1d9 control-plane True
Your cluster frsca.k8s.local is ready
$
After the new cluster is initialized we should be able to access it from the command line:
$ kubectl get no -o wide
NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME
i-02c5243cea5f23694 Ready node 2m26s v1.28.3 172.20.73.45 3.94.102.18 Ubuntu 22.04.3 LTS 6.2.0-1012-aws containerd://1.7.7
i-04784cb25b92c067b Ready node 2m24s v1.28.3 172.20.82.34 18.209.8.72 Ubuntu 22.04.3 LTS 6.2.0-1012-aws containerd://1.7.7
i-0f270bd256adce1d9 Ready control-plane 3m52s v1.28.3 172.20.119.153 54.196.135.210 Ubuntu 22.04.3 LTS 6.2.0-1012-aws containerd://1.7.7
$
Cleaning up AWS resources
If you’ve created a new Kubernetes cluster in AWS just for testing the deployment in a “production-like” environment, you should consider releasing the resources in order to avoid continuous charges.
You can use the kops delete command to validate the resources that will be released (without deleting them!):
$ kops delete cluster --name ${NAME}
I1120 12:47:42.082857 22992 delete_cluster.go:128] Looking for cloud resources to delete
TYPE NAME ID
autoscaling-config control-plane-us-east-1a.masters.frsca.k8s.local lt-057044f271dc66887
autoscaling-config nodes-us-east-1a.frsca.k8s.local lt-0866cba225d45306a
autoscaling-group control-plane-us-east-1a.masters.frsca.k8s.local control-plane-us-east-1a.masters.frsca.k8s.local
autoscaling-group nodes-us-east-1a.frsca.k8s.local nodes-us-east-1a.frsca.k8s.local
dhcp-options frsca.k8s.local dopt-0374c31cc95b5a986
eventbridge-rule frsca.k8s.local-ASGLifecycle frsca.k8s.local-ASGLifecycle
eventbridge-rule frsca.k8s.local-InstanceScheduledChange frsca.k8s.local-InstanceScheduledChange
eventbridge-rule frsca.k8s.local-InstanceStateChange frsca.k8s.local-InstanceStateChange
eventbridge-rule frsca.k8s.local-SpotInterruption frsca.k8s.local-SpotInterruption
iam-instance-profile masters.frsca.k8s.local masters.frsca.k8s.local
iam-instance-profile nodes.frsca.k8s.local nodes.frsca.k8s.local
iam-role masters.frsca.k8s.local masters.frsca.k8s.local
iam-role nodes.frsca.k8s.local nodes.frsca.k8s.local
instance control-plane-us-east-1a.masters.frsca.k8s.local i-0f270bd256adce1d9
instance nodes-us-east-1a.frsca.k8s.local i-02c5243cea5f23694
instance nodes-us-east-1a.frsca.k8s.local i-04784cb25b92c067b
internet-gateway frsca.k8s.local igw-0b70a777f5929b10b
load-balancer api-frsca-k8s-local-j3ama8 arn:aws:elasticloadbalancing:us-east-1:433017611331:loadbalancer/net/api-frsca-k8s-local-j3ama8/353bbfd704f112ac
route-table frsca.k8s.local rtb-03771c9cd5b76a17d
security-group api-elb.frsca.k8s.local sg-0268db53c7d674cad
security-group masters.frsca.k8s.local sg-0c4132730465bf233
security-group nodes.frsca.k8s.local sg-037d8d090ab804418
sqs https://sqs.us-east-1.amazonaws.com/433017611331/frsca-k8s-local-nth https://sqs.us-east-1.amazonaws.com/433017611331/frsca-k8s-local-nth
subnet us-east-1a.frsca.k8s.local subnet-051df654aff53feb1
target-group tcp-frsca-k8s-local-jm7aju arn:aws:elasticloadbalancing:us-east-1:433017611331:targetgroup/tcp-frsca-k8s-local-jm7aju/ce3f4a7107478e44
volume a.etcd-events.frsca.k8s.local vol-0e053c0555c104cc3
volume a.etcd-main.frsca.k8s.local vol-02cab7bfbdc27e616
vpc frsca.k8s.local vpc-02a863199cf902a54
Must specify --yes to delete cluster
$
Running the same command with the --yes flag will actually remove all of the listed resources, for example:
$ kops delete cluster --name ${NAME} --yes
Troubleshooting
At one point after you have your new kOps server installed, especially if you decide to keep it for longer, you might get an error like this one while trying to access it using the kubectl client:
$ kubectl get nodes
I1121 17:31:36.414385 26376 versioner.go:58] the server has asked for the client to provide credentials
E1121 17:31:36.980819 26376 memcache.go:265] couldn't get current server API group list: the server has asked for the client to provide credentials
...
error: You must be logged in to the server (the server has asked for the client to provide credentials)
$
The reason for this error is most probably that the certificates assigned to the kubeconfig file have expired. During the installation, kOps generates certificates with limited validity: 18 hours by default (can be customized) for authentication of the kubectl client to the Kubernetes API service.
To update the certificates and resolve the problem simply run the following command once more to renew your certificates:
$ kops export kubeconfig --admin
Using cluster from kubectl context: frsca.k8s.local
kOps has set your kubectl context to frsca.k8s.local
$
That should solve the problem in most cases.
Conclusion
As we’ve seen, it is really easy to get started using kOps in cloud environments. There are just a few prerequisites that we should meet anyway, no matter the deployment tool we use, and then we can create a new cluster with just a few commands.
In the following blogs, we are going to discuss in detail some of the more advanced cluster lifecycle management tasks and how kOps can help us to address them.