Kubernetes Mastery: From Pods to Production Multi-Tenant Clusters
Kubernetes is the industry-standard platform for running containerized workloads in production. In this guide, you will learn Kubernetes end-to-end: starting with Pods and controllers like Deployments and StatefulSets, then moving on to networking with Services and Ingress, persistent storage with Persistent Volumes and StorageClasses, and security through RBAC and network policies. We also cover availability patterns like Pod Disruption Budgets and multi-tiered scaling, as well as controlling resource usage and cost with CPU/memory requests, quotas, and autoscaling. By the end, you’ll have a well-rounded understanding of what it takes to operate Kubernetes in a production environment. For hands-on practice, see Refonte’s DevOps Engineer Program, and for related DevOps topics, visit our DevOps Hub.
Kubernetes Basics: Pods, Containers, and Controllers
A Pod is the smallest deployable unit in Kubernetes (kubernetes.io). A Pod encapsulates one or more containers that share an IP address and volume mounts. In practice, most Pods contain a single application container, but you can also bundle helper containers (sidecars) in the same Pod to perform logging, proxying, or data conversion tasks. All containers in a Pod share storage and network context, essentially behaving like a single “logical host” (kubernetes.io). Because Pods are volatile entities - the control plane creates or destroys them to match the desired state - you should never depend on an individual Pod’s lifetime (kubernetes.io). For example, a Deployment might create a ReplicaSet that continuously spawns new Pods to replace any that fail. From one moment to the next a Pod’s IP might change, so you must use higher-level abstractions (like Services) to track workloads, not Pod IPs directly (kubernetes.io).
apiVersion: v1
kind: Pod
metadata:
name: example-pod
spec:
containers:
- name: nginx
image: nginx:1.23
ports:
- containerPort: 80
In the manifest above, a single Pod runs one container with the NGINX image. If this Pod crashes, Kubernetes can start a new one, but by itself that Pod could be short-lived. That is why controllers are used.
Kubernetes provides several controllers to manage Pods. A ReplicaSet ensures a fixed number of Pods are running (kubernetes.io). For example, you could create a ReplicaSet that maintains three Pods with a certain label. More commonly, you use a Deployment, which creates and manages a ReplicaSet for you. A Deployment lets you declaratively specify the Pod template, replica count, and update strategy. When you update the Deployment (for example, change the container image), Kubernetes will create a new ReplicaSet and perform a rolling update without downtime. For instance:
apiVersion: apps/v1
kind: Deployment
metadata:
name: webapp-deploy
spec:
replicas: 3 # keep 3 Pods running
selector:
matchLabels:
app: example
template:
metadata:
labels:
app: example
spec:
containers:
- name: nginx
image: nginx:1.23
ports:
- containerPort: 80
In this Deployment manifest, Kubernetes will always maintain 3 replicas of the NGINX container (with label app: example). If one Pod fails, a new one will be created. If you update the image version, Kubernetes will gradually replace the old Pods with new ones, one by one, handling rollbacks if needed. This declarative approach means you say what you want (3 replicas of a given Pod spec) and Kubernetes figures out how to converge to that state (kubernetes.io).
For stateful applications like databases, you use a StatefulSet instead of a Deployment. A StatefulSet assigns each Pod a unique, persistent identity and stable storage. As Kubernetes documentation says, “A StatefulSet runs a group of Pods, and maintains a sticky identity for each of those Pods” (kubernetes.io). This sticky identity is useful when each replica needs its own name or persistent volume. For example, if you have a 3-node Cassandra cluster, each node Pod must have a stable hostname (like cassandra-0, cassandra-1, cassandra-2) and each needs its own PersistentVolume. A StatefulSet creates and deletes Pods in order, ensures their names are stable, and makes sure each replica keeps its associated volume even if rescheduled. By contrast, a Deployment (via ReplicaSet) treats each Pod as identical, with no guarantee of order or persistent identity (kubernetes.io).
Another controller type is a DaemonSet, which ensures one Pod copy runs on every node (or a subset of nodes) in the cluster. DaemonSets are useful for tasks like logging or monitoring: for instance, you might run a log-collector or node-exporter container on all nodes automatically. If you add a new node to the cluster, the DaemonSet controller launches its Pod on that node as well.
Kubernetes also supports Jobs and CronJobs for batch tasks. A Job runs one or more Pods to completion and then stops. A CronJob schedules Jobs on a time-based schedule (similar to a Linux cron). These are useful for periodic or one-off tasks, such as backups, database maintenance, or batch data processing. For example:
apiVersion: batch/v1
kind: CronJob
metadata:
name: db-backup
spec:
schedule: "0 2 * * *"
jobTemplate:
spec:
template:
spec:
containers:
- name: backup
image: backup-tool:latest
args: ["--perform-backup"]
restartPolicy: OnFailure
Below is a summary table of common Kubernetes controllers and when to use them:
| Controller | Purpose | Example Use Case |
|---|---|---|
| Deployment | Maintain a set of identical Pods and roll out updates | Stateless web servers |
| ReplicaSet | Ensures a fixed number of Pods are running | (Used by Deployments) |
| StatefulSet | Pods with stable identity and storage | Databases, Zookeeper |
| DaemonSet | Run one Pod per node (or per selected nodes) | Logging agent, monitoring agent |
| Job / CronJob | Run Pods to completion (once or on a schedule) | Batch processing, backups |
In practice, you rarely manually create Pods without a controller, because controllers provide self-healing. When you write YAML, you choose the right controller kind (apps/v1) based on your application type (stateless vs stateful, daemon, or scheduled job). As a next step in a real workflow, teams often automate the creation of these resources via a pipeline: for example, after building a new container image your CI/CD system could update the image in the Deployment and apply it. See our CI/CD pipelines guide for how to integrate Kubernetes manifests into automated deployment workflows. You might also manage these manifests declaratively using GitOps tools (see our GitOps guide).
Kubernetes Networking: Services, Ingress, and DNS
In Kubernetes, each Pod gets its own IP address within the cluster network. However, Pods are ephemeral and may be created or destroyed at any time (kubernetes.io). To give a stable endpoint to a set of Pods, you use a Service (kubernetes.io). A Service is an abstraction that groups one or more Pods (selected by labels) and exposes them via a virtual IP (ClusterIP) and DNS name. This way, other Pods or clients can reach a Service address, and Kubernetes automatically load-balances traffic to the healthy Pods behind it. As the official docs explain, a Service “expose[s] an application running in your cluster behind a single outward-facing endpoint” (kubernetes.io). You don’t need to modify your app; the Service handles discovery and routing.
A Service is defined like this:
apiVersion: v1
kind: Service
metadata:
name: my-service
spec:
selector:
app: myapp
ports:
- protocol: TCP
port: 80
targetPort: 8080
In this example, any Pod with the label app=myapp will be eligible for this Service. Clients inside the cluster can connect to my-service on port 80, and Kubernetes will forward those requests to port 8080 on one of the Pods matching the selector. Importantly, if the set of Pods changes (e.g., scaling the Deployment up/down), the Service IP and DNS remain constant.
Kubernetes supports several Service types for different exposure levels. Below is a summary:
| Service Type | Description | Use Case |
|---|---|---|
| ClusterIP | Stable internal IP, only reachable within cluster | Service-to-Service communication |
| NodePort | Exposes service on a port on each Node’s IP | Simple external access (not load-balanced) |
| LoadBalancer | Provisions an external load balancer (cloud) | Production external HTTP/HTTPS traffic |
| ExternalName | Maps service to an external DNS name (CNAME) | Access legacy services outside cluster |
- ClusterIP (default) provides a virtual IP inside the cluster. Pods within the cluster can reach the Service at
<IP>:port. This is the most common type for internal communication. - NodePort exposes the Service on a static port on every Node (e.g., port 30001). This allows external traffic if those node ports are accessible, but is not ideal for production load balancing. It’s often used for quick testing or on bare-metal clusters.
- LoadBalancer leverages the cloud provider to create an external load balancer (e.g. AWS ELB, GCP LB) and routes traffic to the Service. This is the usual choice for publicly accessible services in managed Kubernetes. The cloud provider handles assigning an external IP.
- ExternalName makes the Service behave like a DNS alias (a CNAME) to an external host name. This is used to redirect calls to an external service without editing application code.
Beyond Services, an Ingress is the Kubernetes abstraction for HTTP(S) routing at the edge of the cluster. Think of an Ingress as a set of rules for how requests with certain hostnames or paths should be forwarded to Services. An Ingress works in concert with an Ingress Controller (such as NGINX Ingress, Traefik, or the cloud provider’s Ingress) that actually receives traffic and routes it. For example:
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: example-ingress
spec:
tls:
- hosts:
- example.com
secretName: tls-secret
rules:
- host: example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: webapp
port:
number: 80
This Ingress rule says: for HTTP requests to example.com, forward all paths (/) to the Service named webapp on port 80. The tls section says to terminate TLS using the tls-secret (a Kubernetes Secret with certs). The NGINX (or other) Ingress Controller would see this resource and configure its underlying proxy accordingly. In production, you might use an Ingress controller with additional features (rate limiting, WAF, redirects, etc.). Note that an Ingress resource alone does nothing without a controller running in the cluster.
Kubernetes also provides DNS for Services and Pods. By default, a Service named my-service in the namespace default will be discoverable at my-service.default.svc.cluster.local (or just my-service within the same namespace). The CoreDNS add-on automatically creates these records so Pods can simply use the Service name instead of IP. This DNS service is built into the cluster, so as soon as a Service is created, Kubernetes DNS tells other pods how to reach it by name.
In production, you will automate the deployment of these networking components. For example, after building a new frontend image, a CI/CD pipeline might update the Deployment and Service. We discuss pipelines in our CI/CD pipelines guide. You can also use a GitOps workflow where your Ingress and Service YAML live in Git and a controller synchronizes them. Lastly, for monitoring and logging traffic through your Services and Ingress, see our Observability section and free Prometheus guide.
Persistent Storage in Kubernetes: Persistent Volumes and StorageClasses
Pods have only ephemeral storage by default (things like emptyDir are wiped if the Pod restarts). To provide durable storage, Kubernetes uses PersistentVolumes (PV) and PersistentVolumeClaims (PVC). A PersistentVolume is a piece of networked storage in the cluster (for example, an AWS EBS volume or an NFS share) that an admin creates or that gets provisioned. A PersistentVolumeClaim is a request for storage by a user (developer) inside a namespace. When you create a PVC, Kubernetes finds a matching PV and binds them together.
A simple workflow is: first a cluster admin sets up a storage resource. For example, in a cloud you don’t have to pre-create a PV if you define a StorageClass for dynamic provisioning. The official docs note that “cluster administrators can also use StorageClasses to set up dynamic provisioning” (kubernetes.io). A StorageClass defines a provisioner (like kubernetes.io/aws-ebs) and parameters (disk type, IOPS, etc). When a PVC references that StorageClass, Kubernetes will automatically create a PV with those properties (e.g. an EBS disk of the requested size). This means developers don’t have to manually allocate storage in advance.
Here is an example of a PVC using a StorageClass:
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: data-claim
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 10Gi
storageClassName: fast-storage
This claim requests a 10GiB volume with ReadWriteOnce access. Kubernetes (using the fast-storage StorageClass provisioner) will bind it to a new or existing PV. You then mount the PVC into your Pod:
apiVersion: apps/v1
kind: Deployment
metadata:
name: webapp
spec:
template:
spec:
containers:
- name: webapp
image: webapp:latest
volumeMounts:
- name: data
mountPath: /data
volumes:
- name: data
persistentVolumeClaim:
claimName: data-claim
In this Pod template, the volume data-claim is mounted at /data. Now whatever the application writes to /data will persist on the underlying storage, even if the Pod is rescheduled to another node.
PersistentVolumes have several key properties:
- Access Modes specify how the volume can be mounted. Common modes include:
ReadWriteOnce(RWO): mountable read/write by a single node. Most block storage (like AWS EBS or GCE PD) is RWOnce.ReadOnlyMany(ROX): mountable read-only by many nodes (e.g. NFS).ReadWriteMany(RWX): mountable read/write by many nodes (e.g. NFS, GlusterFS, or AWS EFS).
Here is a summary table of access modes:
| Access Mode | Description | Example Storage Type |
|---|---|---|
| ReadWriteOnce | Mounted as read/write by one node | AWS EBS, GCP Persistent Disk |
| ReadOnlyMany | Mounted as read-only by many nodes | NFS |
| ReadWriteMany | Mounted as read/write by many nodes | NFS, GlusterFS, AWS EFS |
When you create a PVC, you choose an access mode. For example, most pods will use RWO. If your application truly needs multi-node RW, you must use RWX on network-attached storage.
-
Reclaim Policy dictates what happens when a PVC is deleted. By default a PV’s reclaim policy may be set to
Delete(the underlying storage is deleted) orRetain(storage is kept after the claim is removed). For production, decide carefully:Retainprevents accidental data loss, whereasDeleteautomates cleanup but will delete data when the claim is deleted. -
StorageClass : You can create multiple StorageClasses for different performance or backing storage. For example, you might have a
fast-ssdclass and astandard-hddclass. Each class has aprovisionerplugin (like CSI drivers). The StorageClass also defines the defaultreclaimPolicy. In a cloud, a default StorageClass is usually installed (e.g.gp2on AWS). You can list available classes withkubectl get sc.
Below is an example StorageClass and PV:
kind: StorageClass
metadata: { name: fast-storage }
provisioner: kubernetes.io/aws-ebs
parameters:
type: gp2
fsType: ext4
reclaimPolicy: Delete
apiVersion: v1
kind: PersistentVolume
metadata: { name: manual-pv }
spec:
capacity:
storage: 20Gi
accessModes: ["ReadWriteOnce"]
storageClassName: manual
awsElasticBlockStore:
volumeID: vol-0123456789abcdef0
fsType: ext4
If dynamic provisioning is enabled (with a StorageClass), you typically don’t create PVs manually. Instead, create a StorageClass and let PVCs trigger PV creation (kubernetes.io). If not using dynamic provisioning, then an admin would create one or more PVs in advance, and Kubernetes will bind PVCs to matching PVs.
In summary, to persist data in Kubernetes: define a PVC in your application spec, optionally create or rely on a StorageClass, and use volumes in Pods. Avoid storing data only in containers; rather, use the PV/PVC mechanism so that the data outlives individual Pod lifecycles. For more on managing infrastructure as code, see our Infrastructure as Code guide, which covers tools like Terraform or Pulumi to automate creation of resources including storage.
Access Control: RBAC and Pod Security
Securing a Kubernetes cluster starts with controlling who can do what. Kubernetes uses Role-Based Access Control (RBAC) as the primary mechanism for authorization. With RBAC, you create Roles or ClusterRoles that list allowed actions (verbs on resources), and then bind those roles to users, groups, or service accounts.
Roles are namespaced: they grant permissions within a specific namespace. ClusterRoles are cluster-wide (and can include non-namespaced resources like nodes or persistent volumes). You grant permissions by creating RoleBindings or ClusterRoleBindings that attach a (Cluster)Role to a subject (user, group, or ServiceAccount). Here is an example of a Role and RoleBinding:
kind: Role
apiVersion: rbac.authorization.k8s.io/v1
metadata:
name: pod-reader
namespace: development
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "watch", "list"]
kind: RoleBinding
apiVersion: rbac.authorization.k8s.io/v1
metadata:
name: read-pods-binding
namespace: development
subjects:
- kind: User
name: alice
apiGroup: rbac.authorization.k8s.io
roleRef:
kind: Role
name: pod-reader
apiGroup: rbac.authorization.k8s.io
In this example, a user alice in the development namespace is granted the pod-reader role, which only allows listing and viewing Pods. She cannot create or delete them, nor can she access other namespaces unless additional roles are bound. This principle - grant the least privileges necessary - is key. You typically create separate Roles for different functions (e.g., one Role for reading logs, one for deploying apps, etc.) and carefully bind them.
Kubernetes documentation notes that “RBAC Roles and Network Policies are namespace-scoped resources. Using RBAC, Users and Service Accounts can be restricted to a namespace” (v1-35.docs.kubernetes.io). In practice, this means you can map each tenant or team to a namespace, give their members Roles in that namespace, and they cannot see or manage resources outside it. To allow something cluster-wide (like creating namespaces or reading nodes), you use a ClusterRole and bind it at the cluster scope.
- Service Accounts: In addition to real users, Pods use ServiceAccounts to authenticate to the API server. By default each namespace has a
defaultServiceAccount, but you can create dedicated ServiceAccounts for your applications. You then bind Roles to those ServiceAccounts so Pods run with only the permissions they need. e.g., if a Pod needs to read ConfigMaps, give its service account a Role for that.
Beyond RBAC, Kubernetes provides Pod Security Standards (the evolution of the old PodSecurityPolicy) to enforce certain security settings on Pods. You can label namespaces to enforce one of the built-in profiles (privileged, baseline, restricted). For production, you typically start at baseline or restricted. These profiles, if set via admission labels like pod-security.kubernetes.io/enforce: restrictive, automatically block Pod specs that request host namespaces, run as root, or use other risky settings. You should also define securityContext options in your Pod specs, for example:
securityContext:
runAsNonRoot: true
runAsUser: 1000
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
These settings mean the container process must not run as root; it cannot gain extra Linux capabilities; and it has a read-only file system by default. You should also explicitly drop capabilities:
securityContext:
capabilities:
drop: ["ALL"]
Most applications run fine without extra capabilities like SYS_ADMIN, so dropping them by default is strong security practice.
Other pod-level security best practices include:
- Use imagePullPolicy: IfNotPresent with immutable image tags to avoid unpredictable image updates.
- Do not mount hostPath volumes or hostNetwork: true unless absolutely necessary, as they can give pods broad access to the node.
- Store sensitive data (credentials, API keys) in Kubernetes Secret objects, not in ConfigMaps or environment variables in plaintext. Moreover, enable encryption-at-rest for Secrets on the cluster (via the encryption provider config file) so that etcd does not store Secrets unencrypted.
- Define a PodSecurity admission configuration at the cluster level to ensure all namespaces have appropriate defaults (for example, labeling production namespaces with pod-security.kubernetes.io/default: "restricted").
Finally, secure the control plane and networking: restrict access to the Kubernetes API server to only approved IP ranges or VPNs; use TLS for all communication; and regularly rotate service account tokens and API keys. Avoid using the built-in Docker shim (use containerd or CRI-O which are more secure), and keep your Kubernetes version up-to-date to get the latest security patches. Periodically run tools like kube-bench or kube-hunter to check for misconfigurations.
To summarize key pod-level security rules:
- Run as non-root user: set runAsNonRoot: true, runAsUser to a non-zero UID, avoid privileged.
- Drop Linux capabilities: by default drop all (capabilities.drop: ["ALL"]) and only add necessary ones.
- Enforce PodSecurity: use namespace labels to enforce baseline/restricted policies cluster-wide.
- Use Secrets: keep sensitive data in Kubernetes Secrets with encryption at rest.
- Scan images: integrate vulnerability scanning in your CI/CD (trivy, clair, etc.) and use minimal base images (e.g. distroless) to reduce attack surface.
With RBAC + Pod Security + network segmentation (next section), you implement a defense-in-depth approach to securing your cluster.
Network Policies: Securing Pod-to-Pod Communication
By default, Kubernetes allows all pods to talk to each other (as long as they have network connectivity). However, in a production or multi-tenant environment you usually want to restrict traffic. NetworkPolicies act like firewall rules at the pod level (OSI layer 3/4). With NetworkPolicies, you can declare exactly which pods are allowed to communicate with which. The Kubernetes documentation explains: “If you want to control traffic flow at the IP address or port level ..., NetworkPolicies allow you to specify rules for traffic flow within your cluster” (kubernetes.io). Note that enforcing NetworkPolicies requires a compatible network plugin (Calico, Cilium, Weave, etc.) - otherwise the rules won’t be applied (kubernetes.io).
A NetworkPolicy has the form of a specification targeting some Pods (via labels) and then listing allowed ingress and/or egress rules. Critically, if no NetworkPolicy selects a Pod, then that Pod accepts all traffic (the default open model). Once you add a policy that selects a Pod, everything not explicitly allowed by a policy is denied for that Pod. A common pattern is to create a default-deny policy to isolate a namespace, then add allow rules for specific traffic.
For example, here’s a default-deny-all ingress policy:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: deny-all-ingress
spec:
podSelector: {}
policyTypes:
- Ingress
ingress: []
This policy applies to all pods in the namespace (podSelector: {}) and specifies no ingress from rules. The effect is to block all incoming connections to those pods. (You could do similarly for egress by listing an empty egress: [] and policyTypes: ["Egress"].)
Once you have a deny-all baseline, you add specific allowances. For example, to allow traffic from any pod in the same namespace, you could add:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-same-namespace
spec:
podSelector: {}
ingress:
- from:
- podSelector: {} # any pod in same namespace
- namespaceSelector: {} # or specify namespace if across namespaces
Or to allow only frontend pods to reach backend pods:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-frontend
spec:
podSelector:
matchLabels:
app: backend
ingress:
- from:
- podSelector:
matchLabels:
app: frontend
This policy selects Pods with app=backend and permits ingress only from Pods with app=frontend. All other traffic to the backend Pods will be blocked.
NetworkPolicies can also control egress. For example, to allow only DNS and HTTP egress, you might use IP blocks or allow specific ports. That tends to be more advanced. In many clusters, internal communication is the focus so mostly ingress rules are written.
NetworkPolicies are key for multi-tenant security: they can enforce network isolation between teams’ workloads. For instance, you might label each namespace and then a top-level policy that denies cross-namespace traffic except for specific allowed cases. Note: if you want isolation by namespace, it’s common to apply a default-deny-all ingress policy in each namespace automatically.
In short, use NetworkPolicies to implement a zero-trust baseline. Start with a broad deny-all policy and then incrementally add allow rules. This way you minimize unexpected lateral traffic in clusters. For more on troubleshooting and understanding traffic flows, see our Observability page which covers monitoring network metrics, and consider logging pod network events as well.
Pod Disruption Budgets and High Availability
High availability is a core production concern. Kubernetes lets you define Pod Disruption Budgets (PDBs) to ensure that voluntary disruptions (like upgrades or maintenance) don’t take down all replicas of a critical service. A PDB is a policy in policy/v1 API that sets the minimum number (or percentage) of Pods that must remain available during evictions. The control plane will block evictions (such as node drains or deletion of replicas) if it would violate any applicable PDB.
For example, if you have a Deployment with 5 replicas, and you want at least 3 up at all times, you could create:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: webapp-pdb
spec:
minAvailable: 3
selector:
matchLabels:
app: webapp
This PDB says that at least 3 Pods matching app=webapp must be available. If you try to drain a node that would reduce the number below 3, Kubernetes will refuse or delay that eviction. Alternatively, you could use maxUnavailable:
spec:
maxUnavailable: 2
which has the same effect (if maxUnavailable is 2, then 3 must stay available when you have 5 replicas). You can specify these values either as raw numbers or percentages (e.g., "50%"). As the official docs illustrate, if you set minAvailable to 10, then 10 Pods must remain running even during disruptions (v1-32.docs.kubernetes.io).
PodDisruptionBudgets only apply to voluntary disruptions - those triggered by an administrator action like kubectl drain or a deployment update. They do not prevent involuntary failures (like node crashes); for that you rely on having multiple replicas across fault domains. When you use a PDB, you strike a balance between availability and maintainability. For truly critical single-instance applications (scale 1), a PDB isn’t helpful because leaving 100% available means you effectively can’t disrupt it at all. In such cases, you may need a manual process or extra logic.
In practice, always define PDBs for your key deployments (especially stateful or critical stateless services). For instance, a multi-node database cluster should have a PDB ensuring majority nodes remain. For stateless web tiers, you might set maxUnavailable: 1 to allow one pod to be updated at a time. Without a PDB, a cluster-wide operation (like a rolling reboot) could unintentionally drain all pods of a service. Remember too that a PDB on a Deployment prevents its pods from ever being kubectldrain`ed if it would violate the budget; to fully drain a node, you might need to temporarily adjust the PDB. These policies should align with your service-level objectives: e.g., if your SLA says 99.9% uptime, configure PDBs so that you never drop below the required number of pods during maintenance.
For concepts around service reliability (service-level indicators, objectives, etc.), see our SRE Fundamentals content. Pod Disruption Budgets map directly to policies about how much planned downtime is tolerated.
Resource Management and Cost Optimization
Kubernetes runs on actual compute resources, so managing them efficiently is both a technical and financial concern. The main knobs are resource requests, limits, quotas, and autoscaling.
Every container can declare CPU and memory requests and limits in its spec. A simple example:
resources:
requests:
cpu: "500m"
memory: "256Mi"
limits:
cpu: "1"
memory: "512Mi"
Here the container requests 0.5 CPU and 256Mi of RAM for scheduling, and is limited to 1 CPU and 512Mi at runtime. The scheduler uses requests to fit Pods on Nodes, whereas limits prevent a container from using more than a set amount (if exceeded, the container may be throttled or OOM-killed). Setting requests appropriately is important: if you under-request, the scheduler may overcommit nodes and they become oversubscribed; if you over-request, the cluster wastes resources underutilized. In practice, analyze historical usage (via monitoring) to set realistic requests. Also set limits so that runaway processes don’t crash a node.
When requests == limits for both CPU and memory, that Pod is in the “Guaranteed” Quality-of-Service class, which gets priority during contention. If only requests are set (no limits or lower limits), it’s “Burstable.” If neither is set, the Pod is “BestEffort,” which should be avoided in production because such pods are most likely to be evicted first under pressure.
At the namespace level, use ResourceQuotas to prevent one team or application from consuming all resources. A ResourceQuota sets aggregate limits for the namespace. For example:
apiVersion: v1
kind: ResourceQuota
metadata:
name: team-quota
namespace: team-a
spec:
hard:
requests.cpu: "4"
requests.memory: 16Gi
limits.cpu: "6"
limits.memory: 24Gi
pods: "10"
This quota ensures that in team-a namespace, the total CPU requests cannot exceed 4 cores, memory requests 16Gi, etc. After 4 cores of requested CPU are used, further Pod creations will fail until resources are freed. Quotas are powerful for multi-tenant cost control and planning.
Kubernetes also provides autoscaling: - Horizontal Pod Autoscaler (HPA) can automatically adjust the replica count of a Deployment (or StatefulSet) based on observed metrics (CPU utilization, memory, or custom metrics). For example:
yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: webapp-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: webapp-deploy
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60
This HPA ensures the webapp-deploy Deployment runs between 2 and 10 pods, scaling up/down to try to keep CPU usage around 60%. Using HPA can save costs by only running more pods when load is high.
- Cluster Autoscaler: In cloud environments, you can enable the cluster autoscaler, which will add or remove nodes based on pending pods (e.g., if pods can’t schedule due to insufficient nodes or if nodes are underutilized). For example, on AWS or GKE, you tag node groups to be scalable. The cluster autoscaler works in conjunction with requests: if no nodes can satisfy a Pod’s requests, it may spin up another node. Note this is separate from HPA, which only scales pods. Managing node counts is important to reduce idle machine cost.
Other cost-saving practices include: - Rightsize images: Smaller container images reduce the CPU/memory footprint. Use Alpine or distroless base images if possible. - Spot instances: On cloud, consider using spot/preemptible instances for batch or stateless workloads, falling back to on-demand if spots aren’t available. - Shutdown dev clusters: If you run non-production clusters, schedule them to scale to zero at night or when idle to avoid billing. - Monitor & adjust: Use monitoring (see Observability) to see actual resource usage. Tools like Prometheus can track usage per namespace/pod; cloud cost dashboards can attribute charges to resources. Assertions like “80% utilization of CPU”. Right-size requests after observing peak and average loads. - Optimize node sizing: Choose appropriate VM sizes. If some pods require heavy CPU or memory, label nodes and use node selectors or taints/tolerations to run them on specific larger nodes.
If you manage your infrastructure via code (see our Infrastructure as Code section), you can define node groups, auto-scaling limits, and resource quotas in Terraform/CloudFormation so that they are reproducible and version-controlled.
In summary, controlling costs in Kubernetes means explicitly specifying resource usage (requests and limits), enforcing quotas to prevent runaway usage, and enabling autoscaling at both Pod and cluster levels. Regularly review metrics and logs to spot waste (e.g., pods with very low utilization but still running) and adjust your settings. Our Observability page discusses how to monitor resource usage and set alerts.
Security Best Practices for Production
Building on RBAC and Pod-level security, here are broader best practices for production clusters:
-
Regular Patching: Always run a supported Kubernetes version and apply updates for the control plane and node operating systems. Older versions can have known vulnerabilities.
-
Network Segmentation: Use NetworkPolicies (covered earlier) to enforce network isolation. Also use infrastructure firewalls to protect Kubernetes API ports (TCP 6443, etc.) from public internet, allowing only trusted IP ranges or VPN access.
-
Multi-Factor Authentication: If your cluster integrates with an identity provider (OIDC, LDAP), enable strong authentication and possibly MFA for admin users.
-
Audit Logging: Enable Kubernetes audit logs and send them to a secure log store (like Elasticsearch or cloud logging). Audit logs will record every sensitive action (like creation of cluster roles or node changes). Review them for unexpected changes and store logs long-term for forensic analysis.
-
Secrets Handling: Ensure that Kubernetes Secrets are encrypted at rest. Configure the API server with an encryption provider or use tools like HashiCorp Vault or Bitnami SealedSecrets to manage secrets outside etcd. Do not store secrets in container images or plain ConfigMaps.
-
Network Policies Recap: Enforce Widget’s "deny all except defined" rule with NetworkPolicies. This is a must for multi-tenant or compliance-sensitive clusters.
-
Minimize Attack Surface: Disable or remove unused add-ons (e.g. the Kubernetes dashboard addon should be removed or heavily secured). Limit privileges of the
kubeletand API server. Do not rely on Docker’s bridging if possible (the docker-shim is deprecated). -
Container Image Security: Build images with up-to-date base OS (no deprecated packages). Scan images in CI for CVEs using open-source tools (Trivy, Clair) or commercial scanners. Maintain a private registry, restrict image pull to that registry, and use
imagePullSecretsas needed. Avoid running images from Docker Hub or other public repos without trust. Sign images (e.g., with cosign) and verify signatures in runtime if possible. -
Pod Security Context: As mentioned earlier, enforce
runAsNonRootand limited capabilities. Also setreadOnlyRootFilesystem: truefor containers that don’t need to write to disk. If a container truly needs extra capabilities, document and approve them carefully. -
Host Hardening: Use lightweight, secure node OS images (like Container-Optimized OS, Bottlerocket, or Ubuntu Minimal). Ensure
ufw/iptableson nodes allow only needed ports. Use SELinux or AppArmor if supported to sandbox processes. -
Role Separation: Do not mix infrastructure/system services with tenants’ apps on the same cluster if possible. Use separate namespaces or clusters. Only give cluster-admin privileges to very few people and use separate accounts for day-to-day tasks.
-
Backup and disaster recovery: Regularly backup etcd data. Test restores! In production you should have a plan for re-creating the cluster from IaC or backups.
-
Security Policies: Consider using tools like OPA Gatekeeper to enforce policies (e.g., “no privileged containers”, “all pods must have resource limits”, etc.) at admission time. This codifies security hygiene.
By implementing these practices, you significantly reduce the risk of compromise. Security is an ongoing process: run periodic penetration tests, stay on top of CVEs, and iterate your defenses. For a portal perspective on reliability and security practices in operations, see SRE Fundamentals and consider formal DevOps certifications if you want to deepen your expertise.
Observability and Monitoring in Kubernetes
No production system is complete without observability. Kubernetes itself provides some metrics and logs, but you will typically run a full monitoring stack. A common choice is Prometheus for metrics and Grafana for dashboards, along with a logging solution (ELK, EFK, or a cloud logging service) and possibly tracing (e.g. Jaeger or OpenTelemetry).
Prometheus can scrape metrics from Kubernetes components and your applications. Deploy Prometheus either in-cluster (often via the Prometheus Operator) or use a hosted monitoring service. Ensure you gather:
- Node Metrics: via the kubelet’s cAdvisor or Node Exporter (CPU, memory, disk usage per node).
- Kubernetes Metrics: via kube-state-metrics (counts of pods, deployments, etc.) and the API server metrics (latency, errors).
- Application Metrics: add application metrics endpoints (/metrics) if your apps have instrumentation for request rate, latency, error counts, etc.
- Cluster Events: Kubernetes logs events but consider forwarding them to Prometheus as alerts.
Set up alerts in Prometheus/Grafana (or Alertmanager) for critical conditions, like pods in CrashLoopBackOff, resource saturation, or HTTP error spikes. For guidance on scraping metrics from Kubernetes, see our free Prometheus guide.
Logging: Typically, you install a log-collector DaemonSet (such as Fluentd or Fluent Bit) so that it runs on every node and tails logs from all containers. These logs can be forwarded to ElasticSearch, Loki, Splunk, or a cloud logging system. Ensure that container stdout/stderr and any application logs are captured. Use structured logging (JSON) if possible so you can query logs (e.g. search for error patterns or correlate with trace IDs).
Tracing: For microservices architectures, use distributed tracing. Tools like Jaeger or OpenTelemetry let you see request flows across services and identify slow points. Instrument client libraries in your code to generate spans.
By combining metrics, logs, and traces, you form a complete view of the cluster health. Dashboards can show node stability, application performance, error rates, etc. This is how you verify Service Level Objectives (SLOs). For introduction to SLOs and reliability concepts, see SRE Fundamentals.
In short: integrate Prometheus/Grafana for metrics, a log aggregation tool for logs, and optionally a tracing system. Don’t forget to monitor Kubernetes itself (API server availability, etcd health, node readiness). For more tips on observability, see our sibling Observability page.
Multi-Tenant Clusters: Namespace Isolation and Beyond
Running multiple teams or customers on the same Kubernetes cluster saves costs, but requires strong isolation controls. There are two broad models: soft multi-tenancy and hard multi-tenancy.
- Soft multi-tenancy (namespace-based): In this model, each tenant (team or project) gets one or more Kubernetes Namespaces. As the Kubernetes docs state, “a Namespace provides a mechanism for isolating groups of API resources within a single cluster” (v1-35.docs.kubernetes.io). Namespaces provide logical separation: pods in different namespaces can have the same names, resource quotas apply per namespace, and RoleBindings can restrict a user to one namespace. To enforce isolation:
- Use RBAC so each tenant’s users only have Roles in their own namespace (so they can’t see other namespaces).
- Apply ResourceQuotas in each namespace so one team cannot use up all CPU/memory. For example, give each namespace a quota of 100 CPUs and 200Gi RAM.
- Use NetworkPolicies to prevent cross-namespace network traffic by default, then allow only what’s needed (or simply do not use a broad namespaceSelector in allow rules).
- Optionally use the Namespace Hierarchy Controller (HNC) if you have related namespaces (HNC can propagate RoleBindings or quotas across a hierarchy).
In practice, you might have a namespace per team or per environment (dev, staging, prod). Each tenant’s CI/CD pipelines apply Kubernetes resources into that namespace only. For shared services (like a common logging namespace), use dedicated namespaces with strict RBAC.
-
Hard multi-tenancy (cluster isolation): If tenants must be completely isolated (for example public customers, or highly regulated applications), you typically give them separate clusters or virtual clusters. Tools like vcluster or Capsule can create “virtual clusters” on one Kubernetes control plane for stronger isolation. Another approach is one Kubernetes cluster per tenant (using automation to provision clusters via Infrastructure as Code). The downside is increased resource overhead (each cluster has control-plane nodes). Choose hard isolation if you do not trust co-tenants at all, or if you need separate API servers for compliance.
-
Storage Isolation: If multiple tenants share storage, ensure they don’t see each other’s data. Use separate StorageClasses or prefixes for their volumes. For example, namespace A and B might get volumes from different NFS shares. Avoid defaulting all tenants onto the same physical disk unless you trust their isolation fully.
-
Network: A common pattern is VLAN or VPC separation underneath, with NetworkPolicies on top. Even if the network fabric is shared, Kubernetes NetworkPolicies (or Calico profiles) can enforce that namespace A’s pods cannot send traffic to namespace B’s pods at all.
Key best practices summary: - Create Kubernetes Namespaces per tenant or environment. - Bound each namespace with a ResourceQuota to prevent overuse. - Assign Roles so only tenant users can modify their own namespace. - Apply default-deny NetworkPolicies, and explicitly allow needed cross-namespace traffic only. - For very strict isolation, use separate clusters or virtual clusters.
Multi-tenancy is a spectrum, not binary. For example, internal teams within one organization might be comfortable with namespace-level isolation, while external SaaS customers would get their own clusters. Evaluate the trust model. For more on designing clusters from an SRE perspective, see SRE Fundamentals, which discusses service isolation and reliability. Also consider integrating these multi-tenant patterns into your cluster provisioning toolset via Infrastructure as Code.
Explore the Kubernetes Silo
- DevOps Hub (Home Page) - A central location for all our DevOps topics and learning paths.
- CI/CD Pipelines (/devops/ci-cd/) - Best practices on continuous integration and delivery pipelines, including how to automate Kubernetes deployments.
- Observability and Monitoring (/devops/observability/) - Techniques and tools for logging, metrics, and tracing in Kubernetes.
- Infrastructure as Code (IaC) (/devops/iac/) - Managing Kubernetes clusters and cloud resources using declarative, version-controlled code.
- GitOps Workflows (/devops/gitops/) - Running Kubernetes using Git-centric deployment workflows for reproducibility and auditability.
- SRE Fundamentals (/devops/sre-fundamentals/) - Reliability engineering principles, SLOs/SLIs, and practices that complement Kubernetes operations.
- DevOps Certifications (/devops/devops-certifications/) - Information on DevOps and Kubernetes certifications for professional development.
FAQ
Q: What is a Kubernetes Pod, and why shouldn't I rely on a Pod's IP address?
A: A Pod is the smallest deployable unit in Kubernetes, grouping one or more containers that share network and storage (kubernetes.io). Pods are designed to be ephemeral: the cluster can kill and restart pods to match the desired state (for example, a Deployment might recreate pods during updates) (kubernetes.io). Each time a pod restarts it may get a new IP. Therefore, you should not rely on a specific pod’s IP. Instead, use a Service which provides a stable virtual IP and name. Services automatically track the set of Pods, so clients always connect to the Service address.
Q: How do I expose my application outside the cluster?
A: Inside the cluster you use a Service of type NodePort or LoadBalancer, or an Ingress for HTTP(S). A LoadBalancer Service (in cloud providers) will provision an external IP/hostname for your service. Otherwise, you can create an Ingress resource and run an Ingress Controller (like NGINX Ingress). An Ingress allows you to route traffic based on domain names and paths to different Services. In short, a Service gives you a single IP or DNS name for one port, and an Ingress gives flexible HTTP routing to multiple Services with TLS support. For step-by-step setup, check the CI/CD pipelines and GitOps guides which cover deployment of Services and Ingresses.
Q: What’s the difference between a Deployment and a StatefulSet?
A: Use a Deployment for stateless applications where each pod is interchangeable (like web servers). Use a StatefulSet for stateful applications that need stable identities (like databases). A StatefulSet ensures each pod has a persistent hostname and volume, and creates them in order (kubernetes.io). For example, the first pod will always be app-0, then app-1, etc., and Kubernetes will keep their volumes attached. A Deployment would just call them app-xxx with random suffixes. So if you need consistent network IDs or stable disks per pod (e.g. for a MySQL master-slave or Kafka cluster), use StatefulSet. Otherwise, for general scaling and rolling updates, a Deployment is simpler.
Q: How does RBAC secure a Kubernetes cluster?
A: RBAC stands for Role-Based Access Control. It lets you define Roles (with allowed actions like get, list, create on resources) and bind them to users or service accounts. For example, you might create a Role that allows reading pods and then bind it to a developer. That developer can then only view pods (and nothing else) in that namespace. Roles are namespace-scoped so you can isolate teams. ClusterRoles apply cluster-wide. As Kubernetes docs say, RBAC Roles can be restricted to namespaces (v1-35.docs.kubernetes.io). By giving each user the smallest Role they need, you prevent unauthorized access. Always avoid giving broad permissions like “cluster-admin” to everyone; instead tailor Roles for each job function.
Q: What is a Pod Disruption Budget (PDB) and when should I use it?
A: A Pod Disruption Budget limits how many Pods of an application can go down at once during voluntary disruptions (like upgrades or node drains). It’s basically a “safety net” for maintenance. You specify minAvailable or maxUnavailable for a set of pods. For instance, if your app has 5 pods, a PDB with minAvailable: 3 ensures at least 3 stay running at any time. If more than 2 were to be evicted, Kubernetes will block it. Without a PDB, a rolling update or node drain could inadvertently take all pods down. Use PDBs on critical Deployments or StatefulSets to guarantee your service stays partially available during upgrades or scaling. The PDB does not protect against sudden failures, but it does prevent planned disruptions from shutting down the entire service.
Q: How does Kubernetes storage work? What is a StorageClass?
A: Kubernetes separates the request for storage (PVC) from the actual disk (PV). You write a YAML claim (PersistentVolumeClaim) asking for a certain size and access mode. Kubernetes matches that to a PersistentVolume. A StorageClass automates this by specifying a provisioner plugin (like AWS EBS, Azure Disk, etc.). When you create a PVC with a StorageClass, the cluster dynamically provisions the storage. Essentially, a StorageClass defines how to create a volume. The PVC is what you request. This abstraction means you never have to manually create a disk; you just say “I need 10Gi” and K8s handles it. Then you mount the PVC into your Pod. For step-by-step examples, see our DevOps Hub DevOps resources.
Q: How do Network Policies improve security?
A: NetworkPolicies let you define a “default deny” network firewall for pods. Without policies, all pods can talk to each other. With NetworkPolicy, you can say things like “pod with label frontend can only talk to pod with label backend” or “only allow TCP port 53 for DNS”. Once a pod is selected by a policy, Kubernetes blocks any traffic not explicitly allowed. This prevents rogue pods (or compromised pods) from indiscriminately reaching other pods. For example, if you apply a deny-all ingress policy for a namespace, no external pod can reach any pods in that namespace unless you add an allow rule. Implementing NetworkPolicies is key for zero-trust in Kubernetes networks. In our Multitenancy discussion we explained how to use policies to isolate namespaces.
Q: How can I control costs and resources in Kubernetes?
A: First, define resource requests and limits on every pod. Requests ensure the scheduler doesn’t pack too many pods on one node beyond capacity; limits prevent a pod from hogging CPU or memory. Second, use ResourceQuotas in namespaces to cap total usage (e.g., “no more than 8 CPUs aggregate”). Third, enable autoscaling. Use a Horizontal Pod Autoscaler to scale pods up/down with load, and a Cluster Autoscaler to add/remove nodes. That way you only pay for needed compute. Also, clean up unused resources: delete stale namespaces, stop idle pods, and schedule non-critical clusters to sleep when not needed. Monitoring actual usage is vital: use Prometheus/Grafana (see Observability) to visualize CPU and memory trends, and adjust your allocations accordingly. Following these steps keeps your Kubernetes deployment efficient and cost-effective.
