Kubernetes
Kubernetes, often abbreviated K8s, is an open-source system for deploying and managing containerized applications across a cluster of machines. Users describe the state they want, such as a number of application replicas, and Kubernetes controllers work to maintain that state. Its responsibilities include workload placement, replacement of failed application instances, service discovery, and integration with storage and networking systems.[1]
For MLOps, Kubernetes supplies infrastructure on which training jobs, data-processing workloads, and model-serving systems can run. It does not define the model, training algorithm, or inference engine. Those run inside workloads or are managed through additional controllers. GPU access, distributed training, and demand-based scaling require configuration beyond simply installing Kubernetes.[3][4]
Origins and governance
Google released Kubernetes as open-source software in 2014. The project drew on experience with Google's internal cluster-management systems. Kubernetes became the Cloud Native Computing Foundation's first graduated project on March 6, 2018.[1][5]
The research literature distinguishes the related systems. The 2015 Borg paper describes Google's production cluster manager, including admission control, task placement, resource sharing, and recovery. The 2013 Omega paper studies parallel schedulers operating on shared state with optimistic concurrency control. These papers concern their respective systems, rather than benchmarks of Kubernetes. The 2016 article "Borg, Omega, and Kubernetes" discusses lessons across all three systems.[6][7][8]
Desired state and the API
A Kubernetes object records configuration and state in the cluster's API. Many objects have a spec, which describes the requested state, and a status, which reports the observed state. A controller repeatedly compares the two and takes actions to reduce the difference. For example, a Deployment requesting four replicas can cause replacement Pods to be created when existing replicas disappear.[2]
Administrators and applications interact with the API through clients such as kubectl. A manifest commonly expresses an object in YAML, including its apiVersion, kind, metadata, and specification. The API version identifies the representation of that resource; it is not the version number of the application being deployed.[2]
Reconciliation is ongoing rather than a one-time installation command. Acceptance of a manifest does not establish that its workload is healthy: scheduling, image startup, storage attachment, and readiness are separate concerns. Pod and workload status provide evidence about whether the requested application actually became available.[2][9]
Cluster components
A cluster has a control plane and worker nodes. A node can host multiple Pods. Production control planes can use multiple machines for availability, while the arrangement in a development cluster can be much smaller.[10]
| Component | Responsibility |
|---|---|
| API server | Exposes the Kubernetes API through which clients and components operate on cluster resources. |
| etcd | Stores the API server's data in a consistent key-value store. |
| Scheduler | Selects a suitable node for a Pod that has not yet been assigned to one. |
| Controller manager | Runs controllers that implement built-in API behavior. |
| Cloud controller manager | Provides optional integration with a cloud provider. |
| kubelet | Runs on a node and manages the execution of its assigned Pods. |
| Container runtime | Starts and manages containers on the node. |
| Service networking implementation | Implements the network behavior of Services; some clusters use kube-proxy, while others use an implementation supplied by their network plugin. |
These components divide cluster coordination from execution. The scheduler assigns work; the node's runtime executes containers. Kubernetes is therefore not itself a container runtime.[10][11]
Pods and workload resources
A Pod is Kubernetes' smallest deployable unit. It contains one or more closely associated containers that run together on the same node and share networking and available Pod volumes. A multi-container Pod is appropriate for components that need this shared context; independent replicas normally use separate Pods.[9]
Pods are replaceable. A replacement Pod is a new object, rather than the original Pod migrating intact to another machine. Applications that need data to survive replacement must store it outside the lifetime of that Pod, for example through persistent storage.[9][12]
Users normally create workload resources that manage Pods instead of managing every Pod individually.[13]
| Resource | Main role | Example in an AI system |
|---|---|---|
| Deployment | Maintains interchangeable replicas of an application. | Replicas of a stateless inference API. |
| StatefulSet | Manages Pods with distinct identities and associations with persistent storage. | A service whose replicas need stable identities. |
| DaemonSet | Runs node-local facilities on all or selected nodes. | A device plugin or node-level collection agent. |
| Job | Runs work to completion, with configurable parallelism and failure handling. | A batch evaluation or data-preparation task. |
| CronJob | Creates Jobs according to a schedule. | Periodic processing of a dataset. |
The examples describe deployment patterns, not a guarantee that any particular model server or database is suitable for a resource without further configuration. A Job can start replacement Pods after failure and can execute work in parallel. Its application must still handle retries and the possibility of repeated execution correctly.[13][14]
Resources and placement
Resource requests tell the scheduler how much capacity a workload needs. Limits constrain resource use at runtime. For CPU and memory, requests and limits have different meanings: a container can use more than its request when capacity permits, subject to its limit. On Linux, CPU limits use throttling; memory-limit enforcement can result in an out-of-memory termination.[15]
Incorrect requests affect placement and efficiency. An inflated request can prevent useful work from fitting on a node, while an understated request does not remove the application's actual memory needs. A Pod can remain pending when the cluster lacks sufficient requested resources. Setting resource fields is capacity accounting, not a measurement of how fast the program will run.[15]
Placement constraints express requirements beyond resource quantities. Node selectors and node affinity match node labels; Pod affinity and anti-affinity express relationships to other Pods. Topology-spread constraints can distribute replicas across specified topology domains. Hardware-specific workloads may need a particular accelerator type or location, rather than any node with a free device count.[16]
These constraints can conflict with each other or with available capacity. A preference allows the scheduler flexibility, whereas a required constraint can make a Pod unschedulable. Application placement therefore depends on the requested policy and the cluster's actual configuration.[16]
GPUs and other devices
Device plugins
Kubernetes device plugins expose hardware requiring vendor-specific setup, including GPUs, network adapters, and FPGAs. A plugin registers with the kubelet and advertises devices under an extended-resource name. Kubernetes can then account for those resources when placing workloads. The device-plugin framework is documented as stable since Kubernetes v1.26.[17]
With the conventional GPU device-plugin path, administrators install the vendor's drivers and plugin. A container can request a resource such as nvidia.com/gpu or amd.com/gpu through its resource limits. If both a request and limit are specified for that GPU resource, they must be equal; specifying only the limit makes it the request as well.[4]
Extended-resource counts are integers and cannot be overcommitted in the ordinary device-plugin resource model.[17] Vendor extensions can nevertheless expose shared access to a physical GPU as multiple resource units. NVIDIA's plugin, for example, has time-slicing options; its documentation warns that this sharing does not provide memory or fault isolation between workloads on the same underlying GPU. A count of advertised GPU resources therefore needs to be interpreted alongside the plugin's configuration.[18]
Dynamic resource allocation
Dynamic Resource Allocation, or DRA, provides a separate resource-claim model for devices. Drivers publish available devices; classes and claims describe which devices a workload can use. Pods reference claims, and allocation and scheduling coordinate so that Pods reach nodes with access to their allocated resources.[19]
DRA supports richer device selection and configuration than a count alone, including attribute-based filtering and claim-based sharing. Sharing a claim does not by itself establish a particular hardware partitioning or isolation guarantee. Device capabilities and driver behavior remain relevant.[19]
The exact DRA API versions, optional subfeatures, and driver requirements depend on the Kubernetes release and the implementation deployed. Configuration written for one cluster should not be assumed to work on another merely because both clusters expose GPUs. Device plugins and DRA are related ways to integrate devices, not interchangeable syntax for every accelerator.[17][19]
Training and inference
Distributed training
Distributed training can involve several worker processes that need one another before making progress. Starting some workers while others wait for capacity can leave allocated accelerators idle. Gang scheduling addresses admission or placement of a group, rather than treating every worker as an unrelated workload.[20]
Kubernetes documents PodGroup-based gang scheduling, but availability and enablement are release-specific. The documented all-or-nothing placement concerns the configured minimum group size at scheduling time. It does not guarantee that every worker remains alive afterward. A running group's size can fall after deletion or eviction.[20]
Other projects add workload-level management. Kueue manages quota, queues, admission, and preemption while leaving responsibilities such as Pod placement and node autoscaling to the corresponding components. Kubeflow Trainer provides training-specific APIs and runtime integration for distributed AI workloads. KubeRay manages Ray clusters and related workloads through Kubernetes custom resources.[21][3][22]
These layers have distinct jobs. A training controller can organize workers, a queue can decide when a job may start, and Kubernetes can place its Pods. The training framework still determines the computation and communication among workers.[3][21] The result of scheduling is a placement, rather than a measurement of training throughput.[10]
Model serving and scaling
A model-serving application can run in replicated Pods behind a Service. Kubernetes maintains the workload and provides network access, while the server performs inference. For example, batching or model-specific execution belongs to the inference application rather than the generic Service abstraction.[1][23]
Horizontal Pod Autoscaling adjusts a workload's replica count based on configured metrics. It can use resource metrics such as CPU utilization or suitable custom metrics. The mechanism requires the relevant metrics APIs and configuration; it does not automatically choose a useful model-serving signal.[24]
Node autoscaling is separate. A node autoscaler can provision machines for Pods that cannot fit on existing nodes, or consolidate nodes when appropriate. Configured limits, incompatible scheduling requirements, and unavailable cloud capacity can prevent provisioning. More desired replicas do not therefore imply that additional GPU capacity is immediately available.[25]
Health checks also need application-specific meaning. A startup probe allows initialization to finish before other probes run. Readiness controls whether a Pod should receive normal Service traffic. Liveness can trigger a restart. A model server that is still loading its weights should not be treated as ready solely because its process has started; an overly aggressive liveness probe can cause repeated restarts.[26]
Networking and storage
A Service provides access to a changing set of endpoints, usually Pods selected by labels. Clients can use the Service instead of tracking each replacement Pod's address. A Service's exposure depends on its type and the networking implementation. External load balancing and HTTP routing require the appropriate integration or controller; creating an object is not sufficient to supply every external networking facility.[23]
PersistentVolumes represent storage with a lifetime independent of an individual Pod. A PersistentVolumeClaim requests storage with requirements such as capacity and access mode. StorageClasses describe provisioning options. Dynamic provisioning depends on a configured class and provisioner; a claim can remain unbound when compatible storage is unavailable.[12]
For AI workloads, this separation allows model files, datasets, and checkpoint files to be stored independently of a worker Pod. Persistence does not mean that Kubernetes creates application checkpoints. The application must write the necessary state, and recovery depends on the storage and application's restart behavior. A volume abstraction also does not remove storage access-mode or performance constraints.[12][14]
Security and multi-tenancy
Namespaces organize groups of API resources, but namespace separation alone is not complete isolation between tenants. Kubernetes' multi-tenancy guidance combines authorization, quotas, networking controls, and workload isolation according to the trust relationship between users. Stronger requirements can justify separate clusters or dedicated hardware.[27]
Role-based access control governs what users and service accounts can do through the API. Resource quotas limit specified resource consumption or object counts. These mechanisms address different risks: a quota does not prevent unauthorized data access, and authorization does not by itself isolate network traffic.[27]
NetworkPolicy can restrict traffic to and from selected Pods at the network layer, but enforcement requires a compatible network plugin. Creating a policy without an implementation that enforces it has no effect. Policies also do not substitute for application authentication.[28]
Secrets store sensitive configuration, but their name does not imply automatic encryption. Kubernetes documents Secret data as unencrypted in etcd by default unless encryption at rest is configured. Base64 encoding is not encryption. A user able to create a Pod that consumes a Secret may be able to obtain its contents even without direct permission to read that Secret through the API.[29]
Extensions and operations
Custom resources and controllers let projects add domain-specific behavior. An operator combines this API extension pattern with control loops for an application's lifecycle, such as deployment, upgrades, or backups. Such behavior comes from the installed operator; it should not be attributed to Kubernetes universally.[30]
Operational visibility requires more than whether a Pod is running. Kubernetes components expose metrics, logs, and, where supported and configured, traces. These signals can be collected into monitoring systems; OpenTelemetry collectors can receive and export tracing data. The resource Metrics API used for basic inspection and autoscaling is not a replacement for a complete monitoring pipeline.[31]
References
- ^1 ^2 ^3Kubernetes. Overview. Documentation accessed September 27, 2026.
- ^1 ^2 ^3Kubernetes. Objects In Kubernetes. Documentation accessed September 27, 2026.
- ^1 ^2 ^3Kubeflow. Overview of Kubeflow Trainer. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes. Schedule GPUs. Documentation accessed September 27, 2026.
- ^Cloud Native Computing Foundation. Cloud Native Computing Foundation announces Kubernetes as first graduated project. March 6, 2018.
- ^Verma, Abhishek, et al. Large-scale cluster management at Google with Borg. EuroSys, 2015.
- ^Schwarzkopf, Malte, et al. Omega: flexible, scalable schedulers for large compute clusters. EuroSys, 2013, pp. 351-364.
- ^Burns, Brendan, Brian Grant, David Oppenheimer, Eric Brewer, and John Wilkes. Borg, Omega, and Kubernetes. ACM Queue, volume 14, 2016, pp. 70-93.
- ^1 ^2 ^3Kubernetes. Pods. Documentation accessed September 27, 2026.
- ^1 ^2 ^3Kubernetes. Cluster Architecture. Documentation accessed September 27, 2026.
- ^Kubernetes. Kubernetes Components. Documentation accessed September 27, 2026.
- ^1 ^2 ^3Kubernetes. Persistent Volumes. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes. Workload Management. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes. Jobs. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes. Resource Management for Pods and Containers. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes. Assigning Pods to Nodes. Documentation accessed September 27, 2026.
- ^1 ^2 ^3Kubernetes. Device Plugins. Documentation accessed September 27, 2026.
- ^NVIDIA. NVIDIA device plugin for Kubernetes, "Shared Access to GPUs." Repository documentation accessed September 27, 2026.
- ^1 ^2 ^3Kubernetes. Dynamic Resource Allocation. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes. Gang Scheduling. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes SIGs. Kueue overview. Documentation accessed September 27, 2026.
- ^Ray. Ray on Kubernetes. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes. Service. Documentation accessed September 27, 2026.
- ^Kubernetes. Horizontal Pod Autoscaling. Documentation accessed September 27, 2026.
- ^Kubernetes. Node Autoscaling. Documentation accessed September 27, 2026.
- ^Kubernetes. Liveness, Readiness, and Startup Probes. Documentation accessed September 27, 2026.
- ^1 ^2Kubernetes. Multi-tenancy. Documentation accessed September 27, 2026.
- ^Kubernetes. Network Policies. Documentation accessed September 27, 2026.
- ^Kubernetes. Good practices for Kubernetes Secrets. Documentation accessed September 27, 2026.
- ^Kubernetes. Operator pattern. Documentation accessed September 27, 2026.
- ^Kubernetes. Observability. Documentation accessed September 27, 2026.
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 2,564 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent AI-assisted editorial review checked the published text against cited primary documentation and research. Version-specific behavior and study limitations are stated in the article; this is not a guarantee of runtime behavior or factual infallibility.
Cite this page: AI Wiki. "Kubernetes." aiwiki.ai, updated 27 Sept 2026, fact-checked 27 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/kubernetes