Kubernetes Roadmap
From Reconciliation Loops to Production Clusters — Understand It, Don't Just Copy the YAML
Your Journey at a Glance
💡 How to use this roadmap
Work through each phase in order. Click on a skill to expand it — you'll find a description and curated resources. Don't rush; understanding beats speed. Complete one phase before moving to the next.
Foundations — The Control Plane & Reconciliation
Kubernetes is a set of control loops driving actual state toward declared state. Understand that one idea and most of the system stops feeling arbitrary.
Workloads, Health & Configuration
Running applications and updating them without dropping requests — including the probes that decide whether traffic reaches you at all.
Networking & Traffic Management
How pods find each other and how external traffic gets in — including the part of the ecosystem that changed most recently and where most existing tutorials are now out of date.
State, Storage & Resource Management
Running things that remember, and teaching the scheduler what your workloads actually need.
Scheduling, Autoscaling & Availability
Placing workloads deliberately, scaling them safely, and surviving the disruptions that are a normal part of cluster life.
Security, Packaging & Operating a Cluster
Locking the cluster down, managing manifests at scale without hand-editing YAML, and building the debugging instinct that separates operators from copy-pasters.
Scenarios — Real Incidents & Interview Discussions
Six incidents told end to end: the symptom, the investigation, the mechanism underneath, and how to discuss each one in an interview. This is where the previous six phases get exercised against real failures.
Roadmap Complete!
You now have the foundations of a production-ready Java engineer. Apply by building real projects.
Deploy and Operate a Multi-Tier Application on Kubernetes
Take a web frontend, an API, an asynchronous worker, a PostgreSQL database and a Redis cache from manifests to a running, observable, deliberately hardened deployment — then break it on purpose and prove you can diagnose it.
What you'll build
- Kustomize base plus per-environment overlays, or a Helm chart with environment values — rendered with helm template / kubectl diff and reviewed before every apply
- Zero-downtime rolling updates demonstrated under continuous load, with readiness probes gating traffic and a preStop hook covering the endpoint-removal race
- Startup, liveness and readiness probes on every service, each testing something different and justified in writing
- Requests and limits on every container, with QoS classes chosen deliberately and an intentional OOM kill triggered so you recognise exit 137
- External traffic through a Gateway API implementation with TLS terminated via cert-manager — explicitly not the archived ingress-nginx controller
- Default-deny NetworkPolicies with verified connectivity tests proving the database is reachable only from the API and worker
- PostgreSQL on a StatefulSet with volumeClaimTemplates and a restore-from-backup you have actually performed, or a managed operator with the tradeoff documented
- HPA on the worker driven by queue depth rather than CPU, with a tuned scaleUp stabilization window
- PodDisruptionBudgets validated by draining a node and watching the application stay up
- RBAC least-privilege ServiceAccounts, the restricted Pod Security Admission profile enforced, and a securityContext that satisfies it
- A written incident runbook produced by deliberately causing CrashLoopBackOff, ImagePullBackOff, a Pending pod and a Service with no ready endpoints, and recording the exact commands that identified each
Tech stack
Key highlights
- ✦Uses the current traffic-management standard rather than the archived controller most tutorials still recommend
- ✦Every reliability claim is demonstrated under load or under a drain, not asserted in a manifest
- ✦Deliberate failure injection produces a debugging runbook you can reuse and interview with