⎈
Learning Path

Kubernetes Roadmap

From Reconciliation Loops to Production Clusters — Understand It, Don't Just Copy the YAML

Updated September 12, 2026
7Phases
7Weeks
24Skills

Your Journey at a Glance

1Foundations — The Control Plane & Reconciliation3 skills
→
2Workloads, Health & Configuration3 skills
→
3Networking & Traffic Management3 skills
→
4State, Storage & Resource Management3 skills
→
5Scheduling, Autoscaling & Availability3 skills
→
6Security, Packaging & Operating a Cluster3 skills
→
7Scenarios — Real Incidents & Interview Discussions6 skills

💡 How to use this roadmap

Work through each phase in order. Click on a skill to expand it — you'll find a description and curated resources. Don't rush; understanding beats speed. Complete one phase before moving to the next.

1

Foundations — The Control Plane & Reconciliation

Kubernetes is a set of control loops driving actual state toward declared state. Understand that one idea and most of the system stops feeling arbitrary.

Week 1
Read the deep-dive guide

2

Workloads, Health & Configuration

Running applications and updating them without dropping requests — including the probes that decide whether traffic reaches you at all.

Week 2
Read the deep-dive guide

3

Networking & Traffic Management

How pods find each other and how external traffic gets in — including the part of the ecosystem that changed most recently and where most existing tutorials are now out of date.

Week 3
Read the deep-dive guide

4

State, Storage & Resource Management

Running things that remember, and teaching the scheduler what your workloads actually need.

Week 4
Read the deep-dive guide

5

Scheduling, Autoscaling & Availability

Placing workloads deliberately, scaling them safely, and surviving the disruptions that are a normal part of cluster life.

Week 5
Read the deep-dive guide

6

Security, Packaging & Operating a Cluster

Locking the cluster down, managing manifests at scale without hand-editing YAML, and building the debugging instinct that separates operators from copy-pasters.

Week 6
Read the deep-dive guide

7

Scenarios — Real Incidents & Interview Discussions

Six incidents told end to end: the symptom, the investigation, the mechanism underneath, and how to discuss each one in an interview. This is where the previous six phases get exercised against real failures.

Week 7
Read the deep-dive guide

🏆

Roadmap Complete!

You now have the foundations of a production-ready Java engineer. Apply by building real projects.

Capstone Project

Deploy and Operate a Multi-Tier Application on Kubernetes

Take a web frontend, an API, an asynchronous worker, a PostgreSQL database and a Redis cache from manifests to a running, observable, deliberately hardened deployment — then break it on purpose and prove you can diagnose it.

What you'll build

  • Kustomize base plus per-environment overlays, or a Helm chart with environment values — rendered with helm template / kubectl diff and reviewed before every apply
  • Zero-downtime rolling updates demonstrated under continuous load, with readiness probes gating traffic and a preStop hook covering the endpoint-removal race
  • Startup, liveness and readiness probes on every service, each testing something different and justified in writing
  • Requests and limits on every container, with QoS classes chosen deliberately and an intentional OOM kill triggered so you recognise exit 137
  • External traffic through a Gateway API implementation with TLS terminated via cert-manager — explicitly not the archived ingress-nginx controller
  • Default-deny NetworkPolicies with verified connectivity tests proving the database is reachable only from the API and worker
  • PostgreSQL on a StatefulSet with volumeClaimTemplates and a restore-from-backup you have actually performed, or a managed operator with the tradeoff documented
  • HPA on the worker driven by queue depth rather than CPU, with a tuned scaleUp stabilization window
  • PodDisruptionBudgets validated by draining a node and watching the application stay up
  • RBAC least-privilege ServiceAccounts, the restricted Pod Security Admission profile enforced, and a securityContext that satisfies it
  • A written incident runbook produced by deliberately causing CrashLoopBackOff, ImagePullBackOff, a Pending pod and a Service with no ready endpoints, and recording the exact commands that identified each

Tech stack

Kubernetes (kind / k3s locally, EKS / GKE / AKS in cloud)Helm 3 and KustomizeGateway API (Envoy Gateway, Istio, Cilium or Traefik)cert-managerPostgreSQL (StatefulSet or operator)Redismetrics-serverPrometheus & Grafana

Key highlights

  • ✦Uses the current traffic-management standard rather than the archived controller most tutorials still recommend
  • ✦Every reliability claim is demonstrated under load or under a drain, not asserted in a manifest
  • ✦Deliberate failure injection produces a debugging runbook you can reuse and interview with

Want to Go Deeper?

Join a live cohort, read in-depth guides, or watch video lessons on the topics in this roadmap.