The first time this change really lands on you is not during an architecture review. It is at 2:17 a.m., when a Kubernetes Service is timing out, pods are healthy, DNS looks fine, and muscle memory sends you straight to the node to inspect iptables rules, NAT chains, and the KUBE-SERVICES path.
I have spent enough years debugging Kubernetes networking to know the old sequence by heart: establish whether the failure is DNS, routing, policy, Service translation, or conntrack; inspect kube-proxy; trace the relevant iptables chains; capture packets; then work inward. That runbook still works on plenty of clusters in 2026, but on an increasingly important class of managed Kubernetes clusters, the assumptions underneath it are simply wrong.
On Azure Kubernetes Service, AKS Automatic now comes preconfigured with Azure CNI Overlay powered by Cilium, and Microsoft states that AKS clusters using Cilium as the network data plane do not use kube-proxy. On Google Kubernetes Engine, GKE Dataplane V2 is enabled by default for every new Autopilot cluster, uses eBPF, is implemented with Cilium technology, and uses Cilium rather than kube-proxy for Kubernetes Services.
AWS is the reason the headline needs the qualifier. Standard EKS still bootstraps Amazon VPC CNI, kube-proxy, and CoreDNS by default, while AWS documents Cilium as an alternate compatible CNI rather than the default cloud-node networking stack. EKS itself does use eBPF in places, including VPC CNI NetworkPolicy enforcement, so “AWS does not use eBPF” would be wrong; what AWS has not done is make a Cilium-based kube-proxy replacement the documented default for its automated cloud-node tier in the way Azure and Google have.
That distinction is why any discussion of Kubernetes networking defaults in 2026 needs a cloud-by-cloud answer rather than a slogan.
For engineers building those fundamentals, the Refonte Learning DevOps Engineer Program covers Linux and scripting, Docker and Kubernetes, AWS/Azure/GCP, Terraform, CI/CD, and monitoring/logging. Those are the layers you need before eBPF networking starts making practical sense. The live curriculum does not currently name eBPF, Cilium, or CNI-level networking specifically, so treat the program as the foundation this article builds on, not as a claim that it already teaches this exact stack.
The Debugging Playbook That Stopped Working
When someone says, “This cluster does not use iptables anymore,” the statement needs qualification. Linux still has netfilter, other components can still create rules, and upstream kube-proxy itself supports iptables, nftables, and IPVS-related deployment histories; what has changed on a Cilium kube-proxy-free data plane is that Kubernetes Service routing is no longer being implemented through the kube-proxy iptables rule set you went looking for. Kubernetes documentation still describes iptables as the default Linux proxy mode for upstream kube-proxy, while Cilium documents a mode that replaces kube-proxy entirely.
That is a much more precise statement than “iptables is dead.”
A traditional kube-proxy investigation often begins with evidence such as:
kubectl -n kube-system get ds kube-proxy
iptables-save | grep KUBE-
iptables -t nat -L KUBE-SERVICES -n -v
EndpointSlice inspection with kubectl get endpointslice
conntrack inspection when translations or stale flows look suspicious
tcpdump on the relevant veth, bridge, node, or physical interface
On an eBPF-based Cilium service data plane, you can spend ten minutes searching for a KUBE-SVC-* chain that is never going to exist because the Service lookup and backend selection are happening through eBPF programs and maps instead. Cilium’s architecture attaches eBPF programs at Linux networking hooks and can perform networking, service load balancing, visibility, and policy decisions inside the kernel rather than building Kubernetes Service behavior from a large sequential ruleset.
That does not make packet captures or Linux fundamentals obsolete. It changes where I look first.
Old assumption | What I verify in 2026 |
“Every Kubernetes cluster has kube-proxy.” | Check whether the managed service uses Cilium or another kube-proxy replacement. |
“ClusterIP translation lives in KUBE-* iptables chains.” | Determine whether Service state lives in eBPF maps instead. |
“NetworkPolicy probably means iptables rules.” | Check the CNI and policy engine; Cilium and AWS VPC CNI can enforce policy using eBPF. |
“No iptables rule means the CNI is broken.” | Absence may be completely normal on a kube-proxy-free data plane. |
“Start with the node firewall.” | Start with cluster dataplane identity, CNI status, endpoint health, flows, then kernel state. |
This is why I would now add one question before almost every Kubernetes network incident: What data plane am I actually standing on?
If you are still orienting yourself around what belongs to a DevOps role versus neighboring infrastructure disciplines, the Platform Engineer vs. DevOps Engineer guide covers that role boundary. This article deliberately stays below that organizational layer and concentrates on Kubernetes packet processing, service routing, policy, and observability.
Where kube-proxy's Scaling Limits Actually Show Up
The old argument that “iptables cannot scale” is often repeated too casually. The real Kubernetes problem is more specific: in iptables mode, kube-proxy creates rules associated with Services and their endpoints, so very large clusters can accumulate tens of thousands of rules and incur meaningful update costs when Services or EndpointSlices change. Kubernetes’ own networking documentation explicitly identifies that behavior and the time kube-proxy can spend updating large rulesets.
Upstream Kubernetes has not ignored that problem. The nftables backend for kube-proxy became stable in Kubernetes 1.33, with the project describing it as addressing longstanding performance and scalability issues in the iptables implementation; compatibility considerations kept iptables as the default at that point.
So the 2026 story is not “Cilium fixed Kubernetes because kube-proxy was unusable.” The more accurate story is that managed Kubernetes platforms now have multiple viable dataplane directions: improved upstream kube-proxy backends on one side, and eBPF-native service implementations that can eliminate kube-proxy on the other.
For an operator, that means the diversity of debugging models is increasing before it decreases.
What eBPF and Cilium Actually Do, and Where the Clouds Default to Them
eBPF lets verified programs execute at defined hooks in the Linux kernel. For Kubernetes networking, that creates an opportunity to make routing, load balancing, policy, and observability decisions closer to the packet path while retaining Kubernetes context that a generic packet-filtering rule does not naturally carry. Google’s Dataplane V2 documentation describes exactly that model: eBPF programs receive packets in the kernel and can use Kubernetes-specific metadata while deciding how those packets should be processed.
Cilium builds a Kubernetes networking platform around that capability. Its node agent consumes desired state from Kubernetes and programs eBPF state for networking, load balancing, policy, and visibility; Hubble sits on top as an observability layer for flows.
The cloud-by-cloud distinction matters most in practice:
Platform / tier | 2026 default or documented position | kube-proxy status | What that means operationally |
AKS Automatic | Azure CNI Overlay powered by Cilium is preconfigured as the default network | Microsoft says Cilium dataplane clusters do not use kube-proxy | Do not assume Kubernetes Service state exists in kube-proxy iptables rules |
AKS Standard | Cilium remains an explicit networking choice | No kube-proxy when Cilium dataplane is selected | Existing Standard runbooks may differ cluster by cluster |
GKE Autopilot | GKE Dataplane V2 is enabled by default on all new clusters | Cilium implements Kubernetes Services instead of kube-proxy | eBPF-aware troubleshooting is baseline knowledge |
GKE Standard | Dataplane V2 is available but is a cluster-creation choice rather than a universal Standard default | Depends on cluster dataplane | Inventory configuration before applying a runbook |
EKS standard EC2-node cluster | VPC CNI, kube-proxy and CoreDNS are default/bootstrap components | kube-proxy remains part of documented default | Traditional kube-proxy debugging remains relevant |
EKS Auto Mode | AWS provides a built-in networking capability and manages networking components rather than asking users to manage standard add-on pods | AWS hides/manages much of the implementation; it does not document a Cilium default | Do not equate “managed” with “Cilium”; inspect EKS-specific tooling |
EKS with alternate CNI | Cilium can be installed on standard EC2 nodes, but AWS classifies it as an alternate compatible CNI; Auto Mode does not support alternate CNIs | Can be replaced depending on architecture | You own more integration and support responsibility |
Microsoft’s Cilium configuration page, reviewed August 19, 2026, identifies AKS Automatic’s Cilium configuration as the default and says AKS Standard requires the Cilium networking path explicitly. The page showed a last-updated date of August 5, 2026, while a separate AKS platform comparison page was updated August 12, 2026.
That date correction does not change the product fact. Microsoft’s documentation says AKS Automatic uses Azure CNI Overlay powered by Cilium without extra setup, while an AKS Standard creation command specifies --network-dataplane cilium.
A representative Standard creation path is:
az aks create \
--name my-cluster \
--resource-group my-rg \
--network-plugin azure \
--network-plugin-mode overlay \
--network-dataplane cilium \
--generate-ssh-keysThe important part is not memorizing the flag. It is realizing that two engineers saying “we run AKS” may be operating materially different packet paths.
Google’s position is even easier to state. Its GKE Dataplane V2 page was last updated August 11, 2026 UTC, says Dataplane V2 is enabled by default for all new Autopilot clusters, states that it is implemented using eBPF and Cilium, and identifies legacy GKE networking as Calico-based.
What Dataplane V2 Actually Replaced
Google’s documentation draws the line cleanly: Dataplane V2 uses Cilium technology and eBPF, while the legacy GKE dataplane uses Calico and Linux iptables for policy behavior. In relevant Dataplane V2 versions, Cilium also implements Kubernetes Services instead of kube-proxy, which removes kube-proxy/iptables Service-routing bottlenecks but introduces eBPF-map limits and its own operational constraints.
That last clause matters. Every data structure has limits.
Google documents a Dataplane V2 Service map capacity of 260,000 endpoint entries across Services, and it warns that high connection churn can drive significant CPU consumption in anetd. It also warns that installing unrelated eBPF software can interfere with GKE Dataplane V2 because multiple programs may compete for or alter critical dataplane behavior.
This is why I would never sell eBPF as “no more networking bottlenecks.” You are trading one architecture for another architecture with different strengths, failure modes, resource ceilings, and observability surfaces.
AWS requires the most careful wording.
EKS documentation still says that standard cluster creation installs VPC CNI, kube-proxy, and CoreDNS by default, and its current kube-proxy documentation publishes supported kube-proxy builds through Kubernetes 1.36. AWS also says VPC CNI is the only CNI supported by Amazon EKS for EC2 nodes in the fully AWS-supported sense, although upstream Kubernetes allows alternate compatible CNIs such as Cilium and AWS documents partner-supported alternatives.
There are two nuances that keep this from becoming an anti-AWS caricature:
Amazon VPC CNI itself uses eBPF for Kubernetes NetworkPolicy enforcement, so AWS absolutely participates in eBPF-based Kubernetes networking.
AWS now supports Cilium directly for EKS Hybrid Nodes, including kube-proxy replacement there, while explicitly distinguishing that support from Cilium on nodes running in AWS Cloud.
EKS Auto Mode also deserves its own footnote. AWS says Auto Mode uses a built-in networking capability, does not support alternate CNI plugins such as Cilium or Calico, and manages networking behavior through EKS constructs such as NodeClass; configuration settings from the traditional open-source VPC CNI do not simply carry across.
Confidence as of August 19, 2026:
High confidence: Azure has made Cilium/eBPF the network default for AKS Automatic, and Google has made Dataplane V2/eBPF the default for GKE Autopilot. AWS has not documented an equivalent Cilium-based kube-proxy replacement default for EKS cloud-node clusters. That describes current documented product configuration, not AWS’s future roadmap.
That is the honest picture of Kubernetes networking defaults in 2026.
For readers working one layer further out, Pulumi, Terraform, and the 2026 IaC decision guide covers infrastructure provisioning choices. Terraform can create your cluster and configure its networking mode, but it does not replace the need to understand what the selected dataplane does once packets start dropping.
Why the Uneven Shift Matters for Multi-Cloud Teams
The operational risk is not that engineers have forgotten Kubernetes networking. It is that the phrase “Kubernetes networking” now hides increasingly different implementations behind the same API objects.
A Service is still a Service. A NetworkPolicy is still a Kubernetes API object. Your Deployment manifest does not tell you whether packets are being steered by kube-proxy-managed iptables, nftables, Cilium eBPF programs, a cloud-specific managed dataplane, or some combination of technologies at different layers. Kubernetes defines the Service abstraction independently of how a particular implementation realizes it.
That creates a runbook problem.
A multi-cloud team can easily have:
AKS Automatic clusters where Cilium replaces kube-proxy.
Older or Standard AKS clusters with a different networking configuration.
GKE Autopilot clusters on Dataplane V2.
GKE Standard clusters where the dataplane choice varies.
Standard EKS clusters using VPC CNI plus kube-proxy.
EKS clusters where VPC CNI network policies are enforced with eBPF even while kube-proxy still handles Services.
Specialist EKS deployments running an alternate CNI such as Cilium.
If your incident guide begins “SSH to node, run iptables-save, find the KUBE chain,” you no longer have a Kubernetes runbook. You have an implementation-specific runbook that forgot to state its preconditions.
My replacement is to put dataplane identification at step zero:
Question | Command or source | Why I care |
What managed Kubernetes tier is this? | Cloud API / cluster metadata | Automatic/Autopilot/Standard can imply different defaults |
Is kube-proxy present? | kubectl -n kube-system get ds,pods | grep kube-proxy | Determines whether kube-proxy state is even relevant |
Is Cilium running? | kubectl -n kube-system get pods | grep cilium or provider-specific components | Identifies likely eBPF dataplane |
What CNI does the cloud support here? | Cloud configuration and vendor docs | Avoids assumptions from another provider |
Is NetworkPolicy implemented by the same dataplane as Service routing? | Provider/CNI documentation | On EKS, for example, the answer can involve both kube-proxy and eBPF |
Are components hidden because the tier manages them? | Provider documentation | Absence from kube-system does not prove absence from architecture |
The AWS NetworkPolicy example is particularly useful because it breaks the simplistic cilium vs kube-proxy framing. AWS can keep kube-proxy for Service networking while enforcing VPC CNI NetworkPolicy with eBPF; one cluster can therefore require both traditional kube-proxy knowledge and eBPF-aware debugging.
That is what multi-cloud competence looks like in 2026: not memorizing one fashionable dataplane, but recognizing which layer owns each behavior.
What Changes in eBPF Observability and Network Debugging
The old networking workflow was often reconstruction by side effects. You saw an iptables counter increment, a conntrack entry, a SYN on one interface and no SYN-ACK on another, and then you inferred which Kubernetes object produced the state.
eBPF-based Kubernetes observability tooling can move some of that context closer to the dataplane event itself. Cilium’s Hubble observability layer is specifically designed to expose communication behavior and networking events at node, cluster, and multi-cluster scope, using information produced by the Cilium/eBPF dataplane.
I still use tcpdump. I just no longer make it my first tool for every policy or Service question.
A useful new hierarchy is:
1. Kubernetes desired state: Service, EndpointSlice, NetworkPolicy, Pod labels, namespaces.
2. Dataplane health: CNI agent/operator status and provider-managed networking health.
3. Flow evidence: allowed, forwarded, dropped, policy verdicts, source/destination identities.
4. eBPF state: service maps, policy maps, endpoints, program attachment.
5. Packet capture: validate what actually traverses a specific interface.
6. Kernel/network fallback: routes, neighbors, conntrack where applicable, MTU, socket state, host firewall.
Cilium’s troubleshooting documentation recommends detailed status inspection through its debugging tooling, and its Hubble CLI is built specifically for inspecting network flows.
Hubble and Flow-Level Visibility Explained
Think of a Hubble network flow as a higher-context answer to the question I used to answer by correlating four terminals.
Instead of beginning with “I saw a TCP SYN leave veth123,” the flow model can tell you that a workload identity in namespace A attempted to communicate with another workload or Service, what protocol was involved, and whether policy or forwarding behavior allowed or dropped that path. Hubble can expose this at node or cluster scope, and Hubble UI can derive a service map from those flows.
Typical Cilium/Hubble investigation commands include:
cilium status
hubble status
hubble observe
hubble observe --verdict DROPPED
hubble observe --from-namespace frontend
hubble observe --to-namespace paymentsExact flags and availability depend on the environment and how the provider exposes Cilium. Managed offerings do not necessarily expose every upstream Cilium knob or CLI surface, so I treat upstream commands as concepts first and provider support boundaries second. Azure, for example, documents specific Cilium capabilities and notes that many cilium-config changes are not supported in its managed implementation.
GKE has another important variation. Google exposes Dataplane V2 observability and built-in NetworkPolicy logging, and new Dataplane V2 clusters include networking logging facilities that can show allowed and denied policy traffic.
AKS likewise distinguishes base Cilium networking from additional Azure Advanced Container Networking Services. Microsoft’s current feature table says base Cilium provides Kubernetes and Cilium L3/L4 policy support, while container network observability and flow logs are associated with ACNS rather than being an automatic promise that every AKS Cilium user receives the complete upstream Hubble experience.
That is an important procurement and troubleshooting detail: “Cilium under the hood” does not automatically mean “every Cilium observability feature is exposed to me.”
What Changes in How You Debug a Networking Issue
Here is the practical scenario.
A checkout pod cannot reach payments.default.svc.cluster.local:8443. DNS resolves. The Service has healthy endpoints. The application times out.
My old playbook would probably move from Service and endpoints toward kube-proxy state and the node’s NAT table.
My 2026 playbook starts by classifying the dataplane.
Is this AKS Automatic? Expect Azure CNI Overlay powered by Cilium.
Is this GKE Autopilot? Expect Dataplane V2.
Is this standard EKS? Expect VPC CNI plus kube-proxy unless the team deliberately changed the bootstrap architecture.
Is this EKS Auto Mode? Use the EKS Auto Mode networking model rather than assuming the standard VPC CNI configuration is directly inspectable.
Then I verify desired state:
kubectl get svc payments -o wide
kubectl get endpointslice \
-l kubernetes.io/service-name=payments -o yaml
kubectl get pod checkout-xxxxx -o wide --show-labels
kubectl get networkpolicy -ANo dataplane can route to an endpoint Kubernetes never published.
Then I ask whether the problem follows the Service VIP or the backend itself. A controlled test to the backend Pod IP, where architecture and policy permit it, can separate Service load-balancing behavior from basic pod-to-pod reachability.
For Cilium-based environments, I then want endpoint and flow evidence before rummaging through generic firewall state. Upstream Cilium exposes service state and detailed dataplane debugging commands, including service maps used by the eBPF kube-proxy replacement.
A conceptual sequence looks like:
cilium status
hubble observe \
--from-pod default/checkout-xxxxx \
--to-pod default/payments-yyyyy
hubble observe \
--verdict DROPPEDIf the source flow is visible and policy verdict says dropped, that is a different incident from “packet never reached the dataplane.”
If the flow is forwarded toward the correct backend but the application still fails, I move down the stack: listening socket, pod network namespace, MTU, return route, node routing, security groups or NSGs, load-balancer path, and packet capture.
New Tools Replacing Old iptables Habits
The migration is not a one-for-one command substitution.
Old habit | eBPF-era equivalent question |
iptables -L | Which program/map implements this behavior? |
Inspect KUBE-SVC-* | What backend does the eBPF Service map select? |
Count ACCEPT/DROP rules | What flow verdict did the dataplane emit? |
Search source/destination IPs only | Which Kubernetes identities, pods, namespaces and policies were involved? |
Assume conntrack explains every Service translation | Check whether the dataplane uses different eBPF state for service handling |
Restart kube-proxy | Verify Cilium/provider agent health and controller reconciliation |
Read generic node firewall rules first | Correlate Kubernetes state with CNI state first |
Cilium’s kube-proxy replacement relies on eBPF features such as socket-level load balancing, so Service handling can occur at points where an operator trained exclusively on node-level DNAT rules may not expect it.
There is another old habit worth retiring: treating Kubernetes-created iptables chains as a stable API. Kubernetes itself warned years ago that implementation chains such as its KUBE-* rules were never intended as an API or ABI contract for outside tools.
The architecture change has now made that warning painfully practical.
My recommended incident checklist for eBPF clusters is:
Confirm the exact managed-cluster mode and CNI/dataplane.
Confirm whether kube-proxy exists.
Verify Service and EndpointSlice state.
Verify pod identities, labels and applicable policies.
Check CNI/agent status.
Inspect network-flow verdicts.
Inspect eBPF Service/policy state where the provider exposes it.
Capture packets only after you know which interface and stage should contain them.
Escalate into routes, MTU, cloud network controls, DNS, application sockets and host networking.
The important shift is epistemic: I do not begin with the kernel implementation I remember. I begin by proving the kernel implementation the cluster is actually using.
Security Implications and Cilium's Maturity in Context
One reason eBPF dataplanes have moved beyond a performance discussion is policy enforcement.
Cilium can attach eBPF logic to workload networking paths and enforce ingress and egress decisions at the node, while supporting Kubernetes NetworkPolicy as well as Cilium-specific policy resources.
That gives operators opportunities for richer policy and identity-aware visibility, but it does not mean “Cilium automatically secures the cluster.”
You still need:
explicit policy design rather than assuming installation equals isolation;
correct namespace and label selectors;
egress controls where required;
DNS allowances;
host-network exceptions understood;
cloud security groups/NSGs aligned with pod-level policy;
observability proving that policy behaves as intended.
Azure’s Cilium documentation demonstrates why implementation details matter. It documents limitations around ipBlock behavior for pod and node IPs, host-networked pods, and which advanced L7/FQDN capabilities require additional Azure networking services.
GKE makes Kubernetes NetworkPolicy continuously available when Dataplane V2 is used, but Google also documents dataplane-specific limitations and warns against arbitrary third-party eBPF programs that can interfere with critical network programs.
AWS illustrates a different architecture: Amazon VPC CNI implements Kubernetes NetworkPolicy enforcement with eBPF while retaining kube-proxy in the standard Service path. That means security tooling and Service-routing tooling can be implemented through different kernel mechanisms inside the same cluster.
For readers thinking more broadly about container security, Refonte’s DevOps security scanning tools (Trivy) article belongs at a different layer. Image and SBOM scanning answer “what vulnerabilities and packages are inside the artifact”; CNI/eBPF policy answers “what may communicate at runtime.” Those controls complement rather than replace one another.
Reading the CNCF Project Data Correctly: It Is Not Fresh 2026 Data
Cilium’s maturity is real, but older project statistics should not be presented as current-year momentum.
The CNCF Cilium Project Journey Report was published November 11, 2024. Its headline snapshot listed 3,195+ contributors and 567+ contributing companies, and it named organizations including Adobe, Alibaba, AWS, Google, DigitalOcean and Bloomberg in its adoption and end-user discussion.
The same report also contains broader cumulative DevStats measurements that cite 4,464 contributors and 1,011 contributing organizations around its 2024 reporting period, illustrating that the report uses multiple snapshots/scopes for different charts. For the exact requested project-snapshot figures, 3,195+ and 567+ are the clearly labelled headline values.
Those are 2024 project-maturity data, not 2026 adoption statistics.
No 2026 refresh of this specific CNCF Project Journey Report was found in the source review. The responsible interpretation is that by late 2024, Cilium had already demonstrated a large contributor ecosystem, graduated CNCF status, and named large-scale production users. The report supports a maturity argument, but it does not prove fresh 2026 growth.
The strongest 2026 evidence comes from the cloud providers themselves, not from extrapolating the CNCF chart. Azure’s Automatic default and Google’s Autopilot default are current product decisions that demonstrate how far the technology has moved into managed Kubernetes.
That distinction between historical maturity evidence and current product-default evidence is essential to accurate technical writing.
Should You Turn Cilium On Manually, and What Skills Matter Now?
I would not enable Cilium merely because “eBPF is the future.”
Changing the Kubernetes dataplane is a foundational architectural decision. It affects service routing, policy, support boundaries, upgrades, observability, integration with cloud-native networking, and the assumptions embedded in your incident procedures.
A sensible decision framework is:
Situation | My default recommendation |
New AKS Automatic cluster | Use the managed default unless a documented limitation conflicts with requirements |
New GKE Autopilot cluster | Dataplane V2 is already the default; train operators accordingly |
Existing stable AKS Standard cluster | Migrate only for a defined benefit and test the support/feature matrix |
Existing GKE Standard cluster | Remember Dataplane V2 cannot simply be enabled in-place on older existing clusters; plan replacement/migration architecture |
Standard EKS using VPC CNI + kube-proxy | Do not replace the supported default just to follow a trend |
EKS requiring Cilium-specific functionality | Evaluate alternate-CNI support ownership, commercial support and migration risk |
EKS Auto Mode | Do not plan around installing Cilium; AWS says alternate CNIs are not supported there |
Multi-cloud platform | Optimize for common Kubernetes abstractions but maintain provider-specific dataplane runbooks |
Google explicitly documents that Dataplane V2 is selected at cluster creation and existing clusters cannot simply be upgraded into it through the same path. AWS advises teams choosing alternate compatible CNIs on EC2 nodes to obtain appropriate commercial support or possess the internal expertise needed to troubleshoot them.
That support boundary matters more to me than a microbenchmark.
The skills I now expect from a senior DevOps or SRE working serious Kubernetes incidents include:
Kubernetes Services and EndpointSlices.
CNI fundamentals and cloud-specific pod IP management.
Linux network namespaces, veths, routes and neighbor state.
iptables and nftables literacy.
kube-proxy architecture.
eBPF concepts: programs, hooks, maps, identities and service maps.
Cilium operational commands.
Hubble network-flow analysis.
Kubernetes NetworkPolicy semantics.
Cloud firewalls, security groups/NSGs and load balancers.
packet capture and socket-level debugging.
That does not mean junior engineers should start by memorizing BPF bytecode.
The right learning order is Linux networking → containers → Kubernetes → Services/CNI/NetworkPolicy → cloud networking → observability → eBPF implementation details. Without that foundation, hubble observe becomes another magic command copied from a runbook rather than evidence you can reason about.
The career market gives that investment some economic context. Indeed’s U.S. DevOps Engineer salary page, updated August 10, 2026, reports an average base salary of $133,481 per year, based on roughly 4,200 salaries from job postings over the preceding 36 months, with a displayed low of $86,500 and high of $205,979.
I would present that as an Indeed market estimate, not as a guaranteed salary for every DevOps engineer in 2026.
It is higher than Refonte Learning’s own program-page marketing figure of $102,500+ starting salary. Refonte also advertises roughly 100,000 annual job openings for the track, but those are the program’s own marketing claims and were not independently validated here.
Glassdoor and ZipRecruiter are not used as corroborating salary sources here because their exact figures could not be independently confirmed. Indeed’s dated page is directly verifiable.
That matters because compensation claims age almost as quickly as cloud defaults.
Building This Foundation with the Refonte Learning DevOps Engineer Program
eBPF should not be your first Kubernetes topic.
Before a cilium status result means anything, you need to understand what a pod is, how a Service finds endpoints, how Linux networking connects namespaces, why a container needs a CNI, what a cloud VPC/VNet is doing underneath the cluster, and how logs and metrics help you separate infrastructure failure from application failure.
That sequence maps naturally onto the verified curriculum of the Refonte Learning DevOps Engineer Program.
The live program page lists:
Duration: 3 months.
Time commitment: 12–14 hours per week.
Linux: Linux fundamentals and scripting.
Version control: Git and GitHub.
Delivery: Continuous Integration and Continuous Deployment.
Containers: Docker and Kubernetes.
Infrastructure as Code: Terraform.
Cloud: AWS, Azure and GCP.
Operations: monitoring and logging tools.
Practical work: a capstone project.
The program identifies MSc Oskar Eriksson as a lead instructor/mentor with more than 10 years of technology experience and a background that includes cloud computing and DevOps.
Admission is currently aimed at learners working toward a bachelor’s degree or higher. The page lists a $300 one-time enrollment cost, or installment amounts of $204 and $98.
Its listed career outcomes include DevOps Engineer, Cloud Engineer and Site Reliability Engineer.
The honest curriculum note is just as important as the positive case: the page names Docker and Kubernetes generically and does not currently list eBPF, Cilium, Hubble, or CNI-level networking as explicit modules.
That does not make the curriculum irrelevant to the subject. It tells you where this skill belongs in the stack.
A sensible progression from the program into advanced Kubernetes networking would look like:
1. Learn Linux processes, interfaces, routing and scripting.
2. Understand containers and network namespaces.
3. Operate Kubernetes Deployments, Pods and Services.
4. Learn AWS, Azure and Google Cloud network primitives.
5. Understand monitoring and logging during failures.
6. Build or operate clusters using different CNI/dataplane choices.
7. Compare kube-proxy, nftables and eBPF service-routing models.
8. Practice Cilium status inspection and Hubble network-flow analysis.
9. Debug the same synthetic failure on AKS, GKE and EKS.
10. Document different runbooks rather than pretending the clouds are interchangeable.
That last exercise is the one I would put in front of any engineer claiming Kubernetes networking depth in 2026.
Create a broken Service on a traditional kube-proxy cluster. Trace the Service through EndpointSlices into kube-proxy state and the node dataplane. Then reproduce the failure on a Cilium-backed cluster and force yourself not to start with the old KUBE iptables chains.
The difference becomes obvious immediately.
The larger lesson in the Cilium vs. kube-proxy comparison is not that one binary won and another binary lost. Kubernetes kept its API abstractions while cloud providers increasingly changed the implementation underneath them.
Azure chose a Cilium-powered default for AKS Automatic. Google chose an eBPF/Cilium implementation for GKE Autopilot. AWS still documents a different default cloud-node architecture, while selectively using eBPF itself and supporting Cilium in other EKS scenarios.
For DevOps engineers, that means your most valuable networking skill in 2026 is not remembering one more command.
It is being able to look at an unfamiliar Kubernetes cluster, determine which component actually owns the packet path, gather evidence from the right layer, and change your debugging model before your old muscle memory sends you searching for rules that were never there.
