Andrew Mercer
on this page

Container Networking: Docker, Podman, and Kubernetes

Everything here reduces to a handful of Linux kernel primitives. Docker, Podman, and Kubernetes are different control planes arranging the same building blocks, so this guide starts there.

Table of Contents

  1. The Linux primitives
  2. Docker networking
  3. Podman networking
  4. Kubernetes networking
  5. Comparison cheat sheet
  6. Debugging playbook
  7. Mental models to keep

Part 1: The Linux primitives

1.1 Network namespaces

A network namespace (netns) is an isolated copy of the network stack: its own interfaces, routing tables, iptables/nftables rules, conntrack table, ARP table, and socket port space. A "container" has no networking of its own; it is a process (tree) placed in a new netns.

ip netns add demo
ip netns exec demo ip addr        # only lo, and it's DOWN
ls /run/netns                     # named netns are bind mounts of /proc/<pid>/ns/net

To enter a container's netns from the host:

PID=$(docker inspect -f '{{.State.Pid}}' mycontainer)   # or podman inspect
nsenter -t $PID -n ip addr
nsenter -t $PID -n ss -tlnp

This is the most useful debugging technique in container networking, because you can use host tools (tcpdump, ss, ip, conntrack) against a container that has none installed.

1.2 veth pairs

A veth pair is a virtual cable with two ends. Packets in one end come out the other. One end goes into the container netns (usually renamed eth0), the other stays in the host netns, attached to a bridge or routed directly.

ip link add veth-host type veth peer name veth-ctr
ip link set veth-ctr netns demo

On the host you'll see interfaces like veth3f2a1b@if7. The @if7 is the interface index of the peer inside the container. That's how you map a host veth to a container.

1.3 Linux bridges

A Linux bridge is a virtual L2 switch. Host-side veth ends are enslaved to it, so containers on the same bridge share an L2 broadcast domain. The bridge device itself (e.g. docker0, cni0, podman0) usually gets an IP and acts as the containers' default gateway.

bridge link            # which veths are on which bridge
bridge fdb show        # learned MACs
ip -d link show docker0

1.4 Routing, NAT, and conntrack

  • IP forwarding (net.ipv4.ip_forward=1) lets the host route between interfaces. Container runtimes enable it.
  • SNAT/MASQUERADE rewrites the source address of container traffic leaving the host so replies return to the host.
  • DNAT rewrites the destination, which is how published ports and Kubernetes Services work.
  • conntrack remembers the translation for each flow so reply packets are un-NATed automatically. Conntrack table exhaustion and races are a classic source of mysterious Kubernetes packet drops.

1.5 Other pieces you'll meet

Primitive Used for
VXLAN L2-over-UDP overlay tunnels (Docker overlay, Flannel, Calico VXLAN, Cilium)
IPIP / WireGuard Alternative encapsulation/encryption between nodes (Calico, Cilium)
macvlan / ipvlan Give a container an address directly on the physical network
eBPF (tc/XDP) Replace iptables for service load balancing and policy (Cilium)
tun/tap Userspace networking (rootless slirp4netns, pasta)

Part 2: Docker networking

2.1 Architecture

dockerd creates the netns and wires it up through libnetwork, which implements the Container Network Model (CNM): Sandbox (the netns), Endpoint (a veth end), and Network (a group of endpoints that can communicate). Drivers implement Networks.

2.2 The default bridge (docker0)

On install, Docker creates docker0 with 172.17.0.0/16. Containers started without --network join it.

Container A (172.17.0.2) ─ veth ─┐
                                  ├─ docker0 (172.17.0.1) ─ iptables NAT ─ eth0 ─ LAN
Container B (172.17.0.3) ─ veth ─┘

Key limitations of the default bridge:

  • No automatic DNS by container name. Only IP access (or the legacy --link).
  • All containers share one flat network, so there is no isolation between unrelated apps.

2.3 User-defined bridges (use these instead)

docker network create --driver bridge --subnet 10.20.0.0/24 appnet
docker run -d --name db --network appnet postgres
docker run -d --name api --network appnet myapi

Differences from the default bridge:

  • Embedded DNS server at 127.0.0.11 inside each container. Docker intercepts queries there, resolves container names and network aliases, and forwards everything else to the host's upstream resolvers.
  • Better isolation, since containers only see networks they're attached to.
  • Containers can attach and detach live, and can be on multiple networks.
  • Each network gets its own Linux bridge (br-<id>).

2.4 How outbound traffic works

Containers reach the outside via a MASQUERADE rule:

iptables -t nat -S POSTROUTING
# -A POSTROUTING -s 172.17.0.0/16 ! -o docker0 -j MASQUERADE

The path is: container → veth → bridge → host routing → SNAT to host IP → out the physical NIC. Replies come back via conntrack.

2.5 How published ports work (-p 8080:80)

Docker creates a DNAT rule in the DOCKER chain of the nat table:

iptables -t nat -S DOCKER
# -A DOCKER ! -i docker0 -p tcp --dport 8080 -j DNAT --to-destination 172.17.0.2:80

Packet flow for an external client hitting host:8080:

  1. Packet arrives on the host NIC, and PREROUTING sends it to the DOCKER chain.
  2. DNAT rewrites the destination to 172.17.0.2:80.
  3. Routing decision: the destination is on docker0, so it goes through the FORWARD chain.
  4. FORWARD passes through DOCKER-USER → DOCKER-ISOLATION-* → DOCKER, which accepts traffic to published ports.
  5. Reply traffic is un-NATed by conntrack.

docker-proxy: for traffic from the host itself to localhost:8080, DNAT in PREROUTING doesn't apply (locally generated traffic hits OUTPUT, and loopback has special handling). Docker historically runs a userland docker-proxy process per published port to handle this hairpin case. You can disable it with "userland-proxy": false in daemon.json and rely on iptables-only handling.

2.6 iptables chains Docker manages

Chain Purpose
nat/DOCKER DNAT for published ports
nat/POSTROUTING MASQUERADE for outbound
filter/FORWARD Entry point for forwarded container traffic
filter/DOCKER-USER Your hook. Evaluated first, and Docker never overwrites it. Put custom filtering here.
filter/DOCKER-ISOLATION-STAGE-1/2 Blocks traffic between different bridge networks
filter/DOCKER Allows traffic to published container ports

Important gotcha: Docker bypasses UFW/firewalld input rules. Published ports are DNATed and then forwarded, so they traverse FORWARD, not INPUT, and your host firewall's INPUT rules never see them. To restrict access, either bind explicitly (-p 127.0.0.1:8080:80) or add rules to DOCKER-USER. Recent Docker releases tightened some of this behavior (e.g. blocking direct access to unpublished container ports from other hosts), and there's ongoing nftables support, so check your version's docs.

2.7 Other Docker network drivers

host (--network host): The container shares the host's netns. No veth, no NAT, no port mapping, and -p is ignored. It has the best performance but no isolation, and port collisions are possible.

none: Only loopback. For fully isolated workloads, or when you'll wire networking up yourself.

container:\<name> (--network container:other): Joins another container's netns. Both share the same interfaces and localhost. This is exactly how Kubernetes pods work.

overlay (Swarm, or standalone with attachable networks): Multi-host L2-like networks using VXLAN (UDP 4789). Control plane uses gossip on 7946 (TCP/UDP) and Swarm management on 2377. Traffic can optionally be encrypted with IPsec (--opt encrypted). Each host has a docker_gwbridge for egress.

macvlan: Each container gets its own MAC and an IP on your physical LAN, so it looks like a separate machine.

docker network create -d macvlan --subnet 192.168.1.0/24 --gateway 192.168.1.1 \
  -o parent=eth0 lan

Caveats: the host can't talk to its own macvlan containers through the parent interface (the kernel forbids it; you need a macvlan sub-interface on the host as a workaround), the switch port must allow multiple MACs, and many Wi-Fi APs won't work.

ipvlan (L2 or L3 mode): Similar to macvlan but all containers share the parent's MAC. It works where MAC limits apply, and L3 mode routes without broadcast.

2.8 Docker DNS details

  • Containers on user-defined networks get 127.0.0.11 in /etc/resolv.conf.
  • Aliases: --network-alias, and Compose service names automatically.
  • With multiple containers sharing an alias, Docker returns them round-robin (DNS RR, not real load balancing).
  • Swarm services get a VIP (IPVS-based) in addition to DNS.
  • Custom DNS: --dns, --dns-search, or daemon.json.

2.9 Compose

Compose creates a project-scoped user-defined bridge (<project>_default) and attaches every service, so service names resolve as hostnames. Use multiple networks: entries to segment tiers (frontend/backend), and internal: true for networks with no external route.


Part 3: Podman networking

Podman is daemonless and rootless-capable, and its networking differs mainly in who sets it up and how rootless works.

3.1 Stack evolution

  • CNI plugins (legacy, removed in Podman 5).
  • Netavark (Rust) sets up interfaces, bridges, and firewall rules. Aardvark-dns is the companion authoritative DNS server for container names. This is the current default.
podman info --format '{{.Host.NetworkBackend}}'    # netavark
podman network ls
podman network inspect podman

3.2 Rootful networking

Similar to Docker: the default podman network is 10.88.0.0/16 on bridge podman0, with MASQUERADE and DNAT for published ports. Netavark can drive iptables, nftables, or firewalld backends.

Notable differences:

  • The default podman network has DNS disabled. Create a user-defined network (podman network create) to get name resolution via aardvark-dns.
  • Multiple networks per container and static IP/MAC assignment are supported.
  • Drivers: bridge (default), macvlan, ipvlan.

3.3 Rootless networking

An unprivileged user can't create bridges or edit host iptables, so rootless Podman uses a userspace network stack:

  1. Podman creates a user namespace and netns.
  2. A userspace helper connects that netns to the host's network by terminating TCP/UDP flows in userspace and re-originating them as normal host sockets.

pasta (from the passt project, the default in Podman 5+) replaced slirp4netns. It's faster, preserves source IPs better, and can map host addresses into the container more transparently. slirp4netns is still selectable.

podman run --network pasta ...
podman run --network slirp4netns ...

Consequences:

  • Rootless containers can't ping arbitrary hosts unless net.ipv4.ping_group_range allows it.
  • Ports below 1024 need sysctl net.ipv4.ip_unprivileged_port_start=80 (or lower).
  • Throughput is lower than kernel-path networking.
  • Rootless user-defined networks still work: netavark builds a bridge inside a rootless netns, and rootless port forwarding is done by a helper (rootlessport, or pasta).
  • Container-to-container traffic on the same rootless network still goes over a bridge, but that bridge lives inside the rootless network namespace.

3.4 Pods

podman pod create makes an infra container (like Kubernetes' pause container) that owns the shared netns. Every container in the pod joins it, so they share localhost and port space. Ports are published on the pod, not on individual containers.

podman pod create --name web -p 8080:80
podman run -d --pod web nginx
podman run -d --pod web myapp     # reaches nginx on localhost:80

This is the direct bridge to Kubernetes: podman kube generate and podman kube play translate between pods and Kubernetes YAML.

3.5 Quadlet / systemd

With Quadlet .container and .network units, Podman networks become declarative systemd units, so Network=appnet.network in a container unit defines the attachment and ordering.

3.6 macOS/Windows

podman machine runs a Linux VM, and gvproxy (gvisor-tap-vsock) handles host-to-VM port forwarding and DNS, which is why published ports on macOS behave slightly differently.


Part 4: Kubernetes networking

4.1 The Kubernetes network model

Kubernetes doesn't implement networking itself; it defines rules that any conforming network must satisfy:

  1. Every pod gets its own unique IP.
  2. Pods can communicate with all other pods, on any node, without NAT.
  3. Agents on a node (kubelet, system daemons) can communicate with all pods on that node.
  4. A pod sees its own IP as the same address others see.

This flat model is a major difference from Docker's default NAT-based model. Four problems need solving:

Problem Solved by
Container ↔ container in a pod Shared netns (localhost)
Pod ↔ pod CNI plugin
Pod ↔ Service kube-proxy (or eBPF replacement)
External ↔ Service NodePort, LoadBalancer, Ingress/Gateway

4.2 The pod network namespace

When the kubelet creates a pod:

  1. The container runtime (containerd/CRI-O) via CRI creates a pause (sandbox) container, which owns the netns and does nothing else.
  2. The runtime invokes the CNI plugin (ADD) with the netns path. The plugin creates the veth pair, assigns an IP from the node's pod CIDR (via IPAM), and sets up routes.
  3. App containers then join the pause container's netns.

This is why a container crash doesn't change the pod's IP: the netns lives in the pause container. Pod IPs are ephemeral, since a recreated pod gets a new IP.

CNI config lives in /etc/cni/net.d/, and binaries in /opt/cni/bin/.

4.3 Pod-to-pod: how CNI plugins move packets

Same node: pod → veth → node bridge or direct routing → other pod's veth.

Cross-node has three main strategies:

Overlay (encapsulation): Pod packets are wrapped in VXLAN/Geneve/IPIP with the node IPs as outer addresses. The underlay only needs node-to-node connectivity. The cost is an MTU reduction (VXLAN overhead ≈ 50 bytes), a small CPU cost, and harder packet inspection. Used by Flannel (VXLAN), Calico (VXLAN/IPIP), Cilium (VXLAN/Geneve).

Routed (no encapsulation): Nodes announce pod CIDRs via BGP (Calico, Cilium BGP control plane, kube-router), or you use cloud route tables. It's fast and easy to debug with tcpdump, but the underlay must know pod routes.

Native VPC/VNet IPs: Pods get real IPs from the cloud network (AWS VPC CNI, Azure CNI classic, GKE VPC-native). No overlay, and pods are first-class citizens on the network, but you can exhaust IP address space.

CNI Dataplane Notes
Flannel VXLAN (or host-gw) Minimal. No NetworkPolicy.
Calico iptables, nftables, or eBPF; BGP/VXLAN/IPIP Strong policy, flexible routing.
Cilium eBPF Can replace kube-proxy, L3–L7 policy, Hubble observability, WireGuard encryption, Gateway API.
AWS VPC CNI Native VPC IPs (ENIs) Pod density limited by ENI/IP quotas.
Azure CNI Several modes (see below)
Weave / others Various Less common now.

AKS specifics: Azure CNI Node Subnet gives pods VNet IPs from the node subnet (IP-hungry). Azure CNI Overlay gives pods overlay IPs from a private CIDR, so it scales without consuming VNet space. Azure CNI Powered by Cilium adds the eBPF dataplane and policy. kubenet is the older route-table-based option and is being retired.

4.5 Services

Pod IPs churn, so a Service provides a stable virtual IP (ClusterIP) and DNS name that load-balances across a set of pods selected by labels.

The ClusterIP is virtual. It isn't assigned to any interface, and no process listens on it. It exists only as rules matched by the dataplane (kube-proxy or eBPF), which DNAT to a backend pod IP.

Endpoint tracking uses EndpointSlices (which replaced the older Endpoints objects for scale). Only Ready pods are included, which is why readiness probes control traffic.

Service types:

Type What it does
ClusterIP Virtual IP reachable only inside the cluster.
NodePort Opens a port (default 30000–32767) on every node and forwards to the Service.
LoadBalancer Provisions an external LB via the cloud controller (or MetalLB on bare metal). It builds on NodePort.
ExternalName DNS CNAME to an external name. No proxying.
Headless (clusterIP: None) No VIP. DNS returns pod IPs directly. Used with StatefulSets and client-side balancing.

4.6 kube-proxy modes

kube-proxy runs on each node, watches Services and EndpointSlices, and programs the dataplane.

iptables mode (long-time default): Per-Service chains with per-endpoint chains and random-probability load balancing.

iptables -t nat -S KUBE-SERVICES | head
# -A KUBE-SERVICES -d 10.96.0.10/32 -p tcp --dport 53 -j KUBE-SVC-XXXX
# KUBE-SVC-XXXX: -m statistic --mode random --probability 0.33333 -j KUBE-SEP-AAAA ...
# KUBE-SEP-AAAA: -j DNAT --to-destination 10.244.1.5:53

Rule updates are O(n) and traversal is linear, so it gets slow with thousands of Services.

IPVS mode: Uses kernel IPVS hash tables, so it scales better and offers real schedulers (rr, lc, wlc, sh...). It still uses some iptables/ipset rules for masquerading.

nftables mode: The modern replacement, generally available in recent Kubernetes releases, with better update performance than iptables.

eBPF replacement (Cilium, Calico eBPF): Skips kube-proxy entirely, does service lookup in a BPF map at the socket or tc layer, and can even DNAT at connect() time so no per-packet NAT is needed for pod-originated traffic.

Masquerading: KUBE-MARK-MASQ marks packets needing SNAT, for example NodePort traffic or traffic leaving the cluster CIDR, so return traffic flows back through the same node.

4.7 Service traffic walk-through

Client pod → my-svc (ClusterIP 10.96.44.10:80):

  1. Pod resolves my-svc.default.svc.cluster.local via CoreDNS and gets 10.96.44.10.
  2. Packet leaves the pod, hits the node's netfilter (PREROUTING or OUTPUT).
  3. KUBE-SERVICES matches the ClusterIP and picks a backend, so DNAT rewrites the destination to 10.244.2.7:8080.
  4. The packet is routed to the pod (locally, or across the overlay or routed fabric).
  5. Replies are un-DNATed via conntrack, so the client sees them coming from the ClusterIP.

4.8 NodePort, LoadBalancer, and source IP

externalTrafficPolicy:

  • Cluster (default): any node accepts and may forward to a pod on another node. The result is even distribution, but an extra hop and SNAT, which loses the client source IP.
  • Local: only nodes with a local endpoint answer, with no cross-node hop and no SNAT, so the client IP is preserved. The trade-off is potentially uneven balancing, and the cloud LB uses a health check node port to know which nodes have pods.

internalTrafficPolicy: Local does the same for in-cluster clients. Topology-aware routing can prefer same-zone endpoints to reduce cross-AZ cost.

4.9 DNS (CoreDNS)

CoreDNS runs as a Deployment behind the kube-dns Service (typically 10.96.0.10). Kubelet injects into each pod's /etc/resolv.conf:

nameserver 10.96.0.10
search default.svc.cluster.local svc.cluster.local cluster.local
options ndots:5

Records:

  • Service: <svc>.<ns>.svc.cluster.local → ClusterIP (headless → pod IPs)
  • Pod (StatefulSet via headless): <pod>.<svc>.<ns>.svc.cluster.local
  • SRV records for named ports

The ndots:5 trap: Any name with fewer than 5 dots is first tried with each search suffix. So looking up api.example.com generates queries like api.example.com.default.svc.cluster.local → NXDOMAIN, and so on, before the real query. That means extra latency and CoreDNS load. Mitigations: use FQDNs with a trailing dot, set dnsConfig.options ndots: 2, or deploy NodeLocal DNSCache.

Related: the conntrack race on UDP DNS (parallel A and AAAA queries) was a historic source of 5-second DNS timeouts. Node-local caching and newer kernels largely fix this.

dnsPolicy options: ClusterFirst (default), Default (inherit node resolv.conf), ClusterFirstWithHostNet, None (fully custom).

4.10 Ingress and Gateway API

Ingress is an L7 HTTP(S) routing object implemented by an ingress controller (ingress-nginx, Traefik, HAProxy, cloud-native ones such as AGIC). The controller pods are typically exposed via a LoadBalancer Service, and they route by host and path to backend Services (or directly to pod IPs).

Gateway API is the successor: role-oriented resources (GatewayClass, Gateway, HTTPRoute, GRPCRoute, TCPRoute...) with better support for multi-tenancy, traffic splitting, header manipulation, and cross-namespace delegation. Prefer it for new work.

4.11 NetworkPolicy

By default, all pods can talk to all pods. NetworkPolicy restricts this, but only if your CNI enforces it (Flannel alone doesn't).

Rules:

  • Policies are additive allow-lists. There are no deny rules in the core API.
  • A pod is unrestricted until a policy selects it. After that, only allowed traffic passes, in the directions the policy declares.
  • Selectors: podSelector, namespaceSelector, ipBlock.

Namespace default-deny for ingress and egress, plus a DNS egress allowance:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: default-deny, namespace: app }
spec:
  podSelector: {}
  policyTypes: [Ingress, Egress]
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: allow-dns, namespace: app }
spec:
  podSelector: {}
  policyTypes: [Egress]
  egress:
    - to:
        - namespaceSelector:
            matchLabels: { kubernetes.io/metadata.name: kube-system }
      ports:
        - { protocol: UDP, port: 53 }
        - { protocol: TCP, port: 53 }

Calico and Cilium add their own CRDs (global policies, explicit deny, L7 and FQDN rules, and so on).

4.12 hostNetwork, hostPort, and friends

  • hostNetwork: true: The pod uses the node's netns, with no pod IP and no isolation. Used by CNI agents, node exporters, and some ingress setups.
  • hostPort: Maps a node port to a pod port through the CNI portmap plugin. It's discouraged because it constrains scheduling.
  • Dual-stack: Pods and Services can have both IPv4 and IPv6 (ipFamilyPolicy: PreferDualStack).

4.13 Other topics worth knowing

  • Service mesh (Istio, Linkerd, Cilium mesh): sidecars or ambient/eBPF data planes add mTLS, retries, and L7 telemetry on top of the model above.
  • Multi-network pods (Multus): attach additional interfaces (SR-IOV, macvlan) to pods.
  • Egress: Pod traffic leaving the cluster is usually SNATed to the node IP, so use egress gateways or NAT gateways (cloud) when you need stable source IPs.
  • Encryption: WireGuard or IPsec in the CNI, or mTLS via a mesh.

Part 5: Comparison cheat sheet

Concern Docker Podman Kubernetes
Unit Container Container / Pod Pod
Shared netns --network container:x Pod infra container Pause container
Default net docker0 bridge, NAT podman0 (rootful) / pasta (rootless) CNI-provided flat network
Name resolution Embedded DNS (user-defined nets) aardvark-dns (user-defined nets) CoreDNS
Outbound MASQUERADE MASQUERADE / userspace Node SNAT (CNI-dependent)
Inbound -p DNAT -p DNAT / pasta Service, NodePort, LB, Ingress
Multi-host Overlay (Swarm) Not native CNI (overlay/BGP/native)
Policy iptables (DOCKER-USER) Firewall backend NetworkPolicy (+ CNI extensions)
Load balancing DNS RR / Swarm VIP DNS RR Service VIP, kube-proxy/eBPF

Part 6: Debugging playbook

Container and host level

# What's this container's netns/IP/veth?
docker inspect -f '{{json .NetworkSettings.Networks}}' NAME | jq
nsenter -t $(docker inspect -f '{{.State.Pid}}' NAME) -n ip -br a

# Map host veth to container: compare ifindex
ip -br link | grep veth

# Watch traffic on the bridge or a specific veth
tcpdump -ni docker0 port 80
tcpdump -ni vethXXXX

# Firewall and NAT state
iptables -t nat -S; iptables -S FORWARD
nft list ruleset
conntrack -L | grep 172.17.0.2
conntrack -S              # look for insert_failed / drops

# Sockets
ss -tulnp

Kubernetes

# Throwaway debug pod with network tools
kubectl run netshoot --rm -it --image=nicolaka/netshoot -- bash

# Attach to an existing pod without tools (ephemeral container)
kubectl debug -it POD --image=nicolaka/netshoot --target=CONTAINER

# Service wiring
kubectl get svc,endpointslices -o wide
kubectl describe svc my-svc          # empty Endpoints => selector/readiness problem

# DNS
kubectl exec -it POD -- cat /etc/resolv.conf
kubectl exec -it POD -- nslookup kubernetes.default
kubectl -n kube-system logs -l k8s-app=kube-dns

# On the node
iptables-save | grep my-svc
ipvsadm -Ln                          # IPVS mode
cilium status; cilium monitor; hubble observe   # Cilium

Common failure patterns

Symptom Likely cause
Service has no endpoints Label selector mismatch, or pods failing readiness
Connects via pod IP but not Service IP kube-proxy or dataplane issue, or NetworkPolicy
Intermittent timeouts on large payloads MTU mismatch (overlay overhead, VPN, cloud). Test with ping -M do -s 1400.
Slow or failing external DNS lookups ndots:5 amplification, CoreDNS overloaded, UDP conntrack races
Pod can't reach the internet Missing SNAT/masquerade, egress NetworkPolicy, or no NAT gateway
Docker port open despite UFW deny Docker DNAT bypasses INPUT. Bind to 127.0.0.1 or use DOCKER-USER.
Container can't reach the host's own service Bridge gateway IP vs localhost confusion. Use host.docker.internal / host-gateway mapping.
Docker network overlaps corporate VPN Change default-address-pools in daemon.json
Rootless Podman can't bind port 80 net.ipv4.ip_unprivileged_port_start
NodePort works on some nodes only externalTrafficPolicy: Local with no local endpoints
Pods stuck ContainerCreating with CNI errors CNI binary/config missing, IPAM exhausted, or the CNI agent is down

Part 7: Mental models to keep

  1. A container's network is just a netns plus a veth. Everything else is plumbing on the host side.
  2. Docker and Podman, by default, hide containers behind NAT. Kubernetes forbids that inside the cluster and makes NAT the exception.
  3. A Service IP is a rule, not a device. If no rule matches (empty endpoints, dataplane broken), the packet goes nowhere.
  4. The CNI decides how packets cross nodes. Overlay vs routed vs native is the biggest performance, debuggability, and IP-planning decision.
  5. Follow the packet. At each hop, ask what the source and destination are now, what rewrote them (SNAT/DNAT), and where conntrack will send the reply.