Andrew Mercer
on this page

IPVS (IP Virtual Server) is the load balancer built into the Linux kernel — the same technology behind the Linux Virtual Server project, and the thing kube-proxy is quietly running when you set its mode to ipvs. Unlike HAProxy or nginx, it operates entirely in kernel space at Layer 4, which is why it can push millions of concurrent connections without breaking a sweat. The tradeoff is that it doesn't know anything about HTTP — no SSL termination, no header inspection, no Layer 7 routing. It just moves packets, very fast.

This guide walks through how it works, how to set it up, and the stuff that actually trips people up in production (mostly ARP, if you're using Direct Routing).

How it fits together

A basic IPVS setup has three moving parts:

  • The director — the box running IPVS, which owns the virtual IP and decides where each new connection goes.
  • The virtual IP (VIP) — the address clients actually connect to.
  • Real servers — the backends doing the actual work.

The director keeps a connection table so that once a client is assigned to a backend, subsequent packets in that connection keep going to the same place. Where things diverge is in how return traffic gets back to the client — that's the forwarding method, covered below.

Installing it

# Debian/Ubuntu
sudo apt-get update && sudo apt-get install ipvsadm

# CentOS/RHEL
sudo yum install ipvsadm

# Arch
sudo pacman -S ipvsadm

Then load the kernel module and whichever schedulers you plan on using:

sudo modprobe ip_vs
lsmod | grep ip_vs

# Schedulers — load what you need, not all of them
sudo modprobe ip_vs_rr    # round-robin
sudo modprobe ip_vs_wrr   # weighted round-robin
sudo modprobe ip_vs_lc    # least-connection
sudo modprobe ip_vs_wlc   # weighted least-connection
sudo modprobe ip_vs_sh    # source hashing
sudo modprobe ip_vs_dh    # destination hashing
sudo modprobe ip_vs_sed   # shortest expected delay
sudo modprobe ip_vs_nq    # never queue

echo "ip_vs" | sudo tee -a /etc/modules

IP forwarding needs to be on for the director to actually route traffic:

echo "net.ipv4.ip_forward = 1" | sudo tee -a /etc/sysctl.conf
sudo sysctl -p

If you're expecting real traffic volumes, bump the conntrack table now rather than after you get paged for it:

sudo tee -a /etc/sysctl.conf << 'EOF'
net.netfilter.nf_conntrack_max = 1048576
net.netfilter.nf_conntrack_tcp_timeout_established = 7200
net.ipv4.vs.conn_reuse_mode = 1
EOF
sudo sysctl -p

Choosing a scheduling algorithm

There are eight schedulers, but in practice you'll use two or three of them almost exclusively.

Round-robin (rr) is the default mental model — even distribution, no state. Fine when every backend is identical and requests are short-lived.

Weighted least-connection (wlc) is what most production setups actually run. It sends new connections to whichever backend has the fewest active connections, adjusted by weight — so a bigger box configured with -w 3 gets roughly three times the traffic of one at -w 1. This is the right default for long-lived connections and mixed-capacity fleets.

Source hashing (sh) routes by client IP, which gives you session affinity without needing cookies or an application-layer proxy. Use it when the backend expects the same client to keep hitting the same server — in-memory session state, local caches, that kind of thing.

The rest — dh, sed, nq — solve narrower problems (destination-based cache routing, minimizing expected delay, avoiding queueing under low latency requirements) and are worth reaching for only if you've measured a specific problem wlc doesn't solve.

# Set the scheduler when creating the service
ipvsadm -A -t 10.0.0.100:80 -s wlc

# Change it later without touching the backend list
ipvsadm -E -t 10.0.0.100:80 -s sh

Forwarding methods: the part that actually matters

This is where most of the real design decisions live, because it determines your network topology, not just your traffic distribution.

NAT

The director rewrites both the destination and source IP on every packet, in both directions. Everything — request and response — flows through it.

ipvsadm -A -t 10.0.0.100:80 -s wlc
ipvsadm -a -t 10.0.0.100:80 -r 192.168.1.10:80 -m
ipvsadm -a -t 10.0.0.100:80 -r 192.168.1.11:80 -m

ip addr add 10.0.0.100/24 dev eth0
iptables -t nat -A POSTROUTING -s 192.168.1.0/24 -j MASQUERADE

Real servers just need their default gateway pointed at the director. Simple to set up, but every byte of response traffic also passes through the director, which caps your throughput at whatever that one box can push. Fine for smaller setups; not what you want at scale.

Direct Routing (DR)

The director only rewrites the destination MAC address and hands the packet off — the real server responds to the client directly, bypassing the director entirely on the way back. This asymmetry is what makes DR the standard choice for anything serious: the director only ever handles inbound requests, which is a fraction of typical HTTP traffic.

The catch is that real servers need the VIP configured on their loopback interface (so they'll accept packets addressed to it), and you have to suppress ARP replies for that address or every real server will start answering ARP requests for the VIP and chaos ensues.

# On the director
ipvsadm -A -t 10.0.0.100:80 -s wlc
ipvsadm -a -t 10.0.0.100:80 -r 192.168.1.10:80 -g -w 10
ipvsadm -a -t 10.0.0.100:80 -r 192.168.1.11:80 -g -w 10
ip addr add 10.0.0.100/32 dev eth0
# On each real server
ip addr add 10.0.0.100/32 dev lo

echo 1 > /proc/sys/net/ipv4/conf/lo/arp_ignore
echo 2 > /proc/sys/net/ipv4/conf/lo/arp_announce
echo 1 > /proc/sys/net/ipv4/conf/all/arp_ignore
echo 2 > /proc/sys/net/ipv4/conf/all/arp_announce

The constraint: director and real servers need to be on the same L2 segment, since DR works by rewriting MAC addresses rather than IP addresses.

IP Tunneling (TUN)

Same asymmetric idea as DR — director handles inbound, real server replies directly — but the director encapsulates packets in an IP-in-IP tunnel instead of rewriting MACs. That means real servers don't need to be on the same LAN, or even the same country. This is how you'd build a load balancer with backends split across regions.

# Real server side
ip tunnel add tunl0 mode ipip
ip link set tunl0 up
ip addr add 10.0.0.100/32 dev tunl0
echo 1 > /proc/sys/net/ipv4/conf/tunl0/arp_ignore
echo 2 > /proc/sys/net/ipv4/conf/tunl0/arp_announce

More moving parts, and there's tunnel encapsulation overhead, but it's the only one of the three that supports geographically distributed backends.

Which one to pick

Throughput ceiling Backend location Setup effort
NAT Limited by director Anywhere behind it Low
DR Very high Same L2 segment Medium (ARP config)
TUN High Anywhere Higher (tunnel config)

Default to DR unless you have a specific reason not to. Reach for NAT only for small internal setups where simplicity beats throughput. Reach for TUN when your backends genuinely need to be geographically spread out.

A working example end to end

#!/bin/bash
# DR-mode web load balancer, three backends, weighted

ipvsadm -C
ipvsadm -A -t 10.0.0.100:80 -s wlc
ipvsadm -a -t 10.0.0.100:80 -r 192.168.1.10:80 -g -w 10
ipvsadm -a -t 10.0.0.100:80 -r 192.168.1.11:80 -g -w 10
ipvsadm -a -t 10.0.0.100:80 -r 192.168.1.12:80 -g -w 5   # smaller box, half the weight

ip addr add 10.0.0.100/32 dev eth0
ipvsadm-save > /etc/ipvsadm.rules

For session-sticky workloads, swap the scheduler and add a persistence timeout so a client keeps landing on the same backend for the duration you specify:

ipvsadm -A -t 10.0.0.100:80 -s sh -p 360   # 6-minute stickiness

For UDP services like DNS, the syntax is nearly identical — just swap -t for -u:

ipvsadm -A -u 10.0.0.100:53 -s lc
ipvsadm -a -u 10.0.0.100:53 -r 192.168.1.10:53 -m

Firewall marks: grouping services under one policy

If you want HTTP and HTTPS to share the same backend pool and scheduling decision (so a client sticks to the same backend across both), tag them with the same mark in iptables first, then build the IPVS service against the mark instead of a port:

iptables -t mangle -A PREROUTING -p tcp --dport 80 -j MARK --set-mark 1
iptables -t mangle -A PREROUTING -p tcp --dport 443 -j MARK --set-mark 1

ipvsadm -A -f 1 -s wlc
ipvsadm -a -f 1 -r 192.168.1.10:0 -g
ipvsadm -a -f 1 -r 192.168.1.11:0 -g

High availability with keepalived

IPVS itself doesn't handle director failover — that's keepalived's job, using VRRP to move the VIP between two directors and health-checking backends via its own config file rather than shell scripts:

vrrp_instance VI_1 {
    state MASTER
    interface eth0
    virtual_router_id 51
    priority 100
    advert_int 1
    authentication {
        auth_type PASS
        auth_pass secret
    }
    virtual_ipaddress {
        10.0.0.100/24
    }
}

virtual_server 10.0.0.100 80 {
    delay_loop 6
    lb_algo wlc
    lb_kind DR
    protocol TCP

    real_server 192.168.1.10 80 {
        weight 1
        TCP_CHECK {
            connect_timeout 3
            connect_port 80
        }
    }
    real_server 192.168.1.11 80 {
        weight 1
        TCP_CHECK {
            connect_timeout 3
            connect_port 80
        }
    }
}

This is the production-realistic setup — keepalived owns health checking and failover, IPVS just does the forwarding. For connection-state sync between the active and standby director (so failover doesn't drop existing connections), there's also ipvsadm --start-daemon, but keepalived config is the more maintainable path for most teams.

Troubleshooting

Checking current state:

ipvsadm -Ln            # services and backends
ipvsadm -Ln --stats    # with traffic counters
ipvsadm -Ln -c         # connection table

A real server isn't getting traffic. First check whether IPVS even considers it active — a weight of 0 means it's configured but excluded from scheduling:

ipvsadm -Ln | grep 192.168.1.10
ipvsadm -e -t 10.0.0.100:80 -r 192.168.1.10:80 -g -w 1   # re-enable if weight was 0

If the weight looks fine, check basic reachability and that the forwarding flag matches your topology (-m for NAT, -g for DR, -i for tunneling — mismatching this is a common copy-paste mistake).

ARP problems in DR mode are the single most common source of "it's configured right but traffic isn't reaching the backend." Confirm the VIP is actually bound to loopback and that ARP is suppressed:

ip addr show lo | grep 10.0.0.100
sysctl -a | grep arp_ignore
arping -I eth0 10.0.0.100   # should get no reply from real servers, only the director

Connections timing out or the conntrack table filling up:

dmesg | grep conntrack
sysctl -w net.netfilter.nf_conntrack_max=1048576
sysctl -w net.ipv4.vs.conn_reuse_mode=1

Kernel-level debug logging, if you need to see scheduling decisions in real time:

echo 2 > /proc/sys/net/ipv4/vs/debug_level   # 0 off, 1 basic, 2 detailed, 3 verbose
dmesg -w

Real-world usage

kube-proxy in IPVS mode. This is probably the IPVS deployment you're most likely to already be running without realizing it. Enabling it just means loading the modules cluster-wide and flipping the mode in the kube-proxy configmap:

cat > /etc/modules-load.d/ipvs.conf << 'EOF'
ip_vs
ip_vs_rr
ip_vs_wrr
ip_vs_sh
nf_conntrack
EOF
modprobe ip_vs ip_vs_rr ip_vs_wrr ip_vs_sh nf_conntrack
apiVersion: v1
kind: ConfigMap
metadata:
  name: kube-proxy
data:
  config.conf: |-
    mode: "ipvs"
    ipvs:
      scheduler: "wlc"
      syncPeriod: 30s
kubectl delete pod -n kube-system -l k8s-app=kube-proxy   # restart to pick up the change
ipvsadm -Ln   # confirm it's actually doing IPVS now

Worth doing on clusters with a lot of Services — IPVS uses hash tables for lookup instead of iterating iptables rules, so it doesn't degrade linearly as your Service count grows the way iptables mode does.

Database read-replica balancing. Long-lived connections make lc a better fit than wlc here, paired with a persistence timeout long enough to outlast typical connection pooling behavior:

ipvsadm -A -t 10.0.0.100:5432 -s lc -p 3600
ipvsadm -a -t 10.0.0.100:5432 -r 192.168.1.10:5432 -g -w 1
ipvsadm -a -t 10.0.0.100:5432 -r 192.168.1.11:5432 -g -w 1
sysctl -w net.ipv4.vs.timeout_established=28800   # 8 hours

Geo-distributed backends via tunneling:

ipvsadm -A -t 10.0.0.100:80 -s wlc
ipvsadm -a -t 10.0.0.100:80 -r 203.0.113.10:80 -i -w 10   # EU
ipvsadm -a -t 10.0.0.100:80 -r 198.51.100.20:80 -i -w 10  # Asia
ipvsadm -a -t 10.0.0.100:80 -r 192.0.2.30:80 -i -w 10     # US

Legacy short-flag syntax

Older references and scripts often use ipvsadm's short single-letter flags rather than the long-form ones used above — functionally identical, just terser:

ipvsadm -C                                              # clear all rules (same as --clear)
ipvsadm -A -t 10.0.0.100:80 -s wlc                       # add service (same as --add-service)
ipvsadm -a -t 10.0.0.100:80 -r 192.168.1.10:80 -m        # add real server, masquerade (same as --add-server --masquerading)
ipvsadm -l                                               # list (same as --list)
ipvsadm -ln                                              # list, numeric (no DNS resolution)

Bottom line

For a straightforward internal load balancer, NAT mode gets you running fastest. For anything production-facing where throughput matters, DR is the default — the ARP setup is a one-time annoyance, not an ongoing cost. Reach for tunneling only when your backends are actually distributed across networks that DR's same-LAN requirement rules out. And if you're already running Kubernetes at any scale, you may already have IPVS underneath you — worth checking before reaching for anything else.