Load Balancing: L4 vs L7, Algorithms, Health Checks
What L4 and L7 balancers actually forward, why least-request beats round robin, how health checks and draining make deploys safe, and why gRPC breaks it.
A load balancer sits behind one address and forwards each unit of work to one of many backends. What that unit is decides almost everything: a layer 4 balancer forwards TCP (Transmission Control Protocol) connections and never looks inside them, a layer 7 balancer terminates the connection, reads each HTTP request and picks a backend per request. On top of that sit an algorithm that chooses the backend (least outstanding requests usually beats round robin) and health checks that stop sending traffic to instances that are dead, starting up or shutting down.
Context
The first web-scale balancing was DNS (Domain Name System) round robin in the mid-1990s: publish several IP addresses for one name and let resolvers rotate them. It spreads load but cannot see a dead server, so hardware appliances followed (Cisco LocalDirector, F5 BIG-IP in 1997), then software: LVS (Linux Virtual Server, 1998) in the kernel, HAProxy (2001) and Nginx (2004) as user-space proxies. Cloud made it a checkbox: AWS ELB (Elastic Load Balancing) in 2009, split in 2016-2017 into ALB (Application Load Balancer, layer 7) and NLB (Network Load Balancer, layer 4). Google published Maglev, its software layer 4 balancer, in 2016, and Lyft open-sourced Envoy the same year, which became the proxy inside Istio and most service meshes.
You have met one every time you wrote an Nginx upstream block, created a Kubernetes Service (kube-proxy balances connections to its pods), or pointed a domain at an ALB target group. The smallest working example is four lines:
upstream api { # the pool
server 10.0.1.11:8080;
server 10.0.1.12:8080;
}
server { location / { proxy_pass http://api; } } # round robin by default- Backend / upstream / target
- One instance that can serve the traffic. Different products, same idea: Nginx says upstream server, AWS says target, Envoy says host.
- VIP (virtual IP)
- The single address clients connect to. It belongs to the balancer, not to any backend.
- L4 / L7
- Layer 4 (transport: TCP, UDP) and layer 7 (application: HTTP, gRPC) of the OSI (Open Systems Interconnection) model. Shorthand for what the balancer can see.
- Health check
- A probe (active) or an observation of real traffic (passive) that decides whether a backend receives new work.
- Draining
- Stopping new work to a backend while letting in-flight requests and connections finish before it is removed.
Why it matters
Every horizontally scaled service has one, and most of its failure modes are balancing failures in disguise. A deploy that drops a few hundred requests is usually a missing drain. One pod at 100% CPU while its siblings idle is round robin meeting uneven requests, or an L4 balancer meeting long-lived HTTP/2 connections. A whole fleet "going down" because the database blipped is a health check that checked too much. And every retry the balancer makes on your behalf is a decision about idempotency you may not know you took.
Layer 4 vs layer 7: what gets balanced
An L4 balancer picks a backend when a connection opens (the first SYN packet for TCP) and then forwards every packet of that connection to the same backend, rewriting the destination address with NAT (network address translation) or, in DSR (direct server return) setups like LVS-DR and Maglev, letting the backend reply to the client directly. It never decrypts anything, so it is fast and protocol agnostic, but its unit of balancing is the connection.
An L7 balancer is a reverse proxy. It completes the TLS (Transport Layer Security) handshake itself, parses HTTP, and opens its own, usually pooled, connections to the backends. Because it sees each request, it can pick a backend per request, route on path, host or header, retry, add X-Forwarded-For, and count real request latency. The cost is CPU for TLS and parsing, and that it must understand the protocol.
| Layer 4 | Layer 7 | |
|---|---|---|
| Unit | Connection (or UDP flow) | Request |
| Sees | IPs, ports, protocol | Host, path, headers, cookies, status codes |
| TLS | Passed through untouched | Terminated (and often re-encrypted to the backend) |
| Can do | Millions of connections, any protocol, static IPs | Path routing, retries, auth, header rewrite, per-request metrics |
| Client IP | Preserved with DSR or PROXY protocol | Passed in X-Forwarded-For / Forwarded |
| Examples | AWS NLB, Maglev, LVS, kube-proxy | AWS ALB, Nginx, HAProxy, Envoy, Cloudflare |
Choosing a backend
Round robin is the default in Nginx, Envoy and ALB because it is stateless and perfectly fair when every request costs the same. Real requests do not: one endpoint returns a cached row in 2 ms, another builds a report for 4 seconds. Round robin keeps handing work to the instance that is already stuck on reports, because it counts requests sent, not requests still running.
Picking the global minimum has its own problem. With ten balancer instances, each with a slightly stale view, all ten see the same "least loaded" backend and pile onto it at once. The fix, from Mitzenmacher's 2001 "power of two choices" result, is P2C (power of two choices): pick two backends at random and send to the less loaded of the two. It needs no global scan, avoids the herd, and in the balls-into-bins model cuts the worst backend's excess load exponentially compared with one random choice. Envoy's LEAST_REQUEST and Nginx's random two least_conn (since 1.15.1) both do exactly this.
type Backend = {url: string; inFlight: number; healthy: boolean}
export function pick(pool: Backend[]): Backend {
const up = pool.filter(b => b.healthy)
if (up.length === 0) throw new Error('no healthy backends')
const a = up[Math.floor(Math.random() * up.length)]
const b = up[Math.floor(Math.random() * up.length)]
return a.inFlight <= b.inFlight ? a : b // the less busy of two
}
export async function forward(pool: Backend[], req: Request) {
const target = pick(pool)
target.inFlight++ // count what is running,
try { // not what was sent
const {pathname, search} = new URL(req.url)
return await fetch(target.url + pathname + search, req)
} finally {
target.inFlight--
}
}| Algorithm | Picks | Good for | Breaks when |
|---|---|---|---|
| Round robin | Next in the list | Uniform, short requests | Request cost varies; instances differ in size |
| Weighted round robin | Next, proportional to weight | Mixed instance sizes, canary at 5% | Weights drift from real capacity |
| Least conn. | Fewest open connections | Long-lived connections (WebSockets, databases) | Connections are pooled, so the count says little |
| Least request / P2C | Less busy of two random picks | The default to reach for with HTTP and gRPC | Only a few backends, where random pairs repeat |
| Latency-aware | Lowest EWMA of response time | Heterogeneous or noisy-neighbour fleets | A fast-failing backend looks attractive |
| Hash | hash(IP, header or key) | Cache locality, session affinity | Hot keys; plain mod N remaps on scale-out |
Health checks, draining and deploys
A balancer is only as good as its list of healthy backends. Active checks probe each backend on a timer (ALB defaults to every 30 seconds and marks a target unhealthy after 2 failures, healthy after 5 successes). Passive checks watch real traffic: open-source Nginx's max_fails/fail_timeout and Envoy's outlier detection eject a backend after consecutive errors. Active checks catch a backend before users hit it; passive checks react within a few requests instead of a few intervals. Production setups use both.
upstream api {
random two least_conn; # P2C
server 10.0.1.11:8080 max_fails=3 fail_timeout=10s;
server 10.0.1.12:8080 max_fails=3 fail_timeout=10s;
server 10.0.1.13:8080 weight=2; # bigger instance
keepalive 32; # pooled upstream conns
}
server {
listen 443 ssl;
location / {
proxy_pass http://api;
proxy_http_version 1.1;
proxy_set_header Connection ""; # needed for keepalive
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_next_upstream error timeout; # retry on another backend
proxy_next_upstream_tries 2; # (not POST, since 1.9.13)
}
}Health checks are also how deploys avoid dropping requests. The balancer must learn that an instance is leaving before the instance stops accepting work, and must not send to a new instance before it is ready. The sequence for one instance in a rolling deploy:
- 1The orchestrator marks the old instance as leaving: Kubernetes removes the pod from the Service endpoints, AWS moves the target to
draining. The balancer stops sending new requests to it. - 2In-flight requests keep running for the drain window: ALB's deregistration delay defaults to 300 seconds. Meanwhile the process gets SIGTERM and should stop accepting new connections but finish the ones it has.
- 3Because endpoint removal propagates asynchronously, a well-behaved server keeps serving for a few seconds after SIGTERM (a
preStopsleep in Kubernetes) so that balancers with a stale list do not hit a closed port. - 4The new instance starts, and receives traffic only once its readiness check passes, meaning it has loaded config, warmed caches and opened its database pool, not merely that the port is open.
- 5Optionally, slow start (ALB, Envoy, HAProxy) ramps its share up over tens of seconds so a cold instance is not handed a full share immediately, which least-request would otherwise do because its queue is empty.
Layers of balancing, stickiness and long-lived connections
"The load balancer" in a large system is a stack of them, each spreading load across the next tier and each redundant in its own way. The top tier is not a box at all: DNS returns several addresses, and anycast plus ECMP (equal-cost multi-path routing) lets routers spread packets for one IP over many identical L4 balancers. That is the real answer to "is the balancer a single point of failure?"; the smaller answer is an active-passive pair sharing a VIP through VRRP (Virtual Router Redundancy Protocol), for example with keepalived.
Long-lived connections defeat L4 balancing
gRPC runs over HTTP/2, which multiplexes every call onto one long-lived connection. Behind an L4 balancer, including a plain Kubernetes ClusterIP Service, each client opens one connection, lands on one pod, and sends it everything for hours; pods added by the autoscaler get nothing. The fixes all move balancing to the request level: an L7 proxy or mesh sidecar (Envoy, Linkerd, Istio) that balances individual streams, client-side balancing over a headless Service, or at least MAX_CONNECTION_AGE on the server so clients reconnect and re-spread periodically. WebSockets have the same shape and usually settle for least-connections plus periodic reconnects.
Sticky sessions
Session affinity pins a client to one backend, by a balancer-issued cookie (ALB AWSALB, HAProxy cookie) or by hashing the client IP. It exists to rescue apps that keep session state in process memory. The price is uneven load, a lost session whenever that backend dies or is deployed, and IP hashing that sends a whole corporate NAT to one instance. Move the state out, into a shared store or a signed token such as a JWT (JSON Web Token), and keep affinity only as a cache-locality optimisation, never as correctness.
Pitfalls
- A deep health check that takes the whole fleet down
If
/healthalso checks the database, one database blip marks every instance unhealthy at once, and the balancer has nowhere to send traffic. Some balancers fail open when everything is unhealthy (ALB routes to all targets; Envoy's panic threshold ignores health below 50% healthy), others return 503 for everything. Health should mean "this instance can serve", not "every dependency is up". - No drain on deploy
Killing a process the moment the orchestrator says so resets its in-flight requests and races the balancer's endpoint update, so a few requests per instance land on a closed port. Handle SIGTERM by finishing in-flight work, keep serving for a few seconds after it, and set a drain window longer than your slowest request.
- Retries that multiply load or duplicate writes
A balancer retrying on another backend is great for a crashed instance and terrible during overload: each layer that retries multiplies traffic on an already struggling fleet. And retrying a POST that timed out may run it twice. Limit retry counts and budgets, and retry non-idempotent requests only with idempotency keys.
- gRPC or HTTP/2 behind an L4 balancer
Balancing happens once per connection, and the connection lives for hours, so a few pods take all the traffic and new pods take none. CPU graphs show it immediately; the fix is L7 or client-side balancing, not more replicas.
- Trusting X-Forwarded-For from anyone
Behind an L7 proxy every request comes from the proxy's IP, so apps read the client IP from
X-Forwarded-For. A client can send that header too. Take the address the trusted proxy appended (the rightmost untrusted hop), or rate limits and IP allow-lists become spoofable.
Interview questions
Q1What is the difference between L4 and L7 load balancing, and when do you pick each?
L4 forwards connections without reading them; L7 terminates the connection, reads each HTTP request and picks a backend per request. Pick L7 for HTTP services where you want path routing, retries, TLS termination and per-request balancing. Pick L4 for non-HTTP protocols, for pass-through TLS, for static IPs, or as the high-throughput tier in front of a fleet of L7 proxies, which is how large deployments combine them.
Q2Which balancing algorithm would you choose for an HTTP API?
Least outstanding requests with power of two choices. Round robin assumes every request costs the same, so a backend stuck on slow requests keeps receiving its share and tail latency climbs. Tracking in-flight requests adapts to real cost, and choosing the better of two random backends avoids every balancer instance herding onto the same "least loaded" backend. Round robin is fine only when requests are uniform.
Q3Walk me through a zero-downtime deploy behind a load balancer.
For each instance: deregister it so it gets no new requests, let in-flight requests finish during a drain window longer than the slowest request, and have the process handle SIGTERM by finishing work rather than exiting, staying up a few seconds longer because endpoint removal propagates asynchronously. Start the new instance and add it only when a readiness check says it can actually serve, ideally with slow start. Keep enough capacity that the fleet absorbs one instance being out, and roll back on error-rate alarms.
Q4What happens when every backend fails its health check?
It depends on the balancer, and you should know which yours does. ALB fails open and routes to all targets; Envoy enters panic mode below 50% healthy and balances across all hosts; Nginx with passive checks returns errors until fail_timeout expires. Usually the root cause is a health check that depends on a shared dependency, so the fix is a shallower check, not a different balancer.
Q5Our gRPC service has five pods but one of them does all the work. Why?
gRPC multiplexes calls over one long-lived HTTP/2 connection, and a Kubernetes ClusterIP Service balances at L4, once per connection. Each client connected once and has been pinned since. Balance per request instead: an L7 proxy or mesh sidecar, client-side balancing over a headless Service, or at minimum a maximum connection age on the server so clients reconnect and spread out.
Q6Isn’t the load balancer a single point of failure?
Only if you run one. Managed balancers are already distributed across zones. Self-hosted ones run as an active-passive pair sharing a virtual IP via VRRP, or as many active instances behind DNS or anycast with ECMP, where routers spread packets for one IP across all of them. At scale the tiers are layered (DNS, L4, L7), each redundant on its own.
Q7When would you use sticky sessions?
As a performance hint, not for correctness. If losing affinity breaks users, the state belongs in a shared store or a signed token, because the pinned backend will eventually die or be redeployed. Affinity is reasonable for cache locality or for WebSocket reconnects to the same node, preferably via consistent hashing on a user or tenant key rather than a client IP, which lumps everyone behind one NAT together.
Q8How does a backend behind a load balancer learn the client’s IP?
Behind an L7 proxy, from the X-Forwarded-For or standard Forwarded header the proxy adds; behind an L4 balancer, from the PROXY protocol header or because DSR preserves the source address. Only trust the value your own proxy wrote: take the rightmost address that is not one of your trusted hops, since clients can send any X-Forwarded-For they like.
- L4 balances connections without reading them; L7 terminates them and balances requests. Large systems stack both behind DNS or anycast.
- Round robin counts requests sent; least-request counts requests running. Power of two choices gets least-request without herding.
- Active checks catch bad backends early, passive checks catch them fast. Keep health shallow, or one dependency blip ejects the whole fleet.
- Zero-downtime deploys need draining, SIGTERM handling, a short post-SIGTERM grace period and readiness gates on the new instance.
- HTTP/2 and gRPC pin all traffic to one connection, so L4 balancing leaves new pods idle; balance per request.
- Sticky sessions and balancer retries are optimisations with costs; never rely on them for correctness.