HTTP/2 and HTTP/3: Multiplexing, QUIC and HOL Blocking
What HTTP/2 multiplexing fixed in HTTP/1.1, why one lost TCP packet still stalls every stream, and how HTTP/3 over QUIC moves the fix into UDP.
HTTP/1.1 sends one request at a time per connection, so browsers open about six connections per host and still queue the rest. HTTP/2 splits every request into frames and interleaves many streams on one TCP (Transmission Control Protocol) connection, but TCP delivers bytes strictly in order, so a single lost packet stalls every stream until it is retransmitted. HTTP/3 keeps the same model and runs it over QUIC on UDP (User Datagram Protocol), which recovers each stream independently and folds the TLS (Transport Layer Security) handshake into connection setup.
Context
HTTP/1.1 (1997, today RFC 9112) served the web for almost two decades, and the web worked around its limits with hacks: domain sharding across static1/static2 hostnames to get more connections, CSS sprites and concatenated bundles to make fewer requests, inlined images to make none. In 2009 Google started SPDY, an experimental binary protocol in Chrome, which the IETF (Internet Engineering Task Force) turned into HTTP/2, published as RFC 7540 in May 2015 and revised as RFC 9113 in 2022. Google then moved the transport itself into user space with QUIC, standardised as RFC 9000 in 2021; the IETF is explicit that QUIC is a name, not an acronym. HTTP/3 on top of it is RFC 9114 (2022).
You use all three every day without choosing: the browser and server agree on the version, and DevTools shows it in the Network panel's Protocol column as http/1.1, h2 or h3. From a terminal:
curl -sI --http2 https://www.cloudflare.com | head -1
# HTTP/2 200
# needs a curl built with HTTP/3 support
curl -sI --http3 https://www.cloudflare.com | head -1
# HTTP/3 200- Stream
- One request and its response inside a connection. HTTP/2 and HTTP/3 carry many at once, each with its own ID.
- Frame
- The unit HTTP/2 and HTTP/3 actually send: a small header naming the stream, then a piece of headers or body.
- HOL (head-of-line) blocking
- Later work waiting behind earlier work that is stuck, either a slow response (HTTP/1.1) or a lost packet (TCP).
- RTT (round-trip time)
- The time for a packet to reach the server and an answer to come back. Connection setup is counted in RTTs.
- ALPN (Application-Layer Protocol Negotiation)
- The TLS extension in which client and server agree on h2 or http/1.1 during the handshake.
Why it matters
The protocol version changes which performance advice is still true. Domain sharding now hurts, because it splits one multiplexed connection into several that each pay for a handshake and a slow start. Bundling everything into one file matters less, while caching granularity matters more. On mobile networks with packet loss, the transport choice decides whether one dropped packet freezes a page's images, scripts and API calls together. And outside the browser, gRPC is built on HTTP/2, so its long-lived multiplexed connections are why it needs request-level load balancing.
HTTP/2: one connection, many streams
HTTP/1.1 is a text protocol in which a connection carries one response at a time: the client sends a request and must read the whole response before the next one on that connection. A slow response blocks everything queued behind it, which is HTTP-level head-of-line blocking. Browsers cope by opening about six connections per origin, and anything beyond that waits.
HTTP/2 keeps the same methods, headers and status codes, so applications did not change, but replaces the wire format. Every message is cut into binary frames tagged with a stream ID, and frames of different streams are interleaved on one TCP connection. A slow response no longer blocks others: the server sends frames of whatever is ready. Headers are compressed with HPACK (RFC 7541), which keeps a table of headers already sent on the connection, so repeated cookies and user agents shrink to a few bytes. Browsers only speak HTTP/2 over TLS, selected through ALPN as h2.
| HTTP/1.1 | HTTP/2 | HTTP/3 | |
|---|---|---|---|
| Transport | TCP | TCP | QUIC over UDP |
| Concurrency | One request at a time per connection; ~6 connections | Many streams on one connection | Many streams on one connection |
| Headers | Plain text, repeated every time | HPACK compression | QPACK compression |
| HOL blocking | Per connection, at the HTTP level | Gone at HTTP level; a lost TCP packet stalls all streams | Only the stream that lost a packet waits |
| New connection | TCP + TLS: 2-3 RTT | TCP + TLS: 2-3 RTT | 1 RTT, or 0 RTT on resumption |
| Encryption | Optional | Required by browsers | Always, TLS 1.3 built in |
What did not survive
Two HTTP/2 features were walked back. Server push, which let the server send resources before the browser asked, rarely helped in practice because the server could not know what the browser already had cached; Chrome disabled it by default in version 106 (2022), and preload hints plus 103 Early Hints replaced it. The original priority tree was implemented inconsistently and was deprecated in RFC 9113; RFC 9218 (2022) replaced it with a simple priority header carrying an urgency and an incremental flag.
The head-of-line blocking HTTP/2 could not fix
TCP promises to deliver a byte stream in order. If one packet is lost, the receiving operating system holds every later packet in its buffer until the retransmission arrives, even packets that belong to completely different HTTP/2 streams. TCP has no idea that streams exist. With HTTP/1.1 and six connections, a lost packet stalled one connection; with HTTP/2, everything shares one connection, so the same loss stalls the whole page. On a clean network multiplexing wins easily; on a lossy mobile link HTTP/2 can do worse than six HTTP/1.1 connections.
HTTP/3 and QUIC
QUIC is a transport protocol that provides what HTTP/2 needed from TCP plus TLS, implemented on UDP: reliable delivery, congestion control and encryption, but with streams as a first-class concept, so loss recovery and ordering are per stream. HTTP/3 maps HTTP requests onto QUIC streams and uses QPACK (RFC 9204) instead of HPACK, because HPACK assumed the in-order delivery QUIC deliberately gives up.
Two more properties come from owning the transport. The TLS 1.3 handshake is part of QUIC's own handshake, so a new connection is ready after one round trip instead of TCP's one plus TLS's one or two. And connections are identified by a connection ID rather than by the IP address and port pair, so a phone that moves from Wi-Fi to cellular keeps its connection instead of starting over.
How a browser finds HTTP/3
A browser cannot know in advance that a server speaks QUIC on UDP port 443, so the first visit usually happens over HTTP/2, and the server advertises HTTP/3 with an Alt-Svc header (RFC 7838). Later requests try QUIC and fall back to TCP if UDP is blocked. Newer clients can learn it before the first request from an HTTPS DNS (Domain Name System) record (RFC 9460, 2023), which also carries the supported protocols.
# HTTP/3 in Nginx: experimental since 1.25.0 (2023)
server {
listen 443 quic reuseport; # UDP, HTTP/3
listen 443 ssl; # TCP, HTTP/2 + 1.1 fallback
http2 on;
ssl_protocols TLSv1.3; # QUIC requires TLS 1.3
# tell browsers HTTP/3 is available on this port for a day
add_header Alt-Svc 'h3=":443"; ma=86400';
}- 1The first request goes over TCP, negotiates
h2through ALPN, and the response carriesAlt-Svc: h3=":443". - 2The browser caches that hint for
maseconds and, for the next requests, opens a QUIC connection to the same host on UDP 443. - 3If the QUIC handshake succeeds, new requests use HTTP/3; if UDP is filtered, the browser quietly keeps using HTTP/2 and may retry QUIC later.
- 4On a later visit with a stored session ticket, the client may send a safe request in 0-RTT data with its first packet, before the handshake completes.
Pitfalls
- Keeping HTTP/1.1 workarounds
Domain sharding forces several connections where one multiplexed connection would do, each paying a handshake and TCP slow start and splitting prioritisation. Extreme bundling and sprites mean one changed line invalidates a large cached file. Serve from one origin and split bundles along cache boundaries instead.
- Expecting HTTP/3 to be faster everywhere
On a clean, low-latency link the gain is small, and QUIC in user space has historically used more CPU per byte than kernel TCP. HTTP/3 pays off on lossy or high-latency networks and on mobile clients that change networks. Measure real-user metrics per protocol before and after.
- Forgetting that UDP is blocked in places
Some corporate firewalls and networks drop UDP 443. Browsers fall back to HTTP/2, but only if you still serve it. Never run an HTTP/3-only endpoint for browsers, and keep TCP listeners and their certificates healthy.
- Assuming the edge protocol reaches your app
A CDN (content delivery network) or load balancer may speak HTTP/3 to the browser and HTTP/1.1 to your origin. That is usually fine, but multiplexing-dependent features like gRPC need HTTP/2 end to end, and long-lived HTTP/2 connections need request-level, not connection-level, load balancing.
- Running unpatched HTTP/2 servers
Multiplexing gives clients cheap ways to create work. The "Rapid Reset" attack (CVE-2023-44487, October 2023) opened and immediately cancelled streams to produce record-size DDoS (distributed denial of service) traffic. Keep servers and proxies updated, and set limits on concurrent streams.
Interview questions
Q1What problem did HTTP/2 solve?
HTTP-level head-of-line blocking and connection overhead. HTTP/1.1 handles one response at a time per connection, so browsers opened about six and queued the rest. HTTP/2 splits messages into frames on numbered streams and interleaves them on one connection, and adds HPACK header compression, so many requests proceed in parallel with one handshake.
Q2If HTTP/2 already multiplexes, why do we need HTTP/3?
Because HTTP/2 moved head-of-line blocking down into TCP. TCP delivers bytes in order, so one lost packet holds back data for every stream on the connection until it is retransmitted. HTTP/3 runs over QUIC, which knows about streams and recovers them independently, and it also cuts connection setup to one round trip and survives network changes.
Q3Walk me through what happens the first time a browser talks to an HTTP/3-capable server.
Unless it has an HTTPS DNS record, it starts over TCP: a TLS handshake with ALPN negotiates h2, and the response includes Alt-Svc advertising h3 on UDP 443. The browser remembers that, opens a QUIC connection for later requests, and switches to HTTP/3 if the handshake succeeds, falling back to HTTP/2 if UDP is blocked. On later visits it can resume with 0-RTT.
Q4What happens to an HTTP/3 connection when a phone switches from Wi-Fi to cellular?
It keeps going. QUIC identifies connections by connection IDs rather than the IP and port four-tuple, so packets from the new address are recognised as the same connection after a path validation. A TCP connection is bound to the old address and has to be re-established, including the TLS handshake, which is where the visible stall on mobile comes from.
Q5Why is QUIC built on UDP instead of being a new transport protocol?
Deployability. A new IP protocol or a changed TCP would be dropped or mangled by the middleboxes and kernels already deployed, and would take years to roll out. UDP passes almost everywhere, QUIC can run in user space and ship with browsers and servers, and encrypting its own headers keeps middleboxes from ossifying it.
Q6Should we still use domain sharding and big bundles with HTTP/2?
No for sharding, mostly no for giant bundles. Sharding splits the single multiplexed connection into several, each with its own handshake and congestion window. Bundling still reduces per-request overhead somewhat, but with cheap requests it is better to split by how often code changes, so a deploy invalidates only the chunks that changed.
Q7What is 0-RTT and what is the risk?
0-RTT lets a client that has talked to the server before send request data in its very first packet, using keys from a previous session, so there is no setup delay at all. The risk is replay: that early data can be captured and resent. Servers therefore accept only idempotent requests in 0-RTT and can reject others with 425 Too Early.
- HTTP/1.1 handles one response at a time per connection, which is why browsers open about six and why sharding and sprites existed.
- HTTP/2 interleaves frames from many streams on one TCP connection with HPACK header compression; server push and the priority tree did not survive.
- TCP's in-order delivery means one lost packet stalls every HTTP/2 stream, so on lossy links HTTP/2 can lose to HTTP/1.1.
- HTTP/3 runs on QUIC over UDP: per-stream loss recovery, TLS 1.3 in a one-RTT handshake, 0-RTT resumption and connection migration.
- Browsers discover HTTP/3 through Alt-Svc or HTTPS DNS records and fall back to HTTP/2, so keep serving both.
- Drop domain sharding, split bundles by cache lifetime, and measure HTTP/3 with real-user data rather than assuming a win.