Topics
Networking & Protocols

DNS Resolution: From Name to IP Address

How a resolver walks root, TLD and authoritative servers to turn a name into an address, why TTL caching explains "propagation", and where DNS breaks.

Intermediate·13 min read·Updated Oct 6, 2026

DNS (Domain Name System) turns a name like www.example.com into an IP address through a chain of caches and servers. Your machine asks one recursive resolver; on a cache miss, the resolver walks the tree itself: a root server points it to the TLD (top-level domain) servers for .com, those point it to the domain's authoritative servers, and those return the record. Every answer carries a TTL (time to live), and every layer caches it for that long, which is why a changed record reaches users gradually rather than instantly.

Context

Before DNS, every host on the ARPANET downloaded one file, HOSTS.TXT, maintained by hand at the Stanford Research Institute. By the early 1980s it was too big and changed too often to copy around, so Paul Mockapetris designed a distributed, hierarchical replacement (RFC 882/883 in 1983, refined as RFC 1034/1035 in 1987, still the core specs). The design has barely changed since: authority is delegated down a tree of names, and answers are cached everywhere on the way back. The modern additions are mostly about trust and privacy: DNSSEC (DNS Security Extensions; the root zone was signed in 2010), DoT (DNS over TLS, RFC 7858, 2016) and DoH (DNS over HTTPS, RFC 8484, 2018).

You meet DNS every time you add an A record in Cloudflare or Route 53, wait for a domain to "propagate", see ENOTFOUND or NXDOMAIN in a log, or point api at a load balancer with a CNAME. The fastest way to see the whole chain is dig +trace, which does the walk itself instead of asking your resolver:

dig-trace.sh
$ dig +trace www.example.com    # trimmed; final address illustrative

.               518400 IN NS a.root-servers.net.
;; Received 239 bytes from 192.168.1.1#53 in 4 ms

com.            172800 IN NS a.gtld-servers.net.
;; Received 1170 bytes from 198.41.0.4#53(a.root-servers.net)

example.com.    172800 IN NS a.iana-servers.net.
;; Received 656 bytes from 192.5.6.30#53(a.gtld-servers.net)

www.example.com.   300 IN A  203.0.113.10
;; Received 60 bytes from 199.43.135.53#53(a.iana-servers.net)
Stub resolver
The minimal client in your OS (getaddrinfo, systemd-resolved). It reads /etc/hosts, then sends one query to a configured resolver and waits.
Recursive resolver
The server that does the work: your ISP, 1.1.1.1, 8.8.8.8 or CoreDNS in a cluster. It walks the tree on a miss and caches every answer.
Authoritative server
A server that holds the actual zone for a domain and answers for it with authority, the end of the walk.
Zone
The part of the namespace one set of authoritative servers is responsible for, such as example.com and everything under it not delegated further.
TTL (time to live)
Seconds an answer may be cached. Set by the zone owner, counted down by every cache that holds it.
NXDOMAIN
The response code for "this name does not exist". It is cached too.

Why it matters

DNS is the first dependency of every request and the least visible one. A slow resolver adds latency before any TCP connection opens; a wrong TTL turns a five-minute migration into a two-day one; a CNAME mistake sends traffic to a stale load balancer; a cluster with default settings quietly multiplies every external lookup. And when DNS itself goes down, everything that depends on names goes with it, including the tools you would use to fix it: in Facebook's six-hour outage on 4 October 2021, a routing change withdrew the paths to their authoritative DNS servers, so facebook.com stopped resolving for the whole internet, and engineers had to restore access physically.

The resolution walk

Two different kinds of question are asked. Your machine sends a recursive query: "give me the final answer". The resolver then sends iterative queries: each server answers with what it knows, usually a referral ("I don't have it, but these servers are responsible for .com"), and the resolver follows the referral itself. The root and TLD servers never do the full lookup for anyone, which is what lets a handful of them serve the whole internet.

recursive: laptop asks onceiterative: resolver walks the treeLaptopstub + cachewww.example.com?203.0.113.10Recursiveresolvercache hit? stopRoot (.)→ .com serversTLD (.com)→ example.com NSAuthoritative→ A 203.0.113.10
On a cache miss the recursive resolver walks the tree: root refers it to the .com servers, those refer it to example.com's authoritative servers, which return the A record. The resolver caches every step, so the next lookup for any .com name skips the root.
  1. 1
    Browser cache. Browsers keep their own short-lived host cache (Chrome shows it at chrome://net-internals/#dns). A hit skips everything below.
  2. 2
    OS stub resolver. It checks /etc/hosts, its own cache if it has one (macOS, systemd-resolved), then sends one recursive query to the resolver from DHCP or the config.
  3. 3
    Recursive resolver. On a cache hit it answers in a millisecond or two. On a miss it starts from the most specific cached delegation, often already knowing the .com servers.
  4. 4
    Root, TLD, authoritative. Each step is a referral down the tree until an authoritative server returns the record. The 13 root server identities (a to m) are served by well over a thousand anycast instances worldwide.
  5. 5
    Back up the chain. The resolver caches each answer for its TTL and returns the address. A modern client asks for A and AAAA (IPv6) in parallel and races the connections (Happy Eyeballs, RFC 8305).

Records you will actually edit

TypeHoldsExampleWatch out for
A / AAAAIPv4 / IPv6 address@ A 203.0.113.10Several records = round-robin across addresses
CNAMEAlias to another namewww CNAME example.com.Cannot coexist with any other record on that name
MXMail servers with priority@ MX 10 mail.example.com.Must point at a name, never a CNAME
TXTFree textSPF, DKIM, domain verificationLong values are split into 255-byte strings
NSDelegation to name servers@ NS ns1.dns.example.Must match what the registrar has at the TLD
SOAZone metadataSerial, refresh, negative TTLIts last field sets how long NXDOMAIN is cached
CAAWhich CAs may issue certificates@ CAA 0 issue "letsencrypt.org"A missing entry blocks issuance by your CA
SRVHost and port for a service_sip._tcp SRV 10 5 5060 sip...Clients must support it; browsers do not for HTTP

CAA (Certification Authority Authorization) records are checked by certificate authorities before they issue. In a zone file the same records look like this; note the trailing dots, which mark a name as absolute:

example.com.zone
$ORIGIN example.com.
$TTL 3600
@    IN SOA ns1.dns.example. hostmaster.example.com. (
          2026100601 ; serial
          7200       ; refresh
          900        ; retry
          1209600    ; expire
          300 )      ; negative-caching TTL
@    IN NS    ns1.dns.example.
@    IN A     203.0.113.10
www  IN CNAME example.com.
api  IN CNAME my-lb-123.eu-west-1.elb.amazonaws.com.
@    IN MX 10 mail.example.com.
@    IN CAA 0 issue "letsencrypt.org"

TTLs, caching and the myth of propagation

Nothing is pushed when you change a record. The new value sits on your authoritative servers, and every resolver that already cached the old one keeps serving it until its copy expires. "Propagation" is just the slowest cache in the world timing out, bounded by the TTL you set before the change. A missing name is cached too: under negative caching (RFC 2308), an NXDOMAIN is remembered for the SOA's negative TTL, so querying a record before you create it can make the new record invisible for that long.

resolver A
cached 11:10 → old IP12:10 → new IPcached just before the change
resolver B
cached 11:59 → old IP12:59 → new IPworst case: a full TTL
resolver C
no cache12:00 → new IPfirst query after the change
After the A record changes at 12:00 with a TTL of 3600, each resolver flips when its own cached copy expires, so users switch over across a full hour.
  1. 1
    A day before the migration (at least one old TTL ahead), lower the record's TTL to 60-300 seconds. Caches pick up the short TTL as their old copies expire.
  2. 2
    Bring the new target up and keep the old one serving; both must work during the switch.
  3. 3
    Change the record. Within the short TTL nearly all resolvers follow; some clients (long-lived processes, old JVMs, misbehaving resolvers) hold on longer, so watch traffic on the old target.
  4. 4
    Once the old target is idle, raise the TTL back to an hour or more to cut query volume and latency.

DNS as a load balancer

Returning different answers per query gives you coarse global balancing: several A records for round-robin, GeoDNS that answers by the resolver's location, or weighted and health-checked records in Route 53 or Cloudflare. Its limits come from caching: clients keep using whatever they got until the TTL ends, a dead server stays in caches, and the location you see is the resolver's, not the user's (EDNS Client Subnet, RFC 7871, passes a truncated client subnet to help). That is why DNS usually picks a region, and a real load balancer behind it picks the server. Public resolvers such as 1.1.1.1 (Cloudflare, 2018) and 8.8.8.8 (Google, 2009) are themselves anycast: one address announced from many locations, routed to the nearest.

DNS inside Kubernetes

Every pod gets a resolver config pointing at the cluster DNS (CoreDNS) with search domains so that orders resolves to the Service in the same namespace. The default ndots:5 means any name with fewer than five dots is tried against each search domain first, so an external lookup fans out into several failing queries, each for A and AAAA, before the real one.

resolv.conf
# /etc/resolv.conf inside a pod in namespace "shop"
search shop.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5

# resolving api.stripe.com (2 dots < 5) tries, in order:
#   api.stripe.com.shop.svc.cluster.local   -> NXDOMAIN
#   api.stripe.com.svc.cluster.local        -> NXDOMAIN
#   api.stripe.com.cluster.local            -> NXDOMAIN
#   api.stripe.com                          -> answer
# fixes: use "api.stripe.com." (trailing dot) or set
#   dnsConfig.options ndots: "2" in the pod spec

Pitfalls

  • Changing a record without lowering the TTL first

    With a 24-hour TTL, some resolvers keep sending users to the old server for a day after the change, and nothing you do on your side can flush their caches. Lowering the TTL only helps if it is done at least one old TTL before the switch.

  • Querying a name before it exists

    Testing a hostname, then creating the record, can leave the NXDOMAIN cached for the zone's negative TTL (RFC 2308). It looks like propagation is broken; it is a cached "does not exist". Create first, then test, and keep the SOA negative TTL modest.

  • A missing trailing dot in a zone file

    Inside a zone, names without a final dot are relative to the origin, so www CNAME example.com means example.com.example.com.. Web dashboards usually add the dot for you; zone files and Terraform records often do not.

  • Clients that cache forever or block on lookups

    Long-lived processes may resolve once and keep the address: older JVMs with a security manager cached positive lookups indefinitely (see networkaddress.cache.ttl), and connection pools keep sockets to the old IP. In Node.js, dns.lookup runs on the libuv thread pool (4 threads by default), so slow DNS can stall unrelated file and crypto work.

  • Treating DNS round-robin as load balancing

    Clients and resolvers cache one answer and often pick the first address, so load is uneven, and a dead server keeps receiving traffic until every cache expires. Use DNS to choose a region or an entry point, and put a health-checked load balancer behind it.

Interview questions

Q1Walk me through the DNS part of typing www.example.com into a browser.

The browser checks its own cache, then asks the OS, which checks /etc/hosts and its cache and sends a recursive query to the configured resolver. On a miss the resolver asks a root server, gets a referral to the .com servers, asks one of those, gets a referral to example.com's authoritative servers, and asks them for the A and AAAA records. It caches every answer for its TTL and returns the addresses, and the browser opens a connection, racing IPv6 and IPv4.

Q2What is the difference between a recursive and an iterative query?

A recursive query asks for the final answer, and the server must go and find it; that is what clients send to their resolver. An iterative query asks for the best the server has, usually a referral to servers further down the tree; that is what resolvers send to root, TLD and authoritative servers. Keeping those servers iterative-only is what lets them scale.

Q3We changed the A record an hour ago. Why do some users still hit the old server?

Because their resolvers cached the old answer and will keep it until its TTL expires; DNS changes are never pushed. If the old TTL was a day, it can take a day, and some clients cache longer than they should. For the next migration, lower the TTL at least one old TTL in advance, keep both targets serving during the switch, and watch the old one drain before turning it off.

Q4Why can you not put a CNAME on the bare domain?

A CNAME says "this name is an alias, look elsewhere for everything", so RFC 1034 forbids other records on the same name, but the zone apex must have SOA and NS records. Providers solve it with ALIAS, ANAME or CNAME flattening: they resolve the target themselves and publish plain A and AAAA records at the apex.

Q5What happens when someone looks up a name before you create the record?

Their resolver gets NXDOMAIN and caches it, for as long as the zone's SOA negative TTL allows (RFC 2308). Creating the record afterwards does not help those users until that cached "does not exist" expires, which looks like the record failing to propagate.

Q6How does DNS-based load balancing work, and what are its limits?

The authoritative servers return different answers per query: several addresses in rotation, the nearest region by GeoDNS, or weighted and health-checked records. It is limited by caching and by what DNS can see: a client keeps its answer for the whole TTL, failed servers linger in caches, and GeoDNS sees the resolver's location unless EDNS Client Subnet is used. So it routes to a region, and a load balancer handles individual servers.

Q7Why are external DNS lookups slow from Kubernetes pods?

Pods use ndots:5 with several search domains, so a name like api.stripe.com, with fewer than five dots, is first tried as api.stripe.com plus each cluster search domain, all of which fail, each for both A and AAAA, before the real query. That multiplies load on CoreDNS and adds latency. Fix it with a trailing dot on external names, a lower ndots in dnsConfig, or a node-local DNS cache.

Q8Do DoH and DNSSEC solve the same problem?

No. DNS over HTTPS and DNS over TLS encrypt the conversation with your resolver, which gives privacy and integrity on the network but still trusts the resolver. DNSSEC signs the records themselves so a validating resolver can prove they came from the zone owner, which gives authenticity but no privacy. They are complementary.

Key takeaways
  • Clients send one recursive query; on a miss the resolver walks root → TLD → authoritative with iterative queries.
  • Every answer is cached for its TTL at every layer. "Propagation" is caches expiring, so lower the TTL one old TTL before a change.
  • NXDOMAIN is cached too (RFC 2308): create records before you test them.
  • No CNAME at the apex; use ALIAS, ANAME or CNAME flattening. Mind trailing dots in zone files.
  • DNS load balancing picks a region, not a healthy server; caches outlive failures.
  • In Kubernetes, ndots:5 multiplies external lookups; use trailing dots or a lower ndots.

Preparing for interviews? DevRecall turns a job description into a prep plan that points at topics like this one.

Start free