Topics
DevOps & Delivery

Graceful Shutdown and Health Checks

What really happens between SIGTERM and SIGKILL, why pods get traffic after they start stopping, and what liveness, readiness and startup probes should check.

Intermediate·13 min read·Updated Oct 6, 2026

When an orchestrator stops a process it sends SIGTERM, waits a grace period (30 seconds by default in Kubernetes), then sends SIGKILL, which cannot be caught. A graceful shutdown uses that window to stop taking new work, finish what is in flight and close resources in order. Health checks are the other half: readiness decides whether an instance gets traffic, liveness decides whether it gets restarted, and mixing the two up turns one slow database into a fleet-wide restart loop.

Context

Signals are as old as Unix: SIGTERM means "please exit" and can be handled; SIGKILL ends the process immediately with no chance to react. For a long-running server that restarted a few times a year, nobody cared much. Containers and continuous deployment changed that: a service with 20 replicas deployed ten times a day is shut down 200 times a day, and every shutdown that drops a few requests shows up as a steady background of 502 errors. Kubernetes made the contract explicit: docker stop waits 10 seconds between the two signals, Kubernetes waits terminationGracePeriodSeconds (default 30), and probes tell the platform when a container is ready and whether it is still alive.

You have met this as the livenessProbe and readinessProbe blocks in a Deployment, as a /health endpoint behind a load balancer, as server.shutdown=graceful in Spring Boot (graceful by default since 3.4) or srv.Shutdown(ctx) in Go (since 1.8), and as a container that takes exactly 30 seconds to stop every time because nothing in it handles SIGTERM. The minimum a process needs:

minimal.js
process.on('SIGTERM', () => {
  server.close(() => process.exit(0))  // stop accepting, finish, exit
})
SIGTERM / SIGKILL
Signal 15 asks a process to exit and can be handled. Signal 9 kills it at once; no handler runs, no buffer is flushed.
Grace period
Time between SIGTERM and SIGKILL: terminationGracePeriodSeconds in Kubernetes, --stop-timeout in Docker, stopTimeout in AWS ECS (Elastic Container Service).
Draining
Finishing in-flight requests and jobs while refusing new ones.
Readiness
Whether the instance should receive traffic now. Failing it removes the pod from the Service, nothing more.
Liveness
Whether the process is stuck beyond recovery. Failing it makes the kubelet kill and restart the container.
PID 1
The first process in a container. The kernel gives it no default signal handlers, so it ignores SIGTERM unless it installs one.

Why it matters

Every rolling deploy, autoscaler scale-down, node drain and spot-instance reclamation is a shutdown. Get it wrong and each one costs some failed requests, half-processed jobs that are retried or lost, and connections reset mid-response. Health checks fail in the opposite direction: a liveness probe that calls the database restarts every pod when the database is slow, which removes all capacity at the moment the system is struggling, and a missing readiness probe sends traffic to pods still warming up. Both are invisible in development, where processes are started and stopped by hand.

What happens when a pod is deleted

Deleting a pod starts two things at the same time, and nothing waits for the other. The kubelet on the node starts stopping the container, and the control plane removes the pod from the Service's endpoints, which kube-proxy, ingress controllers and cloud load balancers then pick up on their own schedule. If the container exits as soon as it receives SIGTERM, proxies that have not caught up yet keep sending it requests, which fail.

Control planemarked Terminatingremoved from endpointsContainerpreStop sleepSIGTERM: stop accepting, drainexit 0SIGKILLProxiesstill route hereupdated, no new traffic0s~1-5sgrace ends (30s)
Pod termination runs on two independent tracks. The preStop sleep keeps the container serving while proxies catch up; SIGTERM handling drains; SIGKILL arrives at the end of the grace period whether or not draining finished.
  1. 1
    The pod is marked Terminating. From here the grace period clock runs, and it covers everything below, including the preStop hook.
  2. 2
    In parallel, the endpoints controller removes the pod from the Service's EndpointSlices. kube-proxy rewrites its rules, and ingress controllers and cloud load balancers update on their own timers, often a few seconds later.
  3. 3
    The kubelet runs the preStop hook if there is one. A short sleep here (5-10 seconds) keeps the container serving normally while the proxies catch up.
  4. 4
    The kubelet sends SIGTERM to PID 1 in the container. The application stops accepting new connections, finishes in-flight requests and jobs, closes its resources and exits.
  5. 5
    If the container is still running when the grace period ends, the kubelet sends SIGKILL. Whatever was in flight is lost.

Making sure the signal arrives

A shutdown handler only runs if the signal reaches the process that installed it. Docker and Kubernetes send SIGTERM to PID 1 in the container, and two things often sit there instead of your application. The shell-form CMD npm start makes /bin/sh PID 1, and sh does not forward signals to its child. Package-manager scripts such as npm start add another wrapper process that has historically not passed signals on reliably. The result is the classic container that ignores SIGTERM, sits for the whole grace period and is killed, on every deploy.

Even when the application is PID 1, the kernel installs no default handlers for it, so a Node.js process without a SIGTERM listener simply ignores the signal. PID 1 also inherits orphaned child processes and must reap them, or they accumulate as zombies. A minimal init such as tini or dumb-init solves both: it runs as PID 1, forwards signals to your process and reaps children. docker run --init injects tini without changing the image.

Dockerfile
# wrong: shell form, /bin/sh is PID 1 and swallows SIGTERM
CMD npm start

# right: exec form, tini is PID 1, forwards signals, reaps zombies
RUN apk add --no-cache tini
ENTRYPOINT ["/sbin/tini", "--"]
CMD ["node", "dist/server.js"]
CMD npm start
sh (PID 1)npmnodesignal stops at sh, killed after grace
CMD ["node", …]
node (PID 1)works only if node handles SIGTERM
tini + node
tini (PID 1)nodeforwarded, zombies reaped
Who receives SIGTERM, depending on how the container starts.

Draining inside the application

Once SIGTERM arrives, the order matters. First stop taking new work: close the listening socket and stop pulling from queues. Then finish what is in flight. Then close what that work depends on, the database pool last, because the requests being drained still use it. Finally exit with code 0, before the grace period ends. A hard timeout inside the process makes sure a stuck request cannot hold the shutdown hostage until SIGKILL.

Keep-alive connections are the usual trap in Node.js. server.close() stops new connections, but an idle keep-alive socket from a proxy stays open and the close callback never fires. Node 18.2 added server.closeIdleConnections(), and since Node 19 close() calls it for you. Sockets that are busy finish their request; sending Connection: close on those responses tells the client not to reuse them.

shutdown.ts
import http from 'node:http'
import {once} from 'node:events'

let shuttingDown = false
app.use((_req, res, next) => {
  if (shuttingDown) res.setHeader('Connection', 'close')
  next()
})
app.get('/readyz', (_req, res) =>
  res.status(shuttingDown ? 503 : 200).end())
app.get('/livez', (_req, res) => res.status(200).end())

const server = http.createServer(app).listen(8080)

async function shutdown(signal: string) {
  if (shuttingDown) return
  shuttingDown = true                  // readiness fails from now on
  const force = setTimeout(() => process.exit(1), 25_000)
  force.unref()                        // hard stop before SIGKILL

  const closed = once(server, 'close')
  server.close()                       // no new connections
  server.closeIdleConnections()        // idle keep-alive sockets
  await consumer.stop()                // finish jobs, take no new ones
  await closed                         // in-flight requests done
  await db.end()                       // pool last: drain needed it
  console.log(`${signal}: drained, exiting`)
  process.exit(0)
}

process.on('SIGTERM', () => void shutdown('SIGTERM'))
process.on('SIGINT', () => void shutdown('SIGINT'))

Liveness, readiness and startup probes

The three probes ask different questions and trigger different actions, so they should check different things. Readiness is about traffic: an instance that is starting, overloaded or shutting down should say "not ready" and be taken out of rotation, without being restarted. Liveness is about the process itself: it should fail only when the process is wedged in a way that only a restart fixes, such as a deadlock or an event loop blocked for minutes. The startup probe, GA (generally available) since Kubernetes 1.20, holds off the other two until a slow starter has booted, so a generous startup allowance does not force a slow liveness probe for the rest of the pod's life.

ProbeQuestionOn failureMust not check
StartupHas the app finished booting?Keeps waiting; restart after failureThresholdSlow external systems with no time limit
ReadinessShould it receive traffic right now?Removed from endpoints, not restartedA shared dependency that, if down, empties every pod at once
LivenessIs the process stuck for good?Container killed and restartedDatabases, caches, downstream APIs: anything outside the process
deployment.yaml
spec:
  terminationGracePeriodSeconds: 45     # preStop + drain + margin
  containers:
    - name: api
      image: registry.example.com/api:1.42
      startupProbe:                     # up to 30 x 2s to boot
        httpGet: {path: /livez, port: 8080}
        periodSeconds: 2
        failureThreshold: 30
      readinessProbe:                   # in or out of the Service
        httpGet: {path: /readyz, port: 8080}
        periodSeconds: 5
        failureThreshold: 2
      livenessProbe:                    # restart only if wedged
        httpGet: {path: /livez, port: 8080}
        periodSeconds: 10
        failureThreshold: 3
      lifecycle:
        preStop:
          sleep: {seconds: 5}           # Kubernetes 1.30+

Failing readiness during shutdown

Inside Kubernetes, a terminating pod is removed from endpoints regardless of its readiness probe, so the preStop sleep is what closes the race. Outside it, a load balancer that only learns about instances through health checks, such as an ALB (Application Load Balancer) target group or an Nginx upstream with active checks, needs the application to start failing readiness on SIGTERM and keep serving for a few check intervals before closing. The handler above does both: readiness returns 503 the moment shutdown begins, while requests are still served.

Pitfalls

  • A liveness probe that checks dependencies

    If /livez pings the database, a slow database fails every pod's liveness at once and Kubernetes restarts all of them. Restarts add connection storms to an already struggling database, so the outage deepens. Liveness should only prove the process can respond; dependency problems belong in readiness, metrics and alerts.

  • Exiting immediately on SIGTERM

    The endpoint removal and the SIGTERM race each other, so a process that quits at once still receives requests from proxies that have not updated, and those requests fail with connection errors. A short preStop sleep, or continuing to serve for a few seconds after SIGTERM, removes the 502s from every deploy.

  • A grace period shorter than the work

    The grace period includes the preStop hook. With a 10-second sleep and a 30-second default, only 20 seconds remain for draining, and a 25-second report request is killed halfway. Measure the slowest request or job and set terminationGracePeriodSeconds above preStop plus that.

  • Shell-form CMD or npm start as the entrypoint

    The signal stops at the shell or the npm wrapper, the application never sees it, and every stop takes the full grace period followed by SIGKILL. Deploys are slow and every in-flight request is cut off. Use exec form with an init such as tini, or docker run --init.

  • Closing the database pool first

    Shutdown code often closes resources in the order they were opened, which shuts the pool while requests that are being drained still query it, turning a graceful drain into a burst of errors. Close dependencies only after the server and consumers report they are idle.

Interview questions

Q1What is the difference between liveness and readiness probes?

Readiness decides whether a pod receives traffic; liveness decides whether it gets restarted. A failing readiness probe removes the pod from the Service until it recovers, which is right for warming up, overload or shutdown. A failing liveness probe kills the container, which is only right when the process is stuck and a restart is the fix, so liveness must never depend on anything outside the process.

Q2Walk me through implementing graceful shutdown for an HTTP service with a queue consumer.

On SIGTERM I set a flag so readiness returns 503, stop the consumer from fetching new messages, and call server.close plus closeIdleConnections so no new connections are accepted and idle keep-alive sockets go away. I wait for in-flight requests and the current message to finish and be acknowledged, then close the database pool and exit 0. A timer shorter than the grace period forces exit if something hangs, and in Kubernetes a preStop sleep of a few seconds keeps the pod serving while proxies drop it.

Q3What happens when a Kubernetes pod is deleted?

It is marked Terminating, and two things start in parallel: the control plane removes it from the Service endpoints, and the kubelet runs the preStop hook, then sends SIGTERM to PID 1. Proxies learn about the endpoint change asynchronously, so for a few seconds traffic can still arrive. If the container has not exited when terminationGracePeriodSeconds runs out, 30 seconds by default, it receives SIGKILL.

Q4Our containers always take exactly 30 seconds to stop. Why?

SIGTERM is not reaching the application, so Kubernetes waits out the whole grace period and sends SIGKILL. Usually the image uses shell-form CMD or npm start, so a shell or npm is PID 1 and does not forward the signal, or the app is PID 1 without a SIGTERM handler, which the kernel then ignores. Exec-form CMD with tini, or --init, plus a handler in the app fixes it.

Q5Should a readiness probe check the database?

Only carefully. If every pod depends on the same database, a database outage makes every pod unready and the Service has no endpoints, so clients get connection errors instead of a fast, meaningful 503 from the app. I prefer readiness that checks the instance's own state (booted, not overloaded, not shutting down), and handle dependency failures in the request path with timeouts, circuit breakers and alerts.

Q6Why does the startup probe exist if liveness has initialDelaySeconds?

Because initialDelaySeconds is fixed for the pod's life. A service that needs up to two minutes to boot would otherwise need a liveness probe that tolerates two minutes of failure forever, so real hangs are detected slowly. The startup probe gives boot its own generous budget, and once it succeeds liveness and readiness take over with tight settings.

Q7How long should terminationGracePeriodSeconds be?

Long enough for the preStop hook plus the slowest request or job plus a margin, because the hook counts against the same budget. For a typical API that is 5 seconds of preStop plus 20-30 seconds of draining; for workers with long jobs it can be minutes, or the jobs must be resumable so a SIGKILL only costs a redelivery.

Key takeaways
  • SIGTERM asks, SIGKILL forces. The grace period between them (30s default in Kubernetes) includes the preStop hook and is your entire budget for draining.
  • Endpoint removal and SIGTERM happen in parallel; a short preStop sleep keeps the pod serving until proxies stop sending traffic.
  • The signal must reach your process: exec-form CMD, an init like tini as PID 1, and an actual SIGTERM handler.
  • Drain in order: stop accepting, finish in-flight work, close the database pool last, exit before a hard internal timeout.
  • Readiness controls traffic, liveness controls restarts, startup protects slow boots. Liveness must never check dependencies.

Preparing for interviews? DevRecall turns a job description into a prep plan that points at topics like this one.

Start free