Topics
DevOps & Delivery

Observability: Logs, Metrics and Traces

What logs, metrics and traces each answer, why label cardinality breaks metrics, how a trace ID crosses services, and how SLOs turn signals into alerts.

Intermediate·15 min read·Updated Oct 6, 2026

A running system tells you what it is doing through three kinds of signal. Logs record individual events, metrics are numbers aggregated over time, and traces follow one request through every service it touched. Each answers a different question at a different cost: metrics are cheap and tell you that something is wrong, traces tell you where, and logs tell you why. A shared trace ID stamped on all three is what lets you move from one to the next.

Context

For a single server, monitoring meant a few graphs (CPU, memory, disk) and tail -f on a log file. Splitting systems into dozens of services broke that: a slow checkout could be caused by any of them, and no single log file held the whole story. Google described its answer, request tracing, in the Dapper paper (2010); Twitter open-sourced Zipkin in 2012 and Uber released Jaeger in 2017. Prometheus (SoundCloud, 2012; the second project in the CNCF, the Cloud Native Computing Foundation, in 2016) made pull-based metrics the default in Kubernetes clusters. The word observability, borrowed from control theory, spread from around 2016 to mean being able to ask new questions of a system without shipping new code. In 2019 the two competing open tracing standards, OpenTracing and OpenCensus, merged into OpenTelemetry, which now covers all three signals.

You have met all three already: a JSON line in your terminal from a logger such as pino, a Grafana dashboard with request rate and latency, and the waterfall view in Jaeger or Datadog that shows which call was slow. The smallest useful example is one log line that carries the ID of the request it belongs to:

log-line.json
{"level":"error","time":"2026-10-06T09:12:44.120Z",
 "service":"payments","msg":"card declined",
 "trace_id":"4bf92f3577b34da6a3ce929d0e0e4736",
 "order_id":"ord_8812","http.status":402,"duration_ms":312}
Monitoring vs observability
Monitoring checks known failure modes against thresholds. Observability is having enough rich signal to investigate failures nobody predicted.
Span
One timed operation inside a trace, such as an HTTP call or a database query, with a parent span, attributes and a status.
Trace ID
A random 128-bit ID shared by every span of one request, passed between services in a header and stamped on logs.
Cardinality
How many distinct values a label can take. Each distinct combination of metric labels is stored as its own time series.
SLI / SLO
A service level indicator measures what users experience (share of fast, successful requests); a service level objective is the target for it, such as 99.9% over 30 days.

Why it matters

Incidents are measured in time to detect and time to resolve, and both depend on signals that existed before the incident started. Without metrics you learn about an outage from users. Without traces you know checkout is slow but not which of eight services is slow. Without correlated logs you find the trace but not the error that explains it. The opposite failure is just as common: a logging bill larger than the compute bill, a metrics database brought down by one label, and alerts so noisy that the team mutes the channel and misses the real page.

Three signals, three questions

The signals differ in what they keep. A log line keeps every detail of one event. A metric throws the details away and keeps a running number per time window, which is why it is cheap to store for months and fast to graph. A trace keeps the timing and causal structure of one request across process boundaries, which no other signal can reconstruct.

SignalAnswersCost driverWeak at
MetricsIs something wrong, and since when? Error rate, p99 latency, queue depthNumber of time series (label cardinality)Explaining one specific failure
TracesWhere in the call graph did this request spend its time or fail?Spans per request × sampling rateLong-term trends; most requests are sampled out
LogsWhat exactly happened in this event, with which inputs?Bytes ingested and retainedAggregation: counting from logs is slow and costly

Structured logs

A log line like Payment failed for order 8812 can only be grepped. The same event as key-value fields can be filtered, grouped and joined: every line from service=payments with http.status=402 in the last hour, or every line of one request. That last query only works if each line carries a correlation ID, and the best one is the trace ID, because the tracing system already propagates it across services for free.

Alertp99 > 800 ms
Dashboardpayments only
Traceslow PSP call
Logsby trace_id
Causecard network timeout
The usual investigation path: a metric alert says something is wrong, an exemplar or a filter leads to a slow or failed trace, and the trace ID finds the log lines that explain it.

Metrics and the cardinality trap

Most metrics are one of three types. A counter only goes up (requests served, errors), and you graph its rate. A gauge goes up and down (memory in use, queue length). A histogram counts observations into buckets (requests under 100 ms, under 250 ms, …), which is what lets you compute percentiles across many instances. You cannot average percentiles from different servers, but you can add their bucket counts and take the percentile of the sum.

p99-latency.promql
# p99 latency per route over the last 5 minutes,
# summed across every instance before taking the quantile
histogram_quantile(0.99,
  sum by (le, route) (
    rate(http_request_duration_seconds_bucket[5m])
  )
)

# error ratio: errors / all requests
sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))

A time-series database (TSDB) such as Prometheus stores one series for every distinct combination of label values. Labels like route, method and status are fine, because they multiply to a few hundred series. Add user_id and every user becomes a new series per route per status: memory grows with users, queries slow down, and the database eventually falls over. Unbounded values belong in logs and trace attributes, never in metric labels.

route × status
40 routes× 6 statuses= 240 seriescheap, keep it
+ user_id
240 series× 500,000 users= 120M seriesTSDB runs out of memory
Series count is the product of each label's distinct values. Bounded labels stay small; one unbounded label multiplies everything by the number of users.

What to measure

Three checklists cover almost every service. The four golden signals from the Google SRE book (2016) are latency, traffic, errors and saturation. RED (rate, errors, duration), named by Tom Wilkie, is the same idea for request-driven services. USE (utilization, saturation, errors), from Brendan Gregg, applies to resources such as CPU, disks, connection pools and queues. RED tells you users are affected; USE tells you which resource ran out.

Following one request: distributed tracing

When a request enters the system, the first service creates a trace ID and a root span. Every outgoing call carries the trace ID and the current span ID in a header; the receiving service creates a child span with that parent and does the same for its own calls. Each service sends its finished spans to a collector independently, and the backend reassembles them by trace ID into a tree with timings.

trace 4bf92f… · 480 msgatewayGET /checkout · 480 msordersPOST /orders · 440 mspostgresINSERT orders · 50 mspaymentscharge · 320 mspayment providerPOST /charges · 280 msinventoryreserve · 15 ms0240 ms480 ms
One checkout request as a trace waterfall. The root span covers 480 ms; nearly all of it is the payments service waiting on the external payment provider, which no single service's logs would show on their own.

The header format was standardised as W3C Trace Context (a W3C Recommendation since 2020), so services written in different languages and traced by different vendors still join one trace. The traceparent header holds a version, the 128-bit trace ID, the caller's span ID and a flags byte whose last bit says whether this trace is being recorded:

propagation.http
GET /charges HTTP/1.1
Host: payments.internal
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
#            ver-trace id (128 bit)            -parent span id -sampled

Sampling

Recording every span of every request is expensive at scale, so traces are sampled. Head sampling decides at the first service (keep 5% of traces) and passes the decision downstream in the flags; it is cheap but blind, so most of the rare slow or failed requests are thrown away. Tail sampling buffers all spans of a trace in a collector and decides after it finishes: keep every error, every request over 1 second and 1% of the rest. It keeps the interesting traces at the price of running and sizing that buffer.

OpenTelemetry ties the signals together

OpenTelemetry (OTel) provides one API and SDK per language for traces, metrics and logs, auto-instrumentation for common HTTP clients, servers and database drivers, and OTLP (the OpenTelemetry Protocol) for shipping data. Its tracing API reached 1.0 in 2021, and metrics and logs followed as stable specifications. Because the API is vendor-neutral, switching from Jaeger to Datadog or Grafana Tempo is a collector configuration change, not a code change.

place-order.ts
import {trace, SpanStatusCode} from '@opentelemetry/api'
import pino from 'pino'

const tracer = trace.getTracer('orders')
const log = pino()

export async function placeOrder(cart: Cart) {
  return tracer.startActiveSpan('placeOrder', async span => {
    span.setAttribute('cart.items', cart.items.length)
    const {traceId, spanId} = span.spanContext()
    try {
      // auto-instrumented driver: becomes a child span
      const order = await db.insertOrder(cart)
      log.info({trace_id: traceId, span_id: spanId,
        order_id: order.id}, 'order placed')
      return order
    } catch (err) {
      span.recordException(err as Error)
      span.setStatus({code: SpanStatusCode.ERROR})
      throw err
    } finally {
      span.end()
    }
  })
}
App + OTel SDKspans, metrics, logs
Collectorbatch · tail-sample
Prometheusmetrics
Tempo / Jaegertraces
Loki / Elasticsearchlogs
A typical pipeline. The app only speaks OTLP to a local collector; the collector batches, samples and routes each signal to whichever backend stores it.

Alerting on what users feel

An alert should mean a human must act now. Alerts on causes (CPU at 90%, a pod restarted) fire constantly without user impact, and miss failures nobody thought to write a rule for. Alerts on symptoms catch both: the share of requests that are slow or failing. An SLO makes that precise and gives the team an error budget, the amount of unreliability it can spend on deploys and experiments before reliability work takes priority.

  1. 1
    Define the SLI from the user's side: the share of checkout requests that return a non-5xx status in under 800 ms, measured at the load balancer.
  2. 2
    Set the SLO: 99.9% of requests over a rolling 30 days. The error budget is the other 0.1%: about 43 minutes of total outage, or a longer period of partial failure.
  3. 3
    Alert on burn rate, not on single bad minutes. The SRE workbook (2018) suggests paging when the last hour burns budget 14.4× faster than sustainable, which spends 2% of the month's budget in one hour, and opening a ticket for slow burns over days.
  4. 4
    When the budget runs out, the agreed policy applies: freeze risky releases and spend the time on reliability until the window recovers.

Pitfalls

  • Unbounded values as metric labels

    User IDs, request IDs, full URLs with IDs in the path or raw error messages as labels create a new time series per value. Memory and query time in the metrics database grow without limit until it falls over, usually during the incident you needed it for. Put those values in logs and span attributes and label metrics by route templates such as /orders/:id.

  • Logs without a trace or correlation ID

    Without a shared ID, the log lines of one request across five services can only be matched by timestamp, which fails under any real load. Inject the trace ID into the logger context once, in middleware, so every line carries it without anyone remembering to add it.

  • Averages instead of percentiles

    A mean latency of 120 ms can hide 1% of requests taking 5 seconds, and those users are often the most valuable (bigger carts, more data). Graph p50, p95 and p99 from histograms, and never average percentiles across instances: aggregate the bucket counts first, then compute the quantile.

  • Breaking context propagation

    A hand-written HTTP client, a message queue or a thread pool that does not pass the trace context splits one request into disconnected traces, so the waterfall ends exactly where the interesting part starts. Use instrumented clients, and carry traceparent in message headers for asynchronous work.

  • Paging on causes

    CPU, memory and restart alerts fire during harmless spikes and stay quiet when a bug returns wrong data at normal resource usage. On-call fatigue sets in, alerts get muted, and real outages are found by customers. Page on SLO burn and keep cause metrics for diagnosis.

Interview questions

Q1What is the difference between logs, metrics and traces?

Metrics are numeric aggregates over time, logs are records of individual events, and traces are the timed tree of operations for one request across services. Metrics are cheap and good for detecting problems and trends, traces locate where a request spent time or failed, and logs carry the detail that explains why. In practice you go from an alert on a metric to a trace to the logs for that trace ID.

Q2Walk me through implementing distributed tracing in a set of microservices.

Add the OpenTelemetry SDK to each service with auto-instrumentation for the HTTP server, HTTP clients and database drivers, so spans and W3C traceparent propagation come for free. Add manual spans around important business steps, propagate context through message queues by putting traceparent in message headers, and inject the trace ID into the logger. Export over OTLP to a collector that does tail sampling (all errors, all slow requests, a small share of the rest) and forwards to the tracing backend.

Q3What happens to your metrics backend if you add user_id as a label?

Every user becomes a separate time series for every combination of the other labels, so series count multiplies by the number of users. The TSDB's memory and index grow with that, ingestion and queries slow down, and eventually it runs out of memory. Per-user detail belongs in logs or trace attributes, where high-cardinality values are expected.

Q4Why can you not average p99 latencies from ten servers?

Because a percentile is not additive: the average of ten p99s is not the p99 of the combined traffic, and it can be far off when one server is slow. Histograms solve this by exporting bucket counts, which can be summed across servers before computing the quantile. That is what histogram_quantile over a summed rate does in Prometheus.

Q5Head or tail sampling: which would you use?

Tail sampling when I can afford a collector tier, because it decides after the trace completes and can keep every error and slow request. Head sampling is simpler and cheaper but random, so at a 1% rate it drops 99% of the rare failures I most want to see. A common setup is head sampling for very high volume endpoints and tail sampling for everything that matters.

Q6What would you alert on for a checkout service?

On burn rate of an SLO such as 99.9% of checkouts succeeding in under 800 ms over 30 days: page on fast burn, ticket on slow burn. That covers any cause of user pain, including ones nobody predicted. CPU, memory and dependency health stay on dashboards for diagnosis rather than paging anyone.

Q7What is the difference between monitoring and observability?

Monitoring watches for known failure modes with predefined checks and dashboards; observability means the system emits rich enough signal to answer questions you did not anticipate. In practice that means high-cardinality, structured events and traces you can slice by any attribute, on top of the monitoring that tells you something is wrong.

Key takeaways
  • Metrics say that something is wrong, traces say where, logs say why. A shared trace ID connects them.
  • Every distinct label combination is a time series. Keep metric labels bounded and put user or request IDs in logs and spans.
  • Use histograms for latency and look at p95 and p99. Aggregate bucket counts across instances, never percentiles.
  • Traces work by propagating context: W3C traceparent between services, message headers across queues. Tail sampling keeps the errors and slow requests.
  • OpenTelemetry gives one vendor-neutral API for all three signals; the collector decides where they are stored.
  • Page on SLO burn rate, not on CPU. The error budget is what turns reliability into a planning decision.

Preparing for interviews? DevRecall turns a job description into a prep plan that points at topics like this one.

Start free