← Interview Preparation

Observability Reference

How you see production systems — structured logging, Prometheus-style metrics, OpenTelemetry tracing, SLOs and burn-rate alerting, incident response, and platform trade-offs (Datadog, Grafana LGTM, Splunk, Quickwit). For backend and distributed-systems senior loops. Assumes Networking and OS are done; wire-level debugging stays in Networking, resilience patterns in Distributed Systems. Anchor books: Site Reliability Engineering (Google) and Observability Engineering (Majors et al.).

Time budget: ≈30h

How to use this reference

  • Work through topics top to bottom — SLOs assume you understand metrics; distributed tracing assumes you understand correlation IDs from logging.
  • Google SRE Book Ch. 6 (monitoring) and Ch. 4–5 (SLOs, alerting) are the anchor — each subtopic points to specific sections, not cover-to-cover reading.
  • We deliberately do not repeat Networking packet capture, OS syscall I/O, or Distributed Systems circuit-breaker theory — only how to measure and alert on those behaviors.
  • Optional tasks include local Prometheus/Grafana or OTel sandbox exercises — skip if you already use Datadog or similar daily at work.

The Reference

  1. 1

    Logs, metrics, and traces — the three pillars, and why cardinality is the hidden cost center.

    1. 1.1Three Pillars & Observability vs Monitoring2/525m
    2. 1.2Cardinality, Sampling & Cost3/530m
  2. 2

    Structured events, pipelines, and search platforms — from stdout to Splunk, Quickwit, and Loki.

    1. 2.1Structured Logging & Pipelines3/530m
    2. 2.2Log Search Platforms3/530m
  3. 3

    RED, USE, Prometheus, histograms, and the SLI math behind percentiles.

    1. 3.1RED, USE & Prometheus Fundamentals3/530m
    2. 3.2Histograms, Percentiles & SLIs4/535m
  4. 4

    OpenTelemetry, context propagation, and sampling — following a request across twenty services.

    1. 4.1OpenTelemetry Fundamentals3/530m
    2. 4.2Context Propagation & Sampling4/535m
  5. 5

    SLIs, SLOs, error budgets, and burn-rate alerting — the language of reliability engineering.

    1. 5.1SLI & SLO Definition3/530m
    2. 5.2Burn-Rate Alerting4/535m
  6. 6

    Alert design that pages humans only when necessary — and incident response that learns.

    1. 6.1Alert Design & Fatigue3/530m
    2. 6.2Incident Response & Postmortems3/530m
  7. 7

    Datadog unified observability vs the Grafana LGTM stack — vendor and OSS trade-offs at scale.

    1. 7.1Datadog Unified Observability3/530m
    2. 7.2Grafana LGTM Stack4/535m
  8. 8

    Dashboards, exemplars, triage workflows, profiling, and eBPF — finding root cause under fire.

    1. 8.1Dashboards, Exemplars & Triage4/535m
    2. 8.2Profiling & eBPF4/535m