← Interview Preparation

AI Systems Reference

Building production-ready, high-scale applications that deeply integrate AI — not training LLMs, not IDE workflows. Think Gemini inside Google products: gateways, retrieval at scale, serving topology, guardrails, offline/online evals, token economics, and safe rollouts. Assumes AI Engineering is done, plus Distributed Systems, Databases, and Observability as prerequisites. Anchor: AI Engineering (Chip Huyen) production chapters. Topic 9 (Deep Cuts) is optional inference-infra material — vLLM, speculative decoding, GPU fleet design, and disaggregated serving.

Time budget: ≈44h

How to use this reference

  • Work through topics top to bottom — evals and rollouts assume you understand serving and degradation patterns from earlier topics.
  • This track treats the LLM as a probabilistic dependency in a distributed system — idempotency, timeouts, and fallbacks are not optional.
  • We deliberately do not repeat AI Engineering (prompting, Cursor, MCP basics), generic SLO/incident playbooks (Observability), or circuit-breaker theory (Distributed Systems) — only how they apply to AI features at scale.
  • Optional tasks are architecture and rollout design drills — paper or whiteboard level, not model training labs.
  • Topic 9 — Deep Cuts (Inference Infra) is optional. Skip it for product and backend system design; do it for ML platform, inference, or AI infrastructure roles (vLLM, GPU fleet sizing, disaggregated serving). Adds ≈10h on top of the core path.

The Reference

  1. 1

    The production lens — why LLM features are distributed systems with probabilistic cores.

    1. 1.1The Production AI Systems Lens3/530m
    2. 1.2Non-Functional Requirements for AI Features4/530m
  2. 2

    Sync, async, streaming, and multi-step pipelines — how requests flow through production AI stacks.

    1. 2.1Sync, Async & Streaming Request Paths3/530m
    2. 2.2Compound AI Pipelines4/535m
  3. 3

    Hosted APIs vs self-hosted inference — throughput, GPU economics, and tail latency at scale.

    1. 3.1Hosted vs Self-Hosted Serving3/530m
    2. 3.2Scaling Latency & Throughput4/535m
  4. 4

    RAG ingestion and vector ops at production scale — freshness, cost, and index lifecycle.

    1. 4.1RAG Ingestion at Scale4/535m
    2. 4.2Vector Index Operations4/530m
  5. 5

    Guardrails, content policy, and graceful degradation when models and dependencies fail.

    1. 5.1Guardrails & Content Policy4/530m
    2. 5.2Failure Modes & Degradation4/535m
  6. 6

    Offline evals, online experiments, and AI-specific observability — measuring what users actually get.

    1. 6.1Offline & Online Evaluations4/535m
    2. 6.2AI System Observability4/530m
  7. 7

    Token economics, budgets, and semantic caching — running AI features without burning the margin.

    1. 7.1Token Economics & Budgeting3/530m
    2. 7.2Semantic Caching & Optimization4/535m
  8. 8

    Model rollouts, feature flags, and enterprise privacy — operating AI like any other critical service.

    1. 8.1Model Rollouts & Feature Flags4/530m
    2. 8.2Privacy, Compliance & Enterprise5/535m
  9. 9

    Optional hard-core material for ML platform, inference, and AI infrastructure loops — skip unless your target role owns GPUs or serving stacks.

    1. 9.1vLLM & Continuous Batching5/530m
    2. 9.2Speculative Decoding & Prefix Caching5/525m
    3. 9.3GPU Parallelism & Fleet Design5/530m
    4. 9.4Quantization for Inference4/525m
    5. 9.5LoRA & Multi-Adapter Serving4/525m
    6. 9.6Disaggregated Inference5/530m