← Interview Preparation

Data Engineering Reference

Building reliable data pipelines at scale — not SQL interview drills, not storage internals alone. Batch and streaming compute (Spark, Flink), orchestration (Airflow), lakehouse ELT (medallion, dbt), quality and governance, and production operations. Assumes SQL, Databases (storage formats), and Distributed Systems as prerequisites. Anchor: Fundamentals of Data Engineering (Reis & Housley) + DDIA Ch. 10–11. Topic 9 (Deep Cuts) covers Kafka, CDC, Beam, Flink internals, data mesh, and warehouse engines.

Time budget: ≈49h

How to use this reference

  • Work through topics top to bottom — Spark assumes you understand shuffle and partitioning; stream processing assumes batch DAG concepts.
  • We focus on concepts and production patterns, not tool certification — Spark API trivia is less important than execution model and tuning.
  • We deliberately do not repeat SQL query writing, B-tree/MVCC storage, TCP/DNS, or generic SLO mechanics — those live in sibling tracks. AI-specific ingestion (RAG, embeddings) lives in AI Systems.
  • Optional tasks are pipeline design drills — partition plans, DAG sketches, backfill strategies. No cluster setup required.
  • Topic 9 — Deep Cuts (Platform) is optional. Skip for general backend or analytics-engineering loops; do it for data platform, streaming, or CDC-heavy roles. Adds ≈10h on top of the core path.

The Reference

  1. 1

    The production lens — what data engineers own, how pipelines differ from apps, and batch vs streaming mental models.

    1. 1.1The Data Engineering Lens2/530m
    2. 1.2Batch, Streaming & Architecture Patterns3/530m
  2. 2

    From MapReduce to DAG engines — how cluster compute works before you tune Spark.

    1. 2.1MapReduce to DAG Execution3/530m
    2. 2.2Shuffle, Skew & Partitioning4/530m
  3. 3

    The default batch engine — architecture, Spark SQL, and production tuning at Databricks/Google scale.

    1. 3.1Spark Architecture & Job Execution4/530m
    2. 3.2Spark SQL, Catalyst & Tuning4/530m
  4. 4

    Event time, windows, and watermarks — plus Spark Structured Streaming and Flink for production.

    1. 4.1Event Time, Windows & Watermarks4/530m
    2. 4.2Spark Structured Streaming & Flink4/530m
  5. 5

    Airflow and beyond — DAG design, dependencies, backfills, and operating scheduled pipelines.

    1. 5.1Airflow: DAG Design & Operations3/530m
    2. 5.2SLAs, Backfills & Pipeline Idempotency4/525m
  6. 6

    Medallion pipelines, dimensional modeling for gold marts, and dbt as the transform layer.

    1. 6.1Medallion & Incremental Processing4/530m
    2. 6.2Dimensional & Analytical Modeling3/535m
    3. 6.3Conformed Dimensions & Mart Design4/530m
    4. 6.4dbt & the Transform Layer3/530m
  7. 7

    Contracts, lineage, catalogs, and schema evolution — trust at scale.

    1. 7.1Data Contracts, Testing & Lineage4/530m
    2. 7.2Schema Evolution & Metadata Catalogs3/525m
  8. 8

    Reliability semantics, failure recovery, cost, and observability for data platforms.

    1. 8.1Delivery Semantics & Failure Recovery4/530m
    2. 8.2Cost Optimization & Pipeline Observability3/525m
  9. 9

    Optional hard-core material for data platform and streaming-heavy roles — Kafka internals, CDC, Beam, Flink state, data mesh, and warehouse engines.

    1. 9.1Apache Kafka for Data Pipelines5/530m
    2. 9.2Change Data Capture & Debezium4/525m
    3. 9.3Apache Beam & the Unified Model5/525m
    4. 9.4Flink State, Checkpoints & Exactly-Once5/530m
    5. 9.5Data Mesh & Platform Boundaries4/525m
    6. 9.6Warehouse Query Engines5/530m