← Every path

H2-CSPD · Specialist path

Streaming and Data Platforms

H2 Certified Streaming Platform Engineer

Kafka in depth, as the backbone a whole company runs on: internals and tuning, change data capture, schema governance, stream processing, exactly-once at scale, security, and multi-region disaster recovery, with PII kept out of the stream.

7 courses · 23 lessons · ~42 h of guided work

Assumes Foundation + T-Shaped core; CSPE recommended

What you leave with

A production-grade streaming platform: a tuned multi-region Kafka, CDC pipelines, a governed schema registry, stream-processing jobs, exactly-once delivery, mTLS and ACLs, and a rehearsed regional failover.

Bus
Kafka
CDC
Debezium
Schemas
Schema Registry
Processing
Kafka Streams / Flink
Security
mTLS / SASL / ACLs

Syllabus

7 courses · every lesson graded · minutes are guided work

  1. Course 01

    Kafka internals and tuning

    What actually happens inside a broker, and the tuning that decides whether the platform holds under load.

    4 lessons · ~8 h

    Tools Kafka, JMX metrics, a load generator

    1. 01The log and the brokerPartitions, segments, the page cache, and replication; where latency and durability trade off.120 min
    2. 02Producer and consumer internalsBatching, acks, fetch sizing, and the settings that quietly lose data.120 min
    3. 03Tuning for the workloadThroughput versus latency versus durability; a cluster tuned to a real SLA.120 min
    4. 04Capacity and partitionsPartition count, rebalancing cost, and the sizing mistake that is hard to undo.90 min

    Course checkGiven a workload and an SLA, tune a cluster to meet it, and prove where the previous configuration would have failed.

    You leave withA tuned Kafka cluster meeting a throughput-and-latency SLA, with the reasoning for every non-default setting written down.

  2. Course 02

    Connect and change data capture

    Get data in and out reliably. Kafka Connect, Debezium CDC from databases, and the outbox pattern for events you cannot afford to lose.

    3 lessons · ~6 h

    Tools Kafka Connect, Debezium, PostgreSQL

    1. 01Kafka ConnectSource and sink connectors; the failure and restart semantics.120 min
    2. 02CDC with DebeziumStreaming Postgres changes; snapshot and streaming phases; the gotchas.120 min
    3. 03The outbox patternEvents guaranteed alongside a database transaction; no dual-write bug.90 min

    Course checkStream a database's changes into topics with no lost or duplicated event across a source restart.

    You leave withCDC pipelines from PostgreSQL into Kafka via Debezium, with the outbox pattern for guaranteed events and a source-restart drill.

  3. Course 03

    Schema governance

    A schema registry that lets a hundred teams evolve without breaking each other. Compatibility, evolution, and governance as policy.

    3 lessons · ~5 h

    Tools Schema Registry, Avro/Protobuf, CI checks

    1. 01Schemas on the busWhy schemas; the cost of their absence at scale.90 min
    2. 02Compatibility and evolutionBackward, forward and full compatibility; a change that is safe and one that is not.120 min
    3. 03GovernanceRegistry rules enforced in CI; ownership and review of schema changes.90 min

    Course checkDeploy a compatible schema change and have an incompatible one rejected; a consumer keeps working through both.

    You leave withA governed schema registry with enforced compatibility, a safe evolution, and an unsafe one rejected before it ships.

  4. Course 04

    Stream processing

    Compute on the stream itself: aggregations, joins and windows with Kafka Streams and Flink, and the state and time problems that make it hard.

    4 lessons · ~7 h

    Tools Kafka Streams, Flink

    1. 01Stateless and statefulMaps and filters versus aggregations and joins; where state lives.120 min
    2. 02Time and windowsEvent time versus processing time; watermarks and late data.120 min
    3. 03JoinsStream-stream and stream-table joins; the correctness traps.90 min
    4. 04State and recoveryManaging state stores; recovery after a crash with no wrong result.90 min

    Course checkBuild a windowed aggregation with a join that produces correct results under out-of-order and late-arriving events.

    You leave withStream-processing jobs with stateful aggregations, joins and windows, correct under late and out-of-order data.

  5. Course 05

    Exactly-once at scale

    The hardest promise in streaming, made real. Idempotent producers, transactions, and consumers that survive a replay without duplicating a side effect.

    3 lessons · ~6 h

    Tools Kafka transactions, idempotent consumer patterns

    1. 01The delivery guaranteesAt-most, at-least, exactly-once; what each really costs.90 min
    2. 02Idempotence and transactionsIdempotent producers and transactional writes; the read-process-write loop done right.120 min
    3. 03Idempotent consumersSide effects that survive a replay; the dedup key and where to keep it.120 min

    Course checkA pipeline must produce exactly-once results through a broker restart, a consumer crash and a deliberate duplicate.

    You leave withAn exactly-once pipeline: idempotent producers, transactional writes, and idempotent consumers proven against replays and crashes.

  6. Course 06

    Securing the platform

    A shared bus is a shared attack surface. mTLS, authentication, per-tenant authorization, encryption, and keeping PII out of the stream.

    3 lessons · ~6 h

    Tools mTLS, SASL, ACLs, field encryption

    1. 01Transport and authenticationmTLS between brokers and clients; SASL for identity.120 min
    2. 02AuthorizationACLs per topic and per tenant; the noisy-neighbor and read-everything problems.120 min
    3. 03PII in streamsEncryption or tokenization of sensitive fields; a scan that fails on plaintext PII.90 min

    Course checkLock down a cluster so one tenant cannot read another's topics, and prove no PII is written to the stream in the clear.

    You leave withA secured Kafka with mTLS, SASL, per-topic ACLs, encryption, and a check that keeps PII out of the stream.

  7. Course 07

    Multi-region and disaster recovery

    Keep the backbone up when a region is not. Mirroring, offset translation, and a regional failover you have actually rehearsed.

    3 lessons · ~6 h

    Tools MirrorMaker / cluster linking, multi-region topology

    1. 01Multi-region topologiesActive-active versus active-passive; the trade-offs and the costs.120 min
    2. 02Mirroring and offsetsCross-region replication and offset translation so consumers resume correctly.120 min
    3. 03Failover rehearsedA regional failover you executed; RPO measured, duplicates checked.90 min

    Course checkFail the platform over to a second region with bounded data loss and no duplicate delivery, live.

    You leave withA multi-region Kafka with mirroring, offset translation, and a rehearsed regional failover with measured RPO.

Certification course

H2 Certified Streaming Platform Engineer

$799 · one price · courses + 90-day labs + exam

48-hour practical: run the platform through a broker loss, a schema-breaking deploy, a poison message, a consumer-group rebalance storm and a regional failover. Graded on no data lost, no duplicate delivered, and every tenant inside its quota.

Opens when Foundation is complete and the 7 course checks are passed. One proctored attempt, plus a free retake if you fail by a margin. The credential is an Open Badges 3.0 credential, signed and verifiable.

Create an account to enrol