← Every path

H2-CSPE · Core certification path

Secure Platform Engineering

H2 Certified Secure Platform Engineer

Build the platform other teams ship on, hardened from the metal up: an immutable OS, a policy-enforced network, a service mesh that survives a gateway failure, stateful data, a streaming backbone and autoscaling, run as an internal product.

6 courses · 27 lessons · ~46 h of guided work

Assumes Foundation

What you leave with

A multi-node platform on Talos with Cilium, a Kuma mesh, CNPG and Kafka, autoscaling on real signals, an internal developer portal, and a gateway you have broken and repaired.

OS
Talos
CNI / policy
Cilium
Mesh
Kuma multizone
Storage
Longhorn
Database
CloudNativePG
Streaming
Kafka
Autoscaling
KEDA

Syllabus

6 courses · every lesson graded · minutes are guided work

  1. Course 01

    Immutable OS and cluster hardening

    A Kubernetes cluster on an OS with no shell and no package manager, managed entirely by API, with the control plane and etcd hardened.

    5 lessons · ~9 h

    1. 01Why immutableReproduce a drifted mutable node, then the same workload on Talos where drift is impossible.90 min
    2. 02Machine configThe whole cluster as declarative config in git; a change applied by a rolling reboot.120 min
    3. 03Control plane and etcdetcd encryption, backups, and a control-plane node lost and rebuilt from config.120 min
    4. 04Benchmarking the clusterA CIS-style scan, the findings, and the config change that clears each one.90 min
    5. 05Upgrades without downtimeA Kubernetes and Talos upgrade, one node at a time, watched.90 min

    Course checkBring up a three-node cluster from machine config, pass a CIS-style benchmark, and recover a lost control-plane node.

    You leave withA Talos cluster you brought up from declarative config, benchmarked, and healed after a node loss.

  2. Course 02

    Cluster networking and policy

    Cilium as the data plane: identity-based network policy, default-deny, and the visibility to prove what is talking to what.

    5 lessons · ~8 h

    1. 01The Cilium data planeInstall Cilium with KubePrism, kubernetes IPAM, and the values that avoid the VPC-CIDR trap.90 min
    2. 02Identity-based policyNetwork policy keyed on identity, not IP; a rule that survives a pod reschedule.120 min
    3. 03Default-denyFlip the namespace to default-deny and restore only what the app needs, flow by flow.90 min
    4. 04Visibility with HubbleWatch flows live, find the one that should not exist, and write the policy that stops it.90 min
    5. 05Egress controlAn egress allow-list to named destinations; the metadata-endpoint trap and how to close it.90 min

    Course checkGiven a multi-service app, write policy so only the required flows work, prove default-deny, and produce the flow evidence.

    You leave withA default-deny cluster where every allowed flow is named in policy, with Hubble showing the rest is blocked.

  3. Course 03

    Service mesh and gateways

    A multizone Kuma mesh: mTLS between services, gateway routes, and the failure modes that produce a 500 you have to read backwards.

    5 lessons · ~9 h

    1. 01Why a meshThe problems a mesh solves and the ones it adds; when not to reach for one.90 min
    2. 02mTLS everywhereAutomatic mTLS between services; a sidecar opt-out by label and why you might.90 min
    3. 03MultizoneGlobal and zone control planes; a service in one zone consumed from another.120 min
    4. 04Gateways and routesRoutes on the global CP referencing zone MeshServices by label; the label mistake that 500s.120 min
    5. 05Debugging the meshThe MeshGatewayInstance tag trap, the TLS-secret format, and reading a 500 to its cause.90 min

    Course checkA mesh with a broken gateway route returning 500s. Diagnose from mesh telemetry and fix without downtime.

    You leave withA multizone mesh with mTLS everywhere, gateway routes to MeshServices, and a documented fix for the unresolved-backend 500.

  4. Course 04

    Storage and databases

    Stateful workloads done carefully: block storage, PostgreSQL on CNPG with real recovery, and the volume traps that eat a cluster.

    4 lessons · ~7 h

    1. 01Block storageLonghorn with replicas; the control-plane-node trap and the Released-PV capacity leak, both reproduced.120 min
    2. 02Postgres on KubernetesCNPG with primary and replicas, connection pooling, and failover you triggered.120 min
    3. 03Backup and recoverybarmanObjectStore backups and a full restore, including the join-deadlock fix.120 min
    4. 04Data on the meshThe database reachable only through policy; no public path, proven.60 min

    Course checkRecover a CNPG cluster from object storage after a simulated data loss, resolving the join deadlock cleanly.

    You leave withA CNPG cluster with barman backups you restored, and a Longhorn setup that survived a control-plane node loss.

  5. Course 05

    Streaming platform operations

    Run Kafka as shared infrastructure: multi-AZ, secured with mTLS and ACLs, quota'd per tenant, and mirrored across regions.

    4 lessons · ~7 h

    1. 01Kafka as a platformMulti-broker across zones; rack awareness; what a broker loss does and does not cost.120 min
    2. 02Securing the clustermTLS between brokers and clients, SASL, and ACLs scoped per topic and per tenant.120 min
    3. 03Quotas and tenancyProducer and consumer quotas so one tenant cannot starve the rest; proven with a noisy client.90 min
    4. 04Mirroring and DRCross-region mirroring, offset translation, and a failover you rehearsed.90 min

    Course checkOperate a multi-AZ Kafka under a broker failure and a noisy tenant; keep every other tenant inside its quota.

    You leave withA multi-AZ Kafka with per-tenant ACLs and quotas, cross-region mirroring, and a broker-loss drill you ran.

  6. Course 06

    Autoscaling and platform as a product

    Scale on the signals that matter, not just CPU, and wrap the whole platform in a self-service portal so teams ship without filing tickets.

    4 lessons · ~7 h

    1. 01Horizontal and vertical scalingHPA and VPA with limits that make sense; the metrics that lie and the ones that do not.90 min
    2. 02Event-driven scalingKEDA scaling on Kafka lag and queue depth; scale-to-zero and the cold-start cost.120 min
    3. 03The internal developer platformA portal, a golden path, and a template a team uses to ship without touching the cluster.120 min
    4. 04Platform as a productTreating the platform's users as customers: SLOs, docs, and a deprecation you communicated.90 min

    Course checkGiven a bursty workload, configure scaling that holds latency without over-provisioning, and expose one self-service action through the portal.

    You leave withKEDA scaling your service on a queue depth, a golden path in an internal portal, and a paved road a new team used without asking you.

Certification course

H2 Certified Secure Platform Engineer

$499 · one price · courses + 90-day labs + exam

Stand up the whole platform, then survive a chaos day: a node lost, a mesh gateway returning 500s, a storage volume faulted and a traffic spike, in one window. Graded on recovery, policy held, and no data lost.

Opens when Foundation is complete and the 6 course checks are passed. One proctored attempt, plus a free retake if you fail by a margin. The credential is an Open Badges 3.0 credential, signed and verifiable.

Counts toward H2-CTSE. Certified T-Shaped Security Expert is the credential of the whole T: hold all seven core credentials and it is awarded automatically, free, with no extra exam.

Create an account to enrol