H2-CSPE · Core certification path
Secure Platform Engineering
H2 Certified Secure Platform Engineer
Build the platform other teams ship on, hardened from the metal up: an immutable OS, a policy-enforced network, a service mesh that survives a gateway failure, stateful data, a streaming backbone and autoscaling, run as an internal product.
6 courses · 27 lessons · ~46 h of guided work
Assumes Foundation
What you leave with
A multi-node platform on Talos with Cilium, a Kuma mesh, CNPG and Kafka, autoscaling on real signals, an internal developer portal, and a gateway you have broken and repaired.
- OS
- Talos
- CNI / policy
- Cilium
- Mesh
- Kuma multizone
- Storage
- Longhorn
- Database
- CloudNativePG
- Streaming
- Kafka
- Autoscaling
- KEDA
Syllabus
6 courses · every lesson graded · minutes are guided work
Course 01
Immutable OS and cluster hardening
A Kubernetes cluster on an OS with no shell and no package manager, managed entirely by API, with the control plane and etcd hardened.
5 lessons · ~9 h
- 01Why immutableReproduce a drifted mutable node, then the same workload on Talos where drift is impossible.90 min
- 02Machine configThe whole cluster as declarative config in git; a change applied by a rolling reboot.120 min
- 03Control plane and etcdetcd encryption, backups, and a control-plane node lost and rebuilt from config.120 min
- 04Benchmarking the clusterA CIS-style scan, the findings, and the config change that clears each one.90 min
- 05Upgrades without downtimeA Kubernetes and Talos upgrade, one node at a time, watched.90 min
Course checkBring up a three-node cluster from machine config, pass a CIS-style benchmark, and recover a lost control-plane node.
You leave withA Talos cluster you brought up from declarative config, benchmarked, and healed after a node loss.
Course 02
Cluster networking and policy
Cilium as the data plane: identity-based network policy, default-deny, and the visibility to prove what is talking to what.
5 lessons · ~8 h
- 01The Cilium data planeInstall Cilium with KubePrism, kubernetes IPAM, and the values that avoid the VPC-CIDR trap.90 min
- 02Identity-based policyNetwork policy keyed on identity, not IP; a rule that survives a pod reschedule.120 min
- 03Default-denyFlip the namespace to default-deny and restore only what the app needs, flow by flow.90 min
- 04Visibility with HubbleWatch flows live, find the one that should not exist, and write the policy that stops it.90 min
- 05Egress controlAn egress allow-list to named destinations; the metadata-endpoint trap and how to close it.90 min
Course checkGiven a multi-service app, write policy so only the required flows work, prove default-deny, and produce the flow evidence.
You leave withA default-deny cluster where every allowed flow is named in policy, with Hubble showing the rest is blocked.
Course 03
Service mesh and gateways
A multizone Kuma mesh: mTLS between services, gateway routes, and the failure modes that produce a 500 you have to read backwards.
5 lessons · ~9 h
- 01Why a meshThe problems a mesh solves and the ones it adds; when not to reach for one.90 min
- 02mTLS everywhereAutomatic mTLS between services; a sidecar opt-out by label and why you might.90 min
- 03MultizoneGlobal and zone control planes; a service in one zone consumed from another.120 min
- 04Gateways and routesRoutes on the global CP referencing zone MeshServices by label; the label mistake that 500s.120 min
- 05Debugging the meshThe MeshGatewayInstance tag trap, the TLS-secret format, and reading a 500 to its cause.90 min
Course checkA mesh with a broken gateway route returning 500s. Diagnose from mesh telemetry and fix without downtime.
You leave withA multizone mesh with mTLS everywhere, gateway routes to MeshServices, and a documented fix for the unresolved-backend 500.
Course 04
Storage and databases
Stateful workloads done carefully: block storage, PostgreSQL on CNPG with real recovery, and the volume traps that eat a cluster.
4 lessons · ~7 h
- 01Block storageLonghorn with replicas; the control-plane-node trap and the Released-PV capacity leak, both reproduced.120 min
- 02Postgres on KubernetesCNPG with primary and replicas, connection pooling, and failover you triggered.120 min
- 03Backup and recoverybarmanObjectStore backups and a full restore, including the join-deadlock fix.120 min
- 04Data on the meshThe database reachable only through policy; no public path, proven.60 min
Course checkRecover a CNPG cluster from object storage after a simulated data loss, resolving the join deadlock cleanly.
You leave withA CNPG cluster with barman backups you restored, and a Longhorn setup that survived a control-plane node loss.
Course 05
Streaming platform operations
Run Kafka as shared infrastructure: multi-AZ, secured with mTLS and ACLs, quota'd per tenant, and mirrored across regions.
4 lessons · ~7 h
- 01Kafka as a platformMulti-broker across zones; rack awareness; what a broker loss does and does not cost.120 min
- 02Securing the clustermTLS between brokers and clients, SASL, and ACLs scoped per topic and per tenant.120 min
- 03Quotas and tenancyProducer and consumer quotas so one tenant cannot starve the rest; proven with a noisy client.90 min
- 04Mirroring and DRCross-region mirroring, offset translation, and a failover you rehearsed.90 min
Course checkOperate a multi-AZ Kafka under a broker failure and a noisy tenant; keep every other tenant inside its quota.
You leave withA multi-AZ Kafka with per-tenant ACLs and quotas, cross-region mirroring, and a broker-loss drill you ran.
Course 06
Autoscaling and platform as a product
Scale on the signals that matter, not just CPU, and wrap the whole platform in a self-service portal so teams ship without filing tickets.
4 lessons · ~7 h
- 01Horizontal and vertical scalingHPA and VPA with limits that make sense; the metrics that lie and the ones that do not.90 min
- 02Event-driven scalingKEDA scaling on Kafka lag and queue depth; scale-to-zero and the cold-start cost.120 min
- 03The internal developer platformA portal, a golden path, and a template a team uses to ship without touching the cluster.120 min
- 04Platform as a productTreating the platform's users as customers: SLOs, docs, and a deprecation you communicated.90 min
Course checkGiven a bursty workload, configure scaling that holds latency without over-provisioning, and expose one self-service action through the portal.
You leave withKEDA scaling your service on a queue depth, a golden path in an internal portal, and a paved road a new team used without asking you.
Certification course
H2 Certified Secure Platform Engineer
$499 · one price · courses + 90-day labs + exam
Stand up the whole platform, then survive a chaos day: a node lost, a mesh gateway returning 500s, a storage volume faulted and a traffic spike, in one window. Graded on recovery, policy held, and no data lost.
Opens when Foundation is complete and the 6 course checks are passed. One proctored attempt, plus a free retake if you fail by a margin. The credential is an Open Badges 3.0 credential, signed and verifiable.
Counts toward H2-CTSE. Certified T-Shaped Security Expert is the credential of the whole T: hold all seven core credentials and it is awarded automatically, free, with no extra exam.
Create an account to enrol