Kubernetes Platform Build
A cluster only we can run — A fixed-scope engagement for teams who need a platform that belongs to them afterwards.
Twelve fixed-scope weeks to turn a Kubernetes environment into a platform your own team runs. It ends with a list of conditions that are either met or not met on the final day — not with a recommendation.
The build is neither a sprint on an existing platform nor an analysis. It builds the platform your applications will run on, and it is not finished when the clusters are up. It is finished when someone on your team can run them without us.
When a build is the right step
- A move to the cloud has been decided, and nobody in the house has built a platform someone else has to run afterwards.
- A cluster appeared at some point, carries production now, and nobody dares rebuild it.
- Every team has built its own route to production, and no two of them look alike.
- Kubernetes is a given — from the group, from a customer, or from a regulator — and you do not want to start from nothing.
- Operations depend on one person who holds it all in their head. You want that resolved before they leave.
- Access is granted by hand, secrets sit in repositories in clear text, and nobody can say who can reach production today.
- You want the platform finished before the first application is migrated, not built alongside the migration.
When it is not
- You already have a platform and only delivery has seized up. The Platform Recovery Sprint is the shorter route.
- You run one application with one team. A managed runtime at your cloud provider is the more honest suggestion than a cluster of your own.
- You do not yet know whether Kubernetes is the right tool for you. An analysis settles that in days; a build takes weeks.
- You are looking for someone to run the platform permanently. We hand over at the end — the whole shape of this is built around that.
- There is nobody to take the platform over after the handover. Then we build something that belongs to no one after twelve weeks.
If the platform already exists and only delivery has seized up, the shorter route is described elsewhere: See the recovery sprint
Duration and scope
- Duration
- 12 weeks with a fixed start and end date, in four sections of three weeks each. Inside them sit 36 days of our time, spread across three days a week. The rest of the time belongs to your team: they decide, review and take over.
- Scope
- Exactly one platform gets built: 2 environments — one below production and production itself — and 2 workloads from your teams running on it. We onboard the first together; your team onboards the second on its own, following the written path. What is in scope is listed below block by block, and what is not is stated just as plainly next to it.
What gets built — and how you check it
"Production-ready" is not a property here, it is a list: 10 conditions, one per building block. Each is true or false on the final day, and each one is demonstrated, rehearsed or measured — none of them counts because a document says so. 9 of the 10 are checked by your team, not by us.
Cluster baseline
demonstratedwe check togetherchecked in Section 1 · weeks 1–3
Result: Both environments entirely as code: network, node pools, ingress, certificates, backups, cluster upgrades. No click anybody can fail to reconstruct.
Production-ready when: A cluster is rebuilt once during the build, from the repository into an empty account — demonstrated together, not asserted.
Identity and access
demonstratedyour team checkschecked in Section 1 · weeks 1–3
Result: Access to clusters and platform tooling through your identity provider: groups rather than personal exceptions, roles per environment, one documented break-glass path.
Production-ready when: Removing someone from a group in the identity provider removes their cluster access. Demonstrated once, on a real person, by you.
GitOps
demonstratedyour team checkschecked in Section 1 · weeks 1–3
Result: One repository the state of both environments is reconciled from, with mandatory review, a readable history, and a way back to the previous state.
Production-ready when: A change made by hand on the cluster is reverted, and a merged change arrives. You demonstrate both of them once, yourselves.
Secret management
rehearsedyour team checkschecked in Section 2 · weeks 4–6
Result: A path for secrets from their source into the pod, with no clear text in the repository, and a rotation procedure that is written down.
Production-ready when: Your team rotates a secret during the build while the workloads keep running — following the procedure, with us away from the keyboard.
Observability
measuredyour team checkschecked in Section 2 · weeks 4–6
Result: Logs, metrics and alerts for the clusters and for the onboarded workloads, with dashboards a named person is responsible for.
Production-ready when: A test alert reaches the named on-call person through the agreed route, and that workload's metrics are findable in the same session.
Security controls
demonstratedyour team checkschecked in Section 2 · weeks 4–6
Result: Baseline rules in the cluster: separation by namespace, network rules, pod security standards, image provenance, regular base-image updates.
Production-ready when: A deployment that violates the baseline rules is rejected by the cluster, not by a review. Shown with a deliberately broken manifest.
Workload onboarding
demonstratedyour team checkschecked in Section 3 · weeks 7–9
Result: The route every workload takes afterwards: repository template, pipeline, deployment manifests, alerts — plus the written path a team can follow without us.
Production-ready when: Your team onboards the second workload along that written path. We watch and do not intervene; whatever is missing gets fixed in the path, not by us.
Runbooks
rehearsedyour team checkschecked in Section 3 · weeks 7–9
Result: Runbooks for what will actually hit you: a node fails, a certificate expires, a deployment has to go back, a cluster upgrade, a restore from backup.
Production-ready when: Every runbook is executed once by your team during the build, the restore from backup against a real environment.
Training
rehearsedyour team checkschecked in Section 4 · weeks 10–12
Result: Four sessions for the people who will run the platform: cluster basics, the route a change takes, alerts and on-call, upgrade and restore — on your platform, not on examples.
Production-ready when: Everyone who will be on call has worked through the exercises on your platform themselves. Watching does not count.
Ownership transfer
demonstratedyour team checkschecked in Section 4 · weeks 10–12
Result: An ownership map: every building block with a named person, a deputy and an escalation path, confirmed by exactly those people — plus the written remaining backlog with the reasoning.
Production-ready when: Our access is revoked on the final day and the platform keeps running. Every building block has a person who confirmed the assignment.
What is not included
This list is as binding as the one above. Every entry says what happens instead, so you know what is still open after the build.
| Not included | Instead |
|---|---|
| Migrating or rewriting your applications. | Two workloads move onto the platform during the build. Your teams take the rest along the same route — accompanied if you want, but as a separate engagement with its own scope. |
| Running the platform after the handover, on-call included. | You run it — that is what the training, the runbooks and the ownership map are for. If interim technical leadership is what you are missing, that is a separate arrangement. |
| Further clusters, further regions, and a disaster-recovery design beyond backup and restore. | Two environments and one rehearsed restore from backup are included. Anything beyond that is a second stage you can decide on after the build. |
| Running databases, message brokers and data warehouses on the cluster. | The onboarded workloads use your cloud provider's managed services. Stateful systems on the cluster are a separate engagement with their own operational cost. |
| A certification, or a passed audit. | We build the controls and document them so an auditor can examine them. The certification itself is issued by an audit body, not by us. |
| The cost of cloud, licences and managed services. | They run through your account and stay visible there. During the build we choose the options with you, so you know what running it will cost afterwards. |
| A bespoke developer portal, and replacing your CI system. | The route to production is provided through templates and a written path. We connect your existing CI system rather than replacing it. |
What you bring
Anything marked "before the start" has to exist before we begin. Arriving later moves the start date — it does not shorten the build.
- before the start
A cloud account that belongs to you, with a spending limit you own, and access for us from day one.
- before the start
An existing identity provider — Entra ID, Google Workspace, Okta or equivalent — and someone allowed to create groups in it.
- before the start
One named person who owns the platform after the handover and works inside the build from day one. Without them we do not start.
- before the start
Decisions on network and names: address ranges, connectivity to your existing systems, and DNS zones you can administer yourselves.
- before the start
Roughly one day a week per involved team member, for decisions, reviews and the exercises. Without that time the result stays with us.
- by section 2
Your security and compliance requirements in writing — or a person allowed to decide them during the second section.
- by section 3
Two workloads, named at scoping, with teams who have time in the third section. The second one is onboarded by that team itself.
- by section 4
A decision on who is on call after the handover. It does not have to exist on day one, but it does before the fourth section.
The twelve weeks
- Section 1 · weeks 1–3
Access, the open decisions, cluster baseline, identity and GitOps. By the end the platform exists as code, and a change only gets in through a pull request.
- Section 2 · weeks 4–6
Secrets, observability and the security controls. By the end what happens on the platform is visible, and what violates the baseline rules no longer gets through.
- Section 3 · weeks 7–9
The first workload together, the second by your team alone — and the runbooks, each one rehearsed rather than only written.
- Section 4 · weeks 10–12
Training, ownership, and the walk through every condition, one after another. On the final day our access is revoked.
Price
A fixed price for a fixed scope, not a day rate: a build billed by time spent no longer has a fixed scope. We do not publish a price list — the price comes with the proposal, together with the scope, before you commit to anything. Cloud and licence costs run separately, through your account.
What comes next
Afterwards there are three honest options: you carry on alone — the normal case, and the purpose of the handover; you onboard further workloads and we accompany the first few; or you bring in interim technical leadership. None of them is part of the build, and none of them is a precondition for it working.
Plainly: there is no figure from an earlier platform build on this page. The offer is newly scoped, and what we built before was never measured in a way that would prove anything here. What we have worked on is in the references — without metrics, because we did not collect them.
Not sure yet whether this is the right engagement? Settle it in a quarter of an hour, with the engineer who does the work. Book a fit check