Kubernetes Consulting
an architecture assessment of your cluster — We get your cluster into a production-ready state and hand it to your team.
Consulting and delivery for teams already running Kubernetes who find the cluster setting the agenda instead of carrying it. We get the environment into a production-ready state and hand it over.
We work on your cluster, in your repositories, through your reviews. The engagement ends when your team runs an upgrade without us — not when an architecture diagram is signed off.
BOOK A FIT CHECKWhy clusters stay half-finished
Kubernetes is rarely adopted wrongly. It is adopted and then never finished: the cluster runs, the applications run, and everything around them was assembled by hand.
What is missing only shows up in operations. A certificate expires and nobody knows who renews it. A node fails and the recovery lives in a chat thread. An upgrade comes due and the answer is a date next quarter.
That state is expensive and quiet at the same time. It rarely costs downtime in one visible block. It costs attention: every change needs the one person who still remembers how it was meant to work.
How it shows up in operations
Six situations we meet again and again in existing Kubernetes environments. The more of them fit, the more the cluster has stopped being a tool and become a project of its own.
- A cluster upgrade has been postponed for several quarters because nobody can say which workloads will stop starting afterwards.
- The cluster was set up by one person who has moved on to other work. It now gets changed only in an emergency.
- What runs in the cluster is not what any repository says. To learn the current state you ask kubectl, not Git.
- The dashboards exist, but nobody opens them during an incident. Alerts go to a channel everyone has muted.
- A new team needs onto the cluster. What that takes is known only to the person setting it up by hand again.
- Access rights grew with the years. Who is allowed to deploy to production is not a question anyone can answer in one sentence.
What “production-ready” means here
On this page “production-ready” is a list, not a promise. We call a cluster production-ready when these eight statements are true — each of them checkable together with your team.
- Upgrading the control plane and the nodes is planned work with a runbook and a way back, not a risk that moves to the next quarter.
- The cluster's desired state lives in Git. What is written there runs, and what runs is written there.
- Losing a node and losing a namespace have both been rehearsed, not only described.
- Access hangs off roles rather than people. Who may do what in production is written down and can be verified.
- Metrics, logs and traces answer the question of what is broken right now, without anyone connecting to a node.
- Every alert has a recipient and a runbook. Alerts with neither are switched off rather than muted.
- A new workload gets its namespace, limits, access and delivery path through a documented procedure.
- At least two people on your team carry out every one of these without us.
This list is our standard, not an industry one. In the fit check we say which points are realistically reachable in your environment and which we consider too expensive to chase.
If you would rather work these eight through against your own cluster before talking to us: the checklist resolves them into twenty-four statements, with a first step beside every one that stays open. Open the Kubernetes checklist
What is included
Eight blocks of work. Which of them an engagement contains depends on the state of your environment — it is almost never all eight.
Platform rescue
When the cluster is actively doing damage, stabilising comes before rebuilding. We take out the acute failure sources and restore a state in which planned work is possible again.
Cluster architecture
How clusters and environments are cut, plus networking, ingress, storage and resource limits — decided along your workloads rather than along a reference picture.
Security
RBAC, network policies, secrets, image provenance and admission rules. The technical work towards ISO 27001 or NIS2 is part of it; the certificate is issued by an accredited body, not by us.
GitOps
The desired state lives in Git and is reconciled from there. Infrastructure changes go through pull requests, exactly like changes to application code.
Observability
Metrics, logs and traces in one place, wired to exactly the alerts somebody actually responds to. The rest is switched off rather than muted.
Operations model
Who owns what, who gets woken at night, what the runbook says. We write it down with you and rehearse it once, instead of asserting it.
Workload onboarding
The documented path a new service takes into the cluster: namespace, limits, access, delivery path. So that the next workload is not another special case.
Knowledge transfer
Pair programming on your own changes, documentation in the repository, and a handover where your team performs the procedures and we watch.
How an engagement runs
Four steps, in this order. You can stop after the second one and still keep a result you can work from.
Fit check
A conversation about the state of your cluster — and about whether Kubernetes is the right choice for your case at all. If it is not, you hear it in that call.
Cluster review
We read the manifests, clusters and pipelines ourselves and write down which of the eight statements above hold, which do not, and what the gap costs in operations.
Delivery inside the team
Every change goes through a pull request somebody on your team reads. We work in your repositories, not alongside them.
Handover
Your team runs an upgrade and a recovery case while we watch. Plus a list of the things we deliberately did not build — with the reasoning.
Engagement models
Four ways into the work. Which one fits is decided in the fit check; the pages behind them state scope, duration and price in full.
Analyse first
When it is not yet clear where the constraint sits: a fixed-scope analysis that ends in a prioritised roadmap.
Stabilise first
When the cluster is actively doing damage and things have to calm down before anything is rebuilt.
Rebuild it properly
When a grown environment should become a platform your team runs on its own — with a fixed start and end date.
Ongoing engagement
Several months with a fixed number of days per week, where operations needs an experienced hand for longer. Scope and term are set case by case.
What you can check
Plainly: there is no figure from a client engagement on this page. The clusters we built before were never measured in a way that would prove anything here, and we will not invent numbers to fill the gap.
A number belongs in this space. There is none: no earlier engagement measured its effect in a way that would prove anything here. As soon as a client releases a measurement, it goes here — not before.
References
What we have worked on, described without metrics, because we never collected any.
Delivery model
Who does the work, who reviews it, and where delivery accountability stays.
Articles
How we think about clusters, operations and handovers — public, and at full length.
Relevant case studies
Case studies for this work, with context, intervention and handover. Each one appears only once the client has released it in writing.
No case study for this work is published yet. Until then, the references index says what we have worked on.
When Kubernetes is the wrong choice
Kubernetes answers operational problems that plenty of teams simply do not have. In these four cases we advise against it — including when the request already arrives phrased as a Kubernetes project.
One service, one team, modest load
A single application with quiet traffic runs cheaper on a managed platform and with less to operate. Kubernetes mostly adds a layer that somebody then has to maintain.
Nobody operates it after the project
A cluster needs continuous attention: upgrades, certificates, capacity. If nobody owns that afterwards, a managed offering is the more honest answer — and you hear that before the proposal.
The constraint is not operational
When releases wait on approvals, coordination or manual testing, a cluster changes none of it. Kubernetes moves those constraints to a new place rather than removing them.
A legacy system is meant to become maintainable
Containerising an application does not make it easier to cut apart. An architecture problem is solved in the code; Kubernetes only repackages it.
If one of these applies, we say so in the fit check and name the alternative. A Kubernetes project nobody needs is a bad project for us too.
What is not included
Four requests that reach us regularly around Kubernetes and that we turn down. They are here so that nobody first meets them in a proposal.
Round-the-clock operations, on-call cover and SLA-backed run responsibility
We build platforms and hand them over ready to run. Running them is a different business — a managed-service provider can commit to that, we cannot.
Certificates, audit sign-off and conformity statements
Engineering work towards ISO 27001, NIS2 or the EU AI Act is in scope. Issuing the attestation is not: that takes an accredited auditor, not an engineer.
Projects with no in-house engineering team
Knowledge transfer is part of every engagement. With no team to hand over to it turns into permanent outsourcing — a model we do not offer.
A fixed price for a whole transformation, quoted before anyone has looked
Pricing an unexamined legacy estate is guesswork. We estimate after a short assessment — and we will say so when the rebuild is not worth it.
Questions we get about cluster operations
Will you operate the cluster for us afterwards?
No. We get the environment into a state your team can run, and hand it over. On-call duty and SLA-bound operational accountability are what a managed service provider offers; we do not.
We are on a managed Kubernetes service. Does this still apply?
Yes. EKS, AKS, GKE or a self-managed install change who operates the control plane, not the eight statements above. Whatever your provider already covers is recorded as covered in the result.
Do we have to rebuild everything?
No. In most engagements a cluster, some manifests and a pipeline already exist. We carry on with them and rebuild only what we can give a reason for rebuilding.
How do you work on an environment that is live?
Changes take the same path as your own: pull request, review, rollout in a pre-production stage. For the interventions that cannot be rehearsed first, the way back is agreed beforehand.
What if the review concludes that Kubernetes is the wrong choice?
Then that is the result you get, with the alternative next to it. The section above is not rhetorical: we have turned down requests because a managed platform was the cheaper answer.
How soon could you start?
We run only a few engagements in parallel, so availability is confirmed before a proposal is written. The fit check gives you a realistic start date, even when it is later than you hoped.
Is your cluster production-ready?
A conversation about the state of your environment — and an honest answer on whether Kubernetes is the right choice for your case.