Production restarting in a loop after a version upgrade.
Back to home
Platform
Kubernetes. Get the cluster back under control.
An unstable cluster is fixed by measurement, not intuition. We instrument, we fix, we document.
What we take over
Diagnosis of recurring incidents: resources, probes, evictions, nodes.
Manifest and chart review, standardised deployments.
Autoscaling and capacity, node sizing backed by numbers.
CI/CD and GitOps pipeline, reproducible environments.
Observability: metrics, logs, alerts that fire on something real.
Typical missions
Manual deployments moved to a reproducible GitOps pipeline.
Seasonal load peak prepared under a constrained budget.
Something to unblock.
Describe the context, the blocker and the date. Same-day reply.
[email protected]