Envion Software
CS-078DevOps ConsultingB2B Analytics (NDA)

Platform Simplification: We Deleted Most of It

A B2B analytics company — 12 engineers, 3 on infrastructure — was spending more time on its platform than its product: three Kubernetes clusters, a service mesh, self-hosted GitLab, Prometheus, Grafana, Loki and Vault, a PostgreSQL operator, and Kafka on Kubernetes for 200 messages a second, serving 40 requests per second at peak. Cloud spend was €47K/month against €340K revenue, and two of three infrastructure engineers had resigned in six months. Envion’s review found 61% of infrastructure effort went to maintaining the platform itself and 22 of 34 production incidents originated in platform components. The recommendation was removal: nine months later cloud spend is €19K/month, incidents down from 34 to 9, deploy duration 22 → 5 minutes, deploys up from 6 to 40+ a week, and the remaining infrastructure engineer works half-time on product.

Platform Simplification: We Deleted Most of It
01

The challenge

Deployment had become slow and frightening. The CTO's description: "we spend more time on our platform than on our product, and nobody is confident about any of it." Cloud spend was €47K a month against €340K monthly revenue, which their CFO had flagged. Two of three infrastructure engineers had resigned in six months.

They expected Envion to tell them what to add. Better GitOps, a service catalogue, tighter policy enforcement.

02

Decision path

What the review found: twelve engineers, three Kubernetes clusters, a service mesh, self-hosted GitLab, self-hosted Prometheus and Grafana, self-hosted Loki, self-hosted Vault, a self-managed PostgreSQL operator with a custom backup CronJob, ArgoCD, a custom Helm chart library, and Kafka running on Kubernetes for a workload of roughly 200 messages per second. Nine microservices. Total production traffic: about 40 requests per second at peak.

None of this was stupid. Every piece had been added by a competent engineer for a reason that made sense at the time — usually "we might need this at scale" or "this is what serious platforms do." What nobody had done was step back and ask what a twelve-person company actually needs.

The measurements: 61% of infrastructure effort was maintaining the platform itself — cluster upgrades, mesh certificate rotation, Vault unsealing, operator version churn, Kafka rebalancing. Of 34 production incidents in six months, 22 originated in platform components rather than application code; the service mesh alone accounted for seven, including two outages caused by mTLS certificate expiry. And €47K/month broke down as roughly €12K of actual application workload and €35K of platform overhead.

The knowledge problem: both departed engineers had built substantial parts of this. The remaining engineer could not confidently operate the Kafka setup or the PostgreSQL operator. That's not a resilience gap you can fix with documentation.

03

Envion contribution

The recommendation was removal. Service mesh: remove — seven incidents, zero requirements it was meeting at 40 rps. Three clusters: consolidate to one — namespace isolation was sufficient. Self-hosted Vault: managed secrets service — about €400/month versus days of operational load. PostgreSQL operator: managed database — backup, failover and patching become someone else's job. Kafka on Kubernetes: managed queue — 200 messages a second does not need self-operated Kafka. Self-hosted GitLab and logging: managed, with 7-day retention matching actual query behaviour. Nine microservices: consolidate to four — the boundaries reflected nothing and deployment coupling was total anyway.

Two things were kept deliberately: ArgoCD and the custom Helm library, which were working well with low overhead. The point isn't that managed services are always right — it's that every operational component has a carrying cost, and that cost has to be paid by the people you have, not the people you imagine having.

Envion also gave the counterfactual honestly: at roughly 10x current traffic, some of these decisions reverse. The client was told which ones, and what signal to watch for. Simplification isn't a permanent state — it's the right state for their current size, and they should expect to revisit it.

04

Delivery

Envion reviewed and planned; the client's team executed. The sequencing moved low-risk components first so the team built confidence before touching the pieces that mattered, and each removal was verified against actual usage before decommissioning — the observability stack's 400GB of daily logs, for example, had been queried beyond seven days by no one.

05

Outcome and evidence

Nine months on: cloud spend fell from €47K to €19K a month, infrastructure time on platform maintenance from 61% to 22%, production incidents from 34 to 9 per six months with platform-originated incidents down from 22 to 2, deploy duration from 22 minutes to 5, deploys from 6 to 40+ a week, and infrastructure engineers required from 3 to 1.5.

The remaining infrastructure engineer moved to half-time on product work. That was the outcome the CTO cared about most.

Results — 9 months on
MetricBeforeAfter
Cloud spend€47K/mo€19K/mo
Infrastructure time on platform maintenance61%22%
Production incidents (6mo)349
Incidents originating in platform222
Deploy duration22 min5 min
Deploys per week640+
Kubernetes clusters31
Services94
Infrastructure engineers required31.5

Client feedback

What the client says about this engagement

CTO

“I brought Alex in expecting a list of things to add and got a list of things to delete. My first reaction was that he'd misunderstood how sophisticated our setup was, and his answer was that he understood it fine and we were twelve people.

The service mesh number was what convinced me — seven outages in six months caused by the thing we'd installed for reliability, at forty requests per second. We're spending twenty-eight thousand euros a month less and shipping about seven times as often. Nobody misses any of it.”

CTO · B2B analytics company (NDA, anonymized)

From the engagement lead

What I’d tell anyone considering this

Alex B.

“Count the operational components you run and divide by the number of people who can operate each one. If any number in that division is one, that's a finding. If several are, you have a platform your team cannot sustain, and it will hurt you the week someone resigns.

Then look at where your incidents originate. If more than a third come from platform infrastructure rather than application code, your reliability machinery is a net negative and you should be removing pieces, not tuning them.

The uncomfortable part is that adding infrastructure is how engineers demonstrate seriousness, and removing it feels like a downgrade. It isn't. Complexity you can't operate is worse than a simpler system you can. Ask what each component would cost you to lose the person who understands it — and then decide whether it's worth keeping.”

Alex B. · DevOps Practice Lead at Envion Software

Evidence gate. This page publishes only what Envion's project records and client disclosure permissions support. Outcomes are added once verified against a baseline, a measurement period, and an approved source.

FAQ

Questions about this case

Facing a similar challenge?

Spending more on your platform than your product? Count what one person can operate — Envion’s operations review tells you what to delete, not what to add.

Discuss a Similar Challenge

Executive Technology Leadership

Support for high-stakes product and AI decisions

Bring senior technology leadership into the business when the roadmap is unclear, delivery is at risk, an AI initiative needs stronger ownership, or the company needs an experienced technical voice before hiring a permanent CTO.

Discuss Interim CTO Support

Core responsibilities

  • Align product and technology priorities with business goals and measurable outcomes.
  • Review architecture, delivery risks, data foundations, security needs, and AI readiness.
  • Lead internal teams and external partners through a practical execution plan.
  • Clarify team structure, ownership, decision rights, and delivery cadence.
  • Support investor, board, partner, and due-diligence conversations with credible technical judgment.

New experience

Prompt-to-Page — try it right here

Describe the landing page you want, in your own words. We turn it into a finished page and email you a private link in 5–10 minutes — no briefs, no calls, $0 to see the result.

  1. Describe what you want to create.
  2. We structure, write, and compose the page.
  3. You receive a private link when it is ready.

Start with a sentence — the interactive builder takes it from there.

Generate My Page

Safe, respectful content only. No obligation.

Start here

Discuss a Similar Challenge

Share your current state, constraints, and desired outcome — a senior specialist will reply with a concrete next step.

Prefer a direct channel?