Envion Software
CS-069Software AuditPayments Processing (NDA)

Quality Audit: The QA Team Wasn’t the Problem

A payments processor (~€8B annual transaction volume) had four severe production incidents in seven months, including a 90-minute settlement outage with contractual penalties and a regulator notification. The internal plan was to double the QA team and add a two-week regression gate. Envion’s five-week audit found only one of the four severe incidents was a testing failure — the others were a missing guardrail, a missing circuit breaker, and a migration never tested at scale — and put a number on the proposed remedy: the regression gate would have grown batch size from 40 to ~90 changes per release, measurably increasing severity. Twelve months on: one severe incident instead of four, detection time 34 → 3 minutes, change failure rate 19% → 5%, QA headcount unchanged at 9.

Quality Audit: The QA Team Wasn’t the Problem
01

The challenge

The client had experienced four severe production incidents in seven months, including a 90-minute settlement outage that triggered contractual penalties and a regulator notification. Board pressure was intense and the internal narrative had settled on a cause: quality assurance was failing.

The proposed response was to double the QA team and add a mandatory two-week regression cycle before every release. The CTO, who suspected this would make things worse without knowing why, commissioned an independent audit before implementing it.

02

Decision path

Envion analysed the four severe incidents plus 180 lower-severity ones, traced each to root cause, and assessed the code, tests and process around them. Only one of four severe incidents was a testing failure. The distribution mattered more than any individual finding:

Incident 1 — a configuration change applied directly to production, outside the deployment pipeline, by an engineer resolving an urgent customer issue. No test would have caught it. There was no mechanism preventing it.

Incident 2 — the settlement outage. A third-party banking API changed a response field. The client's integration had no schema validation and no circuit breaker; the malformed response propagated into the settlement engine and corrupted a batch. A resilience gap, not a testing gap.

Incident 3 — a genuine regression. Test coverage on the affected reconciliation module was 22%.

Incident 4 — a database migration that locked a table under production load. Tested successfully against a dataset 1/400th the size of production.

03

Envion contribution

The test suite was large and pointed the wrong way. 14,000 tests, 71% overall coverage, a 40-minute run. But coverage mapped almost inversely to risk: 94% on the API layer, 22% on reconciliation, 31% on settlement, 18% on the ledger. The tests had been written where they were easy to write. Roughly 3,000 were testing framework behaviour rather than business logic.

The delivery process was creating the incidents it was meant to prevent. Releases were batched into a Thursday window, averaging 40 changes each. When something broke, isolating the cause among 40 changes took hours, and rollback meant reverting all of them — so the team routinely fixed forward under pressure, which is how two of the four incidents escalated in severity.

Envion put a specific number on the proposed remedy: a mandatory two-week regression cycle would grow batch size to roughly 90 changes per release. Every incident would become harder to diagnose and more expensive to reverse. The proposed fix would have measurably increased severity.

Operational readiness was the quietest and most serious finding. Median time to detect a production issue was 34 minutes, and in two of four cases a customer reported it first. There were no runbooks for the settlement engine. On-call rotation had one engineer who had never been through an incident.

04

Delivery

Envion ranked 23 findings by expected severity reduction per unit of effort, not by severity alone — the distinction that makes an audit actionable rather than overwhelming.

Weeks 1–2, immediate risk reduction: block direct production access with a break-glass procedure (logged, mandatory review); schema validation and circuit breakers on all six external financial integrations; migration testing against a production-scale dataset, mandatory in the pipeline.

Weeks 3–8, structural: test investment redirected to reconciliation, settlement and ledger, with a coverage floor by module criticality rather than a global target; delete the ~3,000 low-value tests to get the suite under 10 minutes, enabling per-change runs; replace the weekly release train with continuous deployment behind feature flags; runbooks for the five critical paths and incident simulation for on-call.

Months 3–6, strategic: business-level monitoring — settlement lag, reconciliation break rate — rather than infrastructure metrics alone, and chaos testing on the external integration layer.

Envion recommended against expanding the QA team, and explained the reasoning in terms the board would accept: the evidence located three of four severe incidents outside QA's control entirely.

05

Outcome and evidence

Twelve months on: one severe incident instead of four in seven months, median detection time down from 34 minutes to 3, no customer-reported incidents, releases of 1–3 changes instead of 40, the test suite running in 8 minutes, settlement module coverage at 88%, change failure rate down from 19% to 5% — and QA headcount unchanged at 9 instead of the proposed 18.

Overall coverage fell and quality improved substantially. That single row does more to explain the audit's value than any other number in the table.

The advice that generalizes: root-cause your incidents before choosing a remedy — take your last twenty and classify them honestly, because the distribution tells you where to spend. Coverage percentage is a nearly meaningless target; what matters is coverage weighted by consequence. Batch size drives incident severity, so any process change that increases batch size in the name of safety deserves specific scrutiny. And check how you find out — if your customers tell you first, that's your first finding.

Results — 12 months on
MetricBeforeAfter
Severe (P1) incidents4 in 7 months1 in 12 months
Median time to detect34 min3 min
Customer-reported incidents2 of 40
Changes per release401–3
Test suite runtime40 min8 min
Coverage, settlement module31%88%
Coverage, overall71%68%
Change failure rate19%5%
QA headcount9 (18 proposed)9

Client feedback

What the client says about this engagement

CTO

“We were about to spend a million euros a year on QA headcount and add a two-week regression gate. Envion showed us the gate would have pushed us from forty changes per release to ninety and made every incident worse. That finding alone justified the engagement.

The uncomfortable part was seeing that three of our four serious incidents had nothing to do with testing — they were a missing guardrail, a missing circuit breaker, and a migration nobody had tested at scale. Our overall coverage number went down and our incident rate went down with it. I use that as a teaching example internally now.”

CTO · Payments processor (NDA, anonymized)

Evidence gate. This page publishes only what Envion's project records and client disclosure permissions support. Outcomes are added once verified against a baseline, a measurement period, and an approved source.

FAQ

Questions about this case

Facing a similar challenge?

Reaching for more testing and more process? Root-cause your last twenty incidents first — the distribution tells you where to spend.

Discuss a Similar Challenge

Executive Technology Leadership

Support for high-stakes product and AI decisions

Bring senior technology leadership into the business when the roadmap is unclear, delivery is at risk, an AI initiative needs stronger ownership, or the company needs an experienced technical voice before hiring a permanent CTO.

Discuss Interim CTO Support

Core responsibilities

  • Align product and technology priorities with business goals and measurable outcomes.
  • Review architecture, delivery risks, data foundations, security needs, and AI readiness.
  • Lead internal teams and external partners through a practical execution plan.
  • Clarify team structure, ownership, decision rights, and delivery cadence.
  • Support investor, board, partner, and due-diligence conversations with credible technical judgment.

New experience

Prompt-to-Page — try it right here

Describe the landing page you want, in your own words. We turn it into a finished page and email you a private link in 5–10 minutes — no briefs, no calls, $0 to see the result.

  1. Describe what you want to create.
  2. We structure, write, and compose the page.
  3. You receive a private link when it is ready.

Start with a sentence — the interactive builder takes it from there.

Generate My Page

Safe, respectful content only. No obligation.

Start here

Discuss a Similar Challenge

Share your current state, constraints, and desired outcome — a senior specialist will reply with a concrete next step.

Prefer a direct channel?