European hosting provider · 17-engineer department · Head of Server

Tens of thousands of servers, one department. Culture as infrastructure.

Faster provisioning and fewer repeat incidents across a Europe-wide hosting estate — not through a heroic rewrite, but through automation-first process and a blameless post-mortem culture that changed how the department handled failure.

The situation

A Europe-wide hosting provider ran an estate of tens of thousands of production Linux and Windows servers — real customers, real workloads, on every one of them. Provisioning new servers was too slow, the same categories of incident kept recurring, and catastrophic-incident escalation needed a clear owner rather than whoever happened to be on shift. Nate led the 17-engineer department responsible for the estate as Head of Server, acting as senior escalation for the incidents that mattered most.

Constraints

  • A heterogeneous estate. Linux, Windows, and VMware, side by side, at scale. There's no single playbook that covers all three, and any process change has to work across all of them or it doesn't really work.
  • 24/7 production, real customers, per server. There's no maintenance window big enough to take a meaningful slice of a hosting estate offline to fix it properly. Change has to happen while the lights stay on.
  • Change at department scale is process design, not code. With 17 engineers and tens of thousands of servers, the lever that moves outcomes isn't a clever script one person writes — it's whether the whole department does the right thing by default, every time, without being told.

Architecture decisions — and why

The provisioning pipeline moved to automation-first. Manual runbooks don't scale past a few thousand servers, and worse, they quietly encode tribal knowledge nowhere durable — the fastest way to provision a server correctly lived in a few senior engineers' heads, not in a system anyone else could run. Automating the pipeline did two things at once: it made provisioning faster, and it forced the tribal knowledge out of people's heads and into something versioned, reviewable, and repeatable by anyone on the team.

Blameless post-mortems became the incident-reduction mechanism — deliberately, not as a culture initiative bolted onto the engineering work. The insight driving this: repeat incidents are, overwhelmingly, an information-flow failure before they're a technical one. When post-mortems assign blame, people route around the process — they under-report, they soften the timeline, they leave out the detail that would actually prevent the next occurrence, because that detail is also the detail that makes them look bad. Remove blame from the process and the information starts flowing honestly, which is the actual precondition for a repeat incident rate that goes down instead of staying flat.

Catastrophic incidents got a named senior-escalation ownership model. A named owner changes incident behavior in a way that "whoever's on call figures it out" doesn't — decisions get made faster because there's no ambiguity about who's authorized to make them, and the postmortem has a clear person who can drive the follow-up actions to actually landing, rather than a list of good intentions nobody owns.

Stack

LinuxWindows ServerVMwareAutomation tooling

What shipped

Department-scale process change, not a single project: an automated provisioning pipeline, a blameless post-mortem culture that actually held under pressure, and a named escalation model for catastrophic incidents — rolled out across a 17-engineer organization responsible for a Europe-wide estate.

Measured outcomes

~40% faster provisioningfrom the automation-first pipeline
~25% fewer repeat incidentsfrom the blameless post-mortem culture

At estate scale, culture is infrastructure.

Lessons

Automation pays twice: once in raw speed, and again in the documentation it forces into existence as a side effect. A runbook that only exists as automation code can't quietly go stale in someone's head the way a wiki page can — it either still runs correctly or it visibly breaks, which is a much healthier failure mode for institutional knowledge than "ask Dave, he remembers."

These are, directly, the leadership lessons that now shape how West Fork Digital runs client incidents: blameless by default, a named owner on anything serious, and automation treated as documentation that happens to also be fast.


Running infrastructure at a scale where manual process is starting to break down, or where repeat incidents are eating a department's time? We've built the automation and the culture that fixes both.

Fixed quote within 48 hours — no obligation.