Infrastructure
The full domain, and the other capabilities within it.
Size it for what the workload actually does, not for what the quote assumes.
Servers and storage covers compute sizing, storage architecture, firmware and lifecycle management, and the resilience design underneath both. It is the layer virtualisation, databases and applications all inherit their performance from.
Sizing is where money is lost in both directions. Over-provisioning is visible and merely wasteful; under-provisioning surfaces as an application problem months later and gets diagnosed everywhere except the disk queue. We size from measurement.
If more than one of these is true, this is usually the right place to start.
| Deliverable | What it contains |
|---|---|
| Workload profile | Measured demand per workload with peaks and growth, and the assumptions the sizing rests on. |
| Sizing and design | Compute and storage specification, protection levels, and the failure case each is sized to survive. |
| Firmware baseline | Target levels per platform, current deviation, and the maintenance process that holds the line. |
| Lifecycle plan | Support expiry per asset and a refresh sequence aligned to budget cycles rather than to surprises. |
A reference, not a template. Your estate decides which parts apply and in what order they arrive.
Targets are agreed with you before the work starts, and reported against for its duration.
By measuring the existing workload over a period long enough to include its peaks — month-end, reporting cycles, seasonal load. Sizing from a configurator without that measurement is guesswork with a price attached.
It is the default for anything latency-sensitive and the price gap has narrowed considerably. It is still worth tiering: archival and backup data on flash is money spent for no measurable benefit.
Hyper-converged simplifies operations and scales predictably, which suits a mid-sized estate well. Separate compute and storage still wins where the two need to scale independently, or where storage demand is large relative to compute.
Enough to run through the failure case you have decided to survive. If losing a node must not degrade service, the remaining nodes carry the full load — so the number is set by the resilience decision, not by a percentage rule.
It matters more than most estates treat it. Inconsistent firmware across identical hosts produces failures that appear on one node and not another, and those are among the hardest problems to diagnose because the hosts are supposed to be the same.
Sometimes, and that is a question worth asking before a refresh rather than after. The honest answer depends on workload profile, data gravity and cost over the asset's life — which is a cloud migration assessment, and we would rather run it than assume either way.
The full domain, and the other capabilities within it.
Hypervisor cluster design, resource policy and right-sizing — including the licensing consequence of the design, which is where surprises usually arrive.
Power, cooling, cabling and physical resilience designed together — with the environmental monitoring that turns a facility into something you can operate.
The fastest way to a useful answer is a short, scoped look at what you already have.