Skip to content

Servers & Storage

Size it for what the workload actually does, not for what the quote assumes.

What this is

Servers and storage covers compute sizing, storage architecture, firmware and lifecycle management, and the resilience design underneath both. It is the layer virtualisation, databases and applications all inherit their performance from.

Sizing is where money is lost in both directions. Over-provisioning is visible and merely wasteful; under-provisioning surfaces as an application problem months later and gets diagnosed everywhere except the disk queue. We size from measurement.

When you need it

If more than one of these is true, this is usually the right place to start.

  • Hardware approaching end of support, where a failure becomes a procurement exercise.
  • Storage performance blamed for application slowness, with no measurement either way.
  • A refresh quote sized from a vendor configurator rather than from your workload.
  • Firmware and driver levels that differ across supposedly identical hosts.

What the scope covers

  • Workload profiling: throughput, IOPS, latency and growth measured over a representative period.
  • Compute sizing with headroom for the failure case, not only for average load.
  • Storage architecture: tiering, protection level, snapshots and replication against stated recovery objectives.
  • Firmware and driver baseline, with a process that keeps identical hosts identical.
  • Lifecycle and support plan: what is under support until when, and the refresh sequence that follows.

What you receive

DeliverableWhat it contains
Workload profileMeasured demand per workload with peaks and growth, and the assumptions the sizing rests on.
Sizing and designCompute and storage specification, protection levels, and the failure case each is sized to survive.
Firmware baselineTarget levels per platform, current deviation, and the maintenance process that holds the line.
Lifecycle planSupport expiry per asset and a refresh sequence aligned to budget cycles rather than to surprises.

Reference architecture

A reference, not a template. Your estate decides which parts apply and in what order they arrive.

Servers and storage reference architecture: compute, storage and visibility layersCompute: Rack / Blade Servers, Firmware Baseline, Out-of-band Mgmt. Storage: SAN / NAS, Tiering, Snapshots. Visibility: Capacity Monitoring, Health Alerting, Lifecycle ReportingComputeRack / Blade ServersFirmware BaselineOut-of-band MgmtStorageSAN / NASTieringSnapshotsVisibilityCapacity MonitoringHealth AlertingLifecycle Reporting
Servers and storage reference architecture: compute, storage and visibility layers

How success is measured

Targets are agreed with you before the work starts, and reported against for its duration.

  • Storage latency at peak against the target the applications actually need.
  • Capacity headroom under the single-failure case, not under normal operation.
  • Hosts matching the firmware baseline, counted rather than assumed.

Questions we are asked

  • How do we know what size we need?

    By measuring the existing workload over a period long enough to include its peaks — month-end, reporting cycles, seasonal load. Sizing from a configurator without that measurement is guesswork with a price attached.

  • Is flash storage always the answer?

    It is the default for anything latency-sensitive and the price gap has narrowed considerably. It is still worth tiering: archival and backup data on flash is money spent for no measurable benefit.

  • Should we buy hyper-converged or keep them separate?

    Hyper-converged simplifies operations and scales predictably, which suits a mid-sized estate well. Separate compute and storage still wins where the two need to scale independently, or where storage demand is large relative to compute.

  • How much headroom should we design for?

    Enough to run through the failure case you have decided to survive. If losing a node must not degrade service, the remaining nodes carry the full load — so the number is set by the resilience decision, not by a percentage rule.

  • Does firmware really matter that much?

    It matters more than most estates treat it. Inconsistent firmware across identical hosts produces failures that appear on one node and not another, and those are among the hardest problems to diagnose because the hosts are supposed to be the same.

  • Should this go to the cloud instead?

    Sometimes, and that is a question worth asking before a refresh rather than after. The honest answer depends on workload profile, data gravity and cost over the asset's life — which is a cloud migration assessment, and we would rather run it than assume either way.

Continue reading

  • Virtualization

    Hypervisor cluster design, resource policy and right-sizing — including the licensing consequence of the design, which is where surprises usually arrive.

  • Data Centre

    Power, cooling, cabling and physical resilience designed together — with the environmental monitoring that turns a facility into something you can operate.

Start with an assessment

The fastest way to a useful answer is a short, scoped look at what you already have.