Skip to content

Network Monitoring

Know that something is wrong before the first user calls.

What this is

Network monitoring is discovery, telemetry collection, thresholding, alerting and reporting. Done properly it answers three questions: is it up, is it healthy, and is it getting worse.

The failure mode is not missing data, it is unread alerts. A system that pages on every interface flap trains people to ignore it, and by the time it reports something real nobody is looking. Tuning is the work.

When you need it

If more than one of these is true, this is usually the right place to start.

  • Outages first reported by users rather than by a monitoring system.
  • An alerting channel with enough noise that the team has muted it.
  • No capacity trend, so upgrades are argued from opinion rather than from a graph.
  • Monitoring that covers the data centre but stops at the branch or the cloud.

What the scope covers

  • Discovery and inventory, including the devices that were never added to the existing system.
  • Telemetry design: what is polled or streamed, at what interval, and what it costs to keep.
  • Thresholds derived from an observed baseline rather than from vendor defaults.
  • Alert routing: who is told, by what channel, at what hour, and what an alert is expected to make them do.
  • Dashboards and reporting for two audiences — an engineer diagnosing now, and a manager planning capacity.

What you receive

DeliverableWhat it contains
Monitoring inventoryEvery device and interface in scope, with coverage stated as a figure and the gaps named.
Baseline and thresholdsObserved normal per metric and the thresholds derived from it, including time-of-day variation.
Alert policySeverity, routing, escalation and suppression rules — with what is deliberately not alerted on, and why.
DashboardsAn operational view for diagnosis and a trend view for capacity, each built for its audience.

Reference architecture

A reference, not a template. Your estate decides which parts apply and in what order they arrive.

Network monitoring reference architecture: discovery, control and visibility layersDiscovery: Device Discovery, Topology Mapping, Inventory. Control: Thresholds, Alert Routing, Maintenance Windows. Visibility: Streaming Telemetry, Flow Analysis, DashboardsDiscoveryDevice DiscoveryTopology MappingInventoryControlThresholdsAlert RoutingMaintenance WindowsVisibilityStreaming TelemetryFlow AnalysisDashboards
Network monitoring reference architecture: discovery, control and visibility layers

How success is measured

Targets are agreed with you before the work starts, and reported against for its duration.

  • Share of the estate monitored against the reconciled inventory.
  • Alerts per day per on-call engineer, and the proportion that led to an action.
  • Incidents detected by monitoring before a user reported them.

Questions we are asked

  • We have a monitoring tool already. Why is this needed?

    Most organisations do. The difference is usually coverage and tuning: devices never added, thresholds left at defaults, and alerts routed to a channel nobody watches. The tool is rarely the problem.

  • How do we reduce alert noise?

    Baseline first, then thresholds from the baseline; suppress the dependent alerts that follow a single root cause; and delete alerts nobody has acted on in six months. That last one is unpopular and it is where most of the noise lives.

  • Polling or streaming telemetry?

    Streaming gives higher resolution and scales better on modern platforms; polling still covers the older devices that will be in your estate for years. Most designs are a mix, and the mix is decided by what your hardware supports.

  • Does this cover cloud and branch?

    It should. Monitoring that stops at the data centre leaves the parts of the path users actually complain about unmeasured, which is why branch and cloud connectivity are in scope from the start.

  • Should this feed our SIEM?

    Selectively. Operational telemetry and security telemetry overlap but are not the same, and sending everything to a SIEM is how licence costs get out of hand. The design states what crosses over and why.

  • Can you run it for us?

    Yes, as a managed service. Monitoring is one of the capabilities where consistency matters more than intensity, and where an out-of-hours rota is difficult to sustain with a small internal team.

Continue reading

  • Networking

    The full domain, and the other capabilities within it.

  • Network Infrastructure

    Switching, routing and structured cabling designed for the traffic you will have in five years, not the traffic you had when the building opened.

  • Enterprise Wi-Fi

    RF design, controller policy and validation surveys for wireless that holds up under density — measured on site, not predicted from a floor plan.

Start with an assessment

The fastest way to a useful answer is a short, scoped look at what you already have.