KPIs are only useful when they measure something that matters and when the measurement is accurate enough to be acted on. ISPs often track availability percentages and present them to management without a clear understanding of what the number represents, how it was calculated, or what operational decisions it should drive. The result is KPIs that look like governance but function as decoration.

This article covers the specific KPIs that matter for ISP operations, what they actually measure, how to calculate them correctly, and what decisions they should inform.

Infrastructure Availability

Infrastructure availability is the percentage of time a specific service or network element is operational during the measurement period. For ISPs, the elements worth tracking separately are: upstream link availability (each uplink separately), core network element availability (each critical router), and aggregated subscriber access availability by PoP.

The calculation is: (measurement period - total downtime) / measurement period x 100. Downtime starts when an outage is confirmed (not when it is reported by a subscriber) and ends when service is confirmed restored (not when the NOC believes it is restored). Scheduled maintenance downtime should be tracked separately from unscheduled downtime: the availability KPI that matters operationally is unscheduled downtime.

Track availability monthly and compare against your SLA commitments for corporate clients and your own operational targets for residential service. A monthly availability figure without a target is just a number. A figure compared against a target is a signal for action.

Mean Time to Restore (MTTR)

MTTR is the average time between an incident being confirmed and service being restored. For ISPs, track MTTR separately by incident category: upstream link failures, core hardware failures, access layer failures, and software/configuration incidents. The categories have different expected resolution profiles and different improvement levers.

MTTR is the operational KPI most directly under the NOC team's control: documentation quality, escalation speed, runbook completeness, and engineering skill all affect it. An MTTR trend going in the wrong direction is a signal to investigate whether incident response procedures are being followed, whether the right skills are available on call, or whether infrastructure changes have created new failure modes that runbooks do not cover.

Throughput Utilisation

Peak throughput utilisation on critical links, expressed as a percentage of capacity, is the capacity planning KPI. Track the 95th percentile of peak hourly throughput on each critical link, plotted as a trend over the past 90 days. The trend line, not the current value, is the signal: a link at 70% utilisation but trending toward 80% in 60 days requires an upgrade decision now, not when it reaches 90%.

Latency and Packet Loss

Latency and packet loss measured from the network to reference points (PKIX, upstream provider, major content destinations) indicate network health in ways that availability metrics do not capture. A network that is technically available but experiencing elevated latency or 1-2% packet loss is degraded in ways subscribers notice before the NOC does.

Measure latency and packet loss continuously from your infrastructure toward reference points and alert when they exceed thresholds. For corporate CIR clients with latency SLA commitments, per-circuit latency measurement is required.

Subscriber-Facing Metrics

MTTR from a subscriber perspective, which includes the time between a subscriber reporting an issue and service being restored, is different from the NOC-measured MTTR that starts at detection. The gap between these two metrics indicates how quickly complaints are being routed to the NOC and how accurately the NOC is detecting incidents before subscribers report them. A large gap between subscriber-reported MTTR and NOC-measured MTTR suggests monitoring coverage is insufficient.

Presenting KPIs to Management and the ISSC

For CTDISR-2025 Section 1 governance requirements, the ISSC should receive network performance KPIs as part of its regular reporting. The board-level format should be simple: a traffic-light summary for each KPI against target, the trend direction, and any actions required. Technical detail belongs in the operational KPI report, not the board report.

For the monitoring infrastructure that generates the raw data KPIs are calculated from, NOC Enablement & Monitoring covers the monitoring stack design. For ISSC reporting that integrates network performance alongside security and compliance KPIs, CISO-as-a-Service covers the governance reporting function.