Network Anomaly Detection with Telemetry and Online ML - 夜莺博客

Network Anomaly Detection with Telemetry and Online ML

Static thresholds are the default alerting model in most networks, and they are the reason operators ignore alerts. A 90% link utilisation threshold fires every day at 14:00 during the backup window and never fires for the slow, correlated degradation that actually precedes an outage. Anomaly detection replaces the fixed line with a model of normal behaviour, learned per link, per prefix, per service, and updated continuously.

This article walks through a production-shaped pipeline: which telemetry to collect, how to turn flows and counters into features, which unsupervised algorithms survive streaming data, how to keep false positives down, and how to prove the model works before you let it page anyone.

Start With the Data You Already Have

You almost certainly have three usable sources without buying anything:

  • Flow records (NetFlow v9/IPFIX/sFlow). Rich, already exported, and the format that most published anomaly-detection research is built on - the NF-UNSW-NB15 datasets used in recent work are exactly this shape.
  • Model-driven telemetry (gNMI dial-in/dial-out). High-frequency interface counters, queue drops and BGP state. Ten-second granularity changes what is detectable.
  • Control-plane events. Log messages, BGP updates, configuration commits. Lower volume, very high signal.

The comparison of flow export protocols, sampling rates and collector behaviour is covered in this sFlow vs NetFlow vs IPFIX guide - read it before you assume your exports are complete, because sampling ratios silently destroy anomaly detection. A 1-in-1000 sFlow sample will never see a small scan.

Feature Engineering Decides the Outcome

Raw byte counts are weak features. The features that make flow-based detection work are ratios, rates and entropy measures computed over a window:

# per (src_prefix, dst_prefix, port, proto) over a 5-minute window
flows            = count(records)
bytes_per_flow   = sum(bytes) / flows
packets_per_flow = sum(packets) / flows
avg_flow_duration= sum(duration) / flows
syn_ratio        = count(tcp_flags & SYN) / flows
rx_tx_bytes_ratio= sum(rx_bytes) / sum(tx_bytes)
dst_port_entropy = -sum(p * log2(p) for p in port_hist)
new_dst_ratio    = count(dst_ip not seen in 30d) / flows

Window choice matters more than the algorithm. Too short and normal burstiness looks anomalous; too long and you cannot localise the event. In practice, run two resolutions in parallel: a 1-minute stream for volumetric events and a 15-minute stream for behavioural drift.

Choosing an Algorithm That Survives Production

Supervised classifiers give the best accuracy on a benchmark and the worst experience in production, because you never have labels for what is coming next. Unsupervised online models are the pragmatic choice:

Approach Strength Weakness
Rolling z-score / EWMA per series Trivial to run, interpretable, per-entity baselines Blind to multivariate correlation; struggles with seasonality
Isolation forest / HBOS on sliding window Handles mixed feature scales, batch retraining is cheap Retraining cadence is a compromise; concept drift
Streaming clustering (DenStream, micro-clusters) Detects novel patterns, no labels, adapts continuously Parameter-sensitive; classic outlier streams are alarm-happy
LSTM / autoencoder reconstruction error Captures temporal structure, good on rich counter streams Needs tuning and a stable environment; hard to explain to an auditor

Recent work on unsupervised online learning over NetFlow specifically addresses the label problem by learning normal behaviour continuously and scoring each new flow, with normalisation (MaxAbsScaler) chosen because it behaves well on streaming data without shifting the mean. That is a practical detail worth reusing: methods that assume a global mean and variance break the moment the model's window slides.

The Pipeline

router --(IPFIX/gNMI)--> collector --> Kafka topic
                                          |
                                    stream processor
                                    (window + features)
                                          |
                     +--------------------+--------------------+
                     |                                         |
              feature store (Parquet)                  online scorer
                                                       (per-entity model)
                                                              |
                                                       alert enricher
                                                  (topology + change log)
                                                              |
                                                        pager / ticket

The enrichment step is what makes the output usable. An alert that says "prefix 203.0.113.0/24 to port 443 is 8 sigma above baseline" is ignorable. The same alert annotated with "this is the path through PE2, and a change ticket closed on PE2 eleven minutes ago" is actionable. Feed the model from a state store that already has the schema for telemetry in place; if you are storing counters in ClickHouse, the patterns in this network telemetry schema guide let you query baselines directly instead of maintaining a second pipeline.

If you already run gNMI collection, the deployment details in this gNMI streaming telemetry pipeline guide give you the transport layer, so the only new component is the scorer.

Controlling False Alarms

Detection quality is not the hard part - alarm fatigue is. Four controls do most of the work:

  1. Per-entity suppression windows. Do not raise a second alarm for the same entity and the same feature family within the hold-down period.
  2. Correlation across entities. One interface outlier is a curiosity; twelve interfaces on the same line card are an event. Group alarms by shared topology parent.
  3. Change-window awareness. Raise thresholds during maintenance windows and for a settlement period afterwards, otherwise every deployment teaches the model that change is anomalous and every subsequent change alerts.
  4. Precision tracking. Label a small sample of alarms weekly. If precision is below roughly one in five, users will stop reading them regardless of recall.

Validation Before You Page Anyone

Run the model in shadow mode for at least a month. Compare its alarms against known incidents and against the existing threshold alerts. Measure:

  • Alarm volume per day per entity - if a single device generates more than a handful, the model is under-confident and needs a tighter baseline.
  • Detection latency versus the threshold rule it is meant to beat. A model that fires four minutes later is not an improvement.
  • False-negative review on the incidents that did occur. Unsupervised models are typically good at unknown-unknowns and weak at known signatures - keep your signature rules alongside the model rather than replacing them.

Useful public datasets for confidence-building include NF-UNSW-NB15 and its v2 revision, which contain labelled NetFlow records across nine anomaly categories and let you sanity-check feature choices before touching production data.

Rollout Order

Start with a single feature family on a single role - uplink utilisation on WAN routers, for example - and prove the loop end to end. Then widen to flows, then to control-plane events. Instrument the pipeline itself: if the collector drops records for twenty minutes, every model downstream will report a perfectly anomalous quiet period. Pair the anomaly layer with solid collection health monitoring such as SNMP exporter based device polling so you can distinguish "nothing is happening" from "nothing is arriving".

原文链接:https://arxiv.org/abs/2509.01375