NetFlow vs IPFIX vs sFlow: Choosing Flow Telemetry - 夜莺博客

NetFlow vs IPFIX vs sFlow: Choosing Flow Telemetry

Interface counters tell you a link is full. Flow records tell you which conversation filled it, which is what you actually need during an incident. Three protocols do this job and they are not interchangeable: NetFlow (Cisco lineage, v5/v9), IPFIX (the standards-track successor, RFC 7011) and sFlow (a different, packet-sampling design). Choosing well is about understanding how each one sees the network.

The Fundamental Difference

NetFlow v9 / IPFIX sFlow
Method Flow cache: packets are aggregated into flows per 5-tuple Statistical packet sampling (e.g. 1 in 2000) plus counter sampling
What you get Every flow that matches the cache criteria, with durations and byte counts Sampled packet headers - an estimate of traffic mix, and any protocol it can parse
Overhead Higher: cache memory and CPU on the device Very low: CPU impact is minimal even at high rates
Accuracy Accurate byte counts per flow Statistically representative; good for ratios and trends, poor for exact accounting
L2 / non-IP visibility Limited (IP flows) Yes - samples raw frames, so ARP storms and non-IP traffic are visible
Extensibility Templates carry any field, including MPLS, BGP next-hop, MAC Extensible via community-defined structures

The practical summary: accounting and per-flow forensics favour NetFlow/IPFIX; broad, cheap visibility on many devices favours sFlow.

Configuration Shape

# Cisco IOS: IPFIX-style flexible NetFlow
table-map TM-IPV4-RECORD
 record ipv4
 exporter EXP1
  destination 10.10.10.70
  source Loopback0
  transport udp 4739
  template data timeout 60
 flow monitor MON-IPV4
  exporter EXP1
  record TM-IPV4-RECORD
  cache timeout active 60
interface GigabitEthernet1/0/24
 ip flow monitor MON-IPV4 input
 ip flow monitor MON-IPV4 output
# Arista EOS: sFlow
sflow source-interface Loopback0
sflow destination 10.10.10.70 6343
sflow sample 4000
sflow polling-interval 20
sflow run
interface Ethernet1
 sflow enable

Sampling Rates and What They Buy

  • 1:1000 to 1:2000 on 10G/25G access ports: enough to see top talkers and traffic mix.
  • 1:4000 or higher in a leaf-spine fabric: keeps collector load sane. Remember that a 400G port sampled at 1:4000 still delivers ~100k samples/s if fully loaded.
  • No sampling for NetFlow: correct for billing and capacity accounting, but plan the cache (active timeout 60s, inactive 15s are common starting points) and verify the device's CPU headroom.

Collector Design

  1. Export to a dedicated collector, never to a syslog host. Datagram loss is silent by design in both protocols.
  2. Size for the peak, not the average: during an incident, exports spike exactly when you need them.
  3. Set the source interface to a loopback so the exporter is reachable regardless of interface state.
  4. Enable export over the management VRF if management and production are separated - and remember that some platforms require a route in the correct VRF for the exporter to work.
  5. Retain raw flow data for at least 7-30 days. You will want it long after the incident is closed.

Symptom-to-Tool Mapping

  • Link saturated, unknown cause → NetFlow/IPFIX top talkers by destination AS or prefix.
  • Broadcast storm or unknown unicast flood → sFlow (it samples frames, not just IP flows).
  • Billing or chargeback → IPFIX, with a stable template and documented cache timers.
  • Trace a specific attack pattern across a fabric → sFlow for breadth plus a targeted SPAN capture on the suspect link.

Related reading: sFlow configuration on Dell OS10, sFlow-RT analytics and dashboards, and telemetry-driven anomaly detection for what to do with the data once you have it.

原文链接:https://datatracker.ietf.org/doc/html/rfc7011