sFlow-RT Flow Analytics and Dashboards Guide - 夜莺博客

sFlow-RT Flow Analytics and Dashboards Guide

NetFlow gives you flow records minutes after the fact; sFlow gives you a statistically sampled view of the network in real time. sFlow-RT turns that stream into queryable metrics and dashboards without a database cluster, which is why it shows up in everything from small campus deployments to AI-fabric monitoring. This guide covers installing it, getting switches to export to it, the queries that answer real operational questions, and how to bridge it into Prometheus and Grafana.

Why sFlow rather than NetFlow here

sFlow     packet-level sampling, always-on, low device CPU, includes L2 + L3
NetFlow   flow cache accounting, per-flow accuracy, heavier device load
IPFIX     NetFlow v10, template based, richer fields
Real-time analytics at scale needs sampling: you cannot export every flow
from a 400G fabric without overwhelming the collector.

A direct comparison of the three export formats is in sFlow vs NetFlow vs IPFIX. The short version: monitoring and DDoS detection favour sFlow; billing and per-subscriber accounting favour NetFlow.

Install

wget https://inmon.com/products/sFlow-RT/sflow-rt.tar.gz
tar -xvzf sflow-rt.tar.gz
./sflow-rt/start.sh          # listens on 6343/udp (sFlow) and 8008/tcp (UI/API)

# or containerised
docker run --rm -p 8008:8008 -p 6343:6343/udp sflow/sflow-rt
# install the useful applications
sudo /usr/local/sflow-rt/get-app.sh sflow-rt browse-metrics
sudo /usr/local/sflow-rt/get-app.sh sflow-rt browse-flows
sudo /usr/local/sflow-rt/get-app.sh sflow-rt fabric-metrics
sudo systemctl restart sflow-rt

Java 17+ is required. The Status page confirms telemetry is arriving: the sFlow Agents, sFlow Bytes and sFlow Packets gauges should all be non-zero. If they are zero, the problem is almost always UDP 6343 being blocked or the agent configured with the wrong collector address.

Point devices at the collector

! Arista EOS
sflow sample 2000
sflow polling-interval 30
sflow run
sflow vrf default destination 10.10.20.60 6343

! Cisco IOS-XE
ip flow-export destination 10.10.20.60 6343   ! for NetFlow; sFlow not on IOS
                                             ! use flexible NetFlow or a
                                             ! sampling-capable platform

! ArubaOS-CX
sflow agent-ip 10.10.20.50
sflow collector 10.10.20.60 6343
sflow sampling 1/4096

Sample rate is the main tuning knob. 1:2000 on high-speed core links keeps the record rate sane while still catching elephant flows; 1:512 on access ports gives better visibility into small flows. See the vendor specifics in Arista EOS sFlow configuration.

Asking useful questions

# top talkers in the last minute
curl -s "http://localhost:8008/app/browse-metrics/html/..." >/dev/null   # UI
# via the REST API (most useful form for automation)
curl -s "http://localhost:8008/flow/ips/json?maxFlows=10"
curl -s "http://localhost:8008/flow/tcp/json?maxFlows=10"
curl -s "http://localhost:8008/metric/ALL/max:ifinoctets/json"
curl -s "http://localhost:8008/metric/10.10.0.1/max:ifoutdiscards/json"
Operational question                 Query shape
Which host started the flood?        flow/ips with the highest bps
Is a link saturated?                 metric//max:ifoutoctets
Are drops growing?                   metric//max:ifoutdiscards
Which protocol dominates?            flow/tcp vs flow/udp breakdown
Is the AI fabric healthy?            ai-metrics app (RoCEv2 operations)

Export to Prometheus and Grafana

# install the prometheus app, then scrape it
- job_name: 'sflow-rt'
  metrics_path: /app/prometheus/scripts/export.js/prometheus/txt
  static_configs:
    - targets: ['sflow-rt:8008']

This is the bridge that makes sFlow useful for long-term trending: sFlow-RT keeps seconds of state, Prometheus keeps months. Grafana dashboards for flow analytics and dropped-packet metrics then sit alongside your SNMP or streaming-telemetry data — see Prometheus SNMP exporter for the device-metric side.

Topology matters

Several applications (locate-address, topology verification, path tracing) need a topology file describing the fabric. Without it, sFlow-RT still gives you per-interface and per-IP metrics, but "which switch port is 10.0.0.5 on?" cannot be answered. Topologies can be built from LLDP data, NetBox, or a Graphviz DOT file.

Pitfalls

Sampling too aggressively         collector CPU/disk; start at 1:2000
Collector in the data path        use a management/out-of-band network
No topology loaded                locate/trace apps unusable
Comparing sFlow counts to SNMP    sampling is statistical; small windows diverge
Single collector, no HA           acceptable for visibility, not for security
                                  tooling that must prove traffic existed

Combine sFlow with a packet-based tool for the opposite trade-off — see ntopng traffic monitoring — and keep in mind that sampled evidence is not forensic evidence.

FAQ

Q: Can sFlow replace NetFlow for capacity planning? For trending interface and top-talker volume, yes. For per-subscriber byte accounting with legal precision, no.
Q: How much storage does it need? sFlow-RT itself keeps rolling in-memory state only; long-term retention is Prometheus's job.
Q: Does it work with containerlab? Yes — there are projects that emulate leaf/spine and EVPN fabrics and replay or generate sFlow into sFlow-RT, which is a good way to validate your dashboards before touching production.

原文链接:https://sflow-rt.com/intro.php