SuzieQ: Network State Observability with Python - 夜莺博客

SuzieQ: Network State Observability with Python

Device-by-device troubleshooting answers one query at a time and forgets the answer as soon as the session closes. SuzieQ takes the opposite approach: it polls every device on a schedule, normalises the state (interfaces, routes, BGP, ARP, MAC, VLANs) into a local database, and lets you ask fabric-wide questions with a SQL-like query language or the Python API. That makes questions like which prefixes flapped in the last hour or which interfaces are down but still in a LAG answerable in seconds.

Install and Initialise

python3 -m venv /opt/suzieq && source /opt/suzieq/bin/activate
pip install suzieq
sqPoller --version

# create the inventory template
suzieq-gather --create-config
vi ~/.suzieq/suzieq-cfg.yml
data-directory: /var/lib/suzieq/parquet
rest-server:
  API_KEY: local-only-key
  address: 127.0.0.1
  port: 8000
devices:
  - name: leaf-01
    address: 10.0.0.11
    model: arista
    username: svc-suzieq
    password: ${SUZIEQ_PW}
  - name: leaf-02
    address: 10.0.0.12
    model: cumulus
    username: svc-suzieq
    password: ${SUZIEQ_PW}

Model drivers exist for Arista EOS, Cumulus/Linux, Cisco NX-OS and IOS-XE, Juniper, FRR and SONiC. Use a read-only service account: SuzieQ only issues show commands, but the account should not be able to do anything else.

Poll and Query

sqPoller -c ~/.suzieq/suzieq-cfg.yml --run-once
docker run -d --name suzieq-poller --restart unless-stopped \
  -v /var/lib/suzieq:/var/lib/suzieq -v ~/.suzieq:/root/.suzieq \
  netenglabs/suzieq:latest sqPoller --run-once

In production, run the poller as a container with a schedule (the image includes a suzieq-poller mode that loops), and keep the Parquet data directory on fast local disk.

Queries use a SQL-like grammar against tables named after the state you collected:

# interactive CLI
suzieq# device show
suzieq# bgp show state != Established
suzieq# interface show adminUp and !operUp
suzieq# route show prefix == 10.20.0.0/16 hostname == leaf-01
suzieq# bgp show hostname == leaf-01 state == Established | count

# time-based query: what changed in the last hour
suzieq# interface show start_time = '1 hour ago'

Because state is stored with timestamps, you can answer change questions without a separate configuration archive: when did this interface last go down, and did its BGP session drop at the same time.

Python API

from suzieq.sqobjects import get_sqobject
import pandas as pd

intf = get_sqobject('interfaces')(context=None)
df = intf.get(hostname=['leaf-01', 'leaf-02'])
print(df[['hostname', 'ifname', 'adminState', 'operState']])

bgp = get_sqobject('bgp')(context=None)
bad = bgp.get(state='NotEstd')
if not bad.empty:
    print(bad[['hostname', 'peer', 'state']])

Wrap checks like the one above in a cron job and you have a lightweight fabric assertion test that does not require a full NMS.

Where SuzieQ Fits

  • Use it for state inventory and equality checks (assert that every leaf sees the same EVPN peers, that no interface is errdisabled, that no BGP session is idle).
  • Keep streaming telemetry for rate and latency metrics; SuzieQ is not a metrics engine.
  • Pair it with intent validation: Batfish for config validation covers what the configuration should do, SuzieQ covers what the devices actually report.
  • Combine with TextFSM parsing when you need a table SuzieQ does not model yet, and drive bulk operations through NAPALM getters.

Scheduling and Alerting on SuzieQ Output

Polling is only half the loop; the other half is asserting. A short Python script on a timer turns SuzieQ into an assertion engine:

from suzieq.sqobjects import get_sqobject
import sys

fail = []
bgp = get_sqobject('bgp')(context=None).get()
not_e = bgp[bgp['state'] != 'Established']
if not not_e.empty:
    fail.append(f"{len(not_e)} BGP sessions not established: "
                f"{list(zip(not_e.hostname, not_e.peer))[:5]}")

intf = get_sqobject('interfaces')(context=None).get()
down = intf[(intf['adminState'] == 'up') & (intf['operState'] == 'down')]
if len(down) > 0:
    fail.append(f"{len(down)} interfaces admin-up but oper-down")

if fail:
    print('\n'.join(fail)); sys.exit(2)
print('fabric assertions passed')

Exit codes 0/2 map cleanly onto monitoring checks, so the same script can run in CI before a change and on a schedule afterwards. Keep the assertions few and meaningful - a check that fires every day gets ignored.

原文链接:https://suzieq.readthedocs.io/en/latest/