GitOps for Network Configuration: Pipeline Design - 夜莺博客

GitOps for Network Configuration: Pipeline Design

Network teams have been storing device configurations in Git for years, but storing is not GitOps. The distinction is that in a real GitOps workflow the repository is the source of truth, every change goes through a pipeline that validates it before it touches a device, and the deployed state is continuously reconciled against the declared state. This article lays out a practical design for network GitOps: repository structure, what to validate in CI, how to stage the rollout, and where human approval belongs.

Git vs GitOps vs config backup

Backup tool (Oxidized/RANCID)   device -> Git, read-only history
Config templating (Jinja2)      source generates config text
GitOps                          Git is authoritative; pipeline validates and
                                applies; drift is detected and corrected

The backup comparison is worth internalising because most organisations already have half the tooling: Oxidized setup and RANCID vs Oxidized show the read-only end of the spectrum.

Repository layout

network/
  inventory/
    devices.yaml            # hostname, platform, mgmt ip, tags
    groups.yaml             # site/role groupings
  templates/
    cisco_iosxe/base.j2
    arista_eos/base.j2
  data/
    site-a/leaf-01.yaml     # per-device intent (VLANs, VRFs, peers)
  rendered/                 # CI artefact: full config per device
  policies/
    golden-config.yaml      # must-have lines (NTP, AAA, syslog)
    forbidden.yaml          # must-not-have (telnet, default communities)
  .github/workflows/ci.yml

Separating intent (data/) from syntax (templates/) is what makes the repo reviewable by non-specialists and keeps platform differences out of the business data. The rendering half is plain Jinja2 templating.

Stage 1 — render and lint on every pull request

# .github/workflows/ci.yml (excerpt)
- name: render
  run: python render.py --inventory inventory --data data --out rendered
- name: schema check
  run: python validate_schema.py rendered/  # YAML schema per platform
- name: policy check
  run: python golden_config_check.py policies/ rendered/
- name: syntax check
  run: python batfish_check.py rendered/     # parse + reachability tests

Two classes of check matter: syntactic (will the device accept this?) and semantic (will the forwarding behaviour still be correct?). Batfish covers the second class without hardware — it builds a model of the rendered configs and lets you assert reachability, ACL behaviour and loop freedom. That is the core value of the pipeline: see Batfish config validation.

Stage 2 — staged rollout

1. merge to main          -> rendered artefacts published
2. canary deploy          -> 1 device in 1 site, auto-rollback on failure
3. hold 15 minutes        -> verify with show commands + telemetry
4. wave deploy            -> grouped by role (spines, then leaves)
5. post-checks            -> diff running vs intended, alert on drift

Waves should be defined by blast radius, not by device count. A two-spine datacentre goes one spine at a time even if it only has two devices.

Stage 3 — drift detection

for dev in inventory:
    running  = show running-config
    intended = rendered/.cfg
    if normalize(running) != normalize(intended):
        report_drift(dev, unified_diff(running, intended))

Drift is expected — an operator will always add a description or an ACL entry in a hurry. What matters is that drift is detected and either reconciled or folded back into intent. Suppress known-cosmetic differences (timestamps, encrypted secrets, counters) or the report becomes noise nobody reads.

Where human approval belongs

Auto-approve:   description changes, template-only changes to unused vars
Require review: any change to routing policy, ACLs, or a spine/edge
Require window: changes to firewalls, load balancers, WAN edges
Forbiddden:     direct CLI commits not going through the pipeline

Practical pitfalls

Secrets in Git            -> use SOPS/age or a vault-backed render step
Non-idempotent commands   -> "no ..." lines that flap; test re-render = no diff
Ordering dependencies     -> VLAN must exist before SVI; encode in template
Partial failure handling  -> never leave a device half-committed; use commit
                             confirmed / checkpoint-rollback where available
Rollback plan             -> every wave needs a defined revert commit

The recurring failure mode is a pipeline that renders and pushes correctly but has no revert story. Have the rollback be a commit like any other change, and test it before you need it. The config-compliance pipeline pattern with Jenkins and Ansible is described in Jenkins network configuration compliance.

FAQ

Q: Do I need Kubernetes and Argo CD? No. Argo CD is convenient for continuous reconciliation, but a CI pipeline plus a scheduled drift job delivers the same guarantees for network devices.
Q: What about vendor APIs rather than CLI? Prefer NETCONF/gNMI for rendering fidelity (YANG types), CLI for legacy. The pipeline shape is identical.
Q: How do I convince management? Start read-only: publish diff reports for every change. The first unauthorised change caught by CI usually ends the debate.

原文链接:https://argo-cd.readthedocs.io/en/stable/