Juniper Apstra: Intent-Based DC Fabric Automation - 夜莺博客

Juniper Apstra: Intent-Based DC Fabric Automation

Juniper Apstra is an intent-based networking controller for data center fabrics: instead of pushing per-device CLI, you describe the fabric you want — racks, leaf and spine roles, tenant connectivity — and Apstra generates, validates and deploys the EVPN-VXLAN configuration underneath. This guide walks through the Apstra data model (agents, device profiles, blueprints), the staging-versus-committed workflow that lets you preview every change, and the Intent-Based Analytics (IBA) probes that flag drift long before a customer ticket arrives. It is written for engineers who already understand EVPN-VXLAN and want to move fabric operations from copy-paste to a reviewable pipeline. Every command shown is a reference-node or REST API example you can run against your own Apstra instance.

Why intent beats per-device CLI on a leaf-spine fabric

A 4-spine, 24-leaf EVPN fabric has roughly 600 configuration objects that must agree: ASNs, VTEP loopbacks, BGP group policies, VLAN-to-VNI mappings, anycast gateways, and the underlay IGP. Typed by hand, each change is a chance to create a silent inconsistency — a leaf with a stale VNI, a VRF leaking into the wrong tenant, an underlay link with the wrong MTU. Apstra's model is a graph of intent: you add a rack, claim the devices, define the role (leaf or spine) and let the controller render configuration. Nothing lands on a device until you commit, and every commit is stored as an immutable revision you can diff and roll back.

Architecture: server, agents, devices, blueprints

  • Apstra server — the controller (VM or appliance) holding blueprints, analytics and the API. In 4.x+ it can run as a single server or in a clustered HA pair.
  • Device agents — small VMs per managed switch that proxy telemetry and commands, so the controller never needs SSH reachability from a single subnet.
  • Device profiles — declarative descriptions of each switch model: port counts, breakout options, speeds. Profiles are what let the same logical blueprint land on Juniper, Arista or SONiC hardware.
  • Blueprint — the instance of intent for one fabric: physical layout plus logical elements (virtual networks, routing zones, connectivity templates).

Discover devices and build the blueprint

After the server is up, add managed devices and let Apstra run discovery. The REST API is the fastest way to script this:

# authenticate on the reference node
curl -k -X POST https://apstra.example.com/api/auth/login \
  -H "Content-Type: application/json" \
  -d '{"username":"admin","password":""}' -c cookie.jar

# list systems that have been discovered but not yet managed
curl -k -b cookie.jar https://apstra.example.com/api/system-agent-systems | jq '.items[].id'

# create a managed device from a system id
curl -k -b cookie.jar -X POST https://apstra.example.com/api/managed-devices \
  -H "Content-Type: application/json" \
  -d '{"device_profile":"juniper_vjunos_switch","agent_profile":"vJunos",
       "management_ip":"10.0.0.21","system_id":""}'

Then create the blueprint and describe intent — spine count, leaf roles and link speeds:

curl -k -b cookie.jar -X POST https://apstra.example.com/api/blueprints \
  -H "Content-Type: application/json" \
  -d '{"label":"DC1-Fabric","design":"l3clos"}'

# list blueprints and grab the id
curl -k -b cookie.jar https://apstra.example.com/api/blueprints | jq '.items[] | {id,label}'

The staging workflow: preview, diff, commit

Apstra separates staging from active. You add a rack, assign leaf and spine roles, create a virtual network and refine a connectivity template, then stage the change. The commit dialog shows a per-device diff of exactly what will be pushed — including the underlay, overlay and routing-zone deltas. Two habits matter:

  1. Commit in small batches. Adding a rack touches underlay peering on every spine, so stage it alone instead of bundling it with a tenant change.
  2. Read the diff as a review artifact. A commit that unexpectedly changes a spine's BGP policy is a bug in your intent, not something to click through.

On the CLI you can see the rendered configuration the controller would push:

apstra@server> show blueprint DC1-Fabric staged
apstra@server> show blueprint DC1-Fabric diff
apstra@server> show blueprint DC1-Fabric config-errors

Every commit becomes a numbered revision. Rolling back is a first-class operation, which makes change windows dramatically shorter: if a rack expansion breaks a tenant, restore the previous revision and the controller re-renders the fabric.

Intent-Based Analytics: catching drift automatically

IBA probes are continuous checks against the state Apstra collects from devices and agents. Typical probes you should have on day one:

  • BGP session state — every configured EVPN peer should be Established in the expected address family.
  • Underlay link health — member links up, speeds as designed, MTU consistent across the fabric.
  • VNI/VLAN consistency — the same VNI must not map to two VLANs, and no VLAN should be missing on a leaf that hosts the tenant.
  • Interface error counters — CRC and input errors trending upward on the same port point to a bad optic or patch lead.

Because probes are stateful, an anomaly is raised when the condition is violated and cleared when it returns — a much better signal than a one-shot SNMP poll. Alerting can be forwarded to a webhook, Slack or your existing NMS.

Verification checklist before you call it in production

# on any leaf (Juniper example)
show bgp summary
show evpn database
show ethernet-switching vtep
show route table __default_evpn__.evpn.0 extensive | match "VNI|Route Distinguisher"
# fabric-wide checks from Apstra
curl -k -b cookie.jar https://apstra.example.com/api/blueprints/<id>/anomalies | jq '.items[].type'
curl -k -b cookie.jar https://apstra.example.com/api/blueprints/<id>/probes | jq '.items[] | {name,status}'

Automation beyond the UI

Once the fabric lives in Apstra, treat the blueprint as an API surface. The Terraform provider model used for multi-vendor network automation applies directly: create a virtual network and a connectivity template in code, review the plan, then apply. For teams standardising on Ansible, the same objects can be driven from the Apstra collection, but keep the controller as the source of truth — mixing hand-run CLI with intent-based automation is how drift starts.

Operational best practices

  • Model real racks, not abstractions: one rack object per physical rack makes sparing and replacement predictable.
  • Use ESI-LAG multihoming objects for dual-homed servers rather than hand-built LAGs on individual leaves.
  • Version your intent: export the blueprint and keep it in Git, then promote it between lab and production controllers.
  • Test upgrades in a lab blueprint first — Apstra can render configuration for a newer Junos release against a mirror fabric before any production device reboots.
  • Never disable config validation to push a change through; the error is almost always a missing MTU, ASN or routing-zone object.

原文链接:https://www.juniper.net/documentation/us/en/software/apstra/apstra-user-guide/index.html