Terraform for Network Automation: Cisco, Junos and Arista - 夜莺博客

Terraform for Network Automation: Cisco, Junos and Arista

Ansible made network automation accessible by pushing CLI and NETCONF over SSH. Terraform attacks the same problem from the other end: it keeps state, so it can tell you what changed outside the pipeline. For a fleet of VLANs, interfaces and BGP peers defined in Git, that difference matters more than the syntax. This guide covers the vendor provider landscape, real resource examples, and the drift-detection pattern that turns Terraform into a compliance tool.

The mental model worth internalising before writing any HCL: Ansible is a remote-execution engine that happens to be idempotent if you write your tasks carefully, while Terraform is a state machine that compares a desired graph against a recorded inventory on every run. Ansible asks "is this line in the config?" Terraform asks "does the object I believe exists still match what I declared?" That second question is why Terraform produces a plan, why it refuses to act when state and reality disagree, and why it can tell you that a junior engineer pasted a VLAN by hand at 2 a.m.

Provider cheat sheet

  • Cisco IOS XE — CiscoDevNet/iosxe, speaking RESTCONF or NETCONF to the device.
  • Cisco NX-OS — CiscoDevNet/nxos, RESTCONF.
  • Cisco ACI — CiscoDevNet/aci, APIC API.
  • Juniper Junos — Juniper/junos, NETCONF over SSH.
  • Arista EOS — aristanetworks/ceoslab for lab/veos, aristanetworks/cloudvision for a controller-managed estate.

The pattern to notice: every mature provider is model-driven. Terraform is not parsing show run; it is writing YANG-modelled objects through an API. That is why the providers need RESTCONF/NETCONF/eAPI enabled on the device, and why the resources look like data models rather than CLI lines.

Prerequisites: turn on the model-driven APIs first

Nothing in this article works against a factory-default switch. Before Terraform can talk to a device, the device must expose a programmatic interface, and that interface must have a user with the right privilege level and, on newer code, an authorised RESTCONF profile.

! IOS XE 16.x and later: enable the HTTP server and RESTCONF
ip http secure-server
ip http authentication local
restconf
!
! NETCONF is enabled by default on IOS XE 16.6+, but confirm:
show platform software yang-management process

! Junos: NETCONF over SSH is explicit
set system services netconf ssh
set system login user terraform class super-user authentication ssh-ed25519 "ssh-ed25519 AAAA..."

! Arista EOS: eAPI over HTTPS
management api http-commands
   protocol https
   no shutdown
   vrf MGMT

Resist the temptation to give Terraform a privilege 15 local account on production gear. A better pattern on IOS XE and NX-OS is a dedicated priv-15 role scoped by parser view or the TACACS+ command authorisation set that only allows the write verbs the provider actually needs. For a deeper look at how the transport is enabled and tested by hand, see our walkthrough on enabling and verifying NETCONF/RESTCONF on IOS XE with curl.

Pin your providers

Provider versions move. A breaking change in a minor release can rewrite your interface objects while you sleep. Always pin, and commit the lock file.

terraform {
  required_version = ">= 1.6"
  required_providers {
    iosxe = { source = "CiscoDevNet/iosxe", version = "~> 0.5" }
    junos = { source = "Juniper/junos",    version = "~> 2.0" }
  }
}

# terraform.tfstate.lock.hcl is generated by `terraform init`
# and MUST be committed to Git. Without it, two CI runners
# can resolve different provider versions and disagree.

Cisco IOS XE: VLAN and access port

terraform {
  required_providers {
    iosxe = { source = "CiscoDevNet/iosxe", version = "~> 0.5" }
  }
}

provider "iosxe" {
  username = var.username
  password = var.password
  url      = "https://core-sw-1.lab.example.com"
}

resource "iosxe_vlan" "data" {
  vlan_id = 100
  name    = "DATA"
}

resource "iosxe_interface_ethernet" "gi1_0_1" {
  type        = "GigabitEthernet"
  name        = "1/0/1"
  description = "uplink-A"
  enabled     = true
  switchport_mode_access_vlan = iosxe_vlan.data.vlan_id
}

Two things to notice. First, the interface references iosxe_vlan.data.vlan_id, so Terraform builds a dependency graph and creates the VLAN before the port. Second, enabled = true is not decoration: leaving it out on some provider versions leaves the interface in shutdown state and you will spend an afternoon wondering why a brand-new port carries no traffic.

Scale the pattern with for_each over a map of access ports. Resist count, which produces index-based addresses that renumber everything when you remove a middle element.

locals {
  access_ports = {
    "Gi1/0/5"  = { vlan = 100, desc = "desk-A1" }
    "Gi1/0/6"  = { vlan = 100, desc = "desk-A2" }
    "Gi1/0/7"  = { vlan = 200, desc = "printer" }
  }
}

resource "iosxe_interface_ethernet" "access" {
  for_each = local.access_ports

  type        = "GigabitEthernet"
  name        = replace(each.key, "Gi", "")
  description = each.value.desc
  enabled     = true
  switchport_mode_access_vlan = each.value.vlan
}

Cisco NX-OS: VLAN plus a routed sub-interface

NX-OS resources are similarly model-driven, but the module exposes a different attribute set. The important discipline is the same: reference other resources rather than hard-coding IDs, so the graph stays acyclic and self-documenting.

provider "nxos" {
  username = var.username
  password = var.password
  url      = "https://spine-1.lab.example.com"
  insecure = true   # lab only; use a real cert chain in production
}

resource "nxos_vlan" "web" {
  vlan_id = 300
  name    = "WEB"
}

resource "nxos_physical_interface" "e1_1" {
  interface_id = "eth1/1"
  description  = "to compute rack-1"
  mode         = "access"
  access_vlan  = "vlan-${nxos_vlan.web.vlan_id}"
}

Juniper Junos: security zones and routing instances

provider "junos" {
  alias      = "edge"
  ip         = "edge-fw-1.lab.example.com"
  username   = "terraform"
  sshkey_pem = file("~/.ssh/id_ed25519")
}

resource "junos_security_zone" "trust" {
  provider = junos.edge
  name     = "trust"
}

resource "junos_routing_instance" "vrf_blue" {
  provider            = junos.edge
  name                = "VRF-BLUE"
  type                = "virtual-router"
  route_distinguisher = "65000:100"
}

Junos is the cleanest fit for Terraform of the three vendors, because the entire configuration is already a structured hierarchy in the candidate configuration. The provider stages changes into the candidate database and then commits, so a failed plan cannot leave the device half-configured — the commit either succeeds entirely or not at all. That atomicity is exactly the property IOS XE and NX-OS do not give you for free, and it is worth calling out in any design review.

Bind interfaces into the zone with a separate resource rather than expecting the zone object to do it:

resource "junos_security_zone" "untrust" {
  provider = junos.edge
  name     = "untrust"

  inbound_services = ["ssh", "ping"]
}

resource "junos_interface" "ge0_0_0" {
  provider = junos.edge
  name     = "ge-0/0/0"
  unit {
    name  = "0"
    family_inet {
      address {
        cidr = "203.0.113.1/30"
      }
    }
  }
}

Arista: configuration through CloudVision

resource "cvp_configlet" "site_dns" {
  name = "site-dns"
  config = <<-EOT
    ip name-server 10.0.0.53
    ip name-server 10.0.0.54
  EOT
}

Arista splits into two worlds. In a greenfield lab, aristanetworks/ceoslab manages containerised cEOS instances directly, which is superb for CI: spin up a topology, apply a plan, tear it down. In a controller-managed production estate, CloudVision is the source of truth and Terraform should drive the controller, not the switches. If you are only just getting comfortable with the EOS CLI itself, our Arista EOS CLI cheat sheet is a good companion.

State is the dangerous part

Terraform state is a database, and it contains secrets. Provider credentials, pre-shared keys and SNMP community strings all land in the state file in plaintext. Treat it with the same care as a private key.

  • Store state in a remote backend with locking — an S3 bucket with DynamoDB locking, GitLab-managed state, or Terraform Cloud. Never in Git.
  • Encrypt the backend at rest and restrict who can read it.
  • Mark sensitive variables with sensitive = true so they are redacted from CLI output. Note that this only redacts console output; the state still holds the value.
  • Split the blast radius. One state file per site or per device role means a bad plan takes out a rack, not the whole estate.
variable "password" {
  type      = string
  sensitive = true
}

terraform {
  backend "s3" {
    bucket         = "net-tfstate-prod"
    key            = "site-lon-1/network.tfstate"
    region         = "eu-west-1"
    dynamodb_table = "net-tfstate-lock"
    encrypt        = true
  }
}

Environments with variables and tfvars

The same module should drive lab, staging and production. Keep the topology in code and the environment-specific values in a .tfvars file, so promotion is a review of a diff rather than a rewrite.

# environments/prod.tfvars
username        = "terraform-svc"
vlans = {
  100 = "DATA"
  200 = "VOICE"
  300 = "WEB"
}
bgp_asn = 65001
terraform plan  -var-file=environments/prod.tfvars
terraform apply -var-file=environments/prod.tfvars

The killer use case: drift detection in CI

terraform plan -detailed-exitcode
# exit 0 = no changes expected
# exit 2 = drift: someone changed the device outside the pipeline

Run that nightly for the managed subset of your estate and wire the exit code into Slack or PagerDuty. Exit code 2 with an empty diff usually means an out-of-band CLI change; exit code 1 is a real error such as a device being unreachable. Keep the blast radius small at first — VLANs, loopbacks, NTP and AAA are good starting resources because a wrong value is visible immediately rather than routing blackholing traffic for an hour.

A minimal GitLab CI job that reports drift without ever applying it:

drift-check:
  stage: verify
  image: hashicorp/terraform:1.7
  script:
    - terraform init -input=false
    - |
      terraform plan -detailed-exitcode -input=false -out=plan.tfplan
      rc=$?
      if [ "$rc" = "2" ]; then
        echo "DRIFT DETECTED"
        exit 2
      elif [ "$rc" = "1" ]; then
        echo "PLAN ERROR"
        exit 1
      fi
  rules:
    - if: $CI_PIPELINE_SOURCE == "schedule"

Feed the human-readable plan output into the alert. A message that says "VLAN 300 name changed from WEB to WEB-OLD on core-sw-1" is actionable; "drift detected" is not.

Importing gear you did not create

Nobody starts with a greenfield. The pragmatic path onto an existing fleet is terraform import, resource by resource, plus a strict rule that once a resource is imported it is no longer allowed to be edited by hand. Import the low-risk objects first, confirm the plan is empty after import, then expand.

# import an existing VLAN 100 into the state under a named resource
terraform import iosxe_vlan.data 100
terraform plan      # should be empty if your HCL matches the device

If the plan is not empty after import, your HCL does not describe the device. Fix the HCL until the plan says "no changes", and only then move to the next object. Chasing ten mismatches at once is how teams conclude that Terraform "does not work for networks".

Common pitfalls and how to get unstuck

Four failures account for almost every "Terraform broke my switch" ticket.

  • The provider times out instead of failing. Most NETCONF/RESTCONF providers default to a short connection timeout. If a device is slow to answer, raise the timeout attribute rather than assuming the credentials are wrong.
  • The plan wants to recreate a resource you barely touched. Usually a provider attribute type mismatch — you wrote "100" where the schema wants 100, or vice versa. Read the schema with terraform providers schema -json before guessing.
  • Someone edits the device and Terraform overwrites it. That is Terraform working as designed. The fix is process, not code: no hand edits on managed objects, and a drift alert that pages a human when the plan is non-empty.
  • Two engineers apply at once. Without remote-state locking, you get a corrupted state file and a very bad afternoon. Enable locking on day one.

Debug a single object in isolation with the targeted plan and the provider's own log level — this prints the exact HTTP or NETCONF payload Terraform sent:

terraform plan -target=iosxe_vlan.data
TF_LOG=DEBUG terraform apply -auto-approve 2>&1 | tee tf-debug.log

Reading the raw payload against the vendor's YANG model resolves the argument far faster than reading the provider source. When the debug log shows a 400 or a NETCONF bad-element, you are looking at a schema mismatch, not a network problem.

Where Terraform is not the answer: imperative, ordered changes (software upgrades, a sequence of reloads), and anything the provider models poorly. Keep Ansible for those, and let Terraform own the declarative steady state. The combination — Ansible for execution, Terraform for the intended configuration, a source of truth such as NetBox for the data — is what most mature teams converge on.

Related reading: Ansible network automation playbook examples, NetBox IPAM guide and NAPALM getters and configuration diff.