Cisco IOS Ansible Playbook: Standard + Unique Configs - 夜莺博客

Cisco IOS Ansible Playbook: Standard + Unique Configs

Moving switch configuration to code starts with separating what is identical across the fleet from what is unique per device — and that split maps perfectly onto Ansible's role structure. This tutorial (from GetLabsDone) builds a two-role playbook for Cisco IOS: a cisco_standard role applying banner, DNS, NTP, enable secret and local user to every device, and a cisco_unique role consuming host_vars for per-device hostname and VLANs. The result is a repeatable, template-driven way to deploy new switches and migrate production gear to infrastructure-as-code.

Prerequisites: Collections, ansible.cfg and a Directory Skeleton

Everything runs from a Linux control node. Install Ansible, the cisco.ios collection that supplies ios_banner, ios_system, ios_ntp_global, ios_config, ios_vlans and friends, plus ansible.netcommon for the network_cli connection plugin.

pip install ansible
ansible-galaxy collection install cisco.ios ansible.netcommon
ansible-galaxy collection list | grep -E "cisco|netcommon"

Because the roles live next to the playbooks, point Ansible at the right directories once in ansible.cfg and never pass -i again. The roles_path entry is what lets a task say roles: cisco_standard without a relative path, and host_key_checking = False avoids the interactive fingerprint prompt on freshly imaged switches.

[defaults]
inventory       = inventories/production/cal/inventory.ini
roles_path      = ./roles
host_key_checking = False
gather_facts    = False
stdout_callback = yaml

[persistent_connection]
connect_timeout = 60
command_timeout = 60

Note gather_facts = False. The default setup module cannot log into a network device, so it must be disabled globally rather than per play — forget it and every network playbook throws a connection error on the implicit first task.

Role Layout

roles/
├── cisco_standard/
│   ├── defaults/main.yml    # common config variables
│   └── tasks/main.yml
└── cisco_unique/
    └── tasks/main.yml
host_vars/production/cal/cal-hq-acc-sw-04.yml
inventories/production/cal/inventory.ini
playbooks/cisco_switch_playbook.yml

Two roles, one job each. cisco_standard holds values that must never drift between devices — banner, DNS servers, NTP sources, enable secret, local admin account. cisco_unique holds the parts that legitimately differ: hostname, VLAN database, interface descriptions. Create the skeleton with ansible-galaxy init roles/cisco_standard and ansible-galaxy init roles/cisco_unique, then delete the unused directories so the repo stays readable.

Standard Configuration Variables

# roles/cisco_standard/defaults/main.yml
std_config:
  ios_banner:
    - banner: |
        Welcome to GetLabsDone Network
        Unauthorized access is strictly prohibited.
  dns:
    - fqdn: getlabsdone.local
      dns_1: 8.8.8.8
      dns_2: 4.2.2.2
  ntp:
    - server1: time.google.com
      server2: time1.google.com
      logging: true
  en_password:
    - password: gld_pass
  local_user:
    - name: gldadmin
      password: testpass

Everything is expressed as a list of dictionaries, and each task loops over one key with with_items. That shape is deliberate: adding a second DNS server, a second NTP source or a second local account is a YAML edit, not a code change. Keeping the values in roles/cisco_standard/defaults/ means a child playbook can override any of them through group_vars without touching the role.

Standard Tasks — Use the Right Module

- name: configure the login banner
  cisco.ios.ios_banner:
    banner: login
    text: "{{ item.banner }}"
  with_items: "{{ std_config.ios_banner }}"

- name: configure DNS on the system
  cisco.ios.ios_system:
    lookup_enabled: yes
    domain_name: "{{ item.fqdn }}"
    name_servers: ["{{ item.dns_1 }}", "{{ item.dns_2 }}"]
  with_items: "{{ std_config.dns }}"

- name: setup NTP across the board
  cisco.ios.ios_ntp_global:
    config:
      servers:
        - server: "{{ item.server1 }}"
        - server: "{{ item.server2 }}"
      logging: "{{ item.logging }}"
    state: replaced
  with_items: "{{ std_config.ntp }}"

Prefer purpose-built modules (ios_banner, ios_system, ios_ntp_global) and fall back to ios_config only when no module exists. Use no_log: true for anything containing passwords. The difference is idempotency: ios_ntp_global with state: replaced compares the running NTP block against the desired one and touches nothing when they match, whereas a raw ios_config line is pushed blindly and always reports changed. The remaining two standard tasks follow the same pattern:

- name: set the enable secret
  cisco.ios.ios_config:
    lines:
      - enable secret {{ item.password }}
  with_items: "{{ std_config.en_password }}"
  no_log: true

- name: configure the local admin user
  cisco.ios.ios_config:
    lines:
      - username {{ item.name }} privilege 15 secret {{ item.password }}
  with_items: "{{ std_config.local_user }}"
  no_log: true

Unique Configuration via host_vars

Store per-device data in host_vars/<env>/<site>/<hostname>.yml — hostname, VLANs, interfaces, ACLs. When migrating a production switch, convert its running config to YAML and drop it in host_vars.

# host_vars/production/cal/cal-hq-acc-sw-04.yml
hostname: cal-hq-acc-sw-04
vlans:
  - vlan_id: 10
    name: USERS
  - vlan_id: 20
    name: VOICE
  - vlan_id: 99
    name: MGMT
interfaces:
  - name: GigabitEthernet1/0/1
    description: USERS-ACCESS-A1
    mode: access
    vlan: 10
  - name: GigabitEthernet1/0/24
    description: UPLINK-TO-DIST
    mode: trunk
    allowed_vlans: [10, 20, 99]

Inventory names must match the host_vars filename for Ansible to pick the file up automatically, and mismatched case is the single most common reason per-host variables silently do not apply. The cisco_unique role then does nothing but render those values onto the device:

# roles/cisco_unique/tasks/main.yml
- name: set the device hostname
  cisco.ios.ios_hostname:
    config:
      hostname: "{{ hostname }}"
    state: merged

- name: provision the VLAN database
  cisco.ios.ios_vlans:
    config: "{{ vlans }}"
    state: merged

- name: configure access and trunk interfaces
  cisco.ios.ios_l2_interfaces:
    config:
      - name: "{{ item.name }}"
        access:
          vlan: "{{ item.vlan }}"
        trunk:
          allowed_vlans: "{{ item.allowed_vlans | default(omit) }}"
    state: replaced
  loop: "{{ interfaces }}"

Inventory

[cal]
cal-hq-acc-sw04 ansible_host=10.1.11.7
[cal:vars]
env=production
site=cal
ansible_ssh_user=[SSH_USERNAME]
ansible_ssh_pass=[SSH_PASSWORD]
ansible_network_os=ios
ansible_connection=network_cli
ansible_become_method=enable
ansible_become=yes
ansible_become_password=[ENABLE_PASSWORD]

Two variables do all the heavy lifting. ansible_network_os=ios loads the Cisco IOS prompt-handling logic, and ansible_connection=network_cli tells Ansible to hold one persistent SSH session for the whole play instead of reconnecting for every task. Replace the bracketed placeholders with {{ vault_ssh_user }}, {{ vault_ssh_password }} and {{ vault_enable_password }}, then store the real values in an encrypted group_vars/cal/vault.yml.

The Top-Level Playbook

# playbooks/cisco_switch_playbook.yml
---
- name: standardise and personalise Cisco IOS access switches
  hosts: cal
  gather_facts: false
  connection: network_cli
  serial: 1
  roles:
    - cisco_standard
    - cisco_unique

Role order matters and is the whole point of the split: the standard role runs first so the device has a known baseline (DNS, NTP, admin user) and the unique role then layers the site-specific configuration on top. serial: 1 is not strictly required for two switches, but it is the habit to build — when the play grows to a hundred devices, one-at-a-time rollout with a short verification pause is what keeps an error from becoming an outage.

Playbook and Device Prerequisites

ansible-playbook cisco_switch_playbook.yml -i inventories/production/cal/inventory.ini

Before automating, the device needs: SSH enabled, a local account or RADIUS login, and an IP + default gateway reachable from the Ansible host. Run the playbook once and the switch gains its banner, DNS, NTP, hostname and VLANs in a single, auditable change set.

! On the switch, verify SSH is ready
ip domain-name getlabsdone.local
crypto key generate rsa modulus 2048
username ansible privilege 15 secret <password>
line vty 0 4
 transport input ssh
 login local
!
! From the control node, prove reachability before you run anything
ansible -i inventories/production/cal/inventory.ini cal -m ansible.netcommon.cli_command -a "command='show version' --ask-vault-pass"

Dry Run, Secrets and Verification

Never point a brand-new playbook at production unchecked. The sequence that has kept this workflow safe is: syntax check, then a diff-only run, then a single-host run, and only then the site.

ansible-playbook cisco_switch_playbook.yml --syntax-check
ansible-playbook cisco_switch_playbook.yml --check --diff --ask-vault-pass
ansible-playbook cisco_switch_playbook.yml --limit cal-hq-acc-sw04 --ask-vault-pass

--check --diff prints the exact configuration lines Ansible would add or replace, which is the moment to catch a wrong VLAN or an accidental trunk change. After the real run, confirm the result on the device rather than trusting the recap:

show running-config | include hostname|vlan|ntp|name-server|banner
show vlan brief
show interfaces status

Then run the playbook a second time. A correct playbook reports changed=0 on the repeat run; anything still showing changed is a module that cannot reach a stable state — usually an ios_config line whose format differs from what the device writes back (ordering, capitalisation, blank lines). Fix the desired-state line rather than accepting the noise, because a permanently-changed task destroys the signal that makes the output useful.

Rolling Back and Change Safety

Cisco IOS gives you a built-in undo: reload in combined with a configuration archive. Set archive and path once, and every write memory keeps 14 checkpoints you can roll back to with configure replace flash:<file> force. In the playbook, capture a pre-change backup on every run so the rollback point always exists:

- name: snapshot the running config before any change
  cisco.ios.ios_config:
    backup: yes
    backup_options:
      filename: "{{ inventory_hostname }}-pre.cfg"
      dir_path: "./backups"
  delegate_to: localhost

For the highest-stakes changes, add reload in 10 in a task, apply the configuration, verify, and only then cancel the scheduled reload. If the change breaks management access the device reboots into the last saved configuration on its own — the network equivalent of a dead-man switch.

Troubleshooting

"Timeout waiting for privilege escalation" means the enable password is wrong or missing: check ansible_become_password and confirm ansible_become_method=enable, not sudo. Tasks succeed but nothing changes is usually a host_vars filename that does not match inventory_hostname exactly, so the unique role renders empty values and the modules quietly do nothing. Connection errors on the first task almost always trace back to gather_facts still being enabled, or to ansible_network_os set to cisco.ios.ios in the inventory when the collection registers it as ios — both forms work with a recent collection, but the value must be one of them. Banner task fails when text contains leading whitespace: use the block scalar | as shown in defaults/main.yml and strip trailing spaces. For anything else, -vvv prints the raw CLI dialogue and shows precisely which prompt Ansible was waiting for.

FAQ

Why two roles instead of two task files? Roles give you defaults, handlers, and a documented variable interface. Splitting standard from unique also makes the intent obvious to the next engineer: if it is in the standard role it must be identical everywhere, and if it is in host_vars it is allowed to differ.

Can I use this for a whole campus? Yes — put each switch in [cal], give it a host_vars file, and add sites as new groups. The playbook itself never changes as the fleet grows, which is the entire argument for infrastructure-as-code.

Where should the passwords live? In Ansible Vault, never in defaults/main.yml. The example above uses plaintext values only so the data model is readable; in production every password field becomes a vault variable and every task that renders one gets no_log: true.

Start from zero with our Ansible network automation 101 tutorial, then go deeper with Ansible network modules: ios_config, facts and config backup and the multi-vendor Ansible and Jinja2 guide. For the operations side of the same platforms, see input-drop troubleshooting on IOS XR.

原文链接:https://getlabsdone.com/how-to-get-started-with-cisco-ios-ansible-playbook/