Ansible Cisco IOS First Playbook: Commands and VLANs - 夜莺博客

Ansible Cisco IOS First Playbook: Commands and VLANs

Network automation no longer requires a developer background - Ansible's YAML playbooks let any network engineer turn repetitive CLI work into repeatable, safe automation. This tutorial from Packetswitch walks through the complete first session with Ansible on a Cisco Catalyst 9300 running IOS-XE: installing Ansible in a virtual environment, structuring your project with ansible.cfg, inventory and group_vars, running show commands through ios_command, making configuration changes with ios_vlans and ios_l2_interfaces, and understanding the idempotency and desired-state concepts that make playbooks safe to run again and again.

Why Ansible for Cisco IOS

Most network engineers already know the CLI by heart, so the appeal of Ansible is not that it hides the CLI - it is that it makes the CLI repeatable. The same sequence of configure terminal, vlan 30 and interface commands that you type once by hand becomes a file that runs identically on ten switches, records what it changed, and refuses to re-apply work that is already done.

Three properties make Ansible a good fit for IOS-XE estates:

  • Agentless. Nothing is installed on the switch. Ansible opens an SSH session (or NETCONF/eAPI) from your control node, exactly as you would from a terminal.
  • Declarative. You describe the desired state - "VLAN 30 named server_vlan exists" - and the module decides which commands, if any, are needed.
  • Idempotent. A second run of an already-satisfied playbook reports no changes, so the same playbook is safe in a cron job, a pipeline or a maintenance window.

Installing Ansible

python3 -m venv venv
source venv/bin/activate
pip install ansible

A virtual environment keeps your Python packages isolated from other projects.

Verify the install and the modules that matter for IOS before you go further:

ansible --version
ansible-galaxy collection list | grep cisco.ios

If cisco.ios is missing, install the collection that provides ios_command, ios_config, ios_vlans and the rest of the IOS module family:

ansible-galaxy collection install cisco.ios

For NETCONF-based platforms you would also install ansible.netcommon, but for a first IOS-XE session the network_cli connection plus the IOS collection is everything you need.

Project Structure

ansible-network-101/
├── ansible.cfg
├── inventory/
│   ├── group_vars/
│   │   └── switches.yml
│   ├── host_vars/
│   └── hostfile.ini
└── playbook.yml

ansible.cfg holds global settings: host_key_checking = False and inventory = inventory/hostfile.ini. The inventory groups devices ([switches]), and group_vars/switches.yml defines connection parameters for that group:

---
ansible_connection: ansible.netcommon.network_cli
ansible_network_os: cisco.ios.ios
ansible_become: yes
ansible_become_method: enable
ansible_user: admin
ansible_password: cisco123
ansible_become_password: cisco123

The group name in the inventory must match the file name in group_vars.

The inventory file itself is deliberately boring - a group name in brackets and the device addresses underneath:

[switches]
sw1 ansible_host=192.168.1.101
sw2 ansible_host=192.168.1.102
sw3 ansible_host=192.168.1.103

Because the connection variables live in group_vars/switches.yml, every host in the group inherits them. Add a switch to the group and it is immediately reachable with no other change - that separation of inventory from connection detail is what lets the same playbook scale from one lab device to a production estate.

In a real environment you would not keep ansible_password in a text file. Switch to SSH keys, or use ansible-vault to encrypt the group_vars file, and reference the vault password at run time:

ansible-vault encrypt inventory/group_vars/switches.yml
ansible-playbook playbook.yml --ask-vault-pass

First Playbook: Running Show Commands

---
- name: "Ansible 101"
  hosts: switches
  gather_facts: no
  tasks:
    - name: Show Version
      cisco.ios.ios_command:
        commands: show version
      register: output
    - name: Print output
      debug:
        msg: "{{ output.stdout_lines }}"

Run it with ansible-playbook playbook.yml. To automate more switches, simply add them to hostfile.ini.

A few things are worth noticing in that short playbook. gather_facts: no matters on network devices: the default fact-gathering step assumes a Linux host with Python available over SSH, which a switch is not, so it would either fail or waste time. Setting it to no tells Ansible to skip straight to your tasks.

ios_command runs one or more show commands and hands back three useful attributes: stdout (a list of strings, one per command), stdout_lines (the same output pre-split into lines) and stderr. You can stack several commands in a single task, which saves a round trip:

    - name: Collect operational state
      cisco.ios.ios_command:
        commands:
          - show ip interface brief
          - show vlan brief
          - show ip route summary
      register: state
    - name: Display the routing summary
      debug:
        var: state.stdout[2]

This "gather and report" pattern is the safe entry point to automation: it touches nothing, so it is a good way to prove connectivity, credentials and privilege escalation all work before you make any change. If it returns data cleanly you know the hard part is behind you.

Making Configuration Changes: Creating VLANs

---
- name: "Ansible 101 - Configuration Changes"
  hosts: switches
  gather_facts: no
  tasks:
    - name: VLAN Config
      cisco.ios.ios_vlans:
        config:
          - name: server_vlan
            vlan_id: 30
          - name: user_vlan
            vlan_id: 31

Verify on the switch with show vlan. To assign an access port, use cisco.ios.ios_l2_interfaces to set interface Te1/0/1 to access mode in VLAN 30.

The ios_vlans module takes a declarative list. You do not send the raw CLI - you state the end result and the module generates the minimal command set to reach it. Assigning the port is the same idea applied to an interface:

    - name: Put the access port in VLAN 30
      cisco.ios.ios_l2_interfaces:
        config:
          - name: TenGigabitEthernet1/0/1
            access:
              vlan: 30
        state: merged

The state parameter controls the module's intent and is worth learning properly: merged overlays the supplied configuration on the running config, replaced makes the affected section match exactly what you supplied, overridden makes the whole feature match, and deleted removes what you name. Choosing the wrong state is the most common way a first playbook surprises someone, so start with merged, check the result with --check --diff, and only then move to the stronger states.

ansible-playbook playbook.yml --check --diff

Running in check mode shows you exactly which commands the module would generate without committing anything - the network equivalent of a dry run, and the single most valuable habit to build early.

Idempotency and Desired State

Ansible operates on a declarative model: you define the desired state, and Ansible only applies changes when the current state differs. Running the VLAN playbook a second time reports changed=0 because the VLANs already exist and match. This is what makes network automation with Ansible safe for routine operations and repeatable maintenance.

The recap line tells the story on every run:

sw1 : ok=1 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0

changed=0 means "already in the desired state" - not "nothing happened". That distinction is the whole reason you can schedule the same playbook hourly: on the first run it makes the change, and on every run afterwards it confirms the state is still correct. If someone logs in and removes VLAN 30, the next run puts it back, and the change-count on that run tells you drift occurred.

The same property powers configuration drift detection. Point a playbook at your whole estate, capture changed per host, and any device reporting a change is a device whose configuration had wandered from the intended baseline.

Backing Up Configurations

Before you trust any automation to change devices, it should be able to save them. The ios_config module with a backup flag writes a timestamped copy of the running configuration to the control node every run:

    - name: Back up the running configuration
      cisco.ios.ios_config:
        backup: yes
        backup_options:
          dir_path: ./backups
          filename: "{{ inventory_hostname }}.cfg"

Combined with a git repository under ./backups, this gives you a free change history: every playbook run leaves a commit, and a diff between two commits shows exactly what changed on the device and when.

Collecting Facts and Building Reports

The network resource modules can also read state, which makes simple commissioning and audit reports straightforward. Gather interface facts and assert the ports you expect are up:

    - name: Read interface state
      cisco.ios.ios_interfaces:
        state: gathered
      register: ifs
    - name: Warn about down ports
      debug:
        msg: "{{ item.name }} is down"
      loop: "{{ ifs.gathered | selectattr('enabled', 'equalto', true)
                                 | selectattr('state', 'equalto', 'down') | list }}"

The same gathered state works across the resource modules, so a playbook can produce a model of the whole device - VLANs, interfaces, L3 configuration, routing - without a single screen-scrape. That model is what feeds a source-of-truth database, and it is also how you spot the one switch that is not configured like its peers.

Troubleshooting: libssh Issues

If you see ansible-pylibssh not installed, falling back to paramiko and the wheel build fails (common on Apple Silicon Macs), install with:

CFLAGS="-I $(brew --prefix)/include -I ext -L $(brew --prefix)/lib -lssh" pip install ansible-pylibssh

If the fallback to paramiko still works, the message is only a warning and you can proceed - paramiko is slower but functional. Where you do hit real trouble is usually authentication or privilege escalation, and the fix is to test the connection in isolation:

ansible all -m cisco.ios.ios_command -a "commands='show version'" -vvv

The -vvv output shows the exact SSH handshake, the enable step and the command sent, which turns a vague "unreachable" into a specific cause: wrong password, missing ansible_become_method: enable, or a switch that has not had SSH enabled with transport input ssh on the vty lines.

Running Against Part of the Inventory

Once the estate grows you rarely want to run a playbook against every switch at once. The --limit flag restricts a run to named hosts or groups without editing the inventory, and combining it with a serial batch keeps a big change from hitting everything simultaneously:

ansible-playbook playbook.yml --limit sw1,sw2
ansible-playbook playbook.yml --serial 5

--serial 5 runs the play five devices at a time, so a mistake surfaces on the first batch rather than rolling across the whole access layer. Pair it with --check for a dry run, and any change you make can be rehearsed on one switch, confirmed in show vlan, and only then widened to the fleet. This staged rollout habit - limit, batch, verify, expand - is what separates a lab playbook from one you would trust in a production window.

Common Pitfalls and Best Practices

  • Always run --check --diff first. It costs seconds and prevents the wrong-state mistake.
  • Keep credentials out of git. Use SSH keys or ansible-vault.
  • Remember gather_facts: no. Network devices are not Linux hosts.
  • Prefer network resource modules to ios_config with raw lines. The resource modules understand the feature and give you idempotency for free, whereas pushing raw CLI lines is idempotent only by accident.
  • Back up before every change and keep the backups in version control.
  • Start with reads. Prove connectivity and credentials with show commands before your first write.

Follow those six rules and the jump from a lab playbook to a production one is mostly a matter of adding hosts to hostfile.ini.

Related Reading

原文链接:https://www.packetswitch.co.uk/ansible-network-automation-101