Ansible Network Automation 101: Cisco IOS First Steps - 夜莺博客

Ansible Network Automation 101: Cisco IOS First Steps

Ansible remains the fastest on-ramp to network automation for engineers who have never written code: playbooks are YAML, modules are declarative, and there is no agent to install on the devices. This 101 guide, based on the PacketSwitch tutorial series, walks a complete first project against a Cisco Catalyst 9300 — installing Ansible in a virtualenv, building a clean directory structure, running show commands with ios_command, and making idempotent configuration changes with ios_vlans. You will also learn the two concepts that make automation safe: idempotency and desired state.

If you are a network engineer who has spent a decade typing into a CLI, the idea of "writing code" to manage switches can feel like a career change. It is not. Ansible deliberately keeps the programming to a minimum: the language is YAML, a plain text format of key-value pairs and lists, and the intelligence lives in modules that other people already wrote. Your job is to describe what the network should look like, not to write the algorithms that get it there. By the end of this guide you will have installed the tooling, structured a project the way professionals do, and pushed a real change to a real switch — twice, to prove it is safe.

What You Need Before You Start

  • A control node: your laptop or a Linux jump host. Windows users should use WSL2; Ansible does not run natively on Windows.
  • Python 3.9 or later on the control node. Everything else installs into a virtual environment.
  • SSH reachability to the managed switch, with a privileged account that can enter enable mode.
  • A Cisco IOS switch (Catalyst 9300 in this tutorial) with ip domain-name set, since SSH key generation depends on it.

You do not need to install anything on the switch. Ansible's network connection plugins drive the existing CLI over SSH, so the device stays a stock device — no agent, no daemon, no package to maintain across hundreds of boxes.

Install Ansible

python3 -m venv venv
source venv/bin/activate
pip install ansible

Always work inside a virtual environment. It keeps the Ansible version your playbooks depend on from drifting when your operating system ships an update, and it makes the exact dependency set reproducible on a colleague's machine. Once the environment exists, ansible --version should print a version and a Python path — if it does not, you are still using the system interpreter and the next command will not find the modules.

Networking support lives in collections, and the two you will use constantly are ansible.netcommon (the connection plugins and platform-independent modules) and cisco.ios (the Cisco IOS module family). Recent Ansible versions bundle them; if yours does not, install them explicitly:

ansible-galaxy collection install ansible.netcommon cisco.ios

Project Structure

ansible-network-101/
├── ansible.cfg
├── inventory/
│   ├── group_vars/switches.yml
│   └── hostfile.ini
└── playbook.yml

ansible.cfg sets host_key_checking = False and points to the inventory. In hostfile.ini, group hosts under [switches]. The group name must match the group_vars/switches.yml file name so variables apply automatically:

[switches]
switch_01 ansible_host=10.10.1.50
# group_vars/switches.yml
ansible_connection: ansible.netcommon.network_cli
ansible_network_os: cisco.ios.ios
ansible_become: yes
ansible_become_method: enable
ansible_user: admin
ansible_password: cisco123
ansible_become_password: cisco123

Never store real credentials in plain text — use Ansible Vault.

The directory layout above is worth internalising because every mature Ansible project is a variation on it. The control node reads ansible.cfg from the current working directory first, so a self-contained ansible.cfg inside the project makes the project portable: clone it on another machine and the inventory path, the SSH behaviour and the output formatting all come along. And the group_vars convention is the reason the group is called switches rather than something more creative — the filename is the group name.

A minimal ansible.cfg looks like this:

[defaults]
inventory = inventory/hostfile.ini
host_key_checking = False
retry_files_enabled = False
stdout_callback = yaml
gathering = explicit

Two of those lines are worth explaining. host_key_checking = False stops Ansible from refusing to connect to a device whose SSH host key has not been recorded, which is essential when you are scripting against lab gear that gets reconfigured and re-keyed constantly. gathering = explicit turns off Ansible's default habit of collecting host facts, because on a network device that step is meaningless and only wastes time.

Credentials Done Properly with Vault

ansible-vault create group_vars/switches.yml

Instead of a clear-text password in the file, encrypt the whole group vars file with a vault password of its own. The playbook then runs with --ask-vault-pass or, in automation, with --vault-password-file pointing at a file that only the automation account can read. The credentials never appear in your shell history and never land in Git. Once encrypted, the file is a binary blob to everyone except Ansible, and editing it later goes through ansible-vault edit.

First Playbook: Run Show Commands

---
- name: "Ansible 101"
  hosts: switches
  gather_facts: no
  tasks:
    - name: Show Version
      cisco.ios.ios_command:
        commands: show version
      register: output
    - name: Print output
      debug:
        msg: "{{ output.stdout_lines }}"
ansible-playbook playbook.yml

Read that playbook top to bottom and it reads like a sentence. A play targets a host group, and each task names a module and passes it arguments. register captures the module's return value into a variable, and the debug task prints it. There is no loop, no condition, no error handling to learn on day one.

What makes ios_command especially useful is that it accepts a list of commands and returns a corresponding list of outputs. That means you can collect a batch of show commands in a single task and keep them in sync by index:

- name: Gather several show outputs at once
  cisco.ios.ios_command:
    commands:
      - show version
      - show ip interface brief
      - show vlan brief
  register: facts_run

- name: Print the interface summary
  debug:
    msg: "{{ facts_run.stdout[1] }}"

The stdout list mirrors the commands list one for one, so stdout[1] is the output of the second command. That index-based access is what you will use when parsing information out of a device before Ansible's structured facts modules cover a given platform.

Make a Configuration Change

---
- name: "Ansible 101 - Configuration Changes"
  hosts: switches
  gather_facts: no
  tasks:
    - name: VLAN Config
      cisco.ios.ios_vlans:
        config:
          - name: server_vlan
            vlan_id: 30
          - name: user_vlan
            vlan_id: 31

Run it once → changed=1; run it again → changed=0. That is idempotency: Ansible detects the current state and only applies changes needed to reach the desired state.

Compare that task with the raw CLI you would type by hand. You would enter configuration mode, create the VLAN, name it, and repeat for the second VLAN. Ansible does that for you, but it also does something you would not: before pushing anything, it reads the current VLAN table, computes the difference between what exists and what you asked for, and pushes only the missing pieces. The second run finds nothing to do, so the task reports no change and the switch is left alone.

This is the property that makes automation trustable. A script that blindly re-sends configuration commands is dangerous; a module that reconciles desired state with observed state is not, because re-running it can never make things worse.

Add a Second Play with Validation

- name: Verify the VLANs exist
  cisco.ios.ios_command:
    commands: show vlan brief
  register: vlan_check

- name: Fail loudly if the new VLAN is missing
  assert:
    that:
      - "'server_vlan' in vlan_check.stdout[0]"
    fail_msg: "server_vlan was not found in the VLAN table"

Good automation verifies its own work. An assert task that fails when the expected VLAN is absent converts a silent misconfiguration into a red error in your terminal. Make a habit of ending every change play with a verification task, because "the module returned success" and "the device is in the state I wanted" are two different claims.

Common First-Project Mistakes

Almost everyone new to Ansible makes the same handful of mistakes, and each one has a clean fix. Forgetting to set ansible_network_os is the classic: without it, Ansible assumes the target is a Linux server, tries to copy and execute a Python module over SSH, and fails with a confusing shell error. Storing credentials in the inventory file instead of the vault is the second, and it is the one that eventually becomes a security incident. A third is leaving gather_facts enabled, which adds a pointless delay to every run against a switch because there is no equivalent of a Linux fact-gathering step on a network device. Finally, many people try to run configuration changes without first confirming the device is reachable and the privilege method works — a task that fails halfway through a change leaves the switch in a partial state. Run a read-only playbook first, always.

Fix libssh Issues (macOS)

pip install ansible-pylibssh
# Apple Silicon: build with brew headers
CFLAGS="-I $(brew --prefix)/include -I ext -L $(brew --prefix)/lib -lssh" pip install ansible-pylibssh

On modern macOS, the default connection library can fail to negotiate with older SSH servers on network gear. Installing the ansible-pylibssh transport switches Ansible to a bundled library that behaves like a real OpenSSH client, which resolves the majority of "unable to open shell" and "SSH connection failed" errors on lab equipment. On Apple Silicon, the CFLAGS line above makes the extension build against Homebrew's headers rather than a mismatched system copy.

Desired State Explained

Playbooks declare what the end state should be, not how to get there. Ansible translates intent into the required CLI commands, which reduces human error and keeps every switch consistent. For the deeper playbook design with roles and host_vars, continue with our Cisco IOS Ansible playbook guide; automation pairs well with MLNX-OS configuration management.

The mental shift from imperative to declarative is the hard part, and it is worth stating plainly. An imperative script says "type this, then type that." A declarative playbook says "the end state is a VLAN named server_vlan with ID 30." The module figures out the commands. When you describe fifty VLANs across twenty switches, the difference stops being philosophical: you describe the desired end state once, and the same playbook converges every device toward it, skipping the work that is already done. This is also what makes a playbook safe to run in a pipeline on every commit — it is a convergence step, not a series of blind edits.

To go further, the natural next reads on this site are our Ansible Cisco playbook examples collection and the ios_config, facts and backup module reference, which cover the modules you will reach for once show commands and VLANs feel routine.

原文链接:https://www.packetswitch.co.uk/ansible-network-automation-101