Ansible for Network Engineers: Complete Guide - 夜莺博客

Ansible for Network Engineers: Complete Guide

Network automation is no longer optional in modern IT environments, and Ansible remains the most accessible entry point for engineers managing Cisco, Juniper or Arista fleets. This complete guide from NetOpsHub takes you from zero to productive: why Ansible's agentless architecture fits network devices, how to gather facts with ios_facts, how to build a configuration backup workflow with ios_config, how to protect credentials with Ansible Vault, and the best practices - dynamic inventory, group variables, error handling and NAPALM integration - that separate a lab toy from a production automation platform.

What makes Ansible different from every other automation tool you may have evaluated is that it asks almost nothing of the network. There is no agent to install, no daemon to keep alive, no vendor SDK to fight with, and no controller-side software running on the device. You point Ansible at a management IP address, hand it credentials, and it opens an SSH session the same way you would from a terminal window. That single design decision is why network teams consistently reach for Ansible first, even when the rest of the organisation has standardised on something else for servers.

Why Ansible for Network Automation

Ansible needs no agents on the devices: it uses SSH for Unix/Linux systems and NETCONF/REST APIs for network devices. Playbooks are human-readable YAML that double as living documentation of your network configuration, with native modules for all major vendors and idempotent operations that are safe to rerun.

The practical consequences matter more than the marketing language. Because a playbook is a text file, it goes into Git, gets peer reviewed, and produces a diff when someone changes a VLAN definition. Because modules are idempotent, running the same playbook twice does not produce duplicate configuration or unexpected device reboots. Because the connection is SSH, the tool works on hardware that is five years old and on hardware that shipped last quarter, with the same mental model. And because the modules are shipped as collections that follow the vendor's own release cadence, a new platform feature typically reaches Ansible within weeks rather than years.

There is also a human factor. YAML is readable by a network engineer who has never written a line of Python, so the automation does not become a private asset owned by whoever happens to be on the team when it is written. The blast radius of a mistake is contained by --check mode, by serial limits that touch one device at a time, and by the pre-change backups you will learn to build in later sections.

Ansible Architecture: Control Node, Managed Nodes and Agentless Design

Ansible has exactly two roles in its architecture:

  • Control node - the machine where you install Ansible and from which you run ansible or ansible-playbook. Any Linux, macOS or WSL2 host with Python 3.9 or newer will do. Windows cannot be a control node natively, but it can be a managed node for server workloads.
  • Managed nodes - the devices you target. For network gear this is the device's management interface. No software is installed on it, no port other than SSH (or NETCONF on 830, or the HTTPS API) needs to be open, and no reboot is required to enable automation.

Under the hood, a playbook run works like this: Ansible reads your inventory, builds a task list, and for each host opens an SSH session, converts the module into a small payload the device can consume, executes it, parses the output, and closes the connection. For network platforms the transport is usually network_cli, which wraps the traditional CLI in a structured, persistent connection - a huge improvement over the early days when a naive script would open and close a session for every single command.

Core Concepts You Need Before Writing a Playbook

The vocabulary is small, and knowing it precisely prevents most beginner confusion.

  • Inventory - the list of devices and their connection details. It can be a static INI or YAML file, or generated dynamically from a cloud provider or an IPAM.
  • Playbook - a YAML file containing one or more plays.
  • Play - a mapping of a group of hosts to an ordered list of tasks, plus the connection settings used for the whole play.
  • Task - a single call to a module, with arguments written in YAML.
  • Module - the unit of work: cisco.ios.ios_config, cisco.ios.ios_facts, cisco.ios.ios_command. Collections such as cisco.ios group related modules.
  • Role - a reusable, opinionated bundle of tasks, defaults, templates and handlers with a defined directory layout.
  • Handler - a task that only runs when notified by another task, which is how you make "save configuration" happen once at the end of a run rather than after every change.
  • Fact - data returned by a module about the device, surfaced as variables prefixed ansible_net_ for network platforms.

Installation and Prerequisites

pip install ansible
ansible --version

On a modern system you will also want the vendor collection that contains the platform modules, plus the network connection plugin:

ansible-galaxy collection install cisco.ios
ansible-galaxy collection install ansible.netcommon
ansible-galaxy collection list

Finally, verify reachability and SSH behaviour before you blame your playbook. A device that presents an interactive banner, an unusual prompt, or a TACACS authorization prompt will break automation in ways that have nothing to do with Ansible's configuration:

ssh admin@192.168.1.1
ansible all -m ansible.netcommon.ping -i inventory.ini

Inventory Deep Dive: Static, Dynamic and Group Variables

A static inventory is the right starting point. Group it by vendor, by site and by role so that a single play can target "Cisco switches in the Frankfurt access layer" without any filtering logic in YAML:

[switches]
sw-fra-01 ansible_host=192.168.1.1
sw-fra-02 ansible_host=192.168.1.2

[switches:vars]
ansible_network_os=cisco.ios.ios
ansible_connection=ansible.netcommon.network_cli

[routers]
rtr-fra-01 ansible_host=192.168.2.1

[routers:vars]
ansible_network_os=cisco.iosxr.iosxr

[frankfurt:children]
switches
routers

[all:vars]
ansible_user=automation
ansible_timeout=60

The moment the network grows past a few dozen devices, generated inventory wins. For pure network estates, a small Python script that reads your IPAM or a CSV export and emits an inventory is often simpler and more robust than a cloud plugin. For cloud-connected environments, use the vendor plugins such as amazon.aws.aws_ec2 or azure.azcollection.azure_rm with keyed groups so the inventory stays declarative.

Above the inventory sits the variable hierarchy. group_vars/all/global.yml holds site-wide settings, group_vars/switches/snmp.yml holds group-specific values, and host_vars/sw-fra-01.yml holds the exceptions that always exist in real networks. Ansible merges these by specificity, which means a host variable silently wins over a group variable - a behaviour worth remembering when a device stubbornly ignores a setting you are sure you applied.

Your First Playbook: Gathering Facts

---
- name: Gather Network Facts
  hosts: all
  gather_facts: false
  connection: network_cli
  vars:
    ansible_network_os: ios
  tasks:
    - name: Get device facts
      ios_facts:
        gather_subset: all
    - name: Display hostname
      debug:
        var: ansible_hostname

Pair it with an inventory file:

[switches]
192.168.1.1
192.168.1.2
[all:vars]
ansible_user=admin
ansible_ssh_pass=your_password
ansible_become_pass=enable_password
ansible_connection=network_cli
ansible_network_os=ios

Always dry-run first with ansible-playbook gather_facts.yml --check.

Note that gather_facts: false is deliberate: the default fact gathering is aimed at Linux hosts and will fail or waste time against a network device. Instead, platform fact modules such as ios_facts are called explicitly as a task, and they are far richer than anything the generic gatherer could produce.

Understanding the ansible_net_* Facts

Once ios_facts has run, the returned data lands in variables you can branch on. The most useful ones in day-to-day work:

  • ansible_net_hostname - the configured hostname, ideal for validating that the device you addressed is the device you meant to change.
  • ansible_net_version - the running IOS version, used to gate image-upgrade logic.
  • ansible_net_model and ansible_net_serialnum - hardware identity for asset reconciliation and RMA workflows.
  • ansible_net_interfaces - a dictionary of interfaces with state, MTU and address information.
  • ansible_net_config - the full running configuration when config is included in the gather subset, which is what makes golden-config comparison possible.

Because facts are plain variables, they become the input to decisions: only reload a device whose version differs from the target, only configure a VLAN on devices that are missing it, only open a ticket where drift is detected. That is the difference between a script that pushes text and a policy engine that enforces intent.

Real-World Example: Configuration Backup

---
- name: Network Configuration Backup
  hosts: all
  gather_facts: false
  connection: network_cli
  vars:
    backup_dir: /path/to/backups
  tasks:
    - name: Create backup directory
      file:
        path: "{{ backup_dir }}"
        state: directory
        mode: '0755'
    - name: Fetch running config
      ios_config:
        backup: yes
        backup_options:
          filename: "{{ inventory_hostname }}-{{ ansible_date_time.date }}.cfg"
          dir_path: "{{ backup_dir }}"

Two details make this backup usable rather than merely present. First, the filename includes the date, because the module's default name is fixed and a nightly job would otherwise overwrite yesterday's copy forever. Second, the backup should run in the same play as the change, not in a separate play that a failed earlier task can skip. A backup that exists in a parallel universe to your change is not a rollback plan.

Generating Configuration with Jinja2 Templates

Copying configuration between devices by hand does not scale, and hardcoding per-device commands in task files is barely better. Templates solve this cleanly. Combine a loop over inventory data with the template module, or render a fragment inline with lookup('template', ...), then push the result through ios_config. A typical template builds an access-interface block from variables such as VLAN ID, description and port list, and the same template then serves every switch in a stack - which is precisely how you get consistency without writing 200 nearly identical task files.

Organising Automation with Roles

roles/
  ios_base/
    defaults/main.yml
    tasks/main.yml
    templates/
    handlers/main.yml
    vars/main.yml
playbooks/
  site.yml
  backup.yml

Roles give you exactly the structure that long playbooks lack: a place for defaults that callers can override, a place for templates, and handlers that fire once at the end of a run. A typical decomposition for a campus network is ios_base (hostname, NTP, syslog, AAA), ios_snmp, ios_vlan, and ios_interface, wired together by a site.yml that lists them in dependency order.

Multi-Vendor Playbooks and cli_command

Networks are rarely single-vendor. Two patterns keep multi-vendor estates manageable. The first is conditional dispatch: gather facts first, then include_tasks the platform-specific file based on ansible_net_model or the inventory's ansible_network_os. The second is the escape hatch: cli_command runs an arbitrary CLI command and returns structured output, and cli_config pushes lines without a vendor-specific module. Both are honest about their limits - they are not idempotent and they will not tell you whether a change was needed - so treat them as bridges for platforms lacking first-class modules, and migrate to a real module when one becomes available.

Idempotency, Check Mode and Diff Mode

Idempotency means running the playbook twice leaves the device in the same state as running it once. In practice you evaluate it by running with --check (no changes pushed) and then for real, and confirming that the second real run reports changed=0. Full command forms are essential here: shutdown compares correctly against the running config, while the abbreviation shut is treated as a brand new line on every run. Add --diff to see the exact lines Ansible would add, and get into the habit of reading that diff before approving a production run.

Error Handling and Retries

- name: Apply configuration with a rollback on failure
  block:
    - name: Push interface changes
      cisco.ios.ios_config:
        src: templates/access.j2
        backup: true
        save_when: modified
  rescue:
    - name: Notify that the device needs attention
      ansible.builtin.debug:
        msg: "{{ inventory_hostname }} failed - review the backup and re-run"

- name: Wait for the device to answer again
  ansible.netcommon.cli_command:
    command: show version
  register: result
  retries: 5
  delay: 15
  until: result is not failed

block and rescue let one device's failure be contained and reported instead of aborting the whole run, while retries with until absorbs the transient failures that are normal on network devices - a slow BGP reconvergence, a management-plane hiccup after a reload, a TACACS server that briefly refused authorization.

Securing Secrets with Ansible Vault

ansible-vault create group_vars/all/vault.yml
ansible-vault edit group_vars/all/vault.yml

Store ansible_ssh_pass and ansible_become_pass inside the encrypted vault and reference them as variables - never hardcode passwords in playbooks.

For larger teams, keep the vault file separate from the plain variables file it complements, commit both, and give the vault its own variable prefix so it is obvious which values are secret. Use --vault-password-file pointing at a file with mode 0600, or better, an external secret manager, so that CI runs never need a human to type a password. ansible-vault encrypt_string is the right tool when only one value in an otherwise readable file needs protection.

Best Practices for Production Automation

  • Dynamic inventory - use cloud inventory plugins (e.g. amazon.aws.aws_ec2) for dynamic environments.
  • Organized group variables - split group_vars/all into vault.yml (secrets) and global.yml, with per-group files for switches and routers.
  • Error handling - register task output and branch on failure with when: conditions instead of letting one failure abort the whole run.
  • Vendor modules - use ios_interface, ios_vlan, ios_bgp, ios_acl and ios_command for their respective jobs.
  • Version control everything - playbooks, inventory, templates and vault files all belong in Git, with the vault password held in a secret manager rather than a wiki page.
  • Change windows and serial limits - roll out to one device, then a canary group, then the fleet; never let a first-ever run touch 300 switches simultaneously.
  • Keep a break-glass path - document how to log in and revert manually when automation is unavailable, and test that path occasionally.

Scaling Ansible for Large Networks

Default Ansible forks five hosts at a time. That is fine for a lab and far too slow for a few hundred switches, so tune forks in ansible.cfg alongside the persistent connection settings:

[defaults]
forks = 50
host_key_checking = False
timeout = 60

[persistent_connection]
command_timeout = 120
connect_timeout = 60

Above a few hundred devices, add serial to the play so failures do not cascade, and consider splitting the estate into inventory groups that map to change windows. For genuinely huge fleets, an execution environment running AWX or Ansible Automation Platform gives you scheduling, RBAC, job templates and an audit trail that a cron entry cannot.

Troubleshooting Common Issues

Test SSH with ansible all -m ping; use -vvv for verbose output; raise ansible_timeout=60 in inventory for slow devices; ensure the enable password is set with ansible_become: yes and ansible_become_method: enable.

Three more failure modes are worth knowing. A play that works interactively but fails under cron is usually a missing --vault-password-file or a PATH difference. A device that times out after a hardware reload needs a longer command_timeout and a wait_for style retry loop, not a bigger global timeout. And an inventory group whose ansible_network_os does not match the platform will fail with a family of confusing parsing errors rather than a clear "wrong OS" message.

NAPALM Integration

For vendor-agnostic operational data, NAPALM (Network Automation and Programmability Abstraction Layer with Multivendor support) works with any Ansible-supported device:

- name: Get operational data with NAPALM
  napalm_get_facts:
    hostname: "{{ inventory_hostname }}"
    username: "{{ ansible_user }}"
    password: "{{ ansible_ssh_pass }}"
    driver: "eos"
    register: napalm_facts

The value of NAPALM is the shape of the data: BGP neighbours, LLDP neighbours and interface counters come back as identical dictionaries whether the device is IOS, IOS-XR, Junos or EOS, so a compliance report can be written once. Its other useful primitive is the candidate-config operations - load_merge_candidate, compare_config, discard_config - which give you a genuine two-phase commit on platforms that support it.

Testing, Linting and CI/CD

Treat automation like application code. Run ansible-lint and ansible-playbook --syntax-check in a pre-commit hook so obviously broken YAML never reaches the repo. Structure your pipeline so the test stage runs --check against a lab or a canary device, and the deploy stage runs the real playbook with serial: 1 before widening. Store the backup artifacts your playbooks generate as CI artifacts, which gives you a complete, timestamped record of what device configuration looked like before every change.

A Practical Learning Path

Move in this order: inventory and facts first, then the backup playbook, then one manual change converted into a templated role validated with --check --diff, and only then vault encryption and a CI pipeline. Start small - a backup playbook is the perfect first production automation - then iterate toward CI/CD pipelines that run --check in the test stage and full playbooks in the deploy stage.

Related Reading

原文链接:https://netopshub.com/blog/ansible-network-engineers-complete-guide