Ansible Network Examples: Facts, Backups and cli_command - 夜莺博客

Ansible Network Examples: Facts, Backups and cli_command

Automating network devices with Ansible starts with a few well-understood patterns: inventory groups with connection variables, facts collection, configuration backups, and platform-independent modules that work across vendors. This guide is based on the official Ansible network examples and walks through a facts-and-backup playbook that runs against Arista EOS, Cisco IOS and VyOS, then shows how to simplify multi-vendor playbooks with cli_command and cli_config. It is the fastest way to move from ad-hoc CLI work to repeatable automation.

Prerequisites: Ansible, Collections and Connection Basics

Before a single task runs, three things have to be true. First, Ansible must be installed on a Linux control node (or WSL on Windows) — network modules run on the controller, not on the device, because switches and routers do not ship a Python interpreter you can push code to. Second, the vendor collections must be installed so that module names such as arista.eos.eos_facts resolve. Third, the devices must be reachable over SSH with a usable account.

# Install Ansible and the network-side collections
pip install ansible
ansible-galaxy collection install ansible.netcommon
ansible-galaxy collection install arista.eos
ansible-galaxy collection install cisco.ios
ansible-galaxy collection install vyos.vyos

ansible --version
ansible-galaxy collection list | grep -E "netcommon|eos|ios"

Keep a project-local ansible.cfg next to your inventory so nobody has to remember command-line flags. The two settings that matter most for network gear are the inventory path and the persistent-connection timeout, because network connections are held open in a background process rather than re-established per task.

[defaults]
inventory = ./inventory.ini
host_key_checking = False
gather_facts = False
retry_files_enabled = False

[persistent_connection]
connect_timeout = 60
command_timeout = 60

Setting gather_facts = False globally is deliberate: the standard setup module tries to log into the device with the normal SSH plugin, which fails on network operating systems. You will gather network facts explicitly instead, using the vendor facts modules described below.

Why Network Devices Need a Special Connection

A server playbook connects over SSH, copies a Python script across, and runs it. A network playbook cannot do that. Instead Ansible opens an SSH session to the CLI and drives it as a human would, sending a command and reading the prompt until the next one appears. That is what ansible.netcommon.network_cli does. It is a persistent connection: one socket is reused across every task in the play, which keeps a twenty-task playbook from authenticating twenty times and triggering the device's login rate limiting.

Two variables control the behaviour. ansible_network_os tells Ansible which module set and which prompt-handling logic to load — cisco.ios.ios, arista.eos.eos, vyos.vyos. ansible_become_method=enable tells it that privilege escalation means typing enable and answering the enable prompt, not running sudo. Get these two right and most "connection refused" and "unable to open shell" errors disappear.

Inventory Groups and Variables

Use [group:vars] sections to set connection variables per platform, and Ansible Vault to encrypt passwords. A minimal inventory for network devices groups switches by platform so playbooks can target each type.

[all:vars]
ansible_connection=ansible.netcommon.network_cli
ansible_user=ansible
ansible_ssh_pass={{ vault_ssh_password }}

[switches:children]
eos
ios
vyos

[eos]
veos01 ansible_host=veos-01.example.net

[eos:vars]
ansible_network_os=arista.eos.eos
ansible_become=yes
ansible_become_method=enable
ansible_become_password={{ vault_enable_password }}

[ios]
ios01 ansible_host=ios-01.example.net

[ios:vars]
ansible_network_os=cisco.ios.ios
ansible_become=yes
ansible_become_method=enable

[vyos]
vyos01 ansible_host=vyos-01.example.net

[vyos:vars]
ansible_network_os=vyos.vyos.vyos
ansible_become=no

The switches group is a parent group, so a play with hosts: switches hits every device regardless of vendor while each child keeps its own connection variables. Store the vault variables in group_vars/all/vault.yml and commit only the encrypted file:

ansible-vault create group_vars/all/vault.yml
ansible-vault edit  group_vars/all/vault.yml
ansible-playbook site.yml --ask-vault-pass

Example 1: Collect Facts and Create Backups

- name: "Demonstrate connecting to switches"
  hosts: switches
  gather_facts: no
  tasks:
    - name: Gather facts (eos)
      arista.eos.eos_facts:
      when: ansible_network_os == 'arista.eos.eos'
    - name: Gather facts (ios)
      cisco.ios.ios_facts:
      when: ansible_network_os == 'cisco.ios.ios'
    - name: Display some facts
      debug:
        msg: "The hostname is {{ ansible_net_hostname }} and the OS is {{ ansible_net_version }}"
    - name: Backup switch (eos)
      arista.eos.eos_config:
        backup: yes
      register: backup_eos_location
      when: ansible_network_os == 'arista.eos.eos'

Facts modules populate ansible_net_* variables such as hostname, version, model and serial number; the config backup creates timestamped files you can store in version control. Extend the same block to VyOS and Cisco and add a normalised reporting task so you get one table per run:

    - name: Gather facts (vyos)
      vyos.vyos.vyos_facts:
        gather_subset: ["config", "interfaces"]
      when: ansible_network_os == 'vyos.vyos.vyos'

    - name: Backup switch (ios)
      cisco.ios.ios_config:
        backup: yes
        backup_options:
          filename: "{{ inventory_hostname }}.cfg"
          dir_path: "{{ playbook_dir }}/backups/{{ ansible_date_time.date if ansible_date_time is defined else 'latest' }}"
      when: ansible_network_os == 'cisco.ios.ios'
      register: backup_ios_location

    - name: Build a fleet report
      ansible.builtin.copy:
        dest: "{{ playbook_dir }}/reports/inventory.csv"
        content: |
          hostname,model,version,serial
          {% for h in ansible_play_hosts_all %}
          {{ hostvars[h].ansible_net_hostname | default(h) }},{{ hostvars[h].ansible_net_model | default('n/a') }},{{ hostvars[h].ansible_net_version | default('n/a') }},{{ hostvars[h].ansible_net_serialnum | default('n/a') }}
          {% endfor %}
      delegate_to: localhost
      run_once: true

Example 2: Platform-Independent Modules

With two or more platforms, replace platform-specific modules with ansible.netcommon.cli_command and ansible.netcommon.cli_config:

- hosts: network
  gather_facts: false
  connection: ansible.netcommon.network_cli
  tasks:
    - name: Run cli_command on Arista
      ansible.netcommon.cli_command:
        command: show ip int br
      register: result
      when: ansible_network_os == 'arista.eos.eos'
    - name: Run cli_command on Cisco IOS
      ansible.netcommon.cli_command:
        command: show ip int br
      register: result
      when: ansible_network_os == 'cisco.ios.ios'

Using groups and group_vars by platform, this can be further simplified to a single task with a {{ show_interfaces }} variable per platform group. Define the command once per vendor in group_vars/eos.yml, group_vars/ios.yml and group_vars/vyos.yml:

# group_vars/ios.yml
show_interfaces: "show ip interface brief"
# group_vars/eos.yml
show_interfaces: "show ip interface brief | json"
# group_vars/vyos.yml
show_interfaces: "show interfaces brief"

# tasks/single_task.yml
    - name: Read interface status on every platform
      ansible.netcommon.cli_command:
        command: "{{ show_interfaces }}"
      register: intf
    - name: Print the parsed output
      ansible.builtin.debug:
        var: intf.stdout_lines

Example 3: Pushing Configuration with cli_config

Reading is half the job; cli_config handles the writing side without vendor-specific modules. It accepts a list of configuration lines and a parents stanza for hierarchy, and it supports the same check and diff modes as every other Ansible module.

    - name: Ensure a description exists on uplink1
      ansible.netcommon.cli_config:
        config: |
          interface Ethernet1
            description UPLINK-TO-CORE
        parents: []
        diff_against: intended
      when: ansible_network_os == 'arista.eos.eos'

    - name: Change the syslog server fleet-wide
      ansible.netcommon.cli_config:
        config: "logging host {{ syslog_server }}"

Run configuration tasks with --check --diff first. For platforms that support it, the diff shows exactly which lines would change, which turns a config push into a reviewable change set rather than a leap of faith.

Running the Playbooks and Reading the Output

ansible-playbook -i inventory.ini facts-demo.yml
ansible-playbook -i inventory.ini cli-demo.yml --check --diff
ansible-playbook -i inventory.ini cli-demo.yml --limit ios01

The recap line is the first thing to read. ok=7 changed=2 means five tasks reported no change and two altered state — on a second run you want changed=0 for anything idempotent. A host that reports unreachable=1 never answered over SSH and will not appear in any later task, so triage connectivity before content. Verbose mode is the fastest way to see the raw CLI conversation: -vvv prints every command Ansible sent and every banner the device returned.

Verifying Backups and Building a Change History

Backups are only useful if you can find them. When backup: yes runs, the module returns backup_path, and the file lands in a backup/ directory next to the playbook by default. Copy it somewhere with history:

    - name: Copy backup files
      ansible.builtin.copy:
        src: "{{ backup_eos_location.backup_path }}"
        dest: "/srv/netbackups/{{ inventory_hostname }}/{{ inventory_hostname }}-{{ lookup('pipe','date +%Y%m%d-%H%M') }}.cfg"
      when: backup_eos_location.backup_path is defined
      delegate_to: localhost

    - name: Detect configuration drift
      ansible.builtin.command:
        cmd: "diff -u /srv/netbackups/{{ inventory_hostname }}/latest.cfg /tmp/{{ inventory_hostname }}.cfg"
      delegate_to: localhost
      changed_when: false
      failed_when: false

Once the files are in Git, every change the automation made is attributable, and git log -p on a device's configuration answers the question "what changed on this switch last Tuesday, and who caused it".

Best Practices

Always run ansible-playbook --syntax-check before execution, use --check --diff for dry runs on config modules, store credentials in Vault, and keep playbooks idempotent by checking state before applying changes. Add a few habits that the official examples imply but do not spell out: pin collection versions in a requirements.yml so a controller rebuild does not silently change module behaviour; use --limit for the first run against a new platform; keep per-vendor variables in group_vars rather than in tasks; and never commit a plaintext enable password, even in a private repository.

# requirements.yml
collections:
  - name: ansible.netcommon
    version: ">=6.0.0"
  - name: cisco.ios
    version: ">=8.0.0"
  - name: arista.eos
    version: ">=9.0.0"
  - name: vyos.vyos
    version: ">=6.0.0"
# ansible-galaxy collection install -r requirements.yml

Troubleshooting Common Errors

"Unable to open shell" almost always means the wrong ansible_network_os, a missing SSH server on the device, or an authentication failure hidden behind a generic message. Verify by hand first: ssh user@device from the controller, then check the value in the inventory. "Timed out waiting for privilege escalation" points at ansible_become_method: for IOS and EOS it must be enable, and ansible_become_password must be set. Commands hang forever when the device prompt does not match what the plugin expects -- usually a custom banner or a --More-- pager; add terminal length 0 in a preceding cli_command, or set ansible_terminal_length. Cloud/AAA login prompts returning an unexpected string can be handled by keeping the vault password out of the play log with no_log: true rather than by disabling prompts.

FAQ

Do I need Python on the switch? No. With network_cli, everything runs on the control node and only SSH is required on the device. Some older devices use ansible.netcommon.local with a provider dictionary, but that path is deprecated.

cli_command or vendor modules? Use cli_command for read-only one-liners that work identically everywhere. Use vendor modules (ios_config, eos_facts) when you need structured output, idempotent state management, or config diffs -- they know the syntax of the platform and will not push an unchanged line.

How many devices can one play hit? Fifty to a hundred is comfortable on a single control node because each host holds one persistent SSH session. Use forks to cap concurrency and protect the devices' AAA servers, and run in batches with serial when a change is disruptive.

More automation content: Ansible network automation 101, Cisco IOS Ansible playbook getting started, Ansible network modules: ios_config, facts and config backup, and SONiC troubleshooting.

原文链接:https://docs.ansible.com/projects/ansible/latest/network/user_guide/network_best_practices_2.5.html