Arista EOS Basics: CLI, Sysdb and Merchant Silicon Explained - 夜莺博客

Arista EOS Basics: CLI, Sysdb and Merchant Silicon Explained

Arista EOS is one of the most widely deployed network operating systems in modern data centers, and understanding its architecture explains why it behaves differently from traditional IOS-style platforms. This article covers the essential foundations every engineer needs: the Sysdb in-memory database model, the CLI modes and configuration sessions, powerful pipe filtering, and the implications of running on Broadcom merchant silicon. You will learn the commands used daily on 7050, 7280 and 7500 series switches, with direct comparisons to Junos and IOS-XR for easy translation.

This expanded walkthrough goes deeper into each layer: what Sysdb actually holds and how daemons use it, how configuration sessions and checkpoints give you rollback, the pipe syntax that replaces most manual scrolling, the daily interface and routing commands, and the ways the choice of switching silicon changes what the platform can do.

EOS Architecture: Sysdb and Agent Processes

EOS runs on a standard Linux kernel and decouples the system into independent agent processes that communicate through a central in-memory database called Sysdb. Routing daemons (Bgpd, Ospfd, Isisd), hardware drivers (FwdAgent, AclAgent) and the FastCLI all read and write Sysdb. Because the forwarding plane uses merchant silicon ASICs (Broadcom Trident, Tomahawk, Jericho; Intel/Barefoot Tofino), buffer depth, TCAM scale and feature support are determined by the chip — an important consideration when sizing buffers for spine-leaf fabrics.

The architectural consequence is worth stating plainly: EOS is a set of cooperating Linux processes with a shared, schema-driven state store, not a monolithic image. Each agent subscribes to the Sysdb paths it cares about and is notified when a value changes, so agents do not poll each other and one protocol daemon can restart without disturbing the rest of the system. FastCLI is itself a Sysdb client, which is why show output stays responsive while the control plane reconverges — the CLI is reading committed state, not waiting on a routing daemon to answer. It is also why the same state that the CLI displays is available to eAPI and to streaming telemetry: there is one source of truth.

When something misbehaves, the first question is usually which process owns the failure, so start by listing them:

show version
show agents
show logging | tail 50

show agents lists the running agents and their state; a stopped or restarting agent explains a surprising amount of unexpected behaviour, and show agent <name> logs drills into a specific daemon. Sysdb itself can be inspected from the CLI (show sysdb node, and deeper path queries such as show sysdb path /Sysdb/interface/status), which is mostly useful in TAC cases and when writing automation — the paths are internal interfaces and change between releases, so build on eAPI rather than on Sysdb paths if you want long-lived tooling.

CLI Modes and Configuration Sessions

EOS uses IOS-style modes but adds Junos-like configuration sessions:

localhost(config)# configure session PLANNED-CHANGE
localhost(config-session-PLANNED-CHANGE)# interface Ethernet1
localhost(config-session-PLANNED-CHANGE-if-Et1)# description To-Peer-Router
localhost(config-session-PLANNED-CHANGE-if-Et1)# show session-config diffs
localhost(config-session-PLANNED-CHANGE-if-Et1)# commit

Changes made in a session are not active until commit, and abort discards the session — perfect for staging risky changes.

The mode structure will feel familiar to anyone from IOS: user exec, privileged exec (enable), global configuration (configure), and interface or protocol submodes. What is different is the session layer. A named session is an isolated candidate configuration that other engineers cannot see or break, so two people can prepare changes on the same switch at the same time. Sessions support show session-config diffs to display exactly what will change, commit timer to apply a change automatically unless you cancel it (a poor man's commit-confirmed), and rollback inside the session to discard part of the work. Useful session commands:

show configuration sessions
show session-config named PLANNED-CHANGE diffs
configure session PLANNED-CHANGE
abort
commit
commit timer 00:10:00

Two habits pay for themselves immediately. Always write changes in a session and read the diff before committing, and always save the resulting configuration with write memory (or copy running-config startup-config) once you have confirmed the change works — a commit survives a process restart, but only the startup configuration survives a reload.

Pipe Filtering and Output Control

EOS pipe commands use Unix-style syntax and can be chained:

show ip route | include 192.168 | count
show bgp summary | exclude Idle
show running-config | section router bgp
show ip route | json

The section pipe extracts a full configuration block, and JSON output feeds directly into automation via eAPI.

The pipe is where EOS shows its Linux heritage. include and exclude take regular expressions rather than simple strings, so show interfaces status | include (notconnect|errdisabled) finds every problem port in one line. begin starts output at the first match, which is how you jump to the relevant part of a long configuration. count turns any output into a number, which is what makes verification scripts trivial: show bgp summary | include Established | count should equal the number of configured peers. json converts the output into structured data, and no-more disables paging, which matters for scripts. Chains are unlimited and evaluated left to right, so filter first and count second.

Essential Interface, VLAN and Routing Commands

show interfaces status
show interfaces Ethernet1 counters rates
show vlan
show mac address-table vlan 100
show spanning-tree
show ip ospf neighbor
show bgp summary

For configuration, an L3 interface, SVI and trunk look like:

interface Ethernet1
 no switchport
 ip address 192.0.2.1/30
interface Vlan100
 ip address 10.100.0.1/24
interface Ethernet3
 switchport mode trunk
 switchport trunk allowed vlan 100,200,300

Three details trip people up when they move from IOS. First, no switchport converts a port to routed mode, but the port also needs no shutdown only if it was previously disabled — EOS ports are administratively up by default. Second, an SVI in EOS stays down unless the VLAN exists in the VLAN database and at least one port is up in that VLAN, which is the same autostate behaviour IOS has, so a "Vlan100 down" message usually means no member port is forwarding. Third, trunk configuration in EOS does not require an encapsulation statement: switchport mode trunk implies 802.1Q, and the allowed VLAN list defaults to all VLANs, which means a careless trunk can leak VLANs you never intended to carry. Always set switchport trunk allowed vlan explicitly. More worked examples of VLAN, trunk and SVI configuration are in our EOS VLAN, trunk and SVI configuration examples and the EOS CLI command modes guide.

Merchant Silicon: What the Chip Decides

Because EOS runs on merchant silicon, the ASIC — not the operating system — sets the ceiling on several capabilities:

  • Buffer depth. Broadcom Trident-based 1RU switches have small shared buffers; Jericho-based platforms in the 7280R and 7800R families carry far more packet memory. Shallow buffers are fine for uniformly distributed east-west flows and can drop microbursts from storage or AI workloads.
  • Table scale. Route and host-entry capacity, MAC scale, ACL entries and TCAM width all come from the chip. Arista's platform data sheets list these numbers per SKU, and they are the reason the same EOS version can behave differently on two switches.
  • Feature support. Overlay capabilities such as VXLAN tunnel endpoints, and advanced match fields in ACLs, are hardware-assisted on some families and not on others. A feature that is not in silicon may run in the software forwarding path, which is fine for a handful of packets and disastrous for a production flow.
  • Telemetry and counters. Platform-specific show platform ... commands expose ASIC counters, and their names differ between chip families — which is why troubleshooting notes should always state the platform and EOS version.

The practical rule: capacity planning for an EOS switch begins with the ASIC, not with the port count. Confirm route, host, MAC and ACL scale against the specific SKU before a design depends on them, and check the release notes when a feature seems to behave differently on a new model.

Checkpoints, Rollback and Image Management

Configuration sessions stage a change; checkpoints preserve a known-good state to return to. Both are cheap, and both are far faster than reading a backup from a TFTP server during an outage:

copy running-config checkpoint:before_change
show checkpoint
configure replace checkpoint:before_change
copy running-config startup-config
show boot-config

configure replace restores the whole configuration from a checkpoint, which is the fastest available rollback. On the image side, EOS stores software as a .swi file in flash and the boot image is selected explicitly (boot system flash:EOS-4.29.6M.swi), so an upgrade is a copy, a boot-image change, and a reload — with the previous image still on the device for a rollback reload. Copy the running configuration to startup before any reload, and take a checkpoint before any change you cannot explain in one sentence.

eAPI and Automation

The same Sysdb state the CLI reads is exposed over HTTPS by eAPI, which is the reason EOS is popular in automation stacks. Enable it once:

management api http-commands
   protocol https
   no shutdown

and you can then run any show or configuration command over JSON-RPC:

curl -k -u admin:password https://switch/command-api \
  -d '{"jsonrpc":"2.0","method":"runCmds","params":{"version":1,"cmds":["show version","show ip route summary"],"format":"json"},"id":1}'

The format option is the whole trick: json returns structured data rather than text to parse, so the same command that an engineer types becomes an API call with typed output. Combined with checkpoints and sessions, this makes safe automation practical — stage changes over eAPI, diff them, commit them, and roll back from a checkpoint if verification fails. For configuration templating against EOS, see our network automation guide covering Cisco, Junos and Arista.

Quick Reference: Top Daily Commands

System health checks start with show version, show environment all and show logging | tail 100. Routing checks use show ip route summary, show bgp summary and show ip ospf neighbor. Remember that wr saves the running config, and checkpoints (copy running-config checkpoint:X) give you rollback points before risky changes.

Add a few more commands to the habit list and most incidents become findable in a minute:

show interfaces status | include (notconnect|errdisabled)
show interfaces counters errors
show interfaces transceiver
show lldp neighbors
show ip route 10.20.30.40
show mac address-table count
show agents
show checkpoint

Read them in a fixed order during an incident: physical layer first (show interfaces status, counters, transceivers), then neighbours, then the route lookup for the actual destination, then protocol state. Most "network is down" tickets are a port that is notconnect or a transceiver with alarms, and those two commands answer it before any routing table is examined.

Habits to Unlearn When Coming from IOS or Junos

  • From IOS: there is no write erase equivalent rhythm and no ISL encapsulation choice; trunk ports are 802.1Q only. show running-config and show startup-config are separate views, and wr is a saving action, not a display command.
  • From IOS: the pipe takes regular expressions, not literal strings, and | section has no IOS equivalent — it is the fastest way to view one protocol block.
  • From Junos: sessions and checkpoints exist, but EOS still has IOS-style modes, and configuration is in flat IOS syntax rather than a hierarchy. commit exists in EOS too, which surprises engineers who expect it to be Junos-only.
  • From both: hardware matters more. On EOS the ASIC defines buffer, table and feature limits, so the same command can succeed on one platform and be unsupported on another.

For more Arista practice, see the Arista EOS configuration cheat sheet, the MLAG configuration run book and BGP TCP-AO configuration on EOS.

原文链接:https://forwardingplane.net/configuration-archive/arista-fundamentals