SONiC Network OS Configuration with GNS3: Easy Guide - 夜莺博客

SONiC Network OS Configuration with GNS3: Easy Guide

SONiC is the open-source network operating system that is taking over data center switching, but practicing on real hardware is expensive. The SONiC Virtual Switch (VS) solves that: it runs as a Docker container or KVM image, and with GNS3 you can build a complete multi-switch Clos topology on a laptop. This article, based on PLVision's hands-on SONiC tutorial, walks through downloading and installing the SONiC VS image in GNS3, understanding ConfigDB (the Redis-backed running configuration), and configuring VLANs, LAGs, routed interfaces and BGP using real SONiC CLI commands.

Why Build a SONiC Lab in GNS3

GNS3 is an open-source network emulator that pairs a graphical topology editor with a hypervisor backend. Because SONiC VS ships as a standard x86 KVM image, GNS3 boots it exactly the way it boots routers or firewalls, while still giving you console access, link management and packet capture on every segment. That makes a virtual SONiC fabric the cheapest possible way to learn the platform: a four-node topology costs nothing but roughly 16 GB of host RAM and a few gigabytes of disk.

Before you start, check three things on the workstation:

  • Virtualization — Intel VT-x or AMD-V must be enabled in firmware. On macOS, install the GNS3 VM (VMware Fusion or VirtualBox) so QEMU runs inside a Linux guest; running the SONiC image directly under QEMU on macOS works, but it is slow and the console is fragile.
  • Memory — allocate at least 4 GB per SONiC node. A two-spine / two-leaf Clos therefore needs about 16 GB of RAM available to the GNS3 VM, plus 2 GB for the VM itself.
  • Disk — the VS image is 2–4 GB uncompressed and every node carries its own disk overlay, so budget 20 GB of free space for a small lab.

A good first topology is two spines and two leaves. The leaves hold VLANs and host subnets, the spines only route. Each adjacent pair gets a point-to-point link (a /31 or /24 works fine), and every node gets a unique Loopback0 address that is advertised through BGP.

Downloading and Importing the SONiC VS Image

The SONiC VS platform is implemented as a vslib library linked with the syncd daemon at build time; it stores SAI configuration data while the Linux networking stack (routes, ARP entries) makes forwarding decisions. To use it in GNS3, download a pre-built image and register it as an appliance:

wget https://sonic-jenkins.westus2.cloudapp.azure.com/job/vs/job/buildimage-vs-201911/181/artifact/target/sonic-vs.img.gz
gunzip sonic-vs.img.gz
wget https://raw.githubusercontent.com/Azure/sonic-buildimage/master/platform/vs/sonic-gns3a.sh
./sonic-gns3a.sh -b <path to sonic.img>

The helper script generates a GNS3 appliance definition that already carries the correct QEMU arguments, console type and adapter count. In the GNS3 GUI, open Edit → Preferences → QEMU VMs → New and point the wizard at the generated definition, or import the appliance template directly. Verify these defaults after import:

  • RAM: 4096 MB; raise to 6144 MB if you plan to run telemetry or install a full routing table.
  • Adapters: 8 or more — SONiC VS exposes Ethernet0 through EthernetN in order, plus eth0 for management.
  • Console: telnet is the most reliable option under GNS3; do not enable "use as a linked base VM" if you want persistent per-node configuration.
  • Disk: reference the raw sonic-vs.img once and let nodes clone from it, instead of copying the image per device.

First boot takes five to ten minutes, because each node initializes Redis, loads the default configuration from /etc/sonic/config_db.json and waits for its containers to converge. Log in with the default credentials and change them right away. Taking a GNS3 snapshot shortly after the login prompt appears makes every later boot much faster.

Wiring the Lab Topology

Drag four SONiC VS nodes onto the canvas and cable them as a Clos: leaf-to-spine only, never leaf-to-leaf. Two practical details save a lot of time later:

  • Management reachability — add a GNS3 Cloud or NAT node and connect it to eth0 on each switch, so you can SSH into the lab from the host instead of fighting four console windows. A small Linux container acting as a DHCP server on that segment is enough.
  • Host simulation — attach a lightweight Linux container or the built-in VPCS to each leaf to generate traffic, then use packet capture on the inter-switch links to confirm forwarding.

Once the nodes are up, give each one an identity before touching any other configuration. On a SONiC VS image the router MAC and the BGP ASN live in the DEVICE_METADATA table, and duplicate values across nodes break LACP and BGP in confusing ways:

sudo config hostname leaf01
sudo config interface ip add eth0 192.168.100.11/24
sudo config interface ip add Loopback0 10.1.1.1/32

Keep a spreadsheet or a generated inventory file with hostname, Loopback0, ASN and router MAC for each node. When the fabric misbehaves, the first thing to check is whether two switches are sharing an identity.

ConfigDB: Running vs. Startup Configuration

On boot, configurations load from /etc/sonic/config_db.json into the Redis ConfigDB namespace — think of the file as startup configuration and ConfigDB as the running configuration. CLI changes update ConfigDB but are not written back automatically. Save with sudo config save -y; apply a saved file with sudo config reload -y (restarts SONiC services). For multi-node topologies, give each switch a unique router MAC and BGP ASN in the DEVICE_METADATA section, and unique Loopback0 addresses.

You can inspect the live configuration without leaving the switch, which is invaluable when a CLI command appears to succeed but nothing changes in the data plane:

sonic-db-cli CONFIG_DB keys 'VLAN|*'
sonic-db-cli CONFIG_DB hgetall 'PORT|Ethernet12'
sonic-db-dump -y -n CONFIG_DB -k 'DEVICE_METADATA|localhost'

The distinction matters in a lab: a change that is present in ConfigDB but absent after a config reload was never saved, while a change present in the file but missing from ConfigDB never reached the running system. Treating those two states separately removes most of the guesswork from SONiC troubleshooting. If you want a deeper treatment of this save/reload/replace cycle, see the dedicated walkthrough linked at the end of the article.

Intra-Rack Switching: VLANs

config vlan add 100
config vlan member add -u 100 Ethernet12
config vlan member add -u 100 Ethernet16
config vlan member add -u 100 Ethernet20
show vlan brief

Members added with -u are untagged; omit the flag to make the port a trunk member. Create the corresponding SVI and give it an address to route between VLANs:

config interface ip add Vlan100 10.0.0.1/24
show ip interfaces

From Linux, verify the state that actually exists in the kernel and the ASIC with bridge vlan list. The default VS image contains ebtables rules that block ARP forwarding between L2 interfaces (ebtables --list); remove the ARP drop rule with ebtables --delete FORWARD 2 if your hosts need to communicate within the VLAN. This single rule is responsible for more "the lab is broken" reports than any other cause.

Link Aggregation with PortChannels

config portchannel add PortChannel0001
config portchannel member add PortChannel0001 Ethernet0
config portchannel member add PortChannel0001 Ethernet4
show interfaces portchannel
show lacp neighbor

Run the identical commands on the peer switch with the same member ports, then verify that LACP reached the bundled state on both sides. On SONiC VS, member ports must not be part of any VLAN while you are building the bundle, and both ends must agree on speed; a speed mismatch leaves the port in an individual state forever. The critical check is show interfaces portchannel for the bundle state plus show lacp neighbor for the peer's system ID.

Inter-Rack Routing: RIFs and BGP

config interface ip add Ethernet8 32.0.0.1/24
config interface ip add PortChannel0001 12.0.0.2/24
config interface ip add Vlan100 10.0.0.1/24
show ip interfaces
show ip route

Add BGP neighbors in the BGP_NEIGHBOR section of /etc/sonic/config_db.json (asn, holdtime, keepalive, local_addr), then sudo config reload -y and check show ip bgp summary. A typical two-spine fabric uses ASN 65100 on the leaves and 65001 on the spines, with the loopback addresses as update sources. Redistribution of connected routes must be configured inside the FRR shell, because SONiC's own CLI does not expose redistribute statements:

sonic# configure terminal
sonic(config)# router bgp 64002
sonic(config-router)# address-family ipv4 unicast
sonic(config-router-af)# redistribute connected
sonic(config-router)# end
sonic# write

Remember that write inside FRR only persists the FRR configuration; it does not replace sudo config save -y, which is what stores the ConfigDB state. Do both, or your carefully built fabric will be empty after the next config reload.

Verifying the Fabric End to End

Work bottom-up rather than jumping straight to BGP. A short, repeatable sequence catches almost everything:

show interfaces status
show interfaces counters rates
show lldp neighbors
show mac
show arp
show ip route
show ip bgp summary

Read the output in that order. show interfaces status should report every cabled port as up with the expected speed; a port stuck at 1000 Mb/s instead of 10000 Mb/s points at a cable or speed mismatch. show lldp neighbors confirms the physical cabling matches your documented design — in a lab this is faster than tracing links by hand. show mac and show arp prove L2 and L3 learning respectively, and only once those look right does show ip bgp summary make sense to interpret. Finally, run an end-to-end test between simulated hosts behind two different leaves:

ping -c 5 10.0.0.2
traceroute 10.1.1.4
sudo ip vrf exec mgmt ping 8.8.8.8

A ping that succeeds with latency in the low milliseconds confirms the whole stack — VLAN, SVI, routed port, BGP advertisement and return path. A ping that fails at the first hop points back at the VLAN or the ebtables rule; a ping that reaches the far SVI but not the far host points at the routing adjacency.

Troubleshooting Common GNS3 Lab Issues

  • Nodes ping within a VLAN but not across — check the ARP-blocking ebtables rule and confirm the SVI exists, not just the VLAN.
  • PortChannel never bundles — both peers must use the same member list, the same speed, and must not have the ports assigned to a VLAN.
  • BGP stays in Active — the neighbor address must be reachable, the ASN in DEVICE_METADATA must match the router bgp statement, and duplicate router MACs will silently break sessions.
  • Configuration disappears on reboot — you edited ConfigDB or FRR without running sudo config save -y.
  • Node hangs at boot — insufficient RAM, missing KVM acceleration, or the disk image copied per node instead of referenced.
  • Can't reach the switch from the host — the management Cloud is not attached to eth0, or the default management IP conflicts with the host network.

When something is genuinely broken inside a daemon, the container architecture gives you a direct route in:

docker ps
docker exec -it bgp vtysh -c "show bgp summary"
docker logs syncd --tail 50
show techsupport

show techsupport (or sudo generate_dump) collects logs, ConfigDB state and system information into a single archive — the fastest way to compare a broken node against a known-good one.

Where to Go Next

Once the four-node fabric converges, extend it. Add a third leaf and break the symmetric configuration on purpose to watch BGP reconverge. Enable a gNMI or telemetry container and stream counters to a collector. Import a breakout cable configuration to see how SONiC splits a 100G port into four 25G ports. Each of those exercises teaches more than a chapter of theory, and all of them run on the same laptop.

Related Reading

原文链接:https://plvision.eu/blog/sdn/sonic-network-os-configuration