Arista EOS QoS Configuration: Class Maps to Strict Priority - 夜莺博客

Arista EOS QoS Configuration: Class Maps to Strict Priority

Default Arista EOS QoS treats all unclassified traffic on a first-in, first-out basis, which means a microburst from a backup job or a vMotion event can drop voice packets and ruin call quality. This article shows three escalating approaches to QoS on Arista switches: trusting DSCP/CoS at the edge for quick wins, building class maps and policy maps with the Modular QoS CLI for explicit classification, and enforcing strict-priority queuing so critical traffic always leaves the switch first. Shallow-buffered platforms like the 7050X benefit most from these configurations.

Why QoS Still Matters on a Leaf-Spine Fabric

Arista EOS runs on merchant silicon that delivers enormous forwarding capacity at a very attractive price per port. That economics is exactly why QoS discipline is still required. Most campus and data centre fabrics are built with an oversubscription ratio between leaf and spine, so a single 100G uplink may carry the aggregate of twenty 25G access ports. When several servers respond to the same request at the same instant, the resulting incast burst can exceed the uplink for a few hundred microseconds. The switch port buffers that burst, and whatever does not fit is dropped.

If every packet is treated identically, the drop decision is random from the application's point of view. A storage replication stream losing a packet simply retransmits. A VoIP or video conferencing flow losing a packet produces an audible glitch or a frozen frame. QoS is the mechanism that tells the switch which packets are allowed to be dropped first, and which ones must be protected regardless of congestion.

The three solutions below escalate from a five-second change to a full multi-queue scheduler design. Start with the first one, measure, and only add complexity when you can show that you need it.

How EOS Models QoS: Traffic Classes and Queues

Before touching configuration, understand the two-stage model that EOS uses. Every frame that enters a port is classified into a traffic class (TC), a value from 0 to 7. The traffic class is then mapped to an output queue: a unicast queue (uc-queue) or a multicast queue (mc-queue). The scheduler attached to those queues decides the order in which packets leave the interface.

Classification can come from three sources, in the following order of precedence:

  • An explicit match in a policy map bound to the interface — for example "anything with DSCP 46 becomes traffic class 5".
  • A trust setting on the ingress port, which copies the incoming DSCP or CoS value into the internal traffic class using the QoS maps.
  • The platform default map, which effectively places unmarked traffic in traffic class 0.

Inspect the current state of each building block before you change anything:

show qos map dscp
show qos map cos
show qos map traffic-class
show qos profile
show qos interfaces
show qos status

A common mistake is to assume the switch is already honouring DSCP values. Out of the box, EOS is in a "trust none" state for the internal traffic class on most platforms, so a phone that faithfully marks its RTP stream as EF still lands in the best-effort queue.

Solution 1: Quick Fix — Trust the Edge

interface Ethernet12
   description UPLINK_TO_VOIP_VLAN
   qos trust dscp
   ! qos trust cos   ! for Layer 2 tagging

Only apply trust on ports where you control the device — an untrusted laptop can mark BitTorrent as EF (DSCP 46) and hog bandwidth.

Choose between the two trust modes according to where the marking is generated. qos trust dscp is correct when the upstream device is a router, firewall or IP phone that sets DSCP at Layer 3. qos trust cos is correct when traffic arrives already tagged with 802.1p priority bits and you want the Layer 2 marking to drive the internal traffic class — typical on a trunk from an access layer that marks at the first hop.

Note that trust is only half of the story. The incoming value has to be mapped to something useful. If the platform default maps DSCP 46 to traffic class 5, trust alone is enough. If it maps everything to traffic class 0, add an explicit map:

! Give EF traffic and AF41 video their own traffic classes
qos map dscp 46 to traffic-class 5
qos map dscp 40 to traffic-class 5
qos map dscp 34 to traffic-class 4
qos map dscp 26 to traffic-class 3
!
! Preserve the marking on egress
qos map traffic-class 5 to cos 5
qos map traffic-class 5 to dscp 46
qos map traffic-class 4 to dscp 34

The table below is the marking convention used by most enterprise voice and video deployments and is worth standardising on before you build the policy:

Traffic type DSCP name DSCP value Traffic class
Voice RTP EF 46 5
Call signalling CS3 24 3
Video conferencing AF41 34 4
Transactional data AF21 18 2
Bulk / backup AF11 10 1
Scavenger CS1 8 1
Best effort CS0 / default 0 0

Solution 2: MQC Class Maps and Policy Maps

Trust-based classification stops being enough the moment you have devices that mark incorrectly, traffic you want to reclassify in the core, or a requirement to rate limit a specific application. The Modular QoS CLI (MQC) solves all three. It follows the same three-step pattern as Cisco IOS — class map, policy map, service-policy — but the EOS keywords are subtly different: classes carry a type (type qos, type pbr, type queuing) and the action on a class is usually set traffic-class rather than set dscp.

! Step 1: Class map identifies voice traffic
class-map type qos match-any CLASS-VOICE
   match ip dscp 46
   match ip dscp 40

! Step 2: Policy map maps it to traffic class 5
policy-map type qos POLICY-EDGE-IN
   class CLASS-VOICE
      set traffic-class 5
   class class-default
      set traffic-class 0

! Step 3: Apply to the interface
interface Ethernet48
   description UPLINK_TO_CORE
   service-policy type qos input POLICY-EDGE-IN

match-any is an OR: any single line matching is enough. Use match-all when a class should only trigger on a combination, such as "DSCP 46 and VLAN 100". Class maps can match on DSCP, CoS, IP precedence, VLAN ID, ACL entries, MPLS exp, source and destination addresses, and even on the MAC or protocol. Keep the number of classes small — five to eight is plenty for an enterprise network — because every additional class costs CPU on the ingress pipeline and increases the chance of an ordering bug.

A richer ingress policy that both reclassifies and marks at the trust boundary looks like this:

class-map type qos match-any CLASS-VIDEO
   match ip dscp 34
   match ip dscp 32
class-map type qos match-any CLASS-BULK
   match ip dscp 10
   match ip dscp 8

policy-map type qos POLICY-TRUST-EDGE
   class CLASS-VOICE
      set traffic-class 5
      set dscp 46
   class CLASS-VIDEO
      set traffic-class 4
      set dscp 34
   class CLASS-BULK
      set traffic-class 1
      set dscp 10
   class class-default
      set traffic-class 0
      set dscp 0

interface Ethernet12
   service-policy type qos input POLICY-TRUST-EDGE

Notice the pattern: a policy bound with input reclassifies and marks. A policy bound with output can shape, police and queue. Mixing the two is a common source of confusion, because a policy map bound in the wrong direction is accepted by the parser but silently never fires on traffic that matters.

If a specific flow has to be capped rather than prioritised, add a policer to the class. Policing drops or re-marks above the rate; it does not buffer, which is exactly what you want for abusive or scavenger traffic:

policy-map type qos POLICY-RATE-LIMIT
   class CLASS-BULK
      police rate 200 mbit burst 100 kbyte
      set traffic-class 1
   class class-default
      set traffic-class 0

interface Ethernet48
   service-policy type qos input POLICY-RATE-LIMIT

Solution 3: Strict Priority Queuing

Classification alone does not create priority. Until you change the scheduler, every traffic class ends up in a queue competing on equal terms. Strict priority is the single most effective change for latency-sensitive traffic: the queue is drained completely before any other queue is served.

! Map Traffic Class 5 to output Queue 5
qos map tc 5 to mc-queue 5 uc-queue 5

! Enforce strict priority for voice
interface Ethernet48
   tx-queue 5
      priority strict
   ! Best-effort data gets round robin
   tx-queue 0
      bandwidth percent 50

This overrides the default weighted round-robin behavior so that class 5 traffic is serviced before anything else. It is the right choice for latency-sensitive voice and video on converged networks.

The trade-off is starvation. A strict priority queue with no bound can consume the entire port for as long as it has traffic, leaving best-effort flows to wait indefinitely. Two safeguards are standard practice. First, keep the strict priority traffic class small and well policed at the edge, so it can never legitimately occupy more than a fraction of the link. Second, on platforms that support it, bound the priority queue with an explicit rate so the hardware will not let it run away:

interface Ethernet48
   tx-queue 5
      priority strict
      shape rate 2 gbps
   tx-queue 4
      bandwidth percent 25
   tx-queue 3
      bandwidth percent 15
   tx-queue 2
      bandwidth percent 10
   tx-queue 0
      bandwidth percent 50

bandwidth percent sets the minimum share of the port guaranteed to the queue under contention; it is a floor, not a ceiling, so an idle queue gives its share back to the others. If you need a hard cap instead, use shape rate, which meters the egress queue and is the right tool for reining in a greedy class without dropping it in the scheduler.

A Complete Reference Configuration

Assembling the three solutions gives a full run book. The example below trusts at the access edge, classifies and re-marks on the uplink, maps traffic classes to output queues, and applies a scheduler with one strict priority queue.

! --- Access edge: trust the phone, drop everyone else's marking ---
interface Ethernet12
   description ACCESS_TO_IP_PHONE
   switchport mode access
   switchport access vlan 100
   qos trust dscp

! --- Classification ---
class-map type qos match-any CLASS-VOICE
   match ip dscp 46
class-map type qos match-any CLASS-VIDEO
   match ip dscp 34
class-map type qos match-any CLASS-CRITICAL
   match ip dscp 26

policy-map type qos POLICY-UPLINK-IN
   class CLASS-VOICE
      set traffic-class 5
   class CLASS-VIDEO
      set traffic-class 4
   class CLASS-CRITICAL
      set traffic-class 3
   class class-default
      set traffic-class 0

! --- Traffic class to queue mapping ---
qos map tc 5 to mc-queue 5 uc-queue 5
qos map tc 4 to uc-queue 4
qos map tc 3 to uc-queue 3

! --- Uplink: policy in, scheduler out ---
interface Ethernet48
   description UPLINK_TO_SPINE
   service-policy type qos input POLICY-UPLINK-IN
   tx-queue 5
      priority strict
   tx-queue 4
      bandwidth percent 25
   tx-queue 3
      bandwidth percent 15
   tx-queue 0
      bandwidth percent 50

Verification and Monitoring

QoS is invisible until you look at counters. Apply the configuration inside a configuration session, then confirm that classification, mapping and scheduling are all doing what you intended:

show qos interfaces
show qos interfaces tx-queue Ethernet48
show qos interfaces counters Ethernet48
show qos status
show qos map dscp
show running-config interfaces Ethernet48

Focus on three numbers per queue. Occupancy shows how deep the queue gets, which reveals whether a burst is being absorbed or dropped. Transmit drops show whether the queue overflowed — on a strict priority queue this should stay near zero. Transmit rate compared with the port rate tells you whether the priority class is small enough to be safe. Repeat the counters a few minutes apart; a queue that is always full is a design problem, not a tuning problem.

If you want to export those counters rather than read them by hand, the same values appear in structured JSON output, which is far easier to parse in a monitoring pipeline:

show qos interfaces counters Ethernet48 json

Lossless, PFC and the RoCEv2 Case

Everything above assumes lossy service: when a queue overflows, the switch drops. Storage traffic carried over converged Ethernet needs the opposite guarantee, and that is where Priority Flow Control (PFC) and Explicit Congestion Notification (ECN) enter the picture. PFC pauses a specific priority on a link instead of dropping, which requires a very different configuration and a strict no-drop queue set. If your fabric carries RoCEv2, read the dedicated lossless tuning walkthrough before enabling PFC, because a misconfigured no-drop queue can turn one congested link into a fabric-wide head-of-line block.

Hardware Considerations

Platforms like the 7050X series have shallow buffers and critically need QoS, while deep-buffered 7280R series are more forgiving but still benefit from proactive queue configuration. Monitor queue behavior with show qos interface Ethernet48 and show qos statistics after applying policies. On current EOS releases the equivalent commands are show qos interfaces tx-queue Ethernet48 and show qos interfaces counters Ethernet48, which expose per-queue occupancy and drop counters directly.

Buffer architecture is the variable that changes the answer more than any other. A 7050X with a few megabytes shared across all ports will drop an incast burst that a 7280R with tens of megabytes absorbs without a single loss. If your monitoring shows drops on the 7050X despite correct QoS, the fix may be architectural — rebalance the uplinks or move the storage pair to a deeper-buffered platform — rather than another scheduler tweak.

Common Pitfalls

  • Trusting everywhere. A trust statement on a user-facing port lets any host claim the priority queue. Trust at the edge means trust the phone, not the wall jack.
  • Marking without scheduling. Building class maps and stopping there leaves all traffic in equal queues. Priority only exists once the scheduler is changed.
  • Unbounded strict priority. A strict priority queue with no policer or shaper at ingress can starve every other class on the port.
  • Applying the policy in the wrong direction. An input policy classifies and re-marks; an output policy queues and shapes. Verify with show running-config interface Ethernet48.
  • Forgetting multicast. A traffic class used by multicast applications needs the mc-queue mapping as well, otherwise replication traffic bypasses your carefully built queue plan.
  • Changing QoS in production without a session. Wrap the change in configure session, review the diff, then commit — it is the fastest rollback path on EOS.

Related Articles

QoS sits alongside several other EOS topics worth reading next. The trust boundary and DSCP marking guide for Cisco switches explains where markings should be applied; the IOS XR class map, policy map and shaping reference shows the same MQC pattern on a service provider platform; and the RoCEv2 PFC and ECN tuning guide covers lossless queues in detail.

For the rest of the EOS toolbox, see the MLAG configuration run book and the EOS configuration cheat sheet.

原文链接:https://techresolve.blog/2026/03/09/configuring-arista-qos