Junos Interface Damping: Hold-Time and Backoff - 夜莺博客

Junos Interface Damping: Hold-Time and Backoff

A physical interface that flaps is one of the most disruptive and least glamorous problems in a Junos network. A millisecond-scale bounce during a transport protection switch makes every routing protocol on the box recompute; a link that oscillates every few seconds keeps the control plane permanently busy and can poison BGP, OSPF and MPLS signalling at the same time. Junos ships two distinct mechanisms for this - the static hold-time timer and the dynamic exponential backoff damping profile - and choosing the wrong one is a common and expensive mistake. This article explains which mechanism fits which failure pattern, how to configure both, and how to verify that the configuration is doing what you think it is.

An important framing note before the configuration: interface damping is only useful where a routing protocol is running on the link. On a pure Layer 2 access port, damping buys nothing - the switch still has to process the link state change, and there is no advertisement to suppress. Damping is a control-plane optimisation, not a Layer 1 fix.

中文原文要点(保留原文)

物理接口抖动(flapping)是 Junos 设备最常见的隐患:毫秒级抖动会让路由协议频繁重算,秒级周期性抖动则持续消耗控制面资源。Junos 提供两套接口阻尼机制——静态 hold-time 定时器与动态指数退避阻尼,本文讲解两者的适用场景、配置思路与验证命令,帮助网工按接口角色(核心/客户/对等/边缘)合理设置。

短时抖动:Hold-Time 定时器

冗余链路切换等毫秒级抖动场景,用 hold-time down/up 在接口级别抑制状态翻转上报:hold-time down 期间忽略 down 事件,定时器到期仍为 down 才上报;hold-time up 同理抑制恢复期的快速抖动。hold-time down 默认 0ms,建议保持默认;hold-time up 用于抑制恢复后的连击。

长时抖动:动态阻尼(Exponential Backoff)

秒级周期抖动用类似 BGP 的惩罚机制:每次接口 down 事件累加 1000 惩罚值,惩罚值按 half-life 指数衰减;累计超过 suppress 阈值后接口进入抑制状态,不再向上层协议上报事件,直到衰减到 reuse 阈值以下解除抑制。主要参数:half-life、reuse、suppress、max-suppress-time。

配置思路(示例,原文)

# 物理接口配置静态 hold-time
set interfaces ge-0/0/1 hold-time up 2000
set interfaces ge-0/0/1 hold-time down 0

# 动态阻尼参数(在接口或分组下配置)
set interfaces ge-0/0/1 damping half-life 5
set interfaces ge-0/0/1 damping suppress 2000
set interfaces ge-0/0/1 damping reuse 750
set interfaces ge-0/0/1 damping max-suppress 20

注意事项(原文)

  • 建议链路两端配置相近的阻尼参数,单端配置可能引发对端行为异常;
  • 阻尼仅适用于有路由协议的链路,纯二层场景收益有限;
  • 先用 show interfaces extensive 观察抖动规律(毫秒级还是秒级),再选对应机制。

The two failure patterns, and why they need different tools

Juniper's own documentation splits physical interface flaps into exactly two categories, and the split is the whole design rationale:

  • Nearly instantaneous multiple flaps of short duration (milliseconds). This is what you see when transport equipment switches between redundant paths. During the switch, both router interfaces see several up/down transitions lasting a few milliseconds each. Nothing is physically broken - the link is back before any human could notice - but each transition generates an advertisement to the routing protocols.
  • Periodic flaps of long duration (seconds). An unstable link between a router interface and transport equipment oscillates on a period of roughly five seconds or more, with an up-and-down duration of about one second. These are real, recurring failures, and they will keep happening for as long as the underlying fault survives.

Hold-time is designed for the first pattern and is actively harmful for the second: a fixed timer cannot suppress an oscillation that lasts longer than the timer. Damping is designed for the second pattern and is the wrong tool for the first, because a millisecond bounce does not accumulate enough of a penalty to matter and the backoff profile needs time to build one.

Hold-time: suppressing sub-second flaps statically

The hold timer works by not advertising a transition until the timer has expired and the state is still the new state. When the interface goes from up to down, the down hold-time timer starts; every transition during that window is ignored; if the timer expires and the interface is still down, the router advertises it as down. When the interface comes back up, the up hold-time timer starts and the same logic applies in reverse. The transition is advertised within about 100 milliseconds of the configured value.

The configuration keys are worth knowing precisely, because they behave differently from what most engineers assume:

  • Range: 0 through 4,294,967,295 milliseconds.
  • Default: 0, which means no damping - transitions are advertised immediately.
  • Timer implementation: most Ethernet interfaces use a one-second polling algorithm, so a sub-second hold value is not honoured with sub-second precision on those ports. One-port, two-port and four-port Gigabit Ethernet interfaces with SFP transceivers are interrupt driven and do honour fine-grained values.
  • Not available on controller interfaces, and not applicable to aggregated Ethernet member links - configure hold-time on the ae- interface and leave the members alone.
  • Platform caveat: the QFX10002-72Q and QFX10002-36Q do not support hold-time down below one second on 100G interfaces; the recommended value there is three seconds.

A practical starting point for an edge interface that sits behind a protecting transport layer:

set interfaces ge-0/0/1 hold-time up 2000
set interfaces ge-0/0/1 hold-time down 0

Leaving hold-time down at zero is deliberate. A real fibre cut should be reported immediately so that the routing protocol can start converging; the cost of a delayed down event is a black hole where traffic keeps being forwarded towards a dead link. The up timer is where the value goes, because suppressing the rapid up/down/up chatter of a recovery is what stops the protocol from tearing down and rebuilding adjacencies repeatedly.

Damping: exponential backoff for long-term flapping

Damping uses the same penalty model as BGP route flap damping, applied to the interface rather than to a prefix. Every time the interface goes down, a penalty of 1000 is added to the interface penalty counter. The counter decays exponentially: after one half-life the accumulated penalty is halved, provided the interface has remained stable. When the accumulated penalty exceeds the suppress threshold, the interface enters the suppress state and further up/down transitions are not reported to the upper-level protocols. When the penalty decays below the reuse threshold, the interface is unsuppressed and reporting resumes.

The Junos damping statement is available at the interface level and at the interface-range level, and it was introduced in Junos OS Release 14.1 - on older code the configuration simply does not exist, which is one reason it is less widely deployed than it should be. The statement accepts four parameters plus an enable flag:

Parameter Meaning Range Default
half-life Decay half-life, in seconds, after which the accumulated penalty counter is halved if the interface stays stable 1 - 30 5
max-suppress Maximum hold-down time, in seconds - the interface cannot stay suppressed longer than this no matter how unstable it has been 1 - 20,000 20
reuse Reuse threshold; below this penalty value the interface is unsuppressed 1 - 20,000 1000
suppress Suppression threshold; above this penalty value the interface is suppressed 1 - 20,000 2000

Two consistency rules are enforced by the commit rather than by the operator, and both catch people out:

  • half-life must be less than max-suppress, or the configuration is rejected.
  • max-suppress must be greater than half-life, or the configuration is rejected.

Note also that the penalty counter increments by 1000 for every down event, so the default suppress threshold of 2000 is crossed on the second down event - the defaults are far more aggressive than the numbers suggest. This mirrors the same lesson that the BGP damping community learned with the RIPE-580 revision: the vendor default threshold is too low and penalises normal operation.

Worked penalty arithmetic

As with BGP damping, the four parameters are not independent - the half-life and max-suppress values together define the highest penalty the interface can ever reach, which in turn determines whether the suppress threshold is reachable at all:

ceiling = reuse * 2^(max-suppress / half-life)

With the Junos interface defaults (reuse 1000, half-life 5, max-suppress 20) the ceiling is 1000 * 2^4 = 16,000. A suppress threshold of 2000 is comfortably reachable, which is why the defaults work in the narrow sense. Now consider an operator who wants to be gentler and sets suppress 8000 while leaving everything else alone: the ceiling is still 16,000, so 8000 is reachable but the interface will sit suppressed for a long time. Raise it to suppress 20000, however, and the threshold can never be reached - the interface will never be damped, and the running configuration will look perfectly sensible while doing nothing at all.

A measured starting profile for a transit-facing link that oscillates on a multi-second period:

set interfaces ge-0/0/1 damping half-life 20
set interfaces ge-0/0/1 damping max-suppress 120
set interfaces ge-0/0/1 damping reuse 1000
set interfaces ge-0/0/1 damping suppress 5000

Check the arithmetic: the ceiling is 1000 * 2^(120/20) = 1000 * 64 = 64,000, well above the 5000 suppress threshold, so the profile can actually engage. Once suppressed above 5000, the decay to the 1000 reuse threshold takes 20 * log2(5) = 20 * 2.32 = about 46 seconds, so a flapping link stops poisoning the routing protocols for roughly three quarters of a minute at a time, then is given another chance. If it flaps again immediately, the penalty accumulates from where it decayed to and suppression returns quickly.

Choosing values by interface role

There is no single correct profile. The interface's role determines how much reachability loss you are willing to trade for control-plane stability:

  • Core and transport-facing interfaces: use hold-time only, with hold-time up in the 2000-5000 ms range and hold-time down near zero. A core link's down state must propagate immediately; the recovery chatter is what needs suppressing. Be conservative with damping here - suppressing a core interface hides a real problem from the routing protocols and from your monitoring.
  • Customer-facing edge ports: this is where the damping profile earns its keep. Customer premises equipment flaps for reasons you cannot fix - power problems, cheap optics, CPE reboots - and each bounce otherwise costs you an OSPF adjacency and a BGP session rebuild. A suppress threshold of 5000 with a 20-second half-life absorbs the noise without hiding a genuine outage for long.
  • Peering and IX interfaces: prefer hold-time with a modest up timer and leave damping disabled. Peering faults need to be visible quickly, and damping a peering interface can suppress exactly the transition you need to see to escalate to the peer.
  • Access and edge ports with no routing protocol: none of this applies. Do not configure damping where it cannot help.

Verification: what to run and what to look for

The one command that answers almost every question is show interfaces extensive, which includes a Damping field for the interface showing the configured half-life, max-suppress, reuse and suppress values along with the current penalty state.

show interfaces extensive ge-0/0/1
show interfaces extensive ge-0/0/1 | match -i "damping|penalty|suppress"
show interfaces terse | match "ge-0/0/1"
show interfaces ge-0/0/1 | match "Last flapped"

A verification workflow that actually catches mistakes:

  1. Fault-inject before you trust the profile. If you have a lab port, physically bounce it past the suppress threshold and confirm the Damping field shows the interface in a suppressed state while the routing protocol does not log an adjacency change.
  2. Check the penalty state over time, not just once. Watch the counter decay across at least two half-lives to confirm the configured decay matches your expectation of how long a flapping link should stay suppressed.
  3. Correlate with protocol logs. If the interface is suppressed but OSPF still logs consecutive adjacency changes, the damping profile is not being applied to that interface - check interface ranges and whether the statement landed on the right unit.
  4. Watch for the inverse error: a suppression that never triggers. Compare the configured suppress threshold against the computed ceiling from the formula above.

Pitfalls that bite in production

  • Suppression is invisible to SNMP and OAM. Junos does not indicate whether an interface is down because of damping or because that is the genuine physical state. Consequently SNMP link traps and OAM protocols cannot distinguish a damped link state from the real one, and traps may not behave as your monitoring expects. If your alerting is built on link traps, add an explicit check of the Damping field.
  • Asymmetric configuration. Juniper explicitly recommends similar damping settings on both ends of a physical interface, and warns that configuring damping on one end only can produce undesired behaviour. The undamped end keeps advertising transitions that the damped end is trying to suppress.
  • Using hold-time to hide a real fault. A large hold-time down value means traffic keeps being forwarded to a dead link for that duration. On a core interface this converts a fast reconvergence into a long black hole.
  • Setting half-life above max-suppress. The commit fails with a rejection, but the habit of pasting values in from another platform's profile is what causes it.
  • Configuring damping on ae- member links. Put hold-time on the aggregated interface, not on its members, or the member link state changes will fight the aggregate's state machine.
  • Forgetting the interface polling granularity. On most Ethernet interfaces the hold timers are implemented with a one-second polling loop, so a 200 ms hold-time value does not behave as written. Choose values that are meaningful for the platform you actually run.

Which mechanism, in one table

Symptom Mechanism Why
Millisecond bounce during transport protection switch hold-time Fixed timer ignores transitions inside the window
Link that flaps up and down every few seconds for hours damping Penalty accumulates faster than it decays, so suppression engages
Single clean fibre cut Neither Both mechanisms delay legitimate failure notification
CPE-side port that reboots nightly damping Bounded suppression (max-suppress) caps the exposure

Related reading

Junos 路由层面抖动抑制见本站 Juniper BGP Session and Route Flaps: Prevention Guide;Junos 排障命令大全见 Junos Troubleshooting Commands Reference;MC-LAG 高可用见 Juniper MC-LAG Best Practices。For the penalty-based tuning counterpart on the BGP side, see BGP route flap damping penalty and RIPE-580 tuning.

原文链接:https://juniper.net/documentation/us/en/software/junos/interfaces-fundamentals/topics/topic-map/interfaces-damping-physical.html