BGP Route Hijacking: Detection and Mitigation Playbook - 夜莺博客

BGP Route Hijacking: Detection and Mitigation Playbook

BGP's default behaviour is to believe every announcement it receives. There is no origin authentication, no ownership check, and the most specific path usually wins. That combination means a misconfigured router in a small network can blackhole a large provider's prefix simply by announcing it. Hijacks are not exotic attacks - the majority are accidents - but the effect is identical: traffic goes somewhere it should not.

This playbook covers how hijacks happen, what you should have deployed before one occurs, how to detect one quickly, and what actually restores service.

How Hijacks Happen

Type Mechanism Typical cause
More-specific hijack Announce a longer prefix than the legitimate holder (a /25 inside a /24) Route leak, misconfigured summarisation, accidental redistribution
Exact-prefix, wrong origin Announce the same prefix with a different origin AS BGP hijack or a stale route from a decommissioned link
Path manipulation Prepend or otherwise craft AS path to win best-path Traffic engineering gone wrong, or deliberate
Sub-prefix squatting Announce unallocated space inside a block you do not route Address-planning error

The more-specific case is the one that hurts most, because RPKI validation of the covering prefix alone will not save you.

Prevention: What to Deploy Before the Incident

1. Filter what you announce

Every prefix you advertise should be generated from a router configuration, not typed by hand. On any reasonable platform that means a prefix-list or a generated policy plus an inbound filter on the peer that drops anything you did not intend to send:

! outbound - announce exactly what you own, nothing else
ip prefix-list ANNOUNCE-TO-TRANSIT seq 10 permit 203.0.113.0/24
route-map TO-TRANSIT permit 10
 match ip address prefix-list ANNOUNCE-TO-TRANSIT
 set community 64500:100

! inbound from a peer - RFC 7454 style hygiene
ip prefix-list DENY-BOGONS deny 10.0.0.0/8 le 32
ip prefix-list DENY-BOGONS deny 172.16.0.0/12 le 32
ip prefix-list DENY-BOGONS deny 192.168.0.0/16 le 32
ip as-path access-list 10 deny ^64500_        ! never accept your own AS
ip as-path access-list 10 permit .*
route-map FROM-PEER permit 10
 match ip address prefix-list DENY-BOGONS

2. Filter what you accept from customers

For downstream customers, an IRR-based filter generator is the standard mechanism: build prefix-lists from the registry objects the customer maintains, and refresh them on a schedule. It is only as good as the registry data, which is exactly why RPKI exists.

3. Deploy RPKI origin validation

RPKI turns origin validation into a routing decision. Build a validator, get the VRP cache, and drop or de-prefer invalid announcements:

! Cisco IOS-XE
router bgp 64500
 bgp rpki server tcp 192.0.2.10 port 3323 refresh 600
 ! drop invalid, prefer valid
 bgp bestpath origin-as validity-check
route-map RPKI-FILTER deny 10
 match rpki invalid
route-map RPKI-FILTER permit 20

Validation states are valid, invalid and not-found. Dropping invalid routes wholesale is aggressive and will occasionally break legitimate traffic from networks with sloppy registry data; most operators start by raising the local preference of valid routes and dropping invalid ones only on high-value peering sessions. Whichever policy you choose, decide it deliberately and record why.

4. Session hygiene and limits

Prefix limits stop a misbehaving peer from filling the table, and session protection stops the churn from taking down the control plane. This is the first thing to get right, because a hijack's knock-on effect is often a session reset rather than the hijack itself - the mechanics are covered in this BGP prefix-limit guide, and the cryptographic session protections in this BGP session hardening guide.

5. Peer-lock and ASPA

Peer-lock is a per-neighbour configuration that says "this peer may only send me these prefixes". It is the single most effective defence against a more-specific hijack arriving from a transit peer, because an unexpected /25 is dropped before it reaches the table. ASPA (Autonomous System Provider Authorization) extends the idea to validate the AS path itself; support is still rolling out, but it is the direction the industry is moving.

Detection

You need to know within minutes, not days. Four layers:

  1. BMP to a collector. BGP Monitoring Protocol gives you a per-peer, per-prefix view of what each router receives and what it selects. This is the highest-value telemetry investment for any network running BGP at scale. The router-side configuration is in this BMP configuration guide, and the collector options are compared in this BMP collector comparison.
  2. Unexpected origin detection. Alert when the origin AS for one of your prefixes changes. This is a five-line query against stored BMP data and catches the majority of exact-prefix hijacks.
  3. More-specific alerts. Alert on any announcement more specific than your allocation that you did not originate. These are almost always accidental, and they are the ones that take your traffic.
  4. RPKI state transitions. An alert when a prefix you care about flips from valid to invalid is an early warning that something has changed upstream.
# what is actually in the table right now - origin AS for our space
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast 203.0.113.0/24 bestpath
show ip bgp regexp ^64500_

# RPKI view
show bgp rpki table
show bgp rpki servers
show bgp ipv4 unicast rpki invalid

Public looking glasses and route collectors add an outside-in view, and they matter because a hijack may not be visible from inside your own network at all if the attacker's path is longer than yours.

Mitigation During an Incident

The order matters, and the first step is not technical:

  1. Confirm the hijack and collect evidence. Screenshot the route from multiple vantage points, note the origin AS, the prefix length and the first-seen time. You will need this when talking to the other network's NOC.
  2. Determine whether it is accidental. The overwhelming majority are. A polite, specific email to the offending network's NOC - with the prefix, the origin AS, and the observed path - resolves most cases faster than escalation.
  3. Ask your upstreams to filter. Your transit providers and peers can drop the rogue announcement on request. They can usually do this within minutes, and it is the fastest way to restore traffic.
  4. Announce more specifics where you can. If the hijack is an exact-prefix origin hijack and more-specifics are accepted by your upstreams, a /25 and a /25 (or a /24 and a /24 from a different upstream) can outcompete the rogue route. Validate the specific customer's expectation first - some networks filter anything longer than /24 in IPv4.
  5. Do not panic-advertise. Announcing a wide supernet or withdrawing prefixes during the incident usually makes it worse and can cause collateral outages for your own customers.
  6. Fix the cause if it is yours. If the hijack came from your own redistribution, withdraw it and check every other router with the same configuration template - this class of fault rarely exists on exactly one device.

After the Incident

  • Turn the detection that fired late into an alert that fires early. If you learned about it from a customer, your monitoring has a gap.
  • Verify your RPKI ROAs are correct. A stale ROA with the wrong max-length is functionally the same as no ROA and occasionally worse.
  • Refresh IRR objects and peer filters. A filter built from data that was three years out of date did not defend anything.
  • Write down the escalation path: registrar, upstream NOCs, and the internal owner of the affected prefix, with current contact details. Incidents are a bad time to discover that the previous network engineer's phone number is the only one anyone has.

原文链接:https://www.rfc-editor.org/rfc/rfc7454.html