Routing Protocols and Configuration Questions
How traffic finds its path across networks: route selection basics (administrative distance, longest-prefix match, control plane versus forwarding plane), static and default routing, interior gateway protocols (OSPF, IS-IS, EIGRP, RIP), BGP, and multicast routing. Covers protocol selection and migration, OSPF area and IS-IS level design, EIGRP feasible successors, BGP path selection, communities, route reflectors and policy, route redistribution, summarization and filtering, VRF-Lite and VRF route leaking, convergence tuning with BFD, ECMP path selection, BGP-based inbound and outbound traffic engineering, multi-homed internet edge and full-table scale, RPKI origin validation and containing route leaks, route flap and oscillation analysis, graceful maintenance drains, first-hop gateway redundancy (HSRP, VRRP, GLBP), and verifying and troubleshooting routing on real devices. Excludes network-wide topology and MPLS or segment-routing backbone design, Layer 2 switching and VLANs, general layered fault isolation and packet capture, cloud VPC routing, and firewall or ACL security.
On-premises network peers with both AWS Direct Connect and Azure ExpressRoute. Some prefixes are chosen via Azure resulting in higher latency. Propose a BGP-based traffic-engineering plan (use of local-preference, MED, AS-path prepending, communities, or prefix filters) to prefer Direct Connect for selected prefixes. Discuss implementation steps, how to test changes, and risks such as route oscillation or unintended path shifts.
Sample Answer
Direct answer
First find which BGP decision step is choosing Azure for each bad prefix, then use the lever that sits at that step. If the Azure route is a longer match (for example a /24 inside a /16 learned from AWS), longest-prefix match decides before anything else and local preference cannot override it (for a packet to 172.20.5.9, a /24 route for 172.20.5.0/24 and a /16 route for 172.20.0.0/16 both contain the address, and the router forwards by the /24 because it is the more specific, whatever the local preferences are): fix the advertisement, not the preference. If the lengths are equal, set a higher local preference on the routes learned over Direct Connect (DC) for only the selected prefixes, using an inbound route-map and an exact prefix-list. Local preference is exchanged only inside your AS, so it is fully under your control. Cisco's best-path order is: highest weight (a Cisco-only value local to one router), highest local preference, locally originated, shortest AS_PATH, lowest origin type, lowest MED, eBGP over iBGP, lowest IGP metric to the next hop, then later tie-breakers. Local preference is second, well before AS-path length and MED. Return traffic from each cloud is decided by that cloud, so it needs separate, smaller tweaks.
Terms: Direct Connect is AWS's private circuit service; ExpressRoute is Azure's; local preference (local-pref) is the BGP value, higher wins, that you set on your own routers to rank paths; MED (multi-exit discriminator) is a hint sent to a neighbor AS; prepending repeats your AS number to lengthen the AS_PATH; community is a tag on a route; iBGP is BGP between routers inside your own AS.
Step 1: diagnose before changing anything
On the on-premises edge, for one affected prefix run show ip bgp <prefix> and read both paths: prefix length (also check whether a longer prefix covers the same addresses), local-pref, AS_PATH, origin, MED and next hop. Three outcomes:
- Different prefix lengths: longest match wins. Change what is advertised or filtered.
- Same length, equal local-pref (default is typically 100): the decision fell through to AS_PATH length or later tie-breakers, which is accidental. Local-pref makes it deliberate.
- Azure path already has a higher local-pref: someone set it. Find that route-map.
Step 2: steer your outbound traffic with an inbound route-map (on-premises to cloud)
Cisco IOS XE; neighbor addresses are placeholders from the documentation range.
ip prefix-list PREFER-DX seq 5 permit 172.20.0.0/16
route-map FROM-DX permit 10
match ip address prefix-list PREFER-DX
set local-preference 200
route-map FROM-DX permit 20
router bgp 64500
address-family ipv4 unicast
neighbor 192.0.2.1 route-map FROM-DX in
Sequence 10 raises local-pref to 200 for 172.20.0.0/16 only (no le, so a more-specific is not matched by accident); sequence 20 lets every other route through untouched at its default. The Azure session needs no change, so those routes stay at the default of typically 100. To apply it without resetting the session, run clear ip bgp 192.0.2.1 soft in, which re-applies the inbound route-map to the routes the neighbor already sent instead of tearing the session down. Because local-pref travels across your iBGP, every router in the AS then prefers the DC path, including routers that learned the Azure route locally.
Azure's own routing guidance uses the same tool: a route-map with set local-preference on the neighbor session, and notes the default is typically 100 with higher preferred.
Step 3: the return direction (cloud to on-premises)
Each cloud routes toward you independently, over its own circuit. If an on-premises prefix is advertised on two circuits to the same cloud, steer with that cloud's tools: AWS evaluates the longest prefix first, then local-preference communities on private virtual interfaces (a virtual interface is the logical link, a VLAN plus a BGP session, that carries private-address traffic from your Direct Connect circuit to a VPC) (7224:7100 low, 7224:7200 medium, 7224:7300 high, evaluated before AS_PATH), then AS_PATH length, then MED. Azure does not honor communities you set on routes advertised to it, so use AS-path prepending to deprioritize a circuit.
Step 4: stop your router becoming a transit router (one that carries traffic between two other networks)
Add outbound route-maps toward both clouds that permit only your own on-premises prefixes. Without them, routes learned from one cloud can be re-advertised to the other, and cloud-to-cloud traffic starts crossing your edge. Azure also states ExpressRoute cannot be configured as a transit router. A second reason is capacity: Azure private peering supports up to 4,000 IPv4 prefixes (10,000 with the premium add-on) and the BGP session drops if the limit is exceeded, so an accidental leak of a large table would take the circuit down.
Step 5: testing and rollout
- Record
show ip bgp 172.20.0.0/16before the change. - Put one prefix in the prefix-list, apply the route-map, run the soft refresh.
- Confirm
show ip bgp 172.20.0.0/16shows local-pref 200 on the DC path and marks it best, and that the Azure path is still listed as a backup. - Compare latency to a host in that prefix before and after, with the same probe repeated at several times of day.
- Fail the DC session in a maintenance window and confirm Azure takes over, then restore.
- Add prefixes to the list a few at a time. Rollback is removing the line and running the soft refresh.
Risks
- Route oscillation. Local-pref is static, so it does not oscillate by itself; a flapping DC session does move the prefix between clouds each time it flaps, so treat session stability as a precondition. MED between AWS and Azure is not compared at all by default, because MED is only compared between paths from the same neighboring AS and the two clouds are different ASes (Azure uses AS 12076) unless
bgp always-compare-medis configured, a setting that makes the router compare MED across paths from different neighboring ASes. - Unintended shifts. An overly broad prefix-list line (
le 32means "any length up to /32", so it would match every more-specific inside the range) moves more traffic than intended; always match the exact prefix. - Failover capacity. When the DC path is lost, the Azure path takes the whole load; confirm it can.
- Asymmetric paths. Outbound via DC while returns arrive via Azure is possible if the cloud side has its own preference; confirm with a traceroute from each side.
Design a multi-homed internet edge for an enterprise with two ISPs using BGP. How do you advertise your prefixes, steer inbound and outbound traffic, and prevent your network from becoming a transit path or leaking addresses?
Sample Answer
Direct answer
Terminate one eBGP session (external Border Gateway Protocol, the protocol that exchanges routes between separate organizations, each identified by an autonomous system number or ASN) on each ISP (internet service provider), ideally on two separate edge routers joined by iBGP (internal BGP, between routers of the same ASN). Originate your public prefixes with network statements, and treat the two directions of traffic as two separate problems. Outbound traffic is steered by local preference on what you accept. Inbound traffic is steered by AS-path prepending on what you advertise. Transit and leaks are prevented by an allow-list of exactly your own prefixes on every outbound policy, and a tight allow-list on what you accept.
Design
- R1 peers with ISP-A, R2 peers with ISP-B, and R1 and R2 run iBGP so both learn both exits. By default a router passes an eBGP-learned route to its iBGP peers with the original next hop unchanged (RFC 4271 section 5.1.3), so R2 would be told that ISP-A's address 192.0.2.1 is the next hop and could only use the route if it can reach that address. Either carry the ISP link subnets in your IGP, or have each edge router advertise itself as the next hop with
neighbor <iBGP peer address> next-hop-self. With illustrative iBGP addresses R1 10.255.0.1 and R2 10.255.0.2, R1 would useneighbor 10.255.0.2 next-hop-selfand R2neighbor 10.255.0.1 next-hop-self. Local preference travels across iBGP, so a value set on R1 steers R2 and every other router in your ASN. - Example numbering (documentation-style values): your ASN 64496, ISP-A ASN 64500 at 192.0.2.1, ISP-B ASN 64501 at 192.0.2.5. Your two public /24s: P1 = 203.0.113.0/24 and P2 = 198.51.100.0/24.
- Policy goals: all outbound traffic leaves via ISP-A while it is healthy and via ISP-B otherwise; inbound traffic for P1 arrives mainly via ISP-A and for P2 mainly via ISP-B; each prefix stays reachable through the other ISP if one link dies.
Step 1: advertise your prefixes
BGP only originates a prefix with network if the exact prefix and mask are in the routing table, so back each one with a static route to Null0 (a discard interface). The Null0 route is always present, so the advertisement is not withdrawn when one internal subnet flaps.
ip route 203.0.113.0 255.255.255.0 Null0
ip route 198.51.100.0 255.255.255.0 Null0
router bgp 64496
network 203.0.113.0 mask 255.255.255.0
network 198.51.100.0 mask 255.255.255.0
Step 2: steer outbound traffic (what you accept)
Ask each ISP for a default route only (simple, small routing table). The inbound route-map accepts only 0.0.0.0/0 and sets local preference. A route-map is a list of numbered clauses checked in sequence order, and a route that matches none of them is rejected, which is called the implicit deny (nothing is permitted unless a clause says so). Higher local preference wins, and it is compared before AS-path length in best-path selection. The default value is 100.
R1 (ISP-A):
ip prefix-list DEFAULT-ONLY seq 5 permit 0.0.0.0/0
route-map ISPA-IN permit 10
match ip address prefix-list DEFAULT-ONLY
set local-preference 200
router bgp 64496
neighbor 192.0.2.1 remote-as 64500
neighbor 192.0.2.1 route-map ISPA-IN in
Reading R1's lines in order: the prefix-list DEFAULT-ONLY is a named list of prefixes, and its one entry (seq 5) permits exactly the default route 0.0.0.0/0. route-map ISPA-IN permit 10 opens clause number 10 of the policy, and permit means a route that matches the clause is accepted and has its set lines applied. match ip address prefix-list DEFAULT-ONLY is the test: the clause applies only to routes in that list. set local-preference 200 is the action. The two neighbor lines name ISP-A as the peer and apply the route-map to routes received from it (in).
R2 (ISP-B) uses the same shape with set local-preference 100. Every router now prefers the ISP-A default (200 over 100). If ISP-A fails, its default disappears and the ISP-B default takes over with no operator action. If you need per-destination control, accept full tables and set local preference by destination or neighbor AS instead.
Step 3: steer inbound traffic (what you advertise)
Other networks choose among the paths they hear mostly by AS-path length, so make the path you do not want look longer (the AS path is the list of ASNs a route has crossed) by prepending your own ASN extra times. Advertise each prefix on both links so each stays reachable if one link fails, and prepend where you want less traffic.
R1 to ISP-A prefers P1 (no prepend on P1, three extra copies on P2):
ip prefix-list MY-P1 seq 5 permit 203.0.113.0/24
ip prefix-list MY-P2 seq 5 permit 198.51.100.0/24
route-map ISPA-OUT permit 10
match ip address prefix-list MY-P1
route-map ISPA-OUT permit 20
match ip address prefix-list MY-P2
set as-path prepend 64496 64496 64496
router bgp 64496
neighbor 192.0.2.1 route-map ISPA-OUT out
Reading it: a P1 route matches clause 10 and is sent unchanged. A P2 route fails clause 10's test, moves to clause 20, matches there, and set as-path prepend 64496 64496 64496 adds three extra copies of your ASN to its AS path, so the path length is 1 + 3 = 4. A route that matches neither clause hits the implicit deny and is not sent. The out on the last line applies the route-map to routes sent to ISP-A.
R2 to ISP-B is the mirror image: P2 clean, P1 prepended three times. A remote network that hears P1 as path length 1 via ISP-A and length 4 via ISP-B normally picks ISP-A. Prepending is a hint, not a command: an upstream can override it with its own local preference, so check the result from outside (a looking glass, a public web page some networks run that shows how they see a given route, or a traceroute from another network).
Step 4: do not become a transit path or leak
The classic accident: BGP re-advertises a best path learned from one eBGP peer to the other eBGP peer, so ISP-A traffic could be carried across your network to ISP-B. Prevent it in layers.
- Outbound allow-list (the route-maps in step 3): each outbound route-map permits only your own prefixes. Anything else, including routes learned from the other ISP, hits the implicit deny.
- AS-path filter for locally originated routes only:
ip as-path access-list 10 permit ^$plusneighbor 192.0.2.1 filter-list 10 outpermits only routes with an empty AS path, which means originated in your ASN. If your software applies the route-map's prepend before the filter-list, a prepended path is no longer empty and would be rejected, so confirm the order on your platform (or test on a session that does not prepend) before relying on this layer. - Inbound allow-list: the DEFAULT-ONLY match drops everything else a provider sends, including your own prefixes coming back and private or bogus space.
- Maximum-prefix (a safety limit on how many prefixes a neighbor may send):
neighbor <ip> maximum-prefix <n>shuts the session when the neighbor sends more than n prefixes (addwarning-onlyto log instead, orrestart <minutes>to retry automatically). Size n to what the provider should send, so a provider mistake cannot flood your routers.
Verify
show ip bgp summary: each neighbor should show a prefix count (the State/PfxRcd column) instead of a state name such as Idle or Active.show ip bgp neighbors 192.0.2.1 advertised-routes: only P1 and P2, and the ISP-A list should show P2 with the longer AS path.show ip bgp 0.0.0.0andshow ip route 0.0.0.0: the ISP-A default is best with local preference 200.- Shut the ISP-A link in a maintenance window: the default should flip to ISP-B and P1 should still be reachable from outside via ISP-B.
Trade-offs and pitfalls
- A single edge router with two links is a single point of failure. Two edge routers cost more but remove it.
- Default-only is simple but you cannot tell which ISP has the shorter path to a given destination. Full tables give control and need router memory and a longer convergence.
- Never accept a longer-prefix or private-space announcement from a provider without a reason. Check each provider's own prefix-length and filtering policy before splitting a prefix to steer traffic.
- Forgetting the Null0 route is the most common reason
networkadvertises nothing. - Changing a BGP policy does not apply to routes already exchanged until they are re-evaluated. Use a soft refresh:
clear ip bgp 192.0.2.1 soft inre-applies inbound policy andclear ip bgp 192.0.2.1 soft outre-sends under the outbound policy, both without dropping the session. A plain hardclear ip bgptears down the TCP session and deletes the routes learned from that peer, which on an internet edge is an outage.
That is every published Routing Protocols and Configuration question for Cloud Architect so far. Browse the other topics in this category, or practice this one interactively.