Routing Protocols and Configuration Questions
How traffic finds its path across networks: route selection basics (administrative distance, longest-prefix match, control plane versus forwarding plane), static and default routing, interior gateway protocols (OSPF, IS-IS, EIGRP, RIP), BGP, and multicast routing. Covers protocol selection and migration, OSPF area and IS-IS level design, EIGRP feasible successors, BGP path selection, communities, route reflectors and policy, route redistribution, summarization and filtering, VRF-Lite and VRF route leaking, convergence tuning with BFD, ECMP path selection, BGP-based inbound and outbound traffic engineering, multi-homed internet edge and full-table scale, RPKI origin validation and containing route leaks, route flap and oscillation analysis, graceful maintenance drains, first-hop gateway redundancy (HSRP, VRRP, GLBP), and verifying and troubleshooting routing on real devices. Excludes network-wide topology and MPLS or segment-routing backbone design, Layer 2 switching and VLANs, general layered fault isolation and packet capture, cloud VPC routing, and firewall or ACL security.
A router has routes for 10.0.0.0/8, 10.1.0.0/16 and 10.1.2.0/24, and receives a packet for 10.1.2.5. Which route is used and why? Then what changes if the /24 disappears, and what decides between two routes of the same length from different protocols?
Sample Answer
Direct answer
The router uses 10.1.2.0/24. All three routes contain the address 10.1.2.5, and the rule is longest-prefix match: the most specific matching route (the one with the most leading bits in common with the destination) is used. A /24 names 256 addresses and a /16 names 65,536, so the /24 is the narrower, more specific claim. If the /24 disappears, the packet follows the /16. When two routes have the same length but come from different protocols, administrative distance (AD, the per-source trust number, lower wins) decides; within one protocol the metric (that protocol's own cost number for a path, lower is better) decides.
Why all three match, computed
Mask the destination with each prefix's mask and compare with the network address:
- 10.1.2.5 AND 255.255.255.0 = 10.1.2.0, which equals the /24 network.
- 10.1.2.5 AND 255.255.0.0 = 10.1.0.0, which equals the /16 network.
- 10.1.2.5 AND 255.0.0.0 = 10.0.0.0, which equals the /8 network.
| Route | Address range it covers | Matches 10.1.2.5? | Prefix length |
|---|---|---|---|
| 10.0.0.0/8 | 10.0.0.0 to 10.255.255.255 | yes | 8 |
| 10.1.0.0/16 | 10.1.0.0 to 10.1.255.255 | yes | 16 |
| 10.1.2.0/24 | 10.1.2.0 to 10.1.2.255 | yes | 24 |
The /24 wins because 24 is the longest. RFC 1812 states the requirement as routers using the most specific matching route (the longest matching network prefix).
What changes when the /24 disappears
Only the /8 and /16 remain, both still match, so the /16 (longer) is chosen. Nothing is dropped: traffic simply flows to the /16's next hop. That is the danger. If the /24 was a carve-out (a small block cut out of a larger block so it can be sent somewhere different) that pointed at a special site, its packets now go to wherever the /16 points, and they may be forwarded to the wrong place or looped without any alarm.
Other destinations show the same rule:
| Destination | Routes that match | Chosen |
|---|---|---|
| 10.1.2.5 | /8, /16, /24 | 10.1.2.0/24 |
| 10.1.3.9 | /8, /16 | 10.1.0.0/16 |
| 10.2.0.1 | /8 | 10.0.0.0/8 |
| 192.0.2.1 | none of the three | default route if present, otherwise the packet is dropped |
Same length, different protocols
The decision order is: (1) longest prefix, (2) lowest AD, (3) lowest metric inside the winning protocol, (4) if still tied, equal-cost multipath (the router installs several next hops, the neighbor addresses it hands packets to, and shares traffic across them). With Cisco defaults, an OSPF /24 (AD 110) beats a RIP /24 (AD 120) for the same prefix, and the metrics are never compared. A RIP /24 still beats an OSPF /16, because length is checked before AD. You can confirm the winner with show ip route 10.1.2.5, which prints the entry (Routing entry for 10.1.2.0/24) and the source and distance of the installed route.
Trade-offs and pitfalls
- Specificity beats trust: a more specific route wins even from a less trusted source. That is why internet-facing filters limit how long a prefix a neighbour may announce, since a leaked or hijacked /24 would override a legitimate /16. For example, if the legitimate owner announces 10.1.0.0/16 and someone else announces 10.1.2.0/24, every packet for the 256 addresses in that /24 follows the /24 to the wrong network, even though the /16 is genuine and the same router also holds it; only the other 65,280 addresses (65,536 - 256) still follow the /16.
- Summarization (advertising one broader prefix instead of many narrower ones) relies on this rule. A summary route can be advertised broadly while a specific route steers exceptions, but if the specific route fails the traffic falls back to the summary silently.
- The default route 0.0.0.0/0 is a /0, so it matches everything and loses to every other match.
Where in an enterprise network would you originate or place default routes: edge, distribution or core? Walk through the trade-offs for a site with two internet exits.
Sample Answer
Direct answer. Originate the default route at the edge, on the routers that physically hold the internet exits, and only while the exit is healthy. Let the distribution layer pass it down and, if the access side sits in an OSPF (Open Shortest Path First) stub area, know that its border routers (ABRs, area border routers) hand it a default of their own, which does not follow the exit's health (explained below). Do not originate it in the core or distribution: those layers have no knowledge of whether the internet path beyond them works, so a default originated there can keep attracting traffic to a dead exit. For two exits, originate from both with deliberately different costs (active/backup) or equal costs (hot-potato routing: each router hands traffic to the nearest exit), and make each origination conditional on a probe of that exit.
What "originate" means here
A default route is 0.0.0.0/0: the route a router uses when no more specific route matches. A router either has a real one (a static route toward the ISP, or one learned by BGP (Border Gateway Protocol) from the ISP) or it learns one from its IGP (interior gateway protocol, such as OSPF). In OSPF the router that injects the default is an ASBR (autonomous system boundary router, a router that injects routes learned from outside OSPF). Cisco IOS command:
router ospf 1
default-information originate [always] [metric metric-value] [metric-type type-value] [route-map map-name]
By default OSPF originates the default only if the router already has a default route in its own routing table. The always keyword removes that condition and advertises it regardless. That is dangerous: a router with a dead ISP link and always keeps attracting traffic it cannot deliver.
Layer by layer trade-offs
| Where | Gain | Risk |
|---|---|---|
| Edge (the internet-facing routers) | The default exists only while the router really has an upstream route; failover follows the exit's health; no extra hop to reach the exit | Needs two ASBRs configured identically; the edge becomes part of the IGP's external routing |
| Distribution | Close to the access layer | Distribution routers usually have no internet attachment, so a static default pointing at the core is blind to exit failures, and each distribution block needs its own decision |
| Core | One place to configure | Same blindness; also the core is a high-blast-radius device (one whose mistakes affect many users) to carry static policy; a mistake here affects every site |
For example, if a core router carries a static default toward E1 and E1's ISP link fails, the core keeps advertising that default and keeps sending internet traffic to E1, which has nowhere to deliver it, because nothing in the core ever tested the exit.
Two internet exits: the design
Assume edge routers E1 (ISP A) and E2 (ISP B), both in OSPF area 0. Each uses a static default toward its ISP gated on a probe (an IP SLA (service-level agreement) ICMP echo to a stable address beyond the ISP), so the default leaves the routing table when the exit stops answering:
ip sla 1
icmp-echo 198.51.100.1 source-ip 203.0.113.2
frequency 5
ip sla schedule 1 life forever start-time now
track 10 ip sla 1 reachability
ip route 0.0.0.0 0.0.0.0 203.0.113.1 track 10
router ospf 1
default-information originate metric 10 metric-type 1
Reading the lines: ip sla 1 with icmp-echo 198.51.100.1 source-ip 203.0.113.2 defines a probe that pings 198.51.100.1 (an illustrative address beyond the ISP) from E1's ISP-facing address; frequency 5 repeats it every 5 seconds; ip sla schedule 1 life forever start-time now starts it immediately and keeps it running; track 10 ip sla 1 reachability makes track object 10 follow whether the probe succeeds; ip route 0.0.0.0 0.0.0.0 203.0.113.1 track 10 keeps the default via the ISP only while track 10 is up; default-information originate metric 10 metric-type 1 advertises it into OSPF. E2 mirrors it with its own probe and track object. Cisco's static-route tracking installs the route only while the track object is up; with the route removed and no always, OSPF stops originating the default, so the other edge's default takes over.
Choosing the policy with the metric type. OSPF external routes come in two kinds. For a type 1 external route the total cost is the internal distance to the ASBR (X) plus the advertised metric (Y), X+Y. For type 2 the two parts are compared separately and the advertised metric Y is compared first. Type 1 routes are always preferred over type 2 routes.
- Nearest exit (hot-potato): both edges use
metric-type 1with the same metric. Each router picks the exit with the lower X+Y, so sites close to E1 leave via ISP A and sites close to E2 leave via ISP B. - Primary/backup: E1 uses a lower advertised metric than E2 (for example 10 and 100) with type 2. Every router prefers E1 regardless of where it sits; E2's default is used only after E1's disappears.
Worked case (illustrative distances): router R is 5 from E1 and 20 from E2, and router S is 20 from E1 and 5 from E2. Nearest-exit with type 1 and metric 10 on both edges: R compares 5 + 10 = 15 via E1 against 20 + 10 = 30 via E2 and exits through E1; S compares 20 + 10 = 30 against 5 + 10 = 15 and exits through E2. Primary/backup with type 2, metric 10 on E1 and 100 on E2: both compare 10 against 100 first and exit through E1, even S, which sits 5 from E2.
I would choose nearest-exit when both ISPs have similar capacity and cost, because it spreads load by geography and shortens paths. I would choose primary/backup when one ISP is metered, smaller, or has better routing; then write the policy into the design document.
Return path. Whichever exit outbound traffic uses, return traffic is decided by what each ISP hears from you (for example BGP advertisements that make one path look longer to the ISP's network, known as prepending). If the two exits do stateful NAT (network address translation) or a firewall, nearest-exit can send the reply through the other exit and break the session unless state is shared (a stateful device remembers each connection and drops a reply that arrives on a path where it never saw the request). That asymmetry is the usual reason to prefer primary/backup at a site with stateful firewalls.
Below the edge
Distribution blocks that do not need the full internal table can sit in an OSPF stub area (the ABR sends a default route into the area instead of external routes) or a totally stubby area (area X stub no-summary on the ABR also suppresses inter-area summaries). Per Cisco's area stub documentation, area stub must be configured on all routers in the stub area, and no-summary applies only to the ABR. Stub areas keep the access routers' tables small, but the default inside them is not the edge's default. In a stub area each ABR originates its own default summary LSA into the area, at a cost set per area (Cisco area default-cost, default 1; RFC 2328 section 12.4.3.1 calls the cost StubDefaultCost). Neither RFC 2328 nor Cisco's documentation ties that default to the ABR holding a working internet route, so if both ISP links fail the access routers in a stub area still hold a default toward the ABR and send internet traffic to it, where it is dropped. Where the access side must stop sending traffic when both exits are down, keep that block in a normal (non-stub) area so it learns the edge-originated default as an external route, which disappears when the edge withdraws it.
Pitfalls
alwayson both edges: an exit with a dead ISP link keeps advertising a default.- A probe aimed at the ISP's next hop only proves the first link works; probe a stable address farther out, and understand that a shared probe target is itself a single point of failure.
- Different metric types on the two edges: type 1 is always preferred over type 2, so a type 1 default on E2 beats a type 2 default on E1 whatever the metric values.
- Forgetting that removing the static route is what withdraws the default: without gating, the default stays.
- Assuming a stub area inherits the edge's health gating: the ABR's own default into the stub area stays up while both exits are dead.
What is the difference between the control plane and the forwarding plane on a router? Explain how the routing table, the forwarding table and adjacency information relate when a new route is learned.
Sample Answer
Direct answer
The control plane is the part of the router that decides: routing protocols run there, exchange routes, and build the routing table (the RIB, Routing Information Base). The forwarding plane (also called the data plane) is the part that moves packets, using a compact lookup structure the control plane programs for it. On Cisco routers that structure is the Forwarding Information Base (FIB) used by Cisco Express Forwarding (CEF), plus an adjacency table holding the Layer 2 rewrite information for each next hop. Once a route is programmed, transit packets (packets passing through the router toward some other destination, as opposed to packets addressed to the router) are normally forwarded by the forwarding plane alone, without involving the control plane. The exceptions are packets that cannot be forwarded yet, such as those waiting for an unresolved ARP entry (step 4 below), and packets that need special handling, which are punted (handed up) to the CPU. The control plane builds, then the forwarding plane executes.
What each table holds
- RIB (routing table): every route's source, administrative distance (AD), metric and next hop. AD is a trust number per route source (for example OSPF defaults to 110 and RIP to 120), and the metric is the protocol's own cost for the path. When several protocols offer the same prefix, the lowest AD wins, and within a protocol the lowest metric wins. The RIB is what
show ip routedisplays. - FIB (CEF table): only the winning prefixes, reduced to prefix to next hop and outgoing interface for fast longest-prefix lookup. The CEF table serves as the FIB.
show ip cefdisplays it. - Adjacency table: the Layer 2 header (rewrite string, the ready-made Ethernet header with the next hop's MAC address that gets put on the front of each packet) for each directly connected next hop, which makes it CEF's equivalent of an ARP entry.
show adjacency detaildisplays it. This is a different use of the word from an OSPF or BGP adjacency, which means a relationship between two neighboring routers; here an adjacency is only a stored Layer 2 address for a next hop.
What happens when a new route is learned
- The routing protocol (for example OSPF, via an LSA, a link-state advertisement it just received, or BGP via an UPDATE message) computes a candidate route in the control plane.
- The RIB compares candidates for the prefix and installs the best one by AD, then by metric.
- The change is passed to the forwarding plane. On IOS XE routers the update goes from the control plane process (IOSd, the IOS software that runs on the main CPU) through the forwarding manager (a software layer that translates routing changes into hardware entries) into the forwarding hardware (the line cards or forwarding chips that move packets), which programs its tables. A prefix has a FIB entry, and the entry names a next hop and interface.
- The FIB entry points to an adjacency. If the router already has the next hop's MAC address (from ARP, the Address Resolution Protocol that maps an IP address to a MAC hardware address), the adjacency is complete and packets are rewritten and sent in the forwarding plane. If ARP has not resolved, the adjacency is incomplete: packets that need it are punted to the CPU (handed up from the fast forwarding path to the slower general-purpose processor) so ARP can ask for the address, and until the reply arrives those packets wait or are dropped. During a ping test that shows up as some pings timing out, and as every ping timing out if the next hop never answers ARP. An incomplete adjacency therefore means the router has the route and the FIB entry but no MAC address for the next hop yet.
- For a packet, the forwarding plane does a longest-prefix match in the FIB (the most specific prefix that contains the destination address wins), selects the adjacency, rewrites the Layer 2 header, decrements TTL and sends it out. No routing protocol is consulted.
Worked example
A router learns 10.20.0.0/24 from OSPF via 10.0.0.2 on GigabitEthernet0/0.
show ip route 10.20.0.0shows a line such asKnown via "ospf 1", distance 110, metric <metric>and the next hop 10.0.0.2. In the one-line table form (show ip route ospf) the same pair appears as[110/<metric>]: 110 is the AD and the second number is the OSPF metric.show ip cef 10.20.0.0 detailshows the FIB entry pointing at 10.0.0.2 and GigabitEthernet0/0.show adjacency GigabitEthernet0/0 detailshows the MAC rewrite for 10.0.0.2. If the entry is incomplete, meaning the MAC address for 10.0.0.2 has not been learned, checkshow ip arpfor whether the router has a resolved entry for 10.0.0.2, and check that 10.0.0.2 is reachable at Layer 2 on that interface.
If the RIB and the FIB disagree, the router forwards by the FIB, which is why a route inshow ip routecan still not carry traffic.
Why it matters operationally
- Control-plane overload is not the same as forwarding failure. On a router with dedicated forwarding hardware, the control plane (its CPU) can be too busy to process routing protocol hellos, so adjacencies drop, while the forwarding plane keeps carrying traffic using the entries it already has.
- Convergence has two stages. After a BGP session comes up and the route is in the routing table, the prefixes may not be reachable for a while because the forwarding modules must be programmed. How long that takes depends on the platform and the table size, so measure it rather than assume a figure.
Pitfalls
- Saying "the router looks at the routing table for every packet" is wrong for CEF platforms. The FIB does the per-packet lookup.
- Debugging only
show ip routewhen the symptom is a forwarding problem. - Forgetting the adjacency: a correct route to a next hop whose ARP is unresolved still drops packets.
Two Cisco routers are the default gateways for a user subnet, with virtual IP 10.50.0.1. Configure HSRP so Router1 is preferred and takes back the role after recovering. What role transitions occur when Router1 fails and returns?
Sample Answer
Direct answer
HSRP (Hot Standby Router Protocol) lets two routers share one virtual gateway address, 10.50.0.1, with one router Active (forwarding) and one Standby (ready to take over). To make Router1 preferred and to make it reclaim the role after it recovers, give it a higher priority than Router2 and enable preempt on it. On failure, Router2 goes Standby to Active; on recovery, Router1 goes through Init, Listen and Speak and then takes Active back (with the 60 second preempt delay configured below it normally waits as Standby first), and Router2 returns to Standby.
Configuration
Router1 (preferred):
interface GigabitEthernet0/0/0
ip address 10.50.0.2 255.255.255.0
standby 1 ip 10.50.0.1
standby 1 priority 110
standby 1 preempt delay minimum 60
Router2 (default priority):
interface GigabitEthernet0/0/0
ip address 10.50.0.3 255.255.255.0
standby 1 ip 10.50.0.1
standby 1 preempt
- Priority range is 1 to 255 and the default is 100, so 110 beats Router2's 100.
preemptmakes a router with a higher priority than the current Active take over. Preempt is not implied by a higher priority alone: it has to be configured.delay minimum 60makes Router1 wait 60 seconds before taking over the Active role, so routing and switching have time to settle after a reboot. The delay range is 0 to 3600 seconds and the default is 0.- Hosts use 10.50.0.1 as default gateway. With HSRP version 1, group 1 uses virtual MAC 0000.0C07.AC01 (the Layer 2 address that goes with the virtual IP), and the new Active router takes over that same MAC, so the ARP entry hosts cache (their IP-to-MAC mapping) stays valid through a failover.
- Router2 also gets
preemptso the pair behaves symmetrically: if Router1 is Active but its priority later falls below 100 (for example through tracking, covered in the pitfalls), Router2 can only take the role from it if Router2 has preempt configured. - Defaults: hello 3 seconds, hold time 10 seconds.
Role transitions
HSRP states are Init, Learn, Listen, Speak, Standby and Active.
| Event | Router1 | Router2 |
|---|---|---|
| Both up, stable | Active | Standby |
| Router1 loses power at t=0 | down | no hellos from Router1 for the 10 s hold time, then Standby to Active (takes 10.50.0.1 and the virtual MAC) |
| Router1 shut down cleanly | sends a Resign message | takes over immediately, without waiting for the hold time |
| Router1 boots | Init, then Listen (hears Router2's hellos), then Speak, then Standby once no Standby hello has been heard for the hold time (the RFC 2281 Standby timer expires) | Active, keeps forwarding |
| Preempt delay (60 s) expires, Router1 has priority 110 against 100 | sends a Coup message and becomes Active | on the Coup (or an Active hello from a higher priority router) moves to Speak and sends a Resign, then settles into Standby |
| Stable again | Active | Standby |
What the messages mean (RFC 2281): a Hello says the router is running and can become Active or Standby, a Coup says it wants to become Active, and a Resign says it no longer wants to be Active. Speak is the state where a router sends hellos and takes part in electing Active and Standby. A router that has not yet learned the virtual IP waits in Learn; here the address is configured, so Router1 goes straight from Init to Listen.
Timing of the failure case: Router2 declares Router1 dead 10 seconds after the last hello it received. Hellos arrive every 3 seconds, and Router1 can fail anywhere in that interval, so takeover happens between 7 and 10 seconds after the failure. If you need faster takeover, shorten hello and hold times (Cisco's command reference allows millisecond hello timers from 15 to 999 ms and recommends hold times under 250 ms only on Cisco 7200-class or better platforms and on Fast Ethernet or faster interfaces).
Verification
show standby brief on both routers. Expected stable output on Router1: group 1, priority 110, state Active, virtual IP 10.50.0.1. show standby gives the long form, including Preemption enabled and the Active virtual MAC address. After failing Router1, Router2's state should read Active. Test by pinging 10.50.0.1 from a host during the event.
Pitfalls
- Forgetting
preempton Router1: when it returns it stays Standby and Router2 keeps the Active role indefinitely, which is the opposite of what was asked. - Preempt with no delay: Router1 can become Active before its uplink routing is ready, black-holing traffic (accepting packets it cannot forward, so they vanish). Tracking an object that follows the uplink (
standby 1 track <object> decrement 20, which lowers the priority by 20 while the tracked object is down, taking 110 to 90, below Router2's 100) makes it give up Active when its uplink fails. - HSRP watches the interface it runs on, not the path beyond it; without tracking, hosts keep sending to an Active router that has lost its uplink.
How do HSRP, VRRP and GLBP provide a redundant default gateway? Compare their election, virtual address behavior and timers, and where each fits.
Sample Answer
Direct answer. A first-hop redundancy protocol (FHRP) lets two or more routers share one virtual gateway IP address, so hosts keep a single static default gateway while the routers decide among themselves who forwards. HSRP (Hot Standby Router Protocol, Cisco-originated) and VRRP (Virtual Router Redundancy Protocol, IETF standard in RFC 5798) elect one forwarder per group and keep the others waiting. GLBP (Gateway Load Balancing Protocol, Cisco-originated) elects one coordinator but lets up to four routers forward at once.
How a host sees it. The host ARPs (Address Resolution Protocol) for the virtual IP and gets a virtual MAC address back. When the forwarding router dies, a peer starts answering for the same virtual IP and virtual MAC, so the host's ARP cache (its remembered IP-to-MAC mappings) stays valid and nothing on the host changes.
Election, virtual address and timers
In plain terms the three differ in four things: who forwards, whether a returning higher-priority router takes the role back, which MAC address the host learns, and how fast a silent failure is noticed. The routers find each other with hello messages sent to a multicast address (a group address that only the group's members listen to). The table gives the values.
| HSRP | VRRP | GLBP | |
|---|---|---|---|
| Roles | Active, Standby; others Listen | Master, Backup | AVG (active virtual gateway, the coordinator), up to four AVFs (active virtual forwarders) |
| Election | Highest priority wins; default priority 100; tie goes to the higher primary IP address | Highest priority wins; default 100 for a backup; tie goes to the higher primary IP address | AVG by highest priority (default 100); a standby virtual gateway takes over if the AVG fails |
| Preemption (a better router taking the role back) | Off by default, needs standby preempt | On by default (Preempt_Mode True) | Gateway preempt off by default; forwarder preempt on by default with a 30 s delay |
| Virtual IP | A separate address, not any router's real address | Can be the real address of one router (the owner, priority 255) | A separate address |
| Virtual MAC | v1: 0000.0C07.ACxx (xx = group in hex); v2: 0000.0C9F.Fxxx | 00-00-5E-00-01-{VRID} (VRID = virtual router ID) | 0007.b400.xxyy style, one MAC per forwarder (example 0007.b400.0101) |
| Hellos | 3 s hello, 10 s hold by default; multicast 224.0.0.2 (v1) or 224.0.0.102 (v2) | Advertisement every 1 s by default; multicast 224.0.0.18, IP protocol 112 | Multicast 224.0.0.102, UDP 3222; hello and hold are configurable with glbp timers |
| Failure detection | Hold time expiry | Master_Down_Interval = (3 x advertisement interval) + skew time (a small extra delay, worked out below, that is shorter for higher-priority backups) | Hold time expiry, plus per-forwarder timers |
| Standard | Cisco proprietary | IETF RFC 5798 | Cisco proprietary |
Load behaviour. HSRP and VRRP use one forwarder per group. To use both routers you run two groups (group 1 active on router A and standby on B, group 2 the reverse) and split the VLAN's hosts between the two gateway IPs by DHCP (Dynamic Host Configuration Protocol) option or by hand. GLBP does the splitting itself: the AVG answers each ARP request with the virtual MAC of a different AVF, chosen by host-dependent (the reply depends on which host asked), round-robin (forwarders take turns) or weighted (share proportional to assigned weights) load balancing (glbp 1 load-balancing round-robin). When an AVF fails, the redirect timer (default 600 s) is how long the AVG keeps replying to ARP requests with the failed forwarder's virtual MAC; after it expires the AVG stops using that MAC in ARP replies. The forwarder timeout (default 14,400 s, 4 hours) is the interval before a secondary virtual forwarder becomes invalid.
Worked example: VRRP failover time
Two routers, virtual IP 10.1.10.1, advertisement interval 1 s, router A priority 120, router B priority 100. Router B waits for the Master_Down_Interval:
skew=256(256−100)×1 s=0.609 s,Master_Down=3×1+0.609=3.609 sSo a silent failure costs hosts about 3.6 s of lost gateway. The skew term deliberately makes a higher-priority backup time out sooner (priority 255 would give 3.004 s), so the best backup wins without a race. HSRP at its default 3 s hello and 10 s hold takes up to 10 s for the same event. Both can be tightened (HSRP accepts sub-second values such as standby 1 timers msec 250 msec 800: Cisco's command reference allows a hello of 15 to 999 ms and a hold time of 50 to 3000 ms, normally at least three times the hello, and warns that HSRP state can flap if the hold time is under 250 ms while the processor is busy), but aggressive timers raise the chance that a busy CPU causes a false failover.
interface Vlan10
ip address 10.1.10.2 255.255.255.0
standby version 2
standby 10 ip 10.1.10.1
standby 10 priority 110
standby 10 preempt
standby 10 track 5 decrement 20
interface Vlan10 is the switch virtual interface (SVI), the router's own address on VLAN 10, which hosts use as their gateway. This is HSRP version 2 on router A: priority 110 beats the default 100, preempt lets A reclaim the role after a reload, and track 5 refers to a tracked object, a monitor defined elsewhere in the configuration (illustratively, one that watches the state of A's uplink interface; its definition is not shown here). When the tracked object goes down it lowers A's priority by 20 to 90. B (priority 100) takes over only if B is configured with standby 10 preempt as well: Cisco's HSRP guide states that the device with the higher priority can become active if it has standby preempt configured. Without preempt on B, B would stay Standby while A, still sending hellos, kept the Active role and the black hole. Without tracking, A would stay Active while its uplink was dead and black-hole the VLAN (silently drop its traffic).
Where each fits, and my pick
- VRRP when the gateways are from different vendors or the standard matters: it is the only one of the three that is an open standard.
- HSRP on an all-Cisco campus where the team already knows it; simplest operations, one forwarder per group, easy to reason about.
- GLBP only where you must spread one VLAN's upstream traffic over two or more gateways without managing multiple groups. Its per-host balancing makes troubleshooting harder (different hosts use different MACs), and it is Cisco-only.
My default: HSRP or VRRP with two groups per VLAN pair if load sharing matters, GLBP as the exception.
Pitfalls
- The FHRP active router must also be the Spanning Tree root (or at least on the forwarding path) for that VLAN, otherwise traffic crosses the inter-switch link twice. Spanning Tree is the protocol that blocks redundant switch links to prevent loops, and the root switch is the centre that the open paths lead toward. If the root is switch B but the active gateway is switch A, a host's frame travels toward B, crosses the B-to-A link to reach A, and the reply makes the same extra crossing back.
- No preemption means that after an outage the traffic stays on the backup until something else fails; with preemption on and no delay, a flapping router can cause repeated failovers. Preemption must be on at both ends of a tracking design: the Active router lowers its priority, and the Standby router needs
preemptto act on the difference. - A priority that does not track the real uplink keeps the gateway "alive" for hosts while it has no route out.
- Skipping
standby version 2leaves a group on the version 1 limits (group numbers 0 to 255, multicast 224.0.0.2); version 2 allows groups 0 to 4095 and uses 224.0.0.102.
Unlock Full Question Bank
Get access to all 8 Routing Protocols and Configuration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.