Routing Protocols and Configuration Questions
How traffic finds its path across networks: route selection basics (administrative distance, longest-prefix match, control plane versus forwarding plane), static and default routing, interior gateway protocols (OSPF, IS-IS, EIGRP, RIP), BGP, and multicast routing. Covers protocol selection and migration, OSPF area and IS-IS level design, EIGRP feasible successors, BGP path selection, communities, route reflectors and policy, route redistribution, summarization and filtering, VRF-Lite and VRF route leaking, convergence tuning with BFD, ECMP path selection, BGP-based inbound and outbound traffic engineering, multi-homed internet edge and full-table scale, RPKI origin validation and containing route leaks, route flap and oscillation analysis, graceful maintenance drains, first-hop gateway redundancy (HSRP, VRRP, GLBP), and verifying and troubleshooting routing on real devices. Excludes network-wide topology and MPLS or segment-routing backbone design, Layer 2 switching and VLANs, general layered fault isolation and packet capture, cloud VPC routing, and firewall or ACL security.
You must migrate a large network from one IGP to another (for example an old RIP network, or OSPF to IS-IS) with minimal disruption. Outline the plan: how both run side by side, how metrics map, the cutover sequence, testing and rollback.
Sample Answer
Direct answer. Run the new IGP (interior gateway protocol, the routing protocol inside one organization) next to the old one on the same links, with the new one deliberately losing in the routing table until you choose to flip it. Router by router, then site by site, you move the preference, watch, and keep the old protocol running as a one-line rollback until a soak period (a fixed observation time with no further changes) ends. Only then do you remove it. The same method covers RIP (Routing Information Protocol) to OSPF (Open Shortest Path First), RIP to EIGRP (Enhanced Interior Gateway Routing Protocol), and OSPF to IS-IS (Intermediate System to Intermediate System).
How two protocols coexist: administrative distance
Each protocol builds its own database and offers routes to the routing table. When two protocols offer the same prefix, the lower administrative distance (AD, a trust ranking) wins. Cisco defaults: EIGRP internal 90, OSPF 110, IS-IS 115, RIP 120, EIGRP external 170. That drives the plan:
- RIP to OSPF or EIGRP: both have a lower AD than RIP (110 and 90 against 120), so the new protocol wins the moment it learns a route. Raise the new protocol's AD first so RIP keeps winning during staging; for OSPF, Cisco's form is
distance ospf intra-area 130, with theinter-areaandexternalkeywords given matching values. - OSPF to IS-IS: OSPF (110) beats IS-IS (115) by default, so IS-IS starts out as the silent backup. Cutover is lowering IS-IS's distance below OSPF's with
distance 100underrouter isis(Cisco's generic form,distance <AD-value>under a routing process, changes that process's distance; check the optional keywords your release offers with?), and rollback is removing that line.
Metric mapping
RIP counts hops, with 16 meaning unreachable and 15 the longest usable path (RFC 2453). OSPF and IS-IS use link costs. They rank paths differently, so a straight swap can change real traffic paths. Choose deliberately, in this order:
- List the paths that matter (data center to each site, site to site, any intentionally preferred backup).
- Set the new protocol's link costs so that these paths come out the same as before. OSPF
ip ospf costsets a link's cost explicitly; for OSPF setauto-cost reference-bandwidthto the same value on every router (default 100 Mbps). For IS-IS enablemetric-style widefirst. - Worked case (illustrative): a site reaches the data center either over one hop on a 2 Mbps link or over two hops on 1 Gbps links. RIP picks the one-hop path. With
auto-cost reference-bandwidth 100000OSPF costs are 100000 / 2 = 50,000 for the slow link and 100000 / 1000 = 100 per fast link, so 2 x 100 = 200 for the two-hop path, and OSPF would switch traffic to the fast path. To keep the old path you would setip ospf coston the slow link below 200; to take the better path you leave the computed costs alone and record the change. - Where RIP's hop count picked a worse path than cost would (a one-hop slow link beating a two-hop fast pair), decide whether to keep the old behavior or fix it, and write it down as an intended change.
During partial migration, where some sites are still on the old protocol, a border router redistributes between them, meaning it copies routes learned from one protocol into the other. Route-map the redistribution (a route-map is an ordered rule list that matches routes and permits, denies or changes them) and tag it (a tag is a number attached to a route as a label):
router ospf 1
redistribute rip subnets metric 100 route-map RIP-TO-OSPF
router rip
version 2
redistribute ospf 1 metric 5 route-map OSPF-TO-RIP
!
route-map RIP-TO-OSPF deny 10
match tag 200
route-map RIP-TO-OSPF permit 20
set tag 100
route-map OSPF-TO-RIP deny 10
match tag 100
route-map OSPF-TO-RIP permit 20
set tag 200
Line by line: the redistribute rip subnets metric 100 route-map RIP-TO-OSPF line copies RIP routes into OSPF with a seed metric of 100 (the starting metric given to routes copied into the new protocol), passing each through the route-map; in RIP-TO-OSPF, sequence 10 denies anything tagged 200 and sequence 20 permits everything else and tags it 100; the RIP side repeats this with seed metric 5 and the tags swapped. Routes that came from RIP get tag 100 as they enter OSPF, and OSPF-to-RIP redistribution denies tag 100, so they never return; the reverse holds for tag 200. This stops the loop that appears when two border routers redistribute both ways. The seed metric into RIP must be 15 or less (5 here), because 16 is infinity. The tags only work if they travel inside RIP updates: the route tag field is defined by RIP version 2 (RFC 2453 section 4.2) and a RIPv1 message has only a must-be-zero field in that place, so a RIPv1 network cannot carry the tags and must be moved to version 2 first, or the loop guard has to be done differently (for example with prefix lists at the border). Hold redistribution to the border routers and remove it when the last site moves.
Phases (the 200-router RIP case)
Take 200 routers moving from RIP to OSPF.
- Design. Inventory every router's RIP configuration: network statements, passive interfaces, static routes redistributed, summaries, route filters. Define the OSPF target: area plan (a single area may suffice for 200 routers, or a backbone plus regions), router IDs, link costs from the mapping above, summarization points, authentication. Write the success criteria (every prefix in the old tables is reachable, same next-hop path for the listed critical flows) and the rollback trigger (for example, any critical prefix missing for more than 5 minutes).
- Staging. Rebuild a representative slice in a lab or virtual topology (one site of each type, the data center, the WAN edge). Run the exact configs and the cutover and rollback commands. Capture
show ip routebefore and after for comparison. - Pilot. Enable OSPF on every router of one low-risk site and its upstream with the raised distance (130), so RIP still carries traffic. Check
show ip ospf neighbor(every adjacency FULL) andshow ip ospf database(every router and prefix present). Then flip only the pilot site's distance below RIP, run a full business cycle, and compare route tables and application tests with the staged expectations. - Cutover. In waves of about 20 routers (10 waves for 200), lower OSPF's distance below 120 per wave, validating after each wave with the same checks plus a traceroute set. The change per router is one line and can be pushed by a change tool. Keep RIP configured and running.
- Rollback. At any point, raise OSPF's distance back to 130 on the affected routers: RIP, still running, immediately has the best route again. Because both protocols stayed up, rollback needs no rebuild.
- Decommission. After a soak period with no rollbacks, remove redistribution, then remove RIP router by router (
no router rip), then remove the raised-distance lines, and update monitoring and documentation.
The same sequence applies to OSPF to IS-IS: enable IS-IS on every link while OSPF (distance 110) still wins over IS-IS (115), set the IS-IS link metrics (with metric-style wide) so the critical paths come out as they do under OSPF, check that the IS-IS table holds the same prefixes with the same next hops, then lower IS-IS's distance with distance 100 wave by wave, with rollback being removal of that line.
Testing checklist for each wave
- Adjacencies at expected count; LSDB prefix count equals the inventory.
- Compare route tables before and after: same prefixes, same exit interface for the critical list.
- Failure test on the pilot: shut one link and verify convergence and that RIP's backup role works.
- Monitor for the first full business day, including batch jobs.
Pitfalls
- Leaving the new protocol at its default distance: RIP-to-OSPF flips routing on every router the moment its first OSPF adjacency forms, with no staging.
- Redistributing both ways at two places without tags: it creates a loop that looks like random reachability loss.
- Mismatched OSPF reference bandwidth between routers, so the same path has different costs depending on who computes it.
- Removing the old protocol on the same day as cutover: it removes the rollback.
Two Cisco routers are the default gateways for a user subnet, with virtual IP 10.50.0.1. Configure HSRP so Router1 is preferred and takes back the role after recovering. What role transitions occur when Router1 fails and returns?
Sample Answer
Direct answer
HSRP (Hot Standby Router Protocol) lets two routers share one virtual gateway address, 10.50.0.1, with one router Active (forwarding) and one Standby (ready to take over). To make Router1 preferred and to make it reclaim the role after it recovers, give it a higher priority than Router2 and enable preempt on it. On failure, Router2 goes Standby to Active; on recovery, Router1 goes through Init, Listen and Speak and then takes Active back (with the 60 second preempt delay configured below it normally waits as Standby first), and Router2 returns to Standby.
Configuration
Router1 (preferred):
interface GigabitEthernet0/0/0
ip address 10.50.0.2 255.255.255.0
standby 1 ip 10.50.0.1
standby 1 priority 110
standby 1 preempt delay minimum 60
Router2 (default priority):
interface GigabitEthernet0/0/0
ip address 10.50.0.3 255.255.255.0
standby 1 ip 10.50.0.1
standby 1 preempt
- Priority range is 1 to 255 and the default is 100, so 110 beats Router2's 100.
preemptmakes a router with a higher priority than the current Active take over. Preempt is not implied by a higher priority alone: it has to be configured.delay minimum 60makes Router1 wait 60 seconds before taking over the Active role, so routing and switching have time to settle after a reboot. The delay range is 0 to 3600 seconds and the default is 0.- Hosts use 10.50.0.1 as default gateway. With HSRP version 1, group 1 uses virtual MAC 0000.0C07.AC01 (the Layer 2 address that goes with the virtual IP), and the new Active router takes over that same MAC, so the ARP entry hosts cache (their IP-to-MAC mapping) stays valid through a failover.
- Router2 also gets
preemptso the pair behaves symmetrically: if Router1 is Active but its priority later falls below 100 (for example through tracking, covered in the pitfalls), Router2 can only take the role from it if Router2 has preempt configured. - Defaults: hello 3 seconds, hold time 10 seconds.
Role transitions
HSRP states are Init, Learn, Listen, Speak, Standby and Active.
| Event | Router1 | Router2 |
|---|---|---|
| Both up, stable | Active | Standby |
| Router1 loses power at t=0 | down | no hellos from Router1 for the 10 s hold time, then Standby to Active (takes 10.50.0.1 and the virtual MAC) |
| Router1 shut down cleanly | sends a Resign message | takes over immediately, without waiting for the hold time |
| Router1 boots | Init, then Listen (hears Router2's hellos), then Speak, then Standby once no Standby hello has been heard for the hold time (the RFC 2281 Standby timer expires) | Active, keeps forwarding |
| Preempt delay (60 s) expires, Router1 has priority 110 against 100 | sends a Coup message and becomes Active | on the Coup (or an Active hello from a higher priority router) moves to Speak and sends a Resign, then settles into Standby |
| Stable again | Active | Standby |
What the messages mean (RFC 2281): a Hello says the router is running and can become Active or Standby, a Coup says it wants to become Active, and a Resign says it no longer wants to be Active. Speak is the state where a router sends hellos and takes part in electing Active and Standby. A router that has not yet learned the virtual IP waits in Learn; here the address is configured, so Router1 goes straight from Init to Listen.
Timing of the failure case: Router2 declares Router1 dead 10 seconds after the last hello it received. Hellos arrive every 3 seconds, and Router1 can fail anywhere in that interval, so takeover happens between 7 and 10 seconds after the failure. If you need faster takeover, shorten hello and hold times (Cisco's command reference allows millisecond hello timers from 15 to 999 ms and recommends hold times under 250 ms only on Cisco 7200-class or better platforms and on Fast Ethernet or faster interfaces).
Verification
show standby brief on both routers. Expected stable output on Router1: group 1, priority 110, state Active, virtual IP 10.50.0.1. show standby gives the long form, including Preemption enabled and the Active virtual MAC address. After failing Router1, Router2's state should read Active. Test by pinging 10.50.0.1 from a host during the event.
Pitfalls
- Forgetting
preempton Router1: when it returns it stays Standby and Router2 keeps the Active role indefinitely, which is the opposite of what was asked. - Preempt with no delay: Router1 can become Active before its uplink routing is ready, black-holing traffic (accepting packets it cannot forward, so they vanish). Tracking an object that follows the uplink (
standby 1 track <object> decrement 20, which lowers the priority by 20 while the tracked object is down, taking 110 to 90, below Router2's 100) makes it give up Active when its uplink fails. - HSRP watches the interface it runs on, not the path beyond it; without tracking, hosts keep sending to an Active router that has lost its uplink.
You run a BGP backbone that uses route reflectors. One RR intermittently 'suppresses' or fails to forward routes for a set of prefixes to certain clients, causing reachability issues. Outline concrete debugging steps with exact data and commands you would collect (Adj-RIB-In, Loc-RIB, BGP update logs, cluster-list checks, CPU/heap stats, and packet captures) and how you'd correlate them to find the root cause.
Sample Answer
Direct answer
Work from the RR's own tables outward, in this order: (1) pin down one concrete missing prefix and one client that lacks it, (2) check whether the RR holds the path and whether that path is its best path, because an RR reflects only its best path, (3) check what the RR actually sends to that client, (4) check the loop-prevention attributes and the cluster configuration, (5) check session-level causes such as prefix limits and hold-timer expiry, and only then (6) turn on update debugging and packet capture for the single prefix. Correlate everything on one timeline. The symptom "RR suppresses routes" is almost always one of: the RR's best path differs from what you expected, a policy filter on the RR-to-client session, a CLUSTER_ID or ORIGINATOR_ID drop, or the session flapping.
Terms: Adj-RIB-In is the per-neighbor table of routes as received; Loc-RIB is the table of routes after best-path selection; ORIGINATOR_ID is the router ID of the route's originator inside the AS; CLUSTER_LIST is the list of RR cluster IDs the route has passed through. Normally an iBGP router (a BGP router inside your own AS) never passes a route it learned from one iBGP peer on to another iBGP peer. A route reflector (RR) is the exception (RFC 4456): a route from a client goes to all other clients and to non-clients, a route from a non-client goes to clients, and the RR reflects only its single best path for each prefix. A peer group is a set of neighbors that share one outbound policy, so a policy mistake on the group hits every member. A community is a tag attached to a route that route-maps can match.
Ordered checks (Cisco IOS XE commands)
| # | Check | Command | What the result tells you |
|---|---|---|---|
| 1 | Does the RR have the prefix, how many paths, which is best, and is it advertised? | show ip bgp 203.0.113.0/24 | Several paths with a different best path than expected means clients receive that best path, whose next hop may be unreachable for them. A line Not advertised to any peer means outbound policy or reflection rules stopped it. |
| 2 | What does the RR send to the affected client? | show ip bgp neighbors <client> advertised-routes | Prefix absent here but present in the RR's table: the problem is on the outbound side of that session (policy, reflection rule). |
| 3 | What did the RR receive before inbound policy? | neighbor <peer> soft-reconfiguration inbound (the router keeps an unmodified copy of everything that neighbor sends, at the cost of extra memory per peer), then show ip bgp neighbors <peer> received-routes | Received but not in the table: inbound policy rejected it. Not received: the problem is upstream of the RR. |
| 4 | Loop-prevention attributes | show ip bgp 203.0.113.0/24 on RR and client: read the Originator: and Cluster list: lines | RFC 4456 says a route whose CLUSTER_LIST contains the receiving RR's own CLUSTER_ID should be ignored, and routers drop it. Two RRs in different clusters configured with the same bgp cluster-id discard each other's routes silently. |
| 5 | Session stability and limits | show ip bgp neighbors <peer> and the log for the maximum-prefix message (on IOS XE its tag is MAXPFXEXCEED; confirm the exact text on your release) | A maximum-prefix limit tears the session down with a Cease notification (the BGP message a router sends just before it closes a session on purpose); with the restart option the session returns on its own, which produces exactly an intermittent pattern. |
| 6 | Update-level evidence | debug ip bgp updates limited by a standard access list that matches only the affected prefix (the argument order and the in/out keywords differ by release, so confirm with debug ip bgp updates ? first) | Shows each UPDATE and withdraw seen for that prefix. |
| 7 | Packet capture | Capture on the RR-client link, filter TCP port 179 | Shows whether the UPDATE left the RR, whether the client ACKed it, and TCP retransmissions or a zero receive window (the receiver telling the sender to stop because its buffer is full) pointing at a slow peer. |
What check 1 and check 4 look like on a router that received a reflected route (illustrative addresses; the line layout follows Cisco's documented show ip bgp output):
BGP routing table entry for 203.0.113.0/24, version 41
Paths: (1 available, best #1, table default)
Not advertised to any peer
Local
10.0.0.2 (metric 20) from 10.0.0.100 (10.0.0.100)
Origin IGP, metric 0, localpref 100, valid, internal, best
Originator: 10.0.0.2, Cluster list: 10.0.0.100
Read it top down. Paths: (1 available, best #1) says the router holds one path and it is the best. Not advertised to any peer is printed where an advertised route would instead list the update groups it was sent to, so here nothing has been sent onward. 10.0.0.2 ... from 10.0.0.100 is the next hop and the neighbor that sent it. Originator: 10.0.0.2 is the router that first injected the route into iBGP, and Cluster list: 10.0.0.100 shows the one cluster it passed through. If the receiving router's own cluster ID appeared in that list, it would ignore the route.
Always scope debugging with the access list. An unscoped debug ip bgp updates on a router holding a million prefixes can overload the control plane, which turns a diagnosis into a second outage.
CPU and memory: check the platform's process-level CPU listing for the BGP process and the router's free memory at the time of each flap. The commands differ by platform, so record the platform's own output with a timestamp. High CPU can delay keepalives and expire a hold timer, which looks like suppression from the client's side.
Correlate
Build a table with one row per observation: timestamp, prefix, client, RR best path, advertised yes/no, session state. Look for the common factor:
- Same set of prefixes missing everywhere: a prefix-list, community match or origin AS in a route-map on the RR.
- Same set of clients missing prefixes: a peer group or session-specific policy, or a duplicated cluster ID on that client's RR.
- Timing matches log lines: the maximum-prefix log message, hold timer expiry or a CPU spike.
- Path differs at RR vs expectation: IGP metric to the next hop. BGP prefers the lowest IGP metric to the next hop before it reaches router ID and cluster list length, so the RR's own IGP view decides.
Worked example
Symptom: client PE7 intermittently lacks 203.0.113.0/24 (TEST-NET-3, one of the address blocks RFC 5737 reserves for examples). PE7 peers only with RR2. The prefix is originated by PE2, a client of RR1 in another cluster, and the same customer route is also learned by PE5, a client of RR2, over an eBGP session that flaps. On RR1, show ip bgp 203.0.113.0/24 shows the path as best and advertised-routes toward RR2 lists the prefix, and the reflected route carries RR1's CLUSTER_ID 10.0.0.100 in its CLUSTER_LIST. During an outage of PE5's eBGP session, RR2 has no path for it. RR2's configuration carries the same bgp cluster-id 10.0.0.100 as RR1 although the design puts the two in different clusters, so RR2 ignores every route RR1 reflects. The pattern looks selective because RR2 has two sources of routes. Routes from its own clients arrive without RR1's cluster ID, so RR2 accepts them. Routes reflected by RR1 carry cluster ID 10.0.0.100, which equals RR2's own, so RR2 drops them. PE7 therefore has the prefix only while PE5's copy exists, and loses it whenever PE5's eBGP session flaps, because the copy reflected by RR1 is always dropped. That is why it looks intermittent. Fix: give RR2 its own cluster ID, then clear ip bgp <neighbor> soft out on RR1 and verify with show ip bgp on RR2 and PE7.
Fix and verification
Apply the fix, refresh with clear ip bgp <neighbor> soft out (or soft in), confirm via advertised-routes on the RR and show ip bgp <prefix> on three clients, then watch for a full cycle of the previous flap interval before closing.
Pitfalls
received-routesneeds soft-reconfiguration inbound on that neighbor; without it the command has nothing stored to show.- A client that sees the RR's best path is not "missing" routes; confirm which path the RR holds before declaring suppression.
You must redistribute between OSPF and EIGRP at one router. How do you do it without creating loops, and how do you keep the path preference you want?
Sample Answer
Direct answer
With one router doing the work (here redistributing between OSPF, Open Shortest Path First, and EIGRP, Enhanced Interior Gateway Routing Protocol), loops are unlikely because a route can only be redistributed if it is in the routing table, and the routing table holds one winner per prefix. Do it deliberately anyway: redistribute in both directions with route-maps that tag by origin and allow-list the prefixes each side may receive (a list of the only prefixes permitted through), give EIGRP explicit seed metrics (the starting metric a redistributed route is given, because EIGRP cannot derive one from OSPF), add the subnets keyword for OSPF, and decide which domain owns each prefix so the protocol with the lower administrative distance (AD, the trust ranking between route sources) does not silently choose for you.
Why one router is safe, and what stops being safe
- Cisco redistribution requires the route to be present in the routing table. If prefix X exists in OSPF (AD 110) and in EIGRP as an internal route (AD 90), only the EIGRP copy is installed, so only that copy can be redistributed into OSPF. The router never sends the other copy back.
- The loop risk starts the moment a second router does the same job. A prefix redistributed at R1 can be learned on the far side by R2 and redistributed back. Tagging now, while there is one router, means adding R2 later is a copy-paste instead of a redesign.
Configuration
route-map OSPF-TO-EIGRP deny 10
match tag 20
route-map OSPF-TO-EIGRP permit 20
set tag 10
!
route-map EIGRP-TO-OSPF deny 10
match tag 10
route-map EIGRP-TO-OSPF permit 20
set tag 20
!
router eigrp 100
default-metric 10000 100 255 1 1500
redistribute ospf 1 route-map OSPF-TO-EIGRP
router ospf 1
redistribute eigrp 100 subnets route-map EIGRP-TO-OSPF
default-metric 10000 100 255 1 1500supplies the five EIGRP metric components for redistributed routes, in this order: bandwidth in kilobits per second (10000 is a 10 Mbps link), delay in tens of microseconds (100 is 1 ms), reliability on a 0 to 255 scale (255 means 100 percent reliable), load on a 1 to 255 scale (1 means minimally loaded) and MTU in bytes (1500, the standard Ethernet size). Without a metric, EIGRP does not know what to assign.- OSPF does not need a metric: it assigns 20 by default and marks the routes type 2 external (E2: the metric stays at the value assigned at redistribution and does not grow as the route crosses the OSPF domain).
subnetsis mandatory in practice; without it, only major classful networks (networks at their old class A, B or C boundaries, such as 10.0.0.0/8) are redistributed. - Tag 10 means "originated in OSPF" and tag 20 "originated in EIGRP". Each map refuses the tag of the domain it feeds, which is the guard that matters once a second border router exists.
Trace the prefix 10.50.0.0/24 (illustrative) through the maps, first as a prefix that lives natively in OSPF. OSPF-TO-EIGRP checks it: clause 10 denies only tag 20, the route has no tag, so clause 20 permits it and sets tag 10, and it enters EIGRP as an external route carrying tag 10 and the default-metric values. If a second border router later tried to feed it back with EIGRP-TO-OSPF, clause 10 would match tag 10 and deny it, so the route cannot return to OSPF. A prefix native to EIGRP takes the mirror path: EIGRP-TO-OSPF gives it tag 20 and it enters OSPF as type 2 external with metric 20, and OSPF-TO-EIGRP denies tag 20 on the way back.
Keeping the path preference you want
- Know the AD table. Connected 0, static 1, EIGRP summary 5, eBGP 20, EIGRP internal 90, OSPF 110, IS-IS 115, RIP 120, EIGRP external 170, iBGP 200. Routes from redistributed sources arrive with the destination protocol's external AD (EIGRP external is 170, so a native OSPF route at 110 beats a copy that came back through EIGRP).
- Do not compare metrics across protocols. An OSPF cost and an EIGRP composite metric are different units. AD is what ranks them. If you want a different winner for a prefix, change which domain announces it, not the metric. For example, if 10.50.0.0/24 sits natively in both OSPF and EIGRP on this router, the EIGRP internal copy (AD 90) is installed whatever its metric, even when the OSPF copy has cost 1.
- Allow-list prefixes per direction. During a migration where 10.50.0.0/24 exists in both protocols, permit it in one direction only (
match ip address prefix-list) so there is exactly one owner. - Set the seed metric when you have two exits. In a later two-router design, use
set metricin the route-map on the preferred border so that router's copy is better inside the destination domain. For example, addset metric 20under clause 20 ofEIGRP-TO-OSPFon border router R1 andset metric 100on R2: OSPF routers prefer the lower type 2 metric, so they send traffic for the prefix toward R1 and fall back to R2 if R1 stops advertising it.
Verify
show ip route 10.50.0.0: source and tag (a route from EIGRP inside OSPF shows type 2 external and tag 20).show ip eigrp topologyandshow ip ospf database: the redistributed prefix appears once.show ip protocols: confirms both redistributions and the map names.- Ping and traceroute across the boundary in both directions.
Pitfalls
- Forgetting
default-metricleaves OSPF routes absent from EIGRP. - Omitting
subnetsloses most of your routes. - Writing a route-map with only a deny clause blocks everything.
- A flat redistribution with no filters is the version that creates a loop later when a second border router is added.
Why do iBGP networks use route reflectors and how do they work? Cover client versus non-client behavior, loop prevention and the problems they can introduce.
Sample Answer
Direct answer
iBGP routers do not re-advertise routes learned from one iBGP peer to another iBGP peer (RFC 4271 section 9.2), so without help every router needs a session with every other. A route reflector (RR) is allowed to break that rule, so routers connect only to the reflector. Clients send their routes to the RR, which reflects them to other clients and non-clients. It scales but adds path-hiding, suboptimal-routing and loop risks you must design for.
Why they exist (computed)
A full mesh of n routers needs n(n-1)/2 sessions. With reflectors:
| Routers | Full mesh sessions | 2 RRs + clients (each client peers with both RRs, plus 1 RR-to-RR session) |
|---|---|---|
| 20 | 190 | 18 x 2 + 1 = 37 |
| 100 | 4,950 | 98 x 2 + 1 = 197 |
How reflection works (RFC 4456)
- A client is an iBGP peer that you tell the RR to reflect routes to and from (on Cisco, with
neighbor <ip> route-reflector-client). It needs a session only with its RR, not with other clients. A non-client is an ordinary iBGP peer of the RR that follows the normal iBGP rules. An RR together with its clients forms a cluster, and each cluster has a CLUSTER_ID (a 4-byte identifier). - A route from a non-client is reflected to all clients. A route from a client is reflected to all non-clients and also to the other clients. The non-client peers themselves must be fully meshed, but clients need not be.
- An RR advertises only its best path per prefix, not all paths.
- It should not modify NEXT_HOP (the address to forward traffic to), AS_PATH (the list of ASNs the route crossed), LOCAL_PREF (the preference shared inside the AS) or MED (a hint from a neighbouring AS about its preferred entry point), so clients still need IGP reachability to the real next hop.
Loop prevention
Two attributes replace the "no iBGP to iBGP" rule:
- ORIGINATOR_ID carries the BGP identifier of the route's originator in the AS. A router that sees its own ID ignores the route.
- CLUSTER_LIST records the reflection path. Each RR prepends its CLUSTER_ID, and an RR that finds its own CLUSTER_ID in the list ignores the route. Route selection also prefers the shorter CLUSTER_LIST.
Redundant RRs in one cluster share a CLUSTER_ID so that an RR discards routes from the other RRs in the same cluster.
Problems and mitigations
The first two rows follow directly from reflecting only a best path, so they come with every RR design. The others come from specific design choices.
| Problem | Why it happens | Mitigation |
|---|---|---|
| Path hiding (visibility) | the RR reflects only its best path, so clients see one path per prefix and cannot load-balance or fail over quickly | Cisco best-external (advertise the best external path even when an internal one is best) or BGP Additional Paths (RFC 7911: several paths per prefix, each with a Path Identifier) |
| Suboptimal routing | the RR picks the best path from its own IGP view, which may differ from the client's | place RRs in the topology and keep the reflection topology congruent with the physical one (RFC 4456 recommendation) |
| Single point of failure | one RR | at least two RRs per cluster, same CLUSTER_ID |
| Loops and oscillations (routers keep switching their best path and never settle) | bad RR topology, inconsistent clients | follow a hierarchy (clients, then RRs, then a meshed top level) and keep IGP costs consistent |
| Next hop unreachable | RR does not change next hop | carry next hops in the IGP |
Worked example
Names are illustrative. A 20-router network has two RRs, RR1 and RR2, sharing one CLUSTER_ID, and 18 clients including C1, C2 and C3. Every client peers with both RRs, and RR1 and RR2 peer with each other: 18 x 2 + 1 = 37 sessions instead of 190.
Client route: C1 learns an external route and sends it to RR1 and RR2. RR1 selects its best path and, because the route came from a client, reflects it to all other clients (C2, C3 and the rest) and to its non-client peers. It adds ORIGINATOR_ID = C1's ID and CLUSTER_LIST = its CLUSTER_ID. C2 can now use the route without any session to C1. The copy RR1 sends to RR2 carries the shared CLUSTER_ID, so RR2 finds its own ID in the list and discards it, and still has C1's own advertisement. If a copy ever reached C1, C1 would ignore it because ORIGINATOR_ID is its own.
Non-client route: now suppose a 21st router N1, outside the 20-router example above, is a non-client peer of RR1 and advertises a route to it. RR1 reflects it only to its clients (C1 to C18), not to other non-clients (RR2 included). N1 must therefore be fully meshed with the other non-clients itself, which is why the non-client side of a design stays small.
Pitfalls
- If the client-to-RR design differs from the physical topology, forwarding loops are possible.
- Clients do not see all paths, so a failure needs the RR to re-run best-path selection before the client recovers.
Unlock Full Question Bank
Get access to all Routing Protocols and Configuration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.