CCNP ENARSI 300-410 · 0 · Start here

Structured troubleshooting and the ENARSI toolkit

How professionals troubleshoot: structured methods, the packet-forwarding and control-plane checklists, the show/debug toolkit, documentation and baselines, and how the ENARSI exam tests you.

68 min read12 chapters0 labs15 quiz8 scenarios15 interview Q&A
Jump to chapter (12)
01

What ENARSI is, how it is weighted, and how this track works

What you will learn

  • What the CCNP ENARSI 300-410 v1.1 exam is, where it sits in the CCNP Enterprise certification, and what its four domains weigh.
  • Why ENARSI is a troubleshooting exam first and a configuration exam second, and what that means for how you study.
  • A structured troubleshooting process and six methods, and when to pick each one.
  • Two checklists you will use in every later module: the packet-forwarding checklist and the routing-protocol checklist.
  • The IOS toolkit: filtered show commands, extended ping and traceroute, safe conditional debugging, logging, SNMP, syslog, NetFlow, IP SLA and Embedded Packet Capture.
  • How to read a routing table like an ENARSI engineer, the time-saving commands, and an exam strategy for troubleshooting tickets.

Prerequisites

You should be comfortable with CCNA-level IPv4 and IPv6 addressing and subnetting, basic static routing, and the idea of OSPF, EIGRP and BGP neighbours. If you have done the ENCOR "Start here" module you already know the life of a packet inside a router (RIB, FIB, adjacency table) and a first troubleshooting method. This module goes deeper into the method, because in ENARSI the method is the skill.

The big picture: the network doctor

Think of two doctors. The first one hears "my stomach hurts" and immediately prescribes a pill. Sometimes the pill works. The second one asks when it started, what you ate, checks your temperature, orders one blood test, and only then treats the cause. When the patient is complicated, only the second doctor succeeds reliably. ENARSI trains you to be the second doctor for routed networks. The patient is a network with OSPF, EIGRP, BGP, redistribution, VPNs and security features all interacting, and the symptom is usually a one-line complaint: "branch users cannot reach the payroll server".

What ENARSI is

ENARSI stands for Implementing Cisco Enterprise Advanced Routing and Services. The exam code is 300-410, the current blueprint is version 1.1, and the exam lasts 90 minutes. It is a concentration exam: pass it together with the core exam ENCOR 350-401 and you earn CCNP Enterprise. Pass it on its own and you earn the Cisco Certified Specialist - Enterprise Advanced Infrastructure Implementation certification. ENCOR gives breadth across the enterprise; ENARSI gives depth in routing and the services around it.

ENARSI 300-410 v1.1 domain weights 1.0 Layer 3 Technologies35% 2.0 VPN Technologies20% 3.0 Infrastructure Security20% 4.0 Infrastructure Services25%

Layer 3 is the biggest domain, but services and security together are 45% of the exam.

The four domains

DomainWeightWhat it covers
1.0 Layer 3 Technologies35%Administrative distance, route maps, loop prevention (filtering, tagging, split horizon, route poisoning), redistribution, summarisation, PBR, VRF-lite, BFD, EIGRP, OSPFv2 and OSPFv3, internal and external BGP
2.0 VPN Technologies20%MPLS operations (LSR, LDP, label switching, LSP), MPLS Layer 3 VPN, DMVPN (GRE/mGRE, NHRP, IPsec, dynamic neighbours, spoke-to-spoke)
3.0 Infrastructure Security20%IOS AAA with TACACS+ and RADIUS, IPv4 and IPv6 ACLs, uRPF, control plane policing, IPv6 first-hop security (RA guard, DHCP guard, binding table, ND inspection, source guard)
4.0 Infrastructure Services25%Device management access, SNMP v2c and v3, logging and conditional debugs, IPv4 and IPv6 DHCP, IP SLA and tracking, NetFlow, Catalyst Center (formerly DNA Center) assurance

A troubleshooting-heavy exam

Read the blueprint carefully and you will notice the verb that appears again and again: troubleshoot. Most items say "Troubleshoot redistribution", "Troubleshoot DMVPN", "Troubleshoot SNMP". Cisco is telling you that knowing how to type the configuration is not enough. You must look at output from a broken network and decide what is wrong. Expect multiple-choice questions built around show output, drag-and-drop items that order a process or match commands to results, and simulation-style items (simlets and testlets) where you run show commands on a topology and answer several questions about it. You cannot go back to a previous question on a Cisco exam, so each item must be closed before you move on.

Worked example. Candidates commonly report roughly 55 to 65 items in 90 minutes. Take 60 items: the weights suggest about 21 Layer 3 items (35%), 12 VPN, 12 security and 15 services items. At 90 minutes that is an average of 1.5 minutes each, but a simlet can easily take 8 to 10 minutes. That means the quick multiple-choice items must be answered in under a minute to bank time. Speed comes from a practised method, not from reading faster.

How this track is organised

This module is the "start here" module. It teaches the method and toolkit once, so every later module can simply say "run the routing-protocol checklist". Then the modules follow the blueprint: Layer 3 (redistribution and route maps, EIGRP, OSPF stub/NSSA/virtual links/authentication, BGP, PBR/VRF-lite/BFD), VPN (MPLS, DMVPN), security (AAA, ACLs, uRPF, CoPP, IPv6 first-hop security) and services (syslog, debug, SNMP, DHCP, IP SLA, NetFlow, OSPFv3 details). Each has configuration labs and ticket labs.

Throughout this module we use one small reference topology, shown in the diagram below, so every example uses the same names and addresses.

PC110.1.10.10 R1Lo 10.255.0.1 R2Lo 10.255.0.2 R3Lo 10.255.0.3 SRV110.3.30.10 10.0.12.0/30OSPF area 0 10.0.23.0/30EIGRP AS 100 R2 redistributes OSPF and EIGRP both ways

Reference topology: PC1 in 10.1.10.0/24 behind R1, SRV1 in 10.3.30.0/24 behind R3, R2 joins two routing domains.

Common mistake. Studying ENARSI only by typing configurations from a guide. Labs that always work teach you nothing about failure. For every lab you complete, break it on purpose (wrong area, missing metric, an ACL in the wrong direction) and read the symptoms.

Exam trap. The domain names matter for your study plan: DHCP, SNMP, NetFlow and IP SLA live in Infrastructure Services (25%), not security. AAA, uRPF and CoPP are Infrastructure Security (20%). Candidates who only study routing protocols leave almost half the marks on the table.

The engineer who knew every command

A network engineer with years of configuration experience failed ENARSI on his first attempt. His score report was strongest in configuration-style items and weakest in the troubleshooting simlets: he ran out of time on the second simlet because he checked every device top to bottom instead of isolating the fault. On his second attempt he practised a fixed checklist (define the path, check the routing table on each hop, check neighbours, then filters) on broken labs every evening and passed with time to spare.

Lesson: ENARSI rewards a repeatable method under time pressure far more than memorised configuration.

"How do you usually approach a routing problem you have never seen before?"

A strong candidate describes a method, not a command: define the problem and scope, identify the traffic path, check the routing table and forwarding decision hop by hop, then the control plane (neighbours, what is advertised and received, filters, redistribution), then data-plane policy (ACLs, NAT, PBR). Mention that you compare against a baseline and that you document the root cause.

Key takeaways

  • ENARSI is 300-410 v1.1, 90 minutes, the CCNP Enterprise concentration exam in advanced routing and services.
  • Weights: Layer 3 35%, VPN 20%, Infrastructure Security 20%, Infrastructure Services 25%.
  • Most blueprint items use the verb "troubleshoot"; expect show-output questions, drag-and-drop and simlets.
  • This module teaches the method and toolkit once; later modules apply it to each technology.
  • The reference topology PC1-R1-R2-R3-SRV1 with OSPF and EIGRP is used in every example here.
02

Troubleshooting as a process

A car mechanic hears "it makes a noise". An impatient mechanic replaces the part that was faulty on the last car he fixed. A good mechanic asks when the noise happens (only when braking? only when cold?), drives the car, listens, lifts it, and checks one system at a time. If the brake pads are fine, he crosses brakes off the list and moves on. Troubleshooting a network works the same way. This chapter turns that habit into a written process you can follow under pressure.

Shoot-from-the-hip versus structured

There are two broad styles of troubleshooting.

Shoot-from-the-hip
You jump straight to the likely cause based on experience: "branch cannot reach the server, it must be the ACL again". When the hunch is right it is very fast. When it is wrong you waste time, you may change things that were fine, and you learn nothing about why the network failed.
Structured troubleshooting
You follow defined steps: define the problem, collect facts, and narrow the causes systematically until one remains. It is slower on easy problems but it finishes on hard problems, and anyone else can follow and check your reasoning.

Experienced engineers blend the two. They allow themselves one quick, read-only check of their hunch, time-boxed to a few minutes. If the hunch is not confirmed, they switch to the structured process at once instead of chasing a second and third hunch.

The generic troubleshooting process

1 Define problem 2 Gather info 3 Analyse 4 Eliminate 5 Propose hypothesis 6 Test hypothesis 7 Solve 8 Document hypothesis wrong: roll back, gather more, try again Steps 1 and 8 are the ones people skip, and the ones that matter most later

The generic process. A disproved hypothesis loops back to gathering more information, never to random changes.

  1. Define the problem. A user report is a symptom, often with a wrong diagnosis attached ("the internet is down"). Turn it into a precise statement: who is affected, what exactly fails, from where to where, since when, always or intermittently, and what changed.
  2. Gather information. Collect show output, logs, monitoring graphs, the change record and the topology. Make sure you have access to every device in the path before you need it.
  3. Analyse the information. Compare what you see with how the network should work: your knowledge of the protocols and your baseline (normal state recorded when all was well).
  4. Eliminate potential causes. Cross off what the facts prove is working. If PC1 can ping SRV1 but HTTPS fails, Layer 1 to Layer 3 along the path are proven good for that flow.
  5. Propose a hypothesis. From what remains, pick the most likely cause, and if two are equally likely, the one that is cheaper to test.
  6. Test the hypothesis. Prefer a read-only test (a show command, a packet capture). If the test is a change, make one change at a time and know how to roll it back. If the hypothesis is disproved, roll back and return to gathering.
  7. Solve the problem. Apply the fix, then verify from the user's point of view, not only the router's, and check you did not break anything else.
  8. Document. Record symptom, root cause, fix, verification and how to prevent it. Update the baseline and diagrams if they changed.

Worked example: defining a problem. The ticket says: "Branch cannot reach the intranet." After five questions it becomes: "Since 09:40 today, all users in 10.1.10.0/24 (behind R1) cannot open HTTPS to SRV1 10.3.30.10. Ping from PC1 to SRV1 works. Users in other branches are not affected. Change CHG-2231 on R3 was implemented at 09:35." That single sentence already eliminates the physical path, the routing between PC1 and SRV1, and every other branch. The obvious place to look is the change on R3, and the likely suspect is something that treats TCP differently from ICMP, such as an ACL.

Gathering evidence of change

"What changed" is the most powerful question in troubleshooting, so make IOS answer it for you. The configuration change logger records every configuration command with the user and line:

! Log every configuration command, send it to syslog, hide passwords
R3(config)# archive
R3(config-archive)# log config
R3(config-archive-log-cfg)# logging enable
R3(config-archive-log-cfg)# notify syslog contenttype plaintext
R3(config-archive-log-cfg)# hidekeys
R3# show archive log config all
 idx   sess           user@line      Logged command
    1     1        netops@vty0     |ip access-list extended SRV-PROTECT
    2     1        netops@vty0     | permit icmp any any
    3     1        netops@vty0     | permit tcp any any eq 22
    4     1        netops@vty0     | deny ip any any log
    5     1        netops@vty0     |interface GigabitEthernet0/2
    6     1        netops@vty0     | ip access-group SRV-PROTECT out

Here the evidence shows that at CHG-2231 an outbound ACL was applied toward SRV1 that allows only ICMP and SSH. HTTPS (TCP 443) is denied. The hypothesis is confirmed with a read-only command before any change is made.

Common mistake. Changing several things at once "to save time". If the problem goes away you do not know which change fixed it, and one of the other changes may have created a new problem that appears next week. One change, one test, and roll back if the test fails.

Exam trap. Drag-and-drop items ask for the order of the steps. "Eliminate potential causes" comes before "propose a hypothesis", and "document" is the last step, after the solution is verified. Also remember that the process loops back to gathering information when a hypothesis fails.

The reload that erased the evidence

A distribution router started dropping OSPF adjacencies every few minutes. The on-call engineer suspected a software bug and reloaded the router. The flapping stopped for two hours and then returned. Because the router had only a small console log buffer and no syslog server, the logs from before the reload were gone. The second time, a senior engineer first captured show logging, show processes cpu sorted and show interfaces, which showed input errors and a duplex mismatch on one uplink after a patch-panel move. Fixing the duplex setting solved it for good.

Lesson: gather information before you act, especially before a reload, which destroys evidence.

"What is the difference between shoot-from-the-hip and structured troubleshooting, and which do you use?"

Explain both honestly: shoot-from-the-hip uses experience to jump to a likely cause and is fast when right; structured follows define, gather, analyse, eliminate, hypothesise, test, solve, document and succeeds on unfamiliar problems. Say that you allow a short, read-only check of your hunch and then fall back to the structured process, and that you always document root cause.

Key takeaways

  • Shoot-from-the-hip is fast when right but unreliable; structured troubleshooting finishes on hard problems.
  • The eight steps: define, gather, analyse, eliminate, propose hypothesis, test, solve, document.
  • A precise problem statement (who, what, where, since when, what changed) eliminates most causes before you touch a device.
  • Test hypotheses read-only where possible; otherwise one change at a time with a rollback plan.
  • Use the configuration change logger (archive / log config) so "what changed" has an answer.
03

Structured approaches: choosing where to start

Imagine a water pipe from a tank on the roof to a tap in the kitchen, and no water comes out of the tap. You could start at the tank and walk down (is there water, is the valve open?), start at the tap and walk up, or cut into the pipe halfway and see if water is flowing there. If a neighbour has the same plumbing and their tap works, you could compare the two houses. Or you could simply replace the tap. Each of those is a real troubleshooting approach, and each is best in a different situation. Step 4 of the process in chapter 2 said "eliminate potential causes". The approaches in this chapter are the ways you do that efficiently.

7 Application 6 Presentation 5 Session 4 Transport 3 Network 2 Data link 1 Physical top-down bottom-up start here: ping Divide and conquer starts at Layer 3 and moves up or down

Three layer-based approaches: from the top, from the bottom, or from the middle.

1. Top-down

Start at the application layer and work down. Does the application respond? Is the port reachable? Is there IP reachability? Is the link up? Use top-down when the complaint is specific to one application and other traffic is fine, for example "SSH to SRV1 works but HTTPS does not". You will usually find the fault at Layer 4 to 7: an ACL, a firewall rule, a NAT translation or a service that is not listening.

2. Bottom-up

Start with cables, optics and interface counters, then Layer 2 (VLANs, trunks, STP), then Layer 3. Use bottom-up when the symptoms smell physical: interfaces flapping, CRC and input errors, a new installation or a recent patch-panel move. It is thorough but slow when the fault is high up the stack, because you check a lot of healthy things first.

3. Divide and conquer

Start in the middle, usually with a Layer 3 test such as a ping between the two endpoints. If the ping works, Layers 1 to 3 on that path are healthy and you move up. If it fails, you move down. Divide and conquer is the default choice when you have no strong clue, because one test halves the search space.

4. Follow the traffic path

Start at the source and follow the packet hop by hop toward the destination, checking the forwarding decision on each device (and then the return path back). Routing problems suit this approach perfectly: on each router you ask "what does this router do with a packet to SRV1?" Traceroute is the quick version of it:

R1# traceroute 10.3.30.10 source 10.1.10.1 numeric
Type escape sequence to abort.
Tracing the route to 10.3.30.10
VRF info: (vrf in name/id, vrf out name/id)
  1 10.0.12.2 1 msec 1 msec 1 msec
  2 10.0.23.2 2 msec 1 msec 2 msec
  3  *  *  *

The trace reaches R3 (10.0.23.2) and stops. Either R3 cannot deliver to SRV1, something blocks the traffic after R3, or the replies cannot find their way back. Traceroute only tells you where replies stopped arriving, so the next step is to log in to R3 and check both directions.

5. Compare configurations

Compare the broken device with a working one (another branch router built from the same template), or compare today's configuration with the last known good configuration. It works even when you do not fully understand the technology, which also makes it risky: you may copy a difference that is not the cause. IOS can do the comparison for you:

R3# show archive config differences flash:R3-baseline.cfg system:running-config
!Contextual Config Diffs:
+ip access-list extended SRV-PROTECT
 +permit icmp any any
 +permit tcp any any eq 22
 +deny ip any any log
interface GigabitEthernet0/2
 +ip access-group SRV-PROTECT out

6. Component swapping

Replace a suspected part with a known-good one: a cable, an optic, a line card, a whole router. If the problem moves with the part, you found it. It is useful for hardware faults and when remote hands can do it quickly, but it proves little about configuration faults, it needs spares, and you may introduce a new variable (different software, different config).

ApproachBest whenWeakness
Top-downOne application fails, others workSlow for physical faults
Bottom-upPhysical symptoms, new installSlow for application faults
Divide and conquerNo strong clue; general defaultNeeds an experienced choice of first test
Follow the pathRouting and reachability problemsNeeds access to every hop
Compare configurationsWorking twin or baseline existsCopies differences that are not the cause
Component swappingSuspected hardwareNeeds spares; weak for config faults

Worked example: choosing an approach. Three tickets arrive. Ticket A: "R1 Gi0/0 has 4,812 CRC errors and OSPF flaps." Physical symptom, so bottom-up: check the cable, optic and duplex first. Ticket B: "PC1 cannot reach SRV1 at all, since this morning." No clue about the layer, so divide and conquer with a ping, then follow the path with traceroute and show ip route 10.3.30.10 on each hop. Ticket C: "HTTPS to SRV1 fails, ping works." Layer 3 is proven, so top-down from the application: check ACLs and the service on SRV1.

Common mistake. Following the path only in the forward direction. A large share of routing faults are on the return path: SRV1's replies go to R3, and R3 has no route back to 10.1.10.0/24. Always ask "can the reply come back?"

Exam trap. If a question says the user can ping but not browse, the answer is almost never a routing protocol problem. If a question describes CRC errors or late collisions, the correct first approach is bottom-up, not divide and conquer.

The template that did not match

A new branch router could not bring up its EIGRP neighbour to the hub, while 40 other branches built from the same template worked. The engineer used compare configurations: a diff of the new router against a working branch showed router eigrp 110 instead of router eigrp 100, a typo during staging. Correcting the AS number brought the neighbour up in seconds.

Lesson: when you have a working twin, a configuration diff is often the fastest approach of all, even before reading protocol output.

"Name some troubleshooting approaches and tell me when you would use each."

List top-down, bottom-up, divide and conquer, follow the traffic path, compare configurations and component swapping. Give one example each: top-down for "one app fails", bottom-up for CRC errors, divide and conquer with a ping when you have no clue, follow the path for routing, compare configs against a working twin, swap components for suspected hardware. Mention checking the return path.

Key takeaways

  • Six approaches: top-down, bottom-up, divide and conquer, follow the traffic path, compare configurations, component swapping.
  • Pick the approach from the symptom: application-only failure means top-down, physical errors mean bottom-up, no clue means divide and conquer.
  • Follow the path is the natural approach for routing faults; always check the return path too.
  • show archive config differences compares a saved good configuration with the running one.
  • Component swapping proves hardware faults but says little about configuration faults.
04

The packet-forwarding checklist

A courier carrying a parcel across the city does the same few things at every depot. Can I physically get to the next depot (is the road open)? Do I know exactly which door to knock on (the local address)? Which depot is next on the map? Is there a pre-printed route card that already says so? Will a checkpoint stop me, send me on a detour or relabel the parcel? And when the receiver sends a reply, can that reply find its way back? A router forwarding a packet asks exactly these questions. This chapter turns them into a checklist you run on every hop of the path.

1 Layer 2 reachability 2 ARP / ND 3 Routing table (RIB) 4 CEF: FIB + adjacency 5 ACL / NAT / PBR 6 Return path Run it on every hop: R1, then R2, then R3 then again from SRV1 back toward 10.1.10.0/24 Stop at the first check that fails

Six questions per hop, in the order a router actually uses the information.

1. Layer 2 reachability

Before a router can route anything it needs working interfaces. show ip interface brief must show the ingress and egress interfaces as up/up. "administratively down" means someone typed shutdown; "down/down" is usually physical; "up/down" often means a keepalive, encapsulation or Layer 2 problem. Check show interfaces for input errors, CRC and drops, and on switches the VLAN and trunk state.

2. ARP (or IPv6 neighbour discovery)

To send a frame to the next hop or to the final host on a connected subnet, the router needs the MAC address. ARP resolves IPv4 addresses to MACs; ND does the same for IPv6.

R3# show ip arp 10.3.30.10
Protocol  Address          Age (min)  Hardware Addr   Type   Interface
Internet  10.3.30.10              0   Incomplete      ARPA

"Incomplete" means R3 sent ARP requests and got no reply. The host is off, in the wrong VLAN, has the wrong subnet, or something between R3 and SRV1 is dropping the frames. No routing command will fix that; the fault is on the local segment.

3. The routing table

The routing table (the RIB, routing information base) holds the best route for each prefix, chosen from all routing sources. The fastest check is to ask the router about the exact destination:

R1# show ip route 10.3.30.10
Routing entry for 10.3.30.0/24
  Known via "ospf 1", distance 110, metric 20, type extern 2, forward metric 1
  Last update from 10.0.12.2 on GigabitEthernet0/0, 00:12:41 ago
  Routing Descriptor Blocks:
  * 10.0.12.2, from 10.255.0.2, 00:12:41 ago, via GigabitEthernet0/0
      Route metric is 20, traffic share count is 1

If the answer is % Subnet not in table or % Network not in table, there is no specific route for it (a default route is not shown by this command, so check for one separately with show ip route 0.0.0.0). A missing route sends you to the routing-protocol checklist in chapter 5.

4. CEF: the FIB and the adjacency table

Routers do not forward by reading the RIB. CEF (Cisco Express Forwarding) builds the FIB (forwarding information base) from the RIB, with recursion already resolved, and an adjacency table holding the pre-built Layer 2 header for each next hop. The data plane uses these. Normally they agree with the RIB, but checking them proves what the router will really do, including which link it picks when there are equal-cost paths.

R1# show ip cef 10.3.30.10
10.3.30.0/24
  nexthop 10.0.12.2 GigabitEthernet0/0
R1# show ip cef exact-route 10.1.10.10 10.3.30.10
10.1.10.10 -> 10.3.30.10 =>IP adj out of GigabitEthernet0/0, addr 10.0.12.2

A FIB entry pointing to drop, Null0 or an incomplete (glean) adjacency explains a failure even when the routing table looks fine.

5. Policy: ACLs, NAT and PBR

A router with a perfect route can still refuse to forward. Check three kinds of policy on the ingress and egress interfaces:

  • ACLs: show ip interface GigabitEthernet0/2 | include access list shows what is applied and in which direction; show access-lists shows the hit counters.
  • NAT: show ip nat translations and show ip nat statistics. A wrong inside/outside interface or a missing translation breaks flows silently.
  • PBR (policy-based routing): show ip policy lists interfaces with a route map. PBR is checked before the routing table, so it can send traffic somewhere the RIB never would.
R1# show ip policy
Interface      Route map
Gi0/1          PBR-WEB

6. The return path

Every successful conversation needs two working paths. Repeat checks 1 to 5 from the destination back toward the source: does R3 have a route to 10.1.10.0/24? Does R2? Is there an ACL on the return interfaces? Asymmetric paths are fine for routers, but stateful devices (firewalls, zone-based firewall, NAT) along the way may drop the returning half.

Worked example. From R2, ping 10.3.30.10 succeeds. From PC1, it fails. Why the difference? R2 sources its ping from 10.0.23.1, a subnet that is connected on R3, so SRV1's replies always find their way back. PC1's packets come from 10.1.10.10. On R3, show ip route 10.1.10.10 returns % Subnet not in table. The forward path works; the return path is missing. The fix belongs in the routing protocol on R3 (in this topology, redistribution of OSPF into EIGRP on R2). Testing with ping 10.3.30.10 source 10.1.10.1 from R1 would have exposed it immediately.

Common mistake. Testing with a plain ping from the router instead of an extended ping sourced from the user's subnet. The router uses its egress interface address as the source, which is often reachable from everywhere, and hides return-path problems.

Exam trap. PBR is evaluated before the routing table for packets arriving on the interface where it is applied, but not for traffic the router generates itself unless ip local policy route-map is configured. A question showing a correct RIB but traffic taking another path usually wants PBR as the answer.

The route that existed, the traffic that did not flow

After a WAN migration, one branch could not reach the data centre, although show ip route on the branch router showed the correct route through the new MPLS link. The engineer ran show ip cef exact-route and saw traffic still leaving the old DSL interface. show ip policy revealed a forgotten PBR route map on the LAN interface that set the old DSL next hop for all traffic to the data centre. Removing the stale route map fixed it.

Lesson: the RIB shows what the router believes; CEF and policy show what it actually does.

"A route is in the routing table, but traffic still fails. What do you check next?"

Walk the rest of the checklist: CEF (show ip cef, exact-route, adjacency and ARP for the next hop), ACLs on ingress and egress, NAT, and PBR, which overrides the RIB. Then the return path from the destination, including stateful devices. Mention using an extended ping with the user's source address.

Key takeaways

  • Per hop: Layer 2, ARP/ND, routing table, CEF (FIB and adjacency), ACL/NAT/PBR, then the return path.
  • show ip route x.x.x.x answers "which route will be used for this address" directly.
  • show ip cef exact-route src dst shows the real forwarding decision, including ECMP choices.
  • PBR is applied before the routing table; always check show ip policy.
  • Test with an extended ping sourced from the user subnet to expose return-path faults.
05

The routing-protocol checklist

Think of how news travels between offices in a company. First, two offices must be on speaking terms. Then the sender must actually mention the news. The receiver must hear it and not throw it away as irrelevant. If the receiver hears two versions of the same story, it must decide which source to believe. And if the news has to pass between two departments that speak different languages, someone must translate it. Routing protocols work the same way, and when a route is missing, one of those five steps failed. Chapter 4 sent you here when the routing table had no route or the wrong one; this chapter tells you how to find out why.

1 Neighboursup and stable? 2 Advertisedsent at all? 3 Receivedor filtered? 4 Wins in RIBAD, metric 5 Redistributionbetween domains nbr tables LSDB, topology filters, stub show ip route seed metric Where does the prefix disappear? Find the first router on which it is missing, then ask which step failed there

Five questions for any missing or wrong route, whatever the protocol.

1. Are the neighbours up?

No neighbour, no routes. Check the neighbour table of the protocol involved:

R2# show ip ospf neighbor
Neighbor ID     Pri   State           Dead Time   Address         Interface
10.255.0.1        1   FULL/BDR        00:00:34    10.0.12.1       GigabitEthernet0/0
R2# show ip eigrp neighbors
EIGRP-IPv4 Neighbors for AS(100)
H   Address                 Interface              Hold Uptime   SRTT   RTO  Q  Seq
                                                   (sec)         (ms)       Cnt Num
0   10.0.23.2               Gi0/1                    13 01:02:11    1   100  0  27

If a neighbour is missing, check that the interface is enabled for the protocol and not passive (show ip protocols, show ip ospf interface brief, show ip eigrp interfaces), then the parameters that must match. OSPF: area, subnet, hello and dead timers, authentication, network type, MTU (stuck in ExStart/Exchange), stub flags, unique router IDs. EIGRP: AS number, K values, subnet, authentication. BGP: remote-as, reachability of the peer address, update-source, eBGP multihop/TTL, and TCP port 179 allowed through ACLs. Also watch the uptime: a neighbour that keeps restarting is as bad as a missing one.

2. Is the route being advertised?

Go to the router that originates the prefix. Is it in the protocol at all? The interface must be covered by a network statement or interface command, or the prefix must be redistributed. In BGP, a network statement only advertises a prefix that exists in the routing table with the exact same mask. Look at the protocol database: show ip ospf database (the LSA must exist), show ip eigrp topology, show ip bgp and show ip bgp neighbors 10.0.0.1 advertised-routes. Summarisation also changes what is advertised: a summary hides the specific prefixes behind it.

3. Is it received, or filtered?

On the next router, is the route arriving? Filters can remove it on the way out of the sender or on the way into the receiver: distribute lists, prefix lists, route maps, BGP inbound policy, EIGRP stub (a stub router does not re-advertise learned routes by default), OSPF stub areas (no type 5 external LSAs) and split horizon on hub interfaces. For BGP, show ip bgp neighbors 10.0.0.1 routes shows what was accepted after inbound policy; received-routes shows everything received but needs soft-reconfiguration inbound.

One OSPF detail is heavily tested: a distribute-list in under OSPF filters routes from entering the routing table, not LSAs from the database. So the LSA is present in show ip ospf database while the route is missing from show ip route.

4. Does it win in the routing table?

The route may be received and still not installed. The router first compares prefixes with the same length; among those, the lowest administrative distance (AD) wins, and within one protocol the lowest metric. A static route (AD 1) beats OSPF (110) for the same prefix; an external EIGRP route (170) loses to OSPF. In BGP, show ip bgp marks a best path that could not be installed with r for RIB-failure, usually because a lower-AD source already owns that prefix. Chapter 8 goes deep on route selection.

5. Is redistribution doing its job?

When the prefix must cross from one protocol to another, check the redistributing router. Redistribution only takes routes that are in the RIB via the source protocol (plus connected networks covered by that protocol). Each target protocol needs a sensible seed metric: OSPF assigns 20 by default (1 for routes from BGP) as external type 2; classic EIGRP needs an explicit metric for routes from other routing protocols, or nothing is redistributed. Route maps and tags on the redistribution line may filter prefixes. Mutual redistribution at two or more points can cause loops and suboptimal routing, which the redistribution module covers in depth.

R2# show ip protocols | section eigrp
Routing Protocol is "eigrp 100"
  Outgoing update filter list for all interfaces is not set
  Incoming update filter list for all interfaces is not set
  Default networks flagged in outgoing updates
  Default networks accepted from incoming updates
  Redistributing: ospf 1
  EIGRP-IPv4 Protocol for AS(100)
    Metric weight K1=1, K2=0, K3=1, K4=0, K5=0

Worked example. R3 has no route to 10.1.10.0/24. Step 1: the EIGRP neighbour R2 is up. Step 2: the prefix should be advertised by R2 into EIGRP through redistribution. On R2, show ip eigrp topology 10.1.10.0/24 returns %Entry 10.1.10.0/24 not in topology table. So R2 is not advertising it. Step 5: show run | section router eigrp shows redistribute ospf 1 with no metric and no default-metric. Classic EIGRP needs a seed metric for OSPF routes. The fix:

! Give redistributed OSPF routes a seed metric: bandwidth (kbps) delay (tens of usec) reliability load MTU
R2(config)# router eigrp 100
R2(config-router)# redistribute ospf 1 metric 100000 10 255 1 1500

Afterwards R3 shows D EX 10.1.10.0/24 [170/...] via 10.0.23.1. The external marker and AD 170 confirm it came from redistribution.

Common mistake. Looking only at the router where the symptom appears. Walk back toward the originator and find the first router on which the prefix is missing; the cause is on that router or on its upstream neighbour.

Exam trap. "The route is in the OSPF database but not in the routing table" points to a distribute-list in or to a better route from a lower-AD source. "The route is in the BGP table with r" means RIB-failure. "Redistributed into EIGRP but nothing appears" means a missing seed metric.

The stub that stopped the transit

A company added a second WAN router at a branch and configured both branch routers as EIGRP stubs, copied from the single-router template. The hub could no longer reach the LAN behind the second router when its own WAN link failed, because the first router, being a stub, did not re-advertise routes learned from its neighbour. Neighbours were up and nothing was filtered by a list, so the cause hid in step 3 (received, but not propagated). Adding eigrp stub connected summary leak-map for the needed prefixes restored the backup path.

Lesson: "filtered" includes protocol features like stub and stub areas, not only distribute lists.

"A route is missing on a router. Walk me through how you find out why."

Say you find the first router where it is missing, then check: neighbours up and stable, is the originator advertising it (network statement, redistribution, database), is it received or filtered (distribute lists, prefix lists, route maps, stub, area types, split horizon), does it win in the RIB (longest match, AD, metric, BGP RIB-failure), and whether redistribution has a seed metric and correct filters. Name the commands for each step.

Key takeaways

  • Five steps: neighbours, advertised, received or filtered, wins in the RIB, redistribution.
  • Find the first router on which the prefix is missing; the cause is there or on its upstream neighbour.
  • OSPF distribute-list in filters the RIB, not the LSDB.
  • BGP r means RIB-failure; BGP network needs an exact RIB match.
  • Classic EIGRP needs a seed metric for routes redistributed from other routing protocols; OSPF defaults to 20 as E2.
06

The IOS troubleshooting toolkit

A doctor has a stethoscope for right now, the patient file for history, a blood-pressure log for trends, and a scan when she needs to see inside. A network engineer has the same four kinds of tool: live commands (show, ping, debug), history (logging buffers and syslog), trends (SNMP, NetFlow, IP SLA) and a view inside the packets (Embedded Packet Capture). Knowing which one answers your current question is half of troubleshooting. The services module later teaches configuring SNMP, NetFlow and IP SLA in depth; here you learn to use them as evidence.

Router Liveshow, ping, debug Historylog buffer, syslog TrendsSNMP, NetFlow, IP SLA PacketsEPC capture

Four kinds of evidence: what is happening now, what happened, how things trend, and what is inside the packets.

Show commands and output filters

Big outputs hide the one line you need. IOS filters let you cut them down: | include (lines matching a regular expression), | exclude, | begin (start at the first match), | section (a configuration block and its children) and | count (how many lines match). Examples: show ip route | include ^O|^D lists only OSPF and EIGRP routes; show ip interface brief | exclude unassigned hides unused interfaces; show logging | include OSPF|EIGRP|BGP finds adjacency changes.

Extended ping and traceroute

A plain ping from the router tests the wrong thing surprisingly often. The extended options let you imitate the user's traffic:

! Source from the user LAN, 100 packets, 1400 bytes, do not fragment
R1# ping 10.3.30.10 source GigabitEthernet0/1 repeat 100 size 1400 df-bit
! Trace from the user LAN without DNS lookups
R1# traceroute 10.3.30.10 source 10.1.10.1 numeric

Know the result codes. Ping: ! reply, . timeout, U destination unreachable (a router sent an ICMP unreachable, often because of a missing route or an ACL), M could not fragment (MTU problem with DF set), & TTL expired (a routing loop). Traceroute: * timeout, A administratively prohibited (an ACL), H host unreachable, N network unreachable. Loss that appears only with large packets and the DF bit points to an MTU problem, typical on GRE and DMVPN tunnels.

Debugging safely

Debugs are powerful but they run on the CPU. On a busy router an unfiltered debug ip packet can starve the CPU and cut you off. Rules for safe debugging:

  • Send debug output to the buffer, not the console: no logging console (or logging console warnings) and logging buffered 65536 debugging.
  • Limit it: with an ACL (debug ip packet 101 detail), a neighbour (debug ip bgp 198.51.100.1 updates) or a condition (debug condition interface GigabitEthernet0/0).
  • Prefer event debugs such as debug ip ospf adj or debug eigrp packets hello over packet debugs.
  • Stop quickly: undebug all. Some engineers schedule a safety net with reload in 10 on lab-like devices.
! Only packets between PC1 and SRV1
R2(config)# access-list 101 permit ip host 10.1.10.10 host 10.3.30.10
R2(config)# access-list 101 permit ip host 10.3.30.10 host 10.1.10.10
R2# debug ip packet 101 detail
R2# show logging | include IP:
R2# undebug all

Logging buffers and timestamps

Logs are only useful if you can line them up across devices. Use service timestamps log datetime msec localtime show-timezone and the same for debug, and synchronise every device with NTP. Syslog severities are 0 emergencies, 1 alerts, 2 critical, 3 errors, 4 warnings, 5 notifications, 6 informational, 7 debugging. Choosing a level includes everything more severe. Adjacency changes such as %OSPF-5-ADJCHG and %DUAL-5-NBRCHANGE are level 5. On a VTY session you must type terminal monitor to see log messages live.

SNMP, syslog and NetFlow as evidence

Syslog sent to a central server (logging host) survives reloads and lets you correlate events across routers. SNMP polling produces graphs of interface utilisation, errors, CPU and memory, and traps announce events such as link down. NetFlow records flows (source, destination, ports, protocol, bytes), so it answers "who is filling this link?" and "did this flow ever reach the router?" Each is evidence you can gather without touching the live network.

IP SLA

IP SLA makes the router generate synthetic traffic continuously (ICMP echo, UDP jitter, TCP connect, HTTP and more) and record the results. It catches intermittent problems that you would never see with a single ping, and it can drive object tracking for static routes and HSRP.

R1(config)# ip sla 10
R1(config-ip-sla)# icmp-echo 10.3.30.10 source-interface GigabitEthernet0/1
R1(config-ip-sla-echo)# frequency 10
R1(config)# ip sla schedule 10 life forever start-time now
R1# show ip sla statistics 10
IPSLAs Latest Operation Statistics

IPSLA operation id: 10
        Latest RTT: 2 milliseconds
Latest operation start time: 10:15:02 IST Wed Sep 30 2026
Latest operation return code: OK
Number of successes: 356
Number of failures: 4
Operation time to live: Forever

Embedded Packet Capture

When you need proof of what is on the wire, EPC captures packets on the router itself and exports a pcap file for Wireshark. The IOS XE syntax:

R3# monitor capture CAP interface GigabitEthernet0/2 both
R3# monitor capture CAP match ipv4 host 10.1.10.10 host 10.3.30.10
R3# monitor capture CAP buffer size 10
R3# monitor capture CAP start
R3# monitor capture CAP stop
R3# show monitor capture CAP buffer brief
R3# monitor capture CAP export flash:cap1.pcap

Worked example. Users report the payroll application "freezes a few times a day". Pings during the call are 100% fine. The IP SLA on R1 shows 356 successes and 4 failures, and the failure timestamps match the complaints. The syslog server shows %OSPF-5-ADJCHG on R2 Gi0/0 going down and up at the same times, and SNMP graphs show input errors climbing on that interface. Three kinds of evidence agree: an unstable link, not the application.

Common mistake. Running an unfiltered debug ip packet on a production router. Beyond the CPU risk, it shows only process-switched packets; CEF-switched transit traffic does not appear, so an empty debug does not prove that traffic is absent.

Exam trap. logging trap sets the severity sent to syslog servers, logging buffered the local buffer, logging console the console and logging monitor the VTY sessions. Level 7 debugging includes all levels; level 4 warnings includes 0 to 4 only.

The debug that took down the router

An engineer investigating a NAT issue on a busy internet edge router typed debug ip nat from a console session with console logging at debugging level. The router tried to print thousands of lines per second on a 9600-baud console, the CPU spiked, BGP keepalives were missed and both ISP sessions dropped. The fix was a reload from the power switch. The next time the same problem was investigated with an ACL-limited debug, console logging disabled and output in the buffer.

Lesson: filter debugs, send them to the buffer, and never debug at full volume to the console on production.

"How do you debug safely on a production router?"

Disable console logging, log to the buffer with timestamps, limit the debug by ACL, neighbour or condition, prefer event debugs over packet debugs, check CPU first with show processes cpu sorted, and stop with undebug all. Mention that packet debugs only show process-switched traffic and that EPC or NetFlow is often a better choice.

Key takeaways

  • Filters (include, exclude, begin, section, count) make show output usable.
  • Extended ping and traceroute imitate the user: source, size, DF bit, repeat count; know the result codes.
  • Safe debugging: buffer not console, ACL or condition limited, undebug all ready.
  • Timestamps plus NTP plus central syslog make logs from different routers comparable.
  • SNMP, NetFlow and IP SLA show trends and intermittent faults; EPC proves what is on the wire.
07

Baselines, documentation and change management

When you visit a doctor for the first time, she records your normal blood pressure, weight and pulse. Years later, a reading of 150 means something only because she knows that you are normally 120. Without the record, she cannot tell "high for you" from "normal for you". A network is the same. "R2 CPU is 45%" is meaningless unless you know that R2 normally runs at 5%. The baseline and the documentation are the network's medical record, and change management is how you avoid making the patient ill in the first place.

What good documentation contains

Physical topology
Devices, ports, cables, optics, patch panels, racks. Answers "which port on R2 goes to R3?"
Logical topology
Subnets, VLANs, routing domains, areas and AS numbers, redistribution points, VPN overlays, where ACLs and NAT sit. Answers "which path should traffic from 10.1.10.0/24 take?"
Addressing plan
Every subnet, its purpose and its gateway, including loopbacks and point-to-point links.
Configuration backups
The current and previous configurations of every device, ideally stored automatically.
Change log
Who changed what, when and why, with the change ticket number.
Contacts and contracts
Escalation paths, ISP circuit IDs, support contract numbers.
R2 CPU, five-minute average baseline band 3-8% today 45% MonWedtoday

A baseline turns a number into a finding: 45% is abnormal only because the normal band is known.

What to baseline

Record the normal state when the network is healthy, at both quiet and busy times, and again after every significant change:

  • CPU: show processes cpu sorted and show processes cpu history. The figure "5%/1%" means 5% total, of which 1% is interrupt-level work (packets switched or punted at interrupt level). High interrupt CPU points to traffic hitting the CPU; high process CPU points to a process such as a routing protocol or SNMP.
  • Memory: show processes memory sorted and, on IOS XE, show platform resources.
  • Interfaces: utilisation, errors, drops (show interfaces, SNMP graphs).
  • Control plane: number of neighbours per protocol, number of routes per source (show ip route summary), BGP prefixes received per peer.
  • Performance: normal latency and jitter from IP SLA, normal top talkers from NetFlow.
R2# show ip route summary
IP routing table name is default (0x0)
IP routing table maximum-paths is 32
Route Source    Networks    Subnets     Replicates  Overhead    Memory (bytes)
application     0           0           0           0           0
connected       0           6           0           576         1824
static          0           0           0           0           0
ospf 1          0           4           0           384         1216
  Intra-area: 2 Inter-area: 0 External-1: 0 External-2: 2
  NSSA External-1: 0 NSSA External-2: 0
eigrp 100       0           3           0           288         912
internal        2                                               1208
Total           2           13          0           1248        5160

Automatic configuration backups

IOS can keep its own configuration history with the archive feature. Every write memory, and once a day regardless, it saves a numbered copy:

! Keep up to 10 archived configs in flash, on every save and every 24 hours
R2(config)# archive
R2(config-archive)# path flash:R2-cfg
R2(config-archive)# maximum 10
R2(config-archive)# write-memory
R2(config-archive)# time-period 1440
R2# show archive
The maximum archive configurations allowed is 10.
There are currently 2 archive configurations saved.
The next archive file will be named flash:R2-cfg-3
 Archive #  Name
   1        flash:R2-cfg-1
   2        flash:R2-cfg-2 <- Most Recent
   3
   4

Change management

Many outages are self-inflicted. A simple change process prevents most of them:

  1. Request with a reason and scope.
  2. Risk assessment and peer review of the exact commands.
  3. A maintenance window agreed with the business.
  4. An implementation plan, a verification plan (which show commands prove success) and a rollback plan (the exact commands or file to go back).
  5. After the change: verify, update documentation and the baseline, close the ticket.

In the next chapters you will see IOS tools that make the rollback plan fast and safe: configure replace and the rollback timer.

Worked example. The baseline for R2 records 13 subnets: 6 connected, 4 OSPF (2 intra-area, 2 external type 2) and 3 EIGRP. After a change, show ip route summary shows 11 subnets, with EIGRP at 1. Two EIGRP routes disappeared. Without a baseline, 11 routes looks perfectly normal. With the baseline you know within a minute that EIGRP lost two prefixes, and the next step is the routing-protocol checklist for EIGRP on R2 and R3.

Common mistake. Creating documentation once for a project and never updating it. A wrong diagram is worse than none, because engineers trust it. Make "update the diagram and baseline" a mandatory closing step of every change.

Exam trap. The archive path, maximum, write-memory and time-period commands create configuration backups; archive log config records commands typed. They are different features under the same archive mode, and questions sometimes mix them up.

The "normal" 60% CPU

A NOC engineer saw a branch router at 60% CPU and opened a critical incident. Two engineers spent an hour looking for a routing loop. The baseline, found later on a shared drive, showed that this old router always ran at 55 to 65% because it performed software encryption for a legacy tunnel. Nothing had changed. The real problem the users reported, slow file transfers, was a duplex mismatch on the LAN switch, visible in the interface error counters that were far above their baseline of zero.

Lesson: compare with the baseline before calling something abnormal, and look for what actually differs from normal.

"What is a network baseline and why does it matter for troubleshooting?"

Explain that a baseline is a record of normal behaviour: CPU and memory, interface utilisation and errors, neighbour and route counts, latency and top talkers, taken at quiet and busy times. It lets you tell abnormal from normal quickly, spot what changed, and plan capacity. Add that you refresh it after every significant change and store it with the documentation.

Key takeaways

  • Documentation: physical and logical topology, addressing plan, config backups, change log, contacts.
  • Baseline CPU, memory, interfaces, neighbour and route counts, latency and top talkers, and refresh it after changes.
  • In "5%/1%" CPU output the second number is interrupt-level load.
  • archive with path, maximum, write-memory, time-period keeps automatic config backups.
  • Change management: request, review, window, implementation, verification and rollback plans, then update documentation.
08

Reading the routing table like an ENARSI engineer

A postman sorting letters uses the most specific instruction he has. A note saying "all letters for this city go to the central depot" is overruled by a note saying "letters for Park Street go to depot 7", and that is overruled by "letters for 12 Park Street go to the reception desk". If two notes are equally specific, he trusts the one from the more reliable source. And if the same source gives two routes, he takes the shorter one. That is exactly how a router chooses: longest match, then administrative distance, then metric. An ENARSI engineer reads every routing table with those three rules in mind.

Anatomy of a routing table entry

R1# show ip route | begin Gateway
Gateway of last resort is 10.0.12.2 to network 0.0.0.0

O*E2  0.0.0.0/0 [110/1] via 10.0.12.2, 01:10:22, GigabitEthernet0/0
      10.0.0.0/8 is variably subnetted, 9 subnets, 4 masks
C        10.0.12.0/30 is directly connected, GigabitEthernet0/0
L        10.0.12.1/32 is directly connected, GigabitEthernet0/0
O E2     10.0.23.0/30 [110/20] via 10.0.12.2, 00:12:41, GigabitEthernet0/0
C        10.1.10.0/24 is directly connected, GigabitEthernet0/1
L        10.1.10.1/32 is directly connected, GigabitEthernet0/1
S        10.3.0.0/16 [1/0] via 10.0.12.2
O E2     10.3.30.0/24 [110/20] via 10.0.12.2, 00:12:41, GigabitEthernet0/0
C        10.255.0.1/32 is directly connected, Loopback0
O        10.255.0.2/32 [110/2] via 10.0.12.2, 01:10:22, GigabitEthernet0/0

Take the highlighted line apart: O E2 is the source (OSPF external type 2), 10.3.30.0/24 the prefix and length, [110/20] the administrative distance and metric, via 10.0.12.2 the next hop, 00:12:41 the age (how long since the route last changed; a small age on a stable network is a flapping hint) and GigabitEthernet0/0 the exit interface.

CodeSourceCodeSource
C / LConnected / local host routeDEIGRP internal
SStaticD EXEIGRP external
OOSPF intra-areaBBGP
O IAOSPF inter-areai L1 / L2IS-IS
O E1 / E2OSPF external type 1 / 2*Candidate default
O N1 / N2OSPF NSSA external+ / %Replicated route / next hop override

Rule 1: longest match

Routes for different prefix lengths do not compete; they are all installed. At forwarding time the router picks the route with the longest matching prefix. For destination 10.3.30.10, R1 has three candidates: 10.3.30.0/24 (OSPF), 10.3.0.0/16 (static) and 0.0.0.0/0 (OSPF default). The /24 wins, even though the static route has a far better AD.

1 Longest matchat lookup time 2 Lowest ADsame prefix, many sources 3 Lowest metricinside one protocol /24 beats /16 beats /0 OSPF 110 beats D EX 170 equal metric: ECMP Steps 2 and 3 decide what is installed; step 1 decides which installed route is used

Route selection: installation by AD and metric, forwarding by longest match.

Rule 2: administrative distance

When the same prefix with the same length is offered by several sources, the lowest AD is installed. Default values: connected 0, static 1, EIGRP summary 5, eBGP 20, EIGRP internal 90, OSPF 110, IS-IS 115, RIP 120, EIGRP external 170, iBGP 200, and 255 means "never use". A static route with a raised AD, such as ip route 10.3.30.0 255.255.255.0 10.0.99.2 200, is a floating static: it only appears when the dynamic route disappears. AD is local to the router; it is never advertised.

Rule 3: metric, and protocol-specific preferences

Inside one protocol the lowest metric wins, and equal metrics give equal-cost multipath. Each protocol has its own metric: OSPF cost, EIGRP composite metric from bandwidth and delay, BGP uses its best-path algorithm instead. OSPF adds a rule before the cost: intra-area beats inter-area beats E1/N1 beats E2/N2, whatever the costs. For two E2 routes with the same metric, the lower forward metric (cost to the ASBR) breaks the tie.

Recursive next hops

A next hop that is not on a connected subnet must itself be looked up: that is recursive lookup. iBGP routes normally point to a remote loopback, for example B 203.0.113.0/24 [200/0] via 10.255.0.3; the router then needs an IGP route to 10.255.0.3. If that recursion fails, BGP marks the next hop inaccessible and does not install the route. CEF resolves the recursion in advance, so show ip cef 203.0.113.1 shows the real exit interface. Static routes can recurse too; a static route pointing only to an Ethernet exit interface, without a next-hop address, relies on proxy ARP and should be avoided.

Worked example. Which route does R2 use for 10.3.30.10 when it has: S 10.3.0.0/16 [1/0], O E2 10.3.30.0/24 [110/20] learned from another ASBR, and D EX 10.3.30.0/24 [170/3072] from R3? Step 1 at install time: the two /24s compete; OSPF AD 110 beats EIGRP external 170, so the OSPF /24 is installed and the EIGRP one is not. The /16 static is installed separately. Step 2 at lookup: the /24 is longer than the /16, so traffic follows the OSPF route. If the OSPF route came from a stale redistribution point, this is exactly how suboptimal routing and loops appear after mutual redistribution.

Common mistake. Comparing ADs of routes with different prefix lengths. A static /16 with AD 1 never "beats" an OSPF /24 for addresses inside the /24; longest match decides first.

Exam trap. iBGP has AD 200, so an iBGP route loses to the same prefix from OSPF (110) or EIGRP external (170). EIGRP external is 170, not 90. OSPF intra-area wins over inter-area even with a higher cost. And show ip route 10.3.30.10 does not fall back to the default route; it reports "Subnet not in table".

The backup route that became primary

An engineer added a backup static route to the data centre through a slow DSL link, ip route 10.3.0.0 255.255.0.0 198.51.100.1, but forgot the AD. The primary path was OSPF for the same 10.3.0.0/16 summary. With AD 1 the static route replaced the OSPF route, and all data-centre traffic moved to DSL. The fix was to add an AD of 250 to make it a floating static. If the OSPF route had been more specific (/24s), the static would not have mattered, which is why the problem only appeared for the summary.

Lesson: AD decides between identical prefixes; always set a higher AD on backup statics.

"How does a router choose between routes from different protocols?"

Explain the order clearly: prefixes of different lengths are all installed and the longest match is used when forwarding; for identical prefixes the lowest AD is installed; inside one protocol the lowest metric wins with ECMP on ties. Give the AD values, mention OSPF's route-type preference and the floating static technique, and note that AD is local and not advertised.

Key takeaways

  • Read every entry as: source code, prefix/length, [AD/metric], next hop, age, exit interface.
  • Longest match first; AD only compares identical prefixes; metric only compares within one protocol.
  • AD values: 0, 1, 5, 20, 90, 110, 115, 120, 170, 200, 255.
  • OSPF prefers intra-area, then inter-area, then E1/N1, then E2/N2, regardless of cost.
  • Recursive next hops (BGP, static) must resolve through another route or the route is not usable.
09

Time-savers: the commands that answer the question directly

If you want to know whether a book is in a library, you can walk every shelf, or you can ask the catalogue. Both give the right answer; one takes an hour and the other ten seconds. In the ENARSI exam and in a real outage, time is the scarcest resource you have. This chapter collects the "catalogue" commands: the ones that answer a precise question without making you read hundreds of lines, and the ones that make changes and rollbacks fast and safe.

Ask about one destination: show ip route x.x.x.x

Instead of scanning the whole routing table, ask for one address. IOS performs the longest-match lookup for you and prints the full detail of the winning route, including the source, AD, metric, OSPF route type, the router that advertised it (from 10.255.0.2) and every next hop. Related forms: show ip route 10.3.0.0 255.255.0.0 longer-prefixes lists every more-specific route inside a block (great for summarisation problems), and show ip route ospf, show ip route eigrp or show ip route bgp limit the output to one source.

Ask what the hardware will do: show ip cef exact-route

With equal-cost paths or PBR, the routing table does not tell you which link a particular flow will use. show ip cef exact-route 10.1.10.10 10.3.30.10 gives the exact exit interface and next hop CEF selects for that source and destination pair. Use it to prove which of two ECMP links carries the problem flow before you start capturing packets.

Read only the relevant configuration: | section and friends

R2# show running-config | section router
R2# show running-config | section ^interface GigabitEthernet0/1
R2# show running-config interface GigabitEthernet0/1
R2# show running-config | include ^router|redistribute|distribute-list
R2# show ip protocols | include Routing Protocol|Redistributing

| section prints a block and all its indented children; show running-config interface prints one interface. The include with several patterns separated by | gives you a one-screen overview of routing processes, redistribution and filters. show ip protocols is the single best overview of every routing process: networks, passive interfaces, filters, redistribution, AD changes and neighbours (information sources).

Stay in configuration mode: do

Leaving configuration mode to verify and re-entering it wastes seconds on every step. Prefix any exec command with do: R2(config-router)# do show ip eigrp neighbors. Combine it with filters: do show run | section router eigrp.

configure terminalrevert timer 5 make the change still reachableconfigure confirm locked outno confirm auto rollbackafter 5 minutes Needs an archive path; the rollback uses the saved config

The revert timer: if you lock yourself out, the router puts the old configuration back on its own.

Fast, safe rollback: archive, configure replace and the revert timer

With an archive path configured (chapter 7), IOS gives you two rollback tools. configure terminal revert timer 5 saves the current configuration and starts a timer; if you do not type configure confirm within 5 minutes, the router rolls back by itself. Perfect for remote changes that might cut your own access, such as an ACL on the management interface. configure replace replaces the whole running configuration with a saved file, applying only the differences, so it is not a disruptive reload:

R3# configure replace flash:R3-cfg-2 list
This will apply all necessary additions and deletions
to replace the current running configuration with the
contents of the specified configuration file, which is
assumed to be a complete configuration, not a partial
configuration. Enter Y if you are sure you want to proceed. ? [no]: y
!Pass 1
!List of Commands:
interface GigabitEthernet0/2
 no ip access-group SRV-PROTECT out
no ip access-list extended SRV-PROTECT
end
Total number of passes: 1
Rollback Done

The list keyword prints the commands applied, which doubles as documentation of what the rollback changed. Before replacing, show archive config differences flash:R3-cfg-2 system:running-config shows the same difference without applying it.

Worked example. Question: "Which route and which exit interface does R1 use for SRV1, and who advertised the route?" The slow way: show ip route (40 lines), find the entry, then show ip ospf database external to find the advertising router. The fast way: show ip route 10.3.30.10 shows "Known via ospf 1 ... type extern 2 ... from 10.255.0.2" in four lines, and show ip cef exact-route 10.1.10.10 10.3.30.10 confirms Gi0/0. Two commands, under 20 seconds.

Common mistake. Using | include on the running configuration and concluding that a command is missing. include shows single lines without context; a redistribute line could belong to any routing process. Use | section when the parent matters.

Exam trap. configure replace and configure terminal revert timer need the archive feature (archive with path) for the rollback to have something to return to. configure confirm stops the timer; without it the rollback happens even if the change was good.

The ACL that locked out the engineer

An engineer applied a new inbound ACL to the WAN interface of a remote branch router at 23:00. The ACL was missing a permit for the management subnet, and his SSH session froze. There was no remote hands until morning. His colleague on another site had adopted the habit of always using configure terminal revert timer 10 for remote ACL changes; on her router the same mistake undid itself after ten minutes. After that night the team made the revert timer mandatory in their change template.

Lesson: build the rollback into the change itself, so it works even when you cannot reach the device.

"You are changing an ACL on a remote router with no out-of-band access. How do you protect yourself?"

Mention configure terminal revert timer with an archive path, and configure confirm once access is verified; or the older reload in 10 with the startup configuration unchanged, cancelled with reload cancel. Add that you review the ACL for a management permit, apply it in the right direction, and have configure replace ready with a known good file.

Key takeaways

  • show ip route x.x.x.x does the longest-match lookup and shows full route detail; longer-prefixes shows what is inside a block.
  • show ip cef exact-route src dst shows the exact exit for a flow, including ECMP and PBR effects.
  • | section, show running-config interface and show ip protocols give focused configuration views.
  • do runs exec commands from configuration mode.
  • configure terminal revert timer plus configure confirm, and configure replace ... list, give safe, fast rollback.
10

Exam strategy for troubleshooting simlets and tickets

An emergency-room doctor does not start with a full body scan. She reads the ambulance note, checks the vital signs, and treats the one thing that will kill the patient first, without causing new damage along the way. Troubleshooting items in ENARSI reward the same discipline: read the ticket carefully, isolate the fault quickly with a fixed routine, fix only what is broken, and leave everything else exactly as you found it.

The item types you will meet

Multiple choice
One or more correct answers, often built around a piece of show or debug output. Read the question stem before the output so you know what to look for.
Drag and drop
Order process steps, match commands to their function, or match symptoms to causes.
Simlet
A topology where you can run show commands (not change configuration) and answer several multiple-choice questions about it.
Testlet
A scenario description followed by several questions about the same scenario.
Simulation
A lab-style task where you configure or fix devices to meet stated requirements.

You cannot return to an earlier item, so make each decision before you move on.

Reading the ticket

Most mistakes in troubleshooting items are reading mistakes. Before you type anything, extract four things from the ticket:

  1. The symptom and scope: which source, which destination, which traffic type. Write down the exact addresses.
  2. The expected result: what must work when you are finished, which becomes your verification test.
  3. The constraints: "do not use static routes", "do not modify R2", "use a route map", "do not change the OSPF area design". Violating a constraint can make a technically working fix wrong.
  4. The hints: words like "after a recent change", "intermittently", "only from branch 2" narrow the search at once.
Read: symptom, scope, expected result, constraints Draw the path, pick the checklist Isolate the device and cause Minimal fix + verify

From a wide ticket to one precise, minimal fix.

Isolating fast

Draw the path on your scratch board: source, each router, destination. Then apply the checklists from chapters 4 and 5 in a fixed order, starting at the router nearest the source:

  • show ip route <destination> on each hop until the route is missing or wrong. That router, or its upstream neighbour, is your fault domain.
  • In the fault domain: neighbours, then what is advertised and received, then filters, AD and redistribution.
  • If routing is correct end to end: ACLs, NAT, PBR on the path, and the return path.

Do not read full running configurations first. They are long, and the fault usually hides in one line that a targeted command reveals faster, for example show ip protocols showing a distribute list, or show ip ospf interface brief showing an interface missing.

The ENARSI fault library

Exam faults come from a fairly small family. Train your eyes to spot them:

AreaTypical injected faults
NeighboursPassive interface, wrong area or AS, authentication key mismatch, timer mismatch, MTU mismatch, wrong BGP remote-as or update-source, missing ebgp-multihop
RoutesMissing seed metric, distribute list or route-map deny, stub or area type hiding routes, missing next-hop-self, changed AD, wrong summary
Data planeACL in the wrong direction or missing permit, PBR with wrong next hop, uRPF strict on an asymmetric path, CoPP dropping a protocol
VPNWrong tunnel source or destination, NHRP network ID or mapping mismatch, IPsec profile mismatch, MTU on tunnels
ServicesMissing ip helper-address, SNMP community or ACL, syslog level too low, IP SLA not scheduled, track object wrong

Not breaking other things

  • Make the smallest change that fixes the cause. Removing an ACL completely "fixes" reachability but breaks the security requirement.
  • Respect every constraint in the ticket; if it says "use a route map", a distribute list with a prefix list is wrong.
  • After the fix, verify the ticket's expected result and re-check anything that worked before (other neighbours, other routes).
  • Save the configuration if the task asks you to (copy running-config startup-config).

Worked example: a time budget. Assume 60 items in 90 minutes, with two simlets and one simulation. Budget 10 minutes for each simlet and 12 for the simulation: 32 minutes. That leaves 58 minutes for 57 other items, just over a minute each. If you notice after 45 minutes that you are on item 25, you are behind: answer the next few multiple-choice items quickly to recover. Check the clock at roughly every 15 items.

Common mistake. Fixing the first thing that looks odd and moving on. A simulation often contains two or three faults on the same path; after each fix, run the end-to-end test again. The next fault only becomes visible once the first is gone.

Exam trap. In simlets, the question often asks why something happens, not how to fix it. Answers that describe a correct fix for a different problem are common distractors. Match the answer to the evidence you actually saw in the output.

The fix that broke the security requirement

In a practice lab the ticket said: "Branch users must reach the web server on TCP 443. The server must remain protected: only ICMP, SSH and HTTPS allowed." A student found the outbound ACL on R3 denying TCP 443 and simply removed the ip access-group command. Reachability worked, but the grader failed the task, because all traffic to the server was now permitted. The correct fix was one line in the ACL: 25 permit tcp any host 10.3.30.10 eq 443, placed before the deny.

Lesson: a fix must satisfy the whole ticket, including the parts that were already working.

"You are given a ticket in a live environment. What do you do in the first five minutes?"

Read and restate the problem with exact addresses and expected result, confirm scope and recent changes, note constraints and change windows, draw the path, and run read-only checks from the source hop outward (show ip route, neighbours, show ip protocols). Say you do not change anything until you have a hypothesis backed by evidence, and that you plan verification and rollback before the fix.

Key takeaways

  • Know the item types: multiple choice, drag and drop, simlets, testlets and simulations; there is no going back.
  • Extract symptom, scope, expected result, constraints and hints before touching the CLI.
  • Isolate with show ip route hop by hop, then the protocol and forwarding checklists.
  • Learn the common fault library; exam faults repeat in patterns.
  • Smallest fix, respect constraints, re-test end to end, look for a second fault, save if asked.
11

Troubleshooting workflow: a multi-fault ticket end to end

A detective story rarely has only one clue. The first clue leads to a suspect, the suspect has an alibi, and a second clue appears only after the first is explained. Real tickets are like that, and so are ENARSI simulations: fix one fault and the symptom changes, revealing the next. This chapter walks through one ticket from the first sentence to the documentation, using every tool from this module on the reference topology.

The ticket

INC-7302: "Since last night's maintenance, branch users in 10.1.10.0/24 cannot open the HR portal on SRV1 (10.3.30.10, HTTPS). Constraints: do not add static routes; SRV1 must stay protected (only ICMP, SSH and HTTPS allowed to it)." In this lab R1 has no default route.

PC1 R1 R2 R3 SRV1 1EIGRP to OSPF filter 2OSPF to EIGRP, no metric 3ACL out Gi0/2 Each fault is visible only after the previous one is fixed

The three faults hidden in INC-7302, in the order you will find them.

Steps 1 and 2: define and gather

Restated: "Hosts in 10.1.10.0/24 cannot reach 10.3.30.10 on TCP 443; a maintenance happened last night." PC1 shows Reply from 10.1.10.1: Destination net unreachable. That message comes from R1, its gateway, so R1 has no usable route. Divide and conquer has already told us where to start: Layer 3 on R1.

R1# show ip route 10.3.30.10
% Subnet not in table
R1# show ip ospf neighbor
Neighbor ID     Pri   State           Dead Time   Address         Interface
10.255.0.2        1   FULL/DR         00:00:36    10.0.12.2       GigabitEthernet0/0

Fault 1: the route is not redistributed into OSPF

The OSPF neighbour is Full, so the problem is not the adjacency. Following the routing-protocol checklist, go to where the prefix should be advertised: R2 must redistribute it from EIGRP into OSPF.

R2# show ip route 10.3.30.10
Routing entry for 10.3.30.0/24
  Known via "eigrp 100", distance 90, metric 3072, type internal
  Redistributing via eigrp 100, ospf 1
R2# show run | section router ospf
router ospf 1
 redistribute eigrp 100 subnets route-map EIGRP-TO-OSPF
 network 10.0.12.0 0.0.0.3 area 0
 network 10.255.0.2 0.0.0.0 area 0
R2# show route-map EIGRP-TO-OSPF
route-map EIGRP-TO-OSPF, permit, sequence 10
  Match clauses:
    ip address prefix-lists: DC-NETS
  Set clauses:
    tag 100
  Policy routing matches: 0 packets, 0 bytes
R2# show ip prefix-list DC-NETS
ip prefix-list DC-NETS: 2 entries
   seq 5 permit 10.3.20.0/24
   seq 10 permit 10.3.40.0/24

Hypothesis: the maintenance rewrote the prefix list and typed 10.3.40.0/24 instead of 10.3.30.0/24, so the route map's implicit deny blocks SRV1's subnet. It is a read-only confirmation; now a single, minimal change:

! Add the missing data-centre subnet; keep the route-map design intact
R2(config)# ip prefix-list DC-NETS seq 15 permit 10.3.30.0/24
R2(config)# do show ip ospf database external 10.3.30.0 | include Link State ID
  Link State ID: 10.3.30.0 (External Network Number )

R1 now shows O E2 10.3.30.0/24 [110/20] via 10.0.12.2. Re-test from the user's point of view.

Fault 2: the return path is missing

R1# ping 10.3.30.10 source GigabitEthernet0/1
Type escape sequence to abort.
Sending 5, 100-byte ICMP Echos to 10.3.30.10, timeout is 2 seconds:
Packet sent with a source address of 10.1.10.1
.....
Success rate is 0 percent (0/5)
R3# show ip route 10.1.10.10
% Subnet not in table

The symptom changed: the forward route exists, but R3 cannot send the replies back. R3's EIGRP neighbour R2 is up, so check whether R2 advertises 10.1.10.0/24 into EIGRP.

R2# show ip eigrp topology 10.1.10.0/24
%Entry 10.1.10.0/24 not in topology table
R2# show run | section router eigrp
router eigrp 100
 network 10.0.23.0 0.0.0.3
 redistribute ospf 1

Classic EIGRP needs a seed metric for routes from OSPF. The maintenance removed the old default-metric line. Fix:

R2(config)# router eigrp 100
R2(config-router)# redistribute ospf 1 metric 100000 10 255 1 1500
R3# show ip route eigrp | include 10.1.10
D EX     10.1.10.0/24 [170/28416] via 10.0.23.1, 00:00:09, GigabitEthernet0/1

The metric checks out: the minimum bandwidth of 100,000 kbps gives 10,000,000 / 100,000 = 100; the seed delay of 10 (tens of microseconds) plus 1 for R3's GigabitEthernet interface gives 11; and 256 x (100 + 11) = 28416. Ping from the branch LAN now succeeds.

Fault 3: ping works, HTTPS does not

Layer 3 is proven in both directions, so switch to top-down for the application.

R3# show ip interface GigabitEthernet0/2 | include access list
  Outgoing access list is SRV-PROTECT
  Inbound  access list is not set
R3# show access-lists SRV-PROTECT
Extended IP access list SRV-PROTECT
    10 permit icmp any any (10 matches)
    20 permit tcp any any eq 22
    30 deny ip any any log (38 matches)
R3# show logging | include SRV-PROTECT
%SEC-6-IPACCESSLOGP: list SRV-PROTECT denied tcp 10.1.10.10(51544) -> 10.3.30.10(443), 1 packet

The constraint says SRV1 must stay protected, so removing the ACL is not allowed. Insert one permit before the deny:

R3(config)# ip access-list extended SRV-PROTECT
R3(config-ext-nacl)# 25 permit tcp any host 10.3.30.10 eq 443

Steps 7 and 8: verify and document

From PC1, the HR portal loads; show access-lists SRV-PROTECT shows matches on line 25. Then check that nothing else broke: both neighbours up, show ip route summary on R2 matches the baseline plus the restored routes, and other branches still reach their services.

#Root causeFix
1Prefix list DC-NETS on R2 had 10.3.40.0/24 instead of 10.3.30.0/24ip prefix-list DC-NETS seq 15 permit 10.3.30.0/24
2OSPF redistributed into EIGRP without a seed metricredistribute ospf 1 metric 100000 10 255 1 1500
3SRV-PROTECT ACL did not allow HTTPS25 permit tcp any host 10.3.30.10 eq 443

Worked example: why three faults looked like one. Before any fix, all three faults produced the same user symptom, "HR portal down". Only after fault 1 was fixed did ping change from "net unreachable" to timeouts (fault 2), and only after fault 2 did ping succeed while HTTPS failed (fault 3). Each fix changed the symptom, and each new symptom pointed to a different checklist.

Common mistake. Stopping after fault 1 because "the route is back now". Always re-run the user's test after every fix.

Exam trap. "Destination net unreachable" from the gateway means the gateway has no route; timeouts usually mean the packet was forwarded and something later (often the return path or an ACL without an ICMP reply) failed. An ACL log line proves a deny at a precise place.

The maintenance that "only tidied up"

This ticket is modelled on a real pattern: an engineer "tidied up" redistribution on a border router, replacing a long prefix list with a short one, and removing a default-metric line that looked unused. Separately, another engineer hardened a server ACL. Each change passed its own quick test. Together they cut off one application. The configuration change logger showed both sessions, which turned a three-hour hunt into a thirty-minute fix.

Lesson: small unrelated changes combine into multi-fault outages; change logs and a method untangle them.

"Tell me about a time you fixed a problem that had more than one cause."

Use a structure like INC-7302: the symptom, how you isolated the first fault with evidence, how the symptom changed after the fix, the second and third faults, how you respected constraints (no static routes, keep the ACL), how you verified end to end, and what you documented so it does not happen again.

Key takeaways

  • Restate the ticket with addresses, protocol and constraints before you touch anything.
  • The gateway's "net unreachable" pointed to a missing route; the routing-protocol checklist found a prefix-list error.
  • After each fix, re-test: the changed symptom (timeouts, then HTTPS only) revealed the next fault.
  • Minimal fixes that respect constraints: add a prefix-list line, add a seed metric, add one ACL permit.
  • Verify end to end, check nothing else changed against the baseline, and document every root cause.
12

Summary and exam checklist

A pilot runs the same checklist before every flight, no matter how many thousand hours she has flown. The checklist is short, it is always in the same order, and it catches the mistakes that experience alone misses. This last chapter is your ENARSI pre-flight checklist: the method, the two checklists, the commands, the facts most often tested, and a glossary. Come back to it before every later module and before the exam.

Ticket Processdefine ... document Approachtop-down, path, compare Checklistsforwarding, protocol Toolkit: show, ping, debug, logs, SLA, EPC

The whole module on one page: process, approach, checklists, toolkit.

Can-do checklist

  • I can state the ENARSI 300-410 v1.1 domains and weights: Layer 3 35%, VPN 20%, Infrastructure Security 20%, Infrastructure Services 25%.
  • I can list the eight process steps in order and explain why defining the problem and documenting matter most.
  • I can choose between top-down, bottom-up, divide and conquer, follow the path, compare configurations and component swapping from a symptom.
  • I can run the packet-forwarding checklist on each hop: Layer 2, ARP, RIB, CEF, ACL/NAT/PBR, return path.
  • I can run the routing-protocol checklist: neighbours, advertised, received or filtered, wins in the RIB, redistribution.
  • I can debug safely, set up logging with timestamps, and use SNMP, syslog, NetFlow, IP SLA and EPC as evidence.
  • I can read every field of a routing table entry and predict route selection with longest match, AD and metric.
  • I can roll back safely with configure terminal revert timer and configure replace.
  • I can work a multi-fault ticket to the end without breaking constraints.

Mini glossary

Baseline
A record of normal behaviour (CPU, memory, traffic, routes, neighbours) used to spot abnormal behaviour.
RIB
Routing information base: the routing table of best routes from all sources.
FIB
Forwarding information base built by CEF from the RIB, with recursion resolved.
Adjacency table
CEF table of pre-built Layer 2 headers for each next hop, built from ARP and ND.
Administrative distance
Trust value of a route source; lower wins for identical prefixes; local only.
Seed metric
The starting metric given to routes when they are redistributed into another protocol.
Floating static
A static route with a raised AD so that it is used only when the dynamic route disappears.
Conditional debug
A debug limited by an interface, neighbour, ACL or other condition to protect the CPU.
EPC
Embedded Packet Capture: capturing packets on the router and exporting a pcap file.
Simlet
An exam item with a live topology for show commands and several related questions.

Most tested facts

  • AD: connected 0, static 1, EIGRP summary 5, eBGP 20, EIGRP 90, OSPF 110, IS-IS 115, RIP 120, EIGRP external 170, iBGP 200, unusable 255.
  • Longest match beats AD; AD compares only identical prefixes; metric compares only within one protocol.
  • OSPF preference: intra-area, inter-area, E1/N1, E2/N2, regardless of cost.
  • OSPF distribute-list in filters the routing table, not the LSDB.
  • Classic EIGRP needs a seed metric for routes from other routing protocols; OSPF uses 20 (1 for BGP) as E2 by default.
  • BGP r = RIB-failure; BGP network needs an exact prefix and mask in the RIB.
  • PBR is processed before the routing table on the interface where it is applied.
  • debug ip packet shows only process-switched packets.
  • Syslog levels 0 to 7: emergencies, alerts, critical, errors, warnings, notifications, informational, debugging.
  • Ping codes: ! reply, . timeout, U unreachable, M could not fragment; traceroute A = administratively prohibited.

Command cheat-sheet

! Where does this destination go?
show ip route 10.3.30.10
show ip route 10.3.0.0 255.255.0.0 longer-prefixes
show ip cef exact-route 10.1.10.10 10.3.30.10
! Control plane overview
show ip protocols
show ip ospf neighbor
show ip eigrp neighbors
show ip bgp summary
show ip route summary
! Policy on the path
show ip interface Gi0/2 | include access list
show access-lists
show ip policy
show ip nat translations
! Tests
ping 10.3.30.10 source Gi0/1 repeat 100 size 1400 df-bit
traceroute 10.3.30.10 source 10.1.10.1 numeric
! Evidence and change
show logging | include ADJCHG|NBRCHANGE
show archive log config all
show archive config differences flash:R3-cfg-2 system:running-config
configure terminal revert timer 5
configure confirm
configure replace flash:R3-cfg-2 list
undebug all

Worked example: a two-minute self-test. Cover the page and answer: (1) Which domain is 25%? Infrastructure Services. (2) Route in OSPF database, not in RIB: two likely causes? A distribute list in, or the same prefix from a lower-AD source. (3) PC can ping, cannot browse: which approach? Top-down. (4) Which command shows the exact exit interface for a flow? show ip cef exact-route. If you got all four in under two minutes, you are ready for the technology modules.

Common mistake. Treating this module as "soft" material and skipping to the protocols. Every later lab and ticket assumes you run the checklists automatically. Practise them until you no longer think about the order.

Exam trap. Watch for answers that are true statements but do not answer the question asked. In troubleshooting items, the right answer is the one that explains the specific evidence shown, not a generally good practice.

From checklist to habit

A NOC team adopted the two checklists from this module as a laminated card at every desk. For the first month engineers ticked each line; after that they no longer needed the card. Their mean time to resolve routing incidents fell noticeably, and escalations to senior engineers dropped, because junior engineers now arrived with a precise fault domain and evidence instead of "the network is down".

Lesson: a checklist practised deliberately becomes instinct, and instinct built on method is what ENARSI measures.

"What makes a good network troubleshooter?"

A strong answer mentions method (a structured process and checklists), fundamentals (how routers really forward and select routes), tools (targeted show commands, safe debugging, logs, IP SLA, captures), discipline (one change at a time, rollback, respecting constraints) and communication (a clear problem statement and documentation). Give a short example of a ticket you solved that way.

Key takeaways

  • ENARSI is a troubleshooting exam: Layer 3 35%, VPN 20%, Security 20%, Services 25%.
  • Process: define, gather, analyse, eliminate, propose, test, solve, document.
  • Two checklists: packet forwarding per hop, and the routing-protocol checklist for missing or wrong routes.
  • Targeted commands and safe tools save time and protect the network.
  • Smallest fix, respect constraints, verify end to end, document; every later module builds on this.
🎓 For educational purposes only — all devices are simulationsTerms of UsePrivacy Policy© 2026 Network Kings
CONFIG by Network Kings — an educational IT simulation platform for learning purposes only. It is not Cisco IOS, Junos, FortiOS or PAN-OS and contains no Cisco, Juniper, Fortinet or Palo Alto Networks software. Cisco, IOS, CCNA, CCNP, Juniper, JNCIA, JNCIS, JNCIP, Fortinet, FortiGate, FortiOS, NSE, Palo Alto Networks, PAN-OS and PCNSE are trademarks of their respective owners. Network Kings is not affiliated with or endorsed by Cisco Systems, Inc., Juniper Networks, Inc., Fortinet, Inc. or Palo Alto Networks, Inc.