CCNP ENCOR 350-401 · Architecture · Design & HA

Enterprise design & high availability

Two-tier and three-tier campus designs, spine-leaf fabrics, WAN and branch options, cloud versus on-premises, and the high-availability tools that keep a campus up: redundant links, EtherChannel, first-hop redundancy aligned with spanning tree, and stateful switchover. Maps to ENCOR 1.1 (enterprise network design: 2/3-tier, fabric, cloud) and 1.2 (high availability: redundancy, FHRP, SSO).

68 min read12 chapters2 labs15 quiz8 scenarios15 interview Q&A

This first module is free: read the lesson and take the quiz. Create a free account to run up to 3 hands-on labs.

Log inStart free
Jump to chapter (12)
01

The enterprise network: what you will learn and the big picture

What you will learn in this module. By the end of these chapters you will be able to look at any enterprise network drawing and explain why it is built the way it is: why there are two of every box, where Layer 2 ends and Layer 3 begins, why the spanning-tree root and the default gateway must live on the same switch, how a supervisor can fail without users noticing, which wireless deployment model fits a campus or a branch, when a workload belongs in the cloud, and how a switch actually forwards a frame in hardware. You will also build a redundant collapsed-core block in the lab and review a broken design until traffic stops taking the long way.

Prerequisites. You should be comfortable with CCNA-level VLANs and trunks, spanning tree (root bridge, root port, blocking), EtherChannel with LACP, HSRP basics and single-area OSPF. If any of these feel shaky, revisit the CCNA modules on STP, EtherChannel and FHRP first. This module does not re-teach them; it uses them as building blocks for design.

An analogy: a network is a city's road system

Picture a well-planned city. Every house sits on a quiet local lane. Lanes feed into arterial roads where the traffic lights, speed checks and turn restrictions live. Arterial roads join a fast highway with no traffic lights at all, because its only job is to move a lot of cars quickly between districts. Every district is planned as a copy of the same template, so a new district can be added without redesigning the city. If a lane is dug up for repairs, only that street is affected, and there is always a second road into every district.

An enterprise network works the same way. Local lanes are the access layer where users plug in. Arterial roads are the distribution layer where policy and routing decisions happen. The highway is the core. Districts are modules (a building, a data centre, a WAN edge). The second road is redundancy. The rule that a dug-up lane only affects one street is the idea of a failure domain.

Campus core Campus buildings Data centre Internet edge WAN / SD-WAN edge Branches Public cloud VPN / interconnect

An enterprise is a set of modules joined by a core. Each module is designed on its own and can fail on its own.

Where this module sits in the exam

The ENCOR 350-401 Architecture domain opens with design. This module covers the design-level items; later modules go deep on the technologies they mention.

Blueprint itemWhat you must be able to doChapters
1.1 Design principlesExplain 2-tier, 3-tier, fabric and cloud designs; high availability with redundancy, FHRP and SSO2 to 7
1.2 WLAN deployment modelsCompare centralised, distributed, controller-less, controller-based, cloud and remote-branch designs9
1.3 On-premises vs cloudChoose where infrastructure should live and how it connects8
1.7 Hardware vs software switchingExplain process switching, CEF, RIB, FIB, adjacency table, CAM and TCAM10

SD-WAN and SD-Access, LISP and VXLAN, HSRP object tracking and deep wireless configuration each get their own modules later. Here you learn why they exist and where they fit.

The lab network you will use

Both labs use the same small collapsed-core block. Two distribution/core switches, DS1 and DS2, are joined by two links (Gi0/6 and Gi0/7). One access switch, AS1, has an uplink to each: Gi0/0 to DS1 and Gi0/1 to DS2. PC1 sits in VLAN 10 (10.1.10.0/24) and PC2 in VLAN 20 (10.1.20.0/24). DS1 owns .2 and DS2 owns .3 in each VLAN, and HSRP provides the virtual gateway .1. A WAN router R1 connects to DS1 over 10.0.1.0/30 and to DS2 over 10.0.2.0/30, and reaches the data-centre server SRV (172.16.0.10). OSPF area 0 runs between R1 and the distribution switches.

Worked example. PC1 (10.1.10.11) opens a file on SRV (172.16.0.10). The frame leaves AS1 on the uplink towards the HSRP active gateway, DS1 routes it onto 10.0.1.0/30, R1 routes it to the data-centre LAN. If DS1 dies, HSRP moves 10.1.10.1 to DS2, spanning tree unblocks the DS2 uplink if needed, and OSPF on R1 already has a second equal-cost path through DS2. Three independent mechanisms, each covering a different failure. That is design.

Common mistake. Treating design as drawing boxes. Every design choice has a protocol consequence. "Two distribution switches" immediately means you must decide the STP root, the HSRP active router, the EtherChannel mode between them and how routing sees both paths. Most outages in well-funded networks are caused by redundancy that was bought but never aligned.

Exam trap. ENCOR design questions rarely ask for commands. They describe a situation (many buildings, a small branch, heavy east-west traffic, a supervisor failure) and ask which design or feature fits. Learn the reason behind each option, not just its name.

The campus that grew one switch at a time

A college started with one core switch and added access switches whenever a lab opened, daisy-chaining some of them. After five years a single faulty NIC in one hostel caused a broadcast storm that took down every building, because all 40 switches shared one flat VLAN and one spanning tree. The redesign split the campus into building modules, each with a distribution pair, routed links to a core pair and VLANs that never leave the building.

Lesson: hierarchy and modularity are not paperwork. They are what keeps one bad cable from becoming a campus-wide outage.

"Walk me through how you would design a new enterprise campus."

Start with requirements: number of users and buildings, applications, growth, availability target. Then describe the hierarchy (access, distribution, core, or collapsed core if small), the Layer 2/3 boundary, redundancy at every layer, how STP and FHRP are aligned, how the campus connects to the data centre, WAN and cloud, and how wireless is deployed. Mentioning failure domains and oversubscription shows you think like a designer, not just a configurator.

Key takeaways

  • Enterprise networks are built from modules joined by a core, like districts joined by a highway.
  • Access connects users, distribution applies policy and is the L2/L3 boundary, core moves traffic fast.
  • Redundancy only works when protocols such as STP, FHRP and routing are aligned with it.
  • This module covers ENCOR 1.1, 1.2, 1.3 and 1.7 at a design level; later modules go deep.
  • The labs use DS1, DS2, AS1, R1, SRV, PC1 and PC2 in a collapsed-core block.
02

Design principles: hierarchy, modularity, resiliency, flexibility

Before you pick a single switch model, you need a way to judge whether a design is good. Architects of buildings use principles such as "load-bearing walls line up floor to floor" and "every floor has two exits". Network designers use four: hierarchy, modularity, resiliency and flexibility. Two measurements keep those principles honest: the size of each failure domain and the oversubscription ratio of each layer.

Hierarchy: every layer has one job

Hierarchy means splitting the network into layers with clearly separated roles: access, distribution and core (the next chapter looks at each). Because every device in a layer does the same job, the network behaves predictably. You know where to look for a policy (distribution), where users attach (access) and which devices must never be slowed down (core). Hierarchy also makes route summarisation possible: a distribution pair can advertise one summary such as 10.1.0.0/16 for a whole building, so a flapping user subnet does not ripple through the core.

Modularity: build with repeatable blocks

Modularity means building the enterprise from self-contained blocks: an access-distribution block per building, a data-centre block, a WAN edge block, an internet edge block. Each block has the same internal design and connects to the core in the same way. The benefits are practical. A new building is a copy of a tested template. A change in one block (say a new VLAN in building B) does not touch the others. Troubleshooting starts by finding the block, then working inside it. Modularity also lets different teams own different blocks.

Resiliency: survive failures, normal and abnormal

Resiliency is the ability to keep working when something goes wrong. That includes the obvious failures (a link, a line card, a whole switch, a power feed) and the abnormal ones: a broadcast storm, a loop created by a user's home switch, a traffic flood, a bad software upgrade. Redundancy is the raw material: two uplinks, two distribution switches, dual supervisors and power supplies. But redundancy only becomes resiliency when failover is fast and automatic, which is the job of EtherChannel, FHRP, routing protocols, SSO and NSF, all covered in later chapters.

Availability is measured in "nines". 99.9 % allows about 8.8 hours of downtime a year, 99.99 % about 53 minutes, 99.999 % about 5 minutes. Each extra nine means removing another single point of failure and shortening another convergence time.

Flexibility: ready for what comes next

Flexibility is the ability to adapt without a redesign: adding wireless, IoT, a new cloud region, a merger, a new security policy. Hierarchical and modular designs are naturally flexible because you change one block or add one block. Designs with fixed, hand-crafted exceptions are not. Scalability is closely related: can the design grow from 5 to 50 access switches without changing its shape?

Failure domains: how far does one fault spread?

A failure domain is the part of the network that is affected when one component fails or misbehaves. A Layer 2 VLAN is a failure domain for broadcasts and loops: a storm reaches every port in that VLAN. A spanning-tree instance is a failure domain for topology changes. The good design habit is to keep Layer 2 domains small, ideally inside one access-distribution block, and to use routed links between blocks, because a routing protocol does not forward broadcasts and contains a failure to one subnet.

Stretched VLAN 50 Bldg A Bldg B Bldg C One loop hits all three Routed between blocks Bldg A Bldg B Bldg C A loop stays in one block

Left: one VLAN across buildings is one failure domain. Right: VLANs end at each building's distribution pair.

Oversubscription: honest about bandwidth

Oversubscription is the ratio between the bandwidth that could arrive from below and the bandwidth available upward. Not every user transmits at line rate at the same time, so some oversubscription is normal and saves money. Common starting guidance for a campus is up to about 20:1 from access to distribution and about 4:1 from distribution to core. Data-centre fabrics aim much lower, often 3:1 or better, because servers really do send at high rates.

Worked example. An access switch has 48 user ports at 1 Gbps (48 Gbps down) and two 10 Gbps uplinks (20 Gbps up). If both uplinks forward, the ratio is 48 : 20 = 2.4:1. If spanning tree blocks one uplink, only 10 Gbps is usable and the ratio becomes 4.8:1. Now stack four such switches behind the same two uplinks: 192 Gbps down, 20 Gbps up, 9.6:1, still inside the 20:1 guideline. Count only links that actually forward.

Common mistakes. Counting an STP-blocked uplink as capacity. Stretching a VLAN between buildings "just for one application", which silently merges failure domains. Buying redundant boxes but feeding both from one power strip or one cable tray.

Exam trap. "Redundancy" and "resiliency" are not the same word. A question that says a network has redundant links but takes 50 seconds to recover is describing redundancy without resiliency; the fix is a faster mechanism (RSTP, EtherChannel, routed links, tuned FHRP), not more links.

Cameras everywhere, outage everywhere

A factory stretched VLAN 50 across three buildings so every security camera could sit in one subnet with the recording server. A contractor plugged a small unmanaged switch into two wall ports in building C, creating a loop. The broadcast storm ran through VLAN 50 into buildings A and B, CPU on the distribution switches hit 100 percent and users in all three buildings lost the network. The engineer found the storm with interface counters and show spanning-tree topology change counts, shut the looped ports, and later redesigned: a camera VLAN per building, routed links between buildings, BPDU guard on access ports and storm control.

Lesson: a VLAN stretched for convenience is a failure domain stretched for disaster.

"What is a failure domain, and how do you keep it small in a campus?"

Define it as the set of devices and users affected by one failure. Then list the tools: keep VLANs local to one access-distribution block, make links between blocks and to the core routed, use routed access where the design allows, summarise routes at distribution, and protect the edge with BPDU guard and storm control. A senior touch is mentioning that control-plane failures such as a flapping route are also contained by summarisation.

Key takeaways

  • Hierarchy gives each layer one role; modularity builds the network from repeatable blocks.
  • Resiliency is redundancy plus fast, automatic failover, against both failures and abnormal traffic.
  • Flexibility and scalability come from changing or adding blocks, not redesigning.
  • Keep Layer 2 failure domains small; route between blocks.
  • Oversubscription guideline: about 20:1 access to distribution, 4:1 distribution to core; count only forwarding links.
03

Three-tier, two-tier and the collapsed core

Think of a big railway network. Local trains stop at every small station (access). Several local lines meet at a junction where passengers change trains and tickets are checked (distribution). Junctions are linked by express trains that never stop at small stations (core). A small town, though, does not need an express line: one junction that also acts as the main station is enough. That is the difference between a three-tier campus and a two-tier (collapsed-core) campus.

The three layers and their jobs

LayerMain jobTypical features
AccessConnect end devices: PCs, phones, APs, printers, camerasVLANs, PoE, 802.1X, port security, DHCP snooping, dynamic ARP inspection, QoS trust boundary, PortFast and BPDU guard
DistributionAggregate access switches; the Layer 2/Layer 3 boundarySVIs and FHRP gateways, STP root, route summarisation, ACLs and policy, redundant uplinks to core
CoreFast, reliable transport between distribution blocks, data centre, WAN and internet edgeLayer 3 only, high-speed links, fast convergence, minimal policy, no users attached

The core is deliberately simple. Anything that costs CPU or adds risk (complex ACLs, NAT, packet inspection) belongs at the distribution layer or in dedicated blocks such as the internet edge. A core switch should do one thing: move packets between blocks as fast and as reliably as possible.

Three-tier Core Core Dist Dist Dist Dist Access bldg A Access bldg B Two-tier (collapsed core) DS1 DS2 core + distribution AS1

Three-tier adds a core pair once there are several distribution blocks. Two-tier merges core and distribution into one pair, as in this module's lab.

Why a separate core? The full-mesh problem

Without a core, every distribution block must connect directly to every other block and to the WAN, data centre and internet edge. The number of block-to-block connections grows as n(n-1)/2.

Worked example. Six building distribution pairs fully meshed need 6 x 5 / 2 = 15 pair-to-pair connections. If each connection is four links (every switch to every switch of the other pair), that is 60 fibres, and building seven adds 24 more. With a core pair in the middle, each distribution pair needs only 4 links (two switches, each to both core switches): 6 x 4 = 24 links, and building seven adds just 4. Routing adjacencies drop the same way.

A common rule of thumb is to introduce a dedicated core once you have more than two or three distribution blocks, or when the campus spans several buildings, or when you need one clean meeting point for the data centre, WAN and internet edge.

Two-tier: the collapsed core

In a collapsed-core design the core and distribution functions run on the same pair of switches. Access switches uplink to the pair; the pair also connects to the WAN routers, firewalls or a small server room. This is the right design for a single building or a medium site: fewer boxes, less cabling, less cost, and still fully redundant if built in pairs. In the lab, DS1 and DS2 are a collapsed core: they are the default gateways for VLANs 10 and 20 (distribution role) and route towards R1 and the data centre (core role).

The trade-offs: the collapsed pair carries both user-facing policy and core transport, so a mistake there affects everything; scaling to many buildings eventually brings back the full-mesh problem; and upgrades must be planned more carefully because there is no separate core to carry traffic.

Summarise at the distribution layer

Hierarchy pays off in routing. If each building uses one block of addresses, its distribution pair can advertise a single summary to the core. In OSPF, with the building in its own area and the distribution switches as ABRs:

! On each distribution switch acting as ABR for area 1
router ospf 1
 area 1 range 10.1.0.0 255.255.0.0
CORE1# show ip route ospf
O IA     10.1.0.0/16 [110/3] via 10.0.11.2, 00:05:12, TenGigabitEthernet1/0/1
                     [110/3] via 10.0.12.2, 00:05:12, TenGigabitEthernet1/0/2

A user subnet flapping inside the building no longer triggers SPF recalculation across the whole campus. The multi-area OSPF module covers the details.

Common mistakes. Attaching servers or users directly to core switches "because there were free ports". Putting heavy ACLs or NAT in the core. Building a collapsed core with a single switch, which creates a single point of failure for the whole site.

Exam trap. A collapsed core merges distribution and core, not access and distribution. When a question describes a small site where one pair of switches provides gateways and connects to the WAN, the answer is two-tier/collapsed core.

Building five broke the budget

A company grew from one office building to four, each with its own collapsed-core pair, and connected the pairs in a full mesh. When the fifth building was approved, the plan needed eight new fibre runs, four new routing adjacencies per switch, and changes on every existing pair. The architect proposed a core pair in the main data room instead. New buildings now connect with four links to the core, existing pairs were migrated one at a time, and OSPF summarisation at each building cut the routing table in the core from hundreds of routes to a handful.

Lesson: a collapsed core is perfect until you have several of them; then a dedicated core restores modularity.

"When would you choose a two-tier collapsed core instead of a three-tier design?"

For a single building or a small-to-medium site where one pair of switches can provide the gateways and connect to the WAN, data centre and internet edge. It saves cost and cabling while keeping redundancy. Move to three-tier when there are several distribution blocks or buildings, because a core avoids the full mesh, gives one meeting point for the other modules and keeps each block's failures local. Mention the n(n-1)/2 link growth.

Key takeaways

  • Access connects devices, distribution is the L2/L3 and policy boundary, core is fast Layer 3 transport.
  • Keep the core simple: no users, no heavy policy, fast convergence.
  • A core avoids the n(n-1)/2 full mesh once there are several distribution blocks.
  • Collapsed core (two-tier) merges distribution and core; ideal for one building or a medium site.
  • Summarise each block at the distribution layer to contain routing churn.
04

Where Layer 2 ends: L2 access, routed access and virtual switching

Every campus design must answer one question early: where does Layer 2 stop and Layer 3 begin? Think of a housing society. Inside the gate, neighbours walk freely between houses (Layer 2, same VLAN). At the gate, a guard checks where you are going and sends you on the right road (Layer 3, routing). You can put the gate at the society entrance, or at each building's door. Both work, but they change how far a problem spreads and how far a resident can wander. This chapter compares the three common answers.

Option 1: Layer 2 access (the looped triangle)

Access switches are pure Layer 2. Every access switch has an uplink trunk to each distribution switch, and the two distribution switches are linked by a Layer 2 trunk (usually an EtherChannel). The VLAN's gateway lives on distribution SVIs protected by an FHRP. Each access switch plus the distribution pair forms a triangle, which is a physical loop, so spanning tree blocks one uplink per VLAN.

  • Pros: a VLAN can span several access switches; simple access configuration; works with any access switch.
  • Cons: STP blocks half the uplinks (unless you load-share per VLAN); failover depends on STP and FHRP timers; the VLAN is a larger failure domain; the STP root and FHRP active must be aligned (next chapter).

This is the design used in the lab: AS1 has Gi0/0 to DS1 and Gi0/1 to DS2, and DS1 and DS2 are linked by Gi0/6 and Gi0/7.

Two variants avoid the loop. In a loop-free U, the two access switches of a pair are joined by a Layer 2 link and the distribution link is routed, so VLANs span only that access pair and nothing blocks. In a loop-free inverted U, the distribution link is Layer 2 and each access switch has only one uplink, so nothing blocks but a single uplink failure isolates that switch. Both are niche; the looped triangle and routed access are what you will meet most.

Option 2: Routed access

The Layer 3 boundary moves down to the access switch. Each uplink is a routed point-to-point link, the access switch is the default gateway for its own VLANs, and it runs OSPF or EIGRP towards the distribution pair.

  • Pros: no STP blocking on uplinks, both uplinks forward with equal-cost multipath (ECMP); no FHRP needed because the access switch is the gateway; routing convergence is fast and deterministic; broadcasts and loops stay on one switch, so failure domains are tiny.
  • Cons: a VLAN cannot span access switches; more IP planning (a /30 or /31 per uplink, subnets per access switch); the access switches need Layer 3 features and licensing.
! Routed access switch ACC-R1 (example)
ip routing
interface GigabitEthernet1/0/49
 description UPLINK-DIST1
 no switchport
 ip address 10.0.101.1 255.255.255.252
 ip ospf network point-to-point
interface GigabitEthernet1/0/50
 description UPLINK-DIST2
 no switchport
 ip address 10.0.102.1 255.255.255.252
 ip ospf network point-to-point
interface Vlan110
 ip address 10.1.110.1 255.255.255.0
router ospf 1
 router-id 10.255.0.21
 passive-interface default
 no passive-interface GigabitEthernet1/0/49
 no passive-interface GigabitEthernet1/0/50
 network 10.0.101.0 0.0.0.3 area 1
 network 10.0.102.0 0.0.0.3 area 1
 network 10.1.110.0 0.0.0.255 area 1
ACC-R1# show ip route ospf
O*IA  0.0.0.0/0 [110/2] via 10.0.101.2, 00:02:10, GigabitEthernet1/0/49
                [110/2] via 10.0.102.2, 00:02:10, GigabitEthernet1/0/50

Two equal-cost default routes: both uplinks carry traffic. Making the access area totally stubby (or the access switch an EIGRP stub) keeps its routing table tiny and stops it from ever being used as a transit path.

Option 3: Virtual switching (StackWise Virtual)

The two distribution switches are combined into one logical switch with one control plane (StackWise Virtual on Catalyst 9000, covered in the high-availability chapter). The access switch bundles its two uplinks into a single multichassis EtherChannel (MEC) with LACP. To spanning tree that is one port, so nothing blocks. The gateway is a single SVI on the logical switch, so no FHRP is needed, and VLANs can still span access switches.

L2 access DS1 DS2 AS1 BLK STP + FHRP Routed access DS1 DS2 ACC L3 uplinks, ECMP Virtual switch one logical switch AS MEC, no blocking

L2 access blocks one uplink; routed access and virtual switching use both.

QuestionL2 accessRouted accessVirtual switching
Uplinks forwardingOne per VLANBoth (ECMP)Both (MEC)
STP roleActive, blockingOnly on edge portsLoop protection only
FHRP neededYesNoNo
VLAN spans access switchesYesNoYes
Failure domainWhole VLANOne access switchWhole VLAN

Worked example. A floor has four access switches with two 10 Gbps uplinks each. With L2 access and no per-VLAN load sharing, each switch really has 10 Gbps upward (80 Gbps of uplinks installed, 40 Gbps used). With routed access or MEC, all 80 Gbps forward. Same cables, twice the usable bandwidth.

Common mistakes. Choosing routed access and then discovering a legacy application that needs one subnet on several floors. Mixing models inside one block without a plan. Forgetting that routed access still needs STP on edge ports with BPDU guard, because users can still plug in switches.

Exam trap. In routed access there is no FHRP and no STP blocking on uplinks; the trade-off is that VLANs cannot span access switches. A question asking how to use both uplinks without spanning VLANs across switches points to routed access; one that must keep VLANs spanning points to virtual switching with MEC.

The lab that needed one subnet

A university converted its science block to routed access. Convergence after an uplink failure dropped from several seconds to well under one second and the uplink utilisation doubled. Two weeks later the physics lab complained: an old instrument-control application discovered its devices with broadcasts and needed all benches on one subnet, but the benches were on two floors with different access switches. The engineer kept routed access for the building, placed the instrument benches on one dedicated access switch per lab, and documented the constraint for future moves.

Lesson: check application Layer 2 requirements before moving the L3 boundary to the access layer.

"Compare Layer 2 access with routed access. Which would you choose?"

Explain that L2 access keeps VLANs spanning access switches but relies on STP (one uplink blocked) and an FHRP aligned with the STP root. Routed access moves the gateway to the access switch, uses both uplinks with ECMP, converges quickly and keeps failure domains tiny, but a VLAN cannot leave its access switch. Choose routed access for new builds without VLAN-spanning needs; choose L2 access or virtual switching with MEC when VLANs must span. Mentioning SD-Access, which gives routed access plus stretched subnets through an overlay, is a strong finish.

Key takeaways

  • L2 access: VLANs can span, but STP blocks an uplink and an FHRP is required.
  • Routed access: the access switch is the gateway, both uplinks forward with ECMP, no FHRP, tiny failure domains.
  • Virtual switching (StackWise Virtual + MEC): no blocking, no FHRP, VLANs can still span.
  • Loop-free U and inverted U remove the loop but have scaling or isolation limits.
  • The lab uses L2 access in a looped triangle, so STP and HSRP must be aligned.
05

Aligning STP, HSRP and EtherChannel: build the redundant block

Imagine an office with two reception desks, one at each end of a long corridor. Visitors always enter through the east door, but the only receptionist on duty sits at the west desk. Every visitor walks the whole corridor, then walks back. Nothing is broken, yet everything is slower and the corridor is crowded. That is what happens in an L2 access design when spanning tree sends frames up one uplink while the HSRP active gateway sits on the other distribution switch. This chapter shows how to align them and then walks through the first lab, Build a redundant collapsed-core block.

Why alignment matters

In the looped triangle, spanning tree picks one root bridge per VLAN and blocks the access uplink that leads away from it. HSRP independently picks one active router per VLAN. If DS2 is the STP root for VLAN 10 but DS1 is HSRP active, PC1's frames go AS1 → DS2 (the forwarding uplink), cross the DS1–DS2 link and only then reach the gateway on DS1. The inter-switch link carries all of VLAN 10's routed traffic and every packet takes an extra hop.

Misaligned (VLAN 10) DS1HSRP active DS2STP root AS1 BLK AS1 to DS2 to DS1: extra hop Aligned (VLAN 10) DS1root + active DS2standby AS1 BLK AS1 straight to the gateway

Make the same switch STP root and HSRP active for each VLAN.

The design rules

  1. Per VLAN, one owner. The same distribution switch is STP root and HSRP active; the other is secondary root and HSRP standby.
  2. Load-share by VLAN. DS1 owns VLAN 10, DS2 owns VLAN 20, so both uplinks carry traffic.
  3. Preempt on both switches. Without standby preempt, a recovered owner stays standby and alignment is lost after the first failure. In production add standby 10 preempt delay minimum 60 so the switch waits for routing to converge before taking over.
  4. Bundle the inter-switch link with LACP. Two links in one Port-channel look like one link to STP, so a member failure causes no topology change, and LACP detects mis-cabling that mode on would hide.
  5. Both distribution switches in the routing protocol, so the WAN router has two equal-cost paths back to every user subnet.

Lab walkthrough: build a redundant collapsed-core block

Start state: addressing, SVIs and HSRP virtual IPs (10.1.10.1 and 10.1.20.1) exist; nothing is aligned, the DS1–DS2 links are separate, and DS2 is missing from OSPF.

Task 1: LACP bundle between DS1 and DS2

! On DS1 and on DS2
interface range g0/6 - 7
 channel-group 1 mode active
interface port-channel 1
 switchport mode trunk
 switchport trunk allowed vlan 10,20
DS1# show etherchannel summary
Group  Port-channel  Protocol    Ports
------+-------------+-----------+-----------------------------------------------
1      Po1(SU)         LACP      Gi0/6(P)    Gi0/7(P)

SU means Layer 2 and in use; P means bundled. Anything else (I stand-alone, s suspended, D down) means the bundle is not healthy.

Task 2: STP roots per VLAN

DS1(config)# spanning-tree vlan 10 root primary
DS1(config)# spanning-tree vlan 20 root secondary
DS2(config)# spanning-tree vlan 20 root primary
DS2(config)# spanning-tree vlan 10 root secondary

The root primary macro writes priority 24576, or 4096 less than the current root if that is already lower; root secondary writes 28672. It is a one-time calculation saved as a number in the configuration.

AS1# show spanning-tree vlan 10
  Root ID    Priority    24586
             Cost        4
             Port        1 (GigabitEthernet0/0)
Interface           Role Sts Cost      Prio.Nbr Type
Gi0/0               Root FWD 4         128.1    P2p
Gi0/1               Altn BLK 4         128.2    P2p

24586 is 24576 plus the VLAN number (extended system ID). Gi0/0 towards DS1 is the root port; for VLAN 20 the roles reverse.

Task 3: HSRP owners with preempt

DS1(config)# interface vlan 10
DS1(config-if)# standby 10 priority 110
DS1(config-if)# standby 10 preempt
DS1(config)# interface vlan 20
DS1(config-if)# standby 20 preempt
! DS2 mirrors it: priority 110 + preempt for group 20, preempt for group 10
DS1# show standby brief
Interface   Grp  Pri P State   Active          Standby         Virtual IP
Vl10        10   110 P Active  local           10.1.10.3       10.1.10.1
Vl20        20   100 P Standby 10.1.20.3       local           10.1.20.1

Task 4: DS2 into OSPF

DS2(config)# router ospf 1
DS2(config-router)# router-id 1.1.1.3
DS2(config-router)# network 10.0.2.0 0.0.0.3 area 0
DS2(config-router)# network 10.1.0.0 0.0.255.255 area 0
R1# show ip route ospf
O        10.1.20.0/24 [110/2] via 10.0.2.2, 00:00:21, GigabitEthernet0/1
                      [110/2] via 10.0.1.2, 00:00:21, GigabitEthernet0/0

Two equal-cost paths: return traffic can reach the users through either distribution switch.

Tasks 5 and 6: failover test and the virtual MAC

Shut interface vlan 10 on DS1. DS2 logs %HSRP-5-STATECHANGE: Vlan10 Grp 10 state Standby -> Active and PC1 still pings 172.16.0.10. Bring the SVI back with no shutdown; after the hold and preempt timers DS1 logs Speak -> Active or Standby -> Active and owns VLAN 10 again. Finally, show standby vlan 10 shows the virtual MAC.

DS1# show standby vlan 10
Vlan10 - Group 10
  State is Active
  Virtual IP address is 10.1.10.1
  Active virtual MAC address is 0000.0c07.ac0a (MAC In Use)
  Preemption enabled
  Priority 110 (configured 110)

Worked example. HSRP version 1 builds its virtual MAC as 0000.0c07.acXX, where XX is the group number in hex. Group 10 is 0x0a, so the answer is 0000.0c07.ac0a. Group 20 would be 0000.0c07.ac14. Because the MAC moves with the active role, PC1's ARP entry for 10.1.10.1 never changes during failover.

Common mistakes. Configuring priority 110 but forgetting preempt, so alignment breaks after the first reboot. Using channel-group 1 mode on on one side and active on the other. Setting STP roots with the macro, then later adding a switch with a lower priority that silently steals the root.

Exam trap. STP root and FHRP active are chosen by different, independent protocols. Nothing aligns them automatically. If a question shows traffic crossing the inter-switch link to reach the gateway, the answer is to align root and active for that VLAN, not to change HSRP timers.

The inter-switch link that was always at 90 percent

A retail head office saw its 2 x 1 Gbps distribution interconnect run near 90 percent every afternoon, while one access uplink per switch sat idle. show spanning-tree root showed DS-B as root for every VLAN (default priorities, lowest MAC), while HSRP was active on DS-A for all VLANs. Every routed packet crossed the interconnect. The engineer set root primary/secondary per VLAN to match the HSRP design, split odd VLANs to DS-A and even VLANs to DS-B with matching HSRP priorities and preempt, and interconnect load fell below 20 percent.

Lesson: default STP elects the switch with the lowest MAC, which is almost never the switch you intended.

"Why should the STP root bridge and the HSRP active router be on the same switch?"

Explain that STP decides which uplink forwards and HSRP decides where the gateway is. If they differ, traffic climbs to the root, crosses the inter-switch link to the gateway, and the link becomes a bottleneck with an extra hop. Align them per VLAN, load-share VLANs across the pair, use preempt so alignment survives failures, and bundle the interconnect with LACP. Mention that routed access or StackWise Virtual removes the problem entirely.

Key takeaways

  • Per VLAN, make one switch both STP root and HSRP active; load-share VLANs across the pair.
  • Use preempt on both switches; add a preempt delay in production.
  • Bundle the inter-switch link with LACP active on both sides; verify Po1(SU) and members (P).
  • Put both distribution switches in OSPF so R1 has two equal-cost return paths.
  • HSRPv1 virtual MAC = 0000.0c07.acXX; group 10 = 0000.0c07.ac0a.
06

Device-level high availability: SSO, NSF, NSR, graceful restart and StackWise

A plane has two pilots. If the captain falls ill, the first officer is already in the cockpit, already knows the altitude, speed and route, and simply takes the controls. The engines never stop. Air-traffic control keeps talking to the same flight number. Network devices have exactly this idea inside one chassis: a second supervisor (the pilot) that knows the full state, hardware that keeps forwarding (the engines), and neighbours that are told not to panic (air-traffic control). This chapter covers the tools that make a single box, or a pair acting as one box, survive its own failures.

Layers of redundancy

So far you have protected against link failures (EtherChannel, dual uplinks), gateway failures (HSRP, VRRP, GLBP) and path failures (OSPF ECMP). Device-level HA protects against failures inside a device: a supervisor engine crash, a software fault, a stack member dying. It matters most where a single box is a single point of failure, such as a modular core chassis or an access switch with users attached.

Supervisor redundancy modes: RPR, RPR+ and SSO

A modular switch (for example a Catalyst 9400 or 9600) can hold two supervisors: one active, one standby. How ready the standby is depends on the redundancy mode:

ModeStandby stateSwitchover impact
RPRPartially booted, not synchronisedLine cards reset; minutes of outage
RPR+Fully booted, configuration synchronisedLinks stay up but state is rebuilt; tens of seconds
SSOFully booted, configuration and state synchronised (STANDBY HOT)Layer 2 state and links kept; typically around a second or less

SSO (stateful switchover) continuously copies the running configuration and protocol state (interface state, MAC tables, STP, LACP, and more) to the standby. On failure the standby takes over with no link flaps and no spanning-tree reconvergence.

redundancy
 mode sso
CORE1# show redundancy states
       my state = 13 -ACTIVE
     peer state = 8  -STANDBY HOT
           Mode = Duplex
Redundancy Mode (Operational) = sso
Redundancy Mode (Configured)  = sso

NSF: keep forwarding while routing restarts

SSO synchronises Layer 2 state, but routing protocol adjacencies traditionally restart on the new supervisor. NSF (non-stop forwarding) lets the hardware keep forwarding packets using the last known FIB while the new active supervisor rebuilds its routing table. To stop neighbours from tearing down adjacencies and rerouting around the switch, the routing protocol uses graceful restart (GR): the restarting router signals that it is restarting, and NSF-aware (helper) neighbours keep its routes and adjacency for a grace period while databases resynchronise.

  • NSF-capable: the device that can restart gracefully (dual supervisors with SSO).
  • NSF-aware / helper: a neighbour that understands graceful restart and keeps forwarding towards the restarting device. Most modern IOS XE devices are NSF-aware by default.
router ospf 1
 nsf                        ! Cisco NSF; "nsf ietf" for the RFC version
router bgp 65001
 bgp graceful-restart

NSR: neighbours never notice

NSR (non-stop routing) goes one step further. The standby supervisor keeps a live copy of the routing protocol state itself: adjacencies, databases, BGP sessions and TCP state. After a switchover the new active continues the conversation as if nothing happened, so no helper is needed and neighbours never see a restart. NSR is useful when neighbours are not NSF-aware or belong to another organisation, such as a service provider.

router ospf 1
 nsr
Active sup fails SSO: standby active GR: helpers wait Routes resynced NSF: hardware keeps forwarding with the last FIB With NSR, the GR step is not needed: state is already on the standby

SSO keeps Layer 2 state, NSF keeps packets moving, graceful restart keeps neighbours from rerouting.

StackWise: many switches, one brain

StackWise (for example StackWise-480 on Catalyst 9300) joins up to eight access switches with stack cables into one logical switch with one management IP and one configuration. One member is active, one is standby (SSO between them), the rest are members. You can build a cross-stack EtherChannel with uplinks on different members, so losing a member does not isolate the stack. The member with the highest stack priority (1 to 15) becomes active.

ACC-STACK# show switch
Switch#   Role    Mac Address     Priority Version  State
-------------------------------------------------------------
*1       Active   00aa.bb00.1100     15     V01     Ready
 2       Standby  00aa.bb00.2200     14     V01     Ready
 3       Member   00aa.bb00.3300     1      V01     Ready

StackWise Virtual: two chassis, one logical switch

StackWise Virtual (SVL) joins two distribution or core switches (Catalyst 9400, 9500 or 9600) over a StackWise Virtual link of one or more Ethernet ports. The pair has one control plane (active plus hot standby with SSO) and two data planes, both forwarding. Access switches connect with a multichassis EtherChannel, so no STP blocking and no FHRP are needed. A separate dual-active detection (DAD) link prevents both switches from becoming active if the SVL fails.

stackwise-virtual
 domain 10
interface range TenGigabitEthernet1/0/47 - 48
 stackwise-virtual link 1
interface TenGigabitEthernet1/0/46
 stackwise-virtual dual-active-detection
! save and reload both switches to form the pair; verify with show stackwise-virtual

Worked example. A core chassis handles 40 Gbps. Without SSO, a supervisor crash takes about three minutes to recover: 40 Gbps x 180 s of dropped traffic and every OSPF neighbour reroutes twice. With SSO and NSF, Layer 2 state survives, the ASICs keep forwarding with the last FIB, and NSF-aware neighbours keep their routes; the users see at most a sub-second blip. Add ISSU (in-service software upgrade, which relies on SSO) and even planned upgrades avoid an outage window.

Common mistakes. Buying dual supervisors but leaving the mode at RPR, or running different software versions that force a lower mode. Combining NSF with very aggressive hello timers or BFD, so neighbours declare the router dead before graceful restart can begin. Forgetting the DAD link on StackWise Virtual.

Exam trap. SSO keeps state on the standby; NSF keeps forwarding during the switchover and needs graceful restart with NSF-aware helpers; NSR keeps routing state on the standby so no helper is needed. HSRP/VRRP protect the gateway between two boxes; SSO protects inside one box.

The switchover that still dropped routes

A bank enabled SSO and NSF on its new core chassis and tested a supervisor failover in a change window. Layer 2 stayed up, but the two old WAN routers dropped their OSPF adjacencies and rerouted all branch traffic for about 40 seconds. show ip ospf neighbor history and the logs showed the WAN routers ran old software that was not NSF-aware, so they ignored the grace signal. The team upgraded the WAN routers to NSF-aware software and enabled NSR on the core for the BGP sessions with the service provider, whose routers they did not control. The next test showed no routing change at all.

Lesson: NSF works only when neighbours help; NSR works even when they cannot.

"Explain the difference between SSO, NSF and NSR."

SSO synchronises configuration and state to a hot standby supervisor so a switchover keeps links and Layer 2 state. NSF lets the data plane keep forwarding with the last FIB during that switchover, and uses graceful restart so NSF-aware neighbours keep adjacencies and routes while the new active rebuilds. NSR keeps the routing protocol state itself on the standby, so neighbours do not need to help and never see a restart. Add StackWise and StackWise Virtual as the switch-level equivalents that use SSO between members.

Key takeaways

  • SSO: standby supervisor is STANDBY HOT with synchronised config and state; configure with redundancy / mode sso.
  • NSF: forwarding continues on the last FIB; needs graceful restart and NSF-aware helper neighbours.
  • NSR: routing state lives on the standby too; no helper required.
  • StackWise joins up to eight access switches into one; StackWise Virtual joins two chassis with SVL and DAD.
  • With StackWise Virtual and MEC, access uplinks do not block and no FHRP is needed.
07

Fabric designs, spine-leaf and the WAN and branch edge

A metro rail network with a ring line and radial lines has a nice property: from any station you can reach any other station with at most one change, and if one ring segment closes, trains simply use the other direction. Compare that with an old tree-shaped bus network where every trip goes to the central depot first. Data-centre and modern campus fabrics are the metro: every edge switch is exactly one hop from every other edge switch through a set of equal paths. This chapter covers fabric designs (ENCOR 1.1 lists "fabric" next to 2-tier and 3-tier) and the WAN and branch designs that connect an enterprise's sites.

Why the data centre moved to spine-leaf

Traditional three-tier data centres were built for north-south traffic: clients outside talking to servers inside. Modern applications are split into many services that talk to each other, so most traffic is now east-west, server to server. In a three-tier design with spanning tree, east-west traffic often climbs to the aggregation or core layer and back down, half the links are blocked, and latency differs depending on where two servers sit.

A spine-leaf (Clos) fabric fixes this with two strict rules:

  • Every leaf (top-of-rack switch where servers, firewalls and routers attach) connects to every spine.
  • Spines never connect to each other, and leaves never connect to each other (apart from special pairs such as vPC or MLAG peers).

So any server reaches any other server on another leaf in exactly leaf → spine → leaf: the same number of hops and predictable latency. The links are routed and every spine offers an equal-cost path, so traffic is spread with ECMP and nothing is blocked.

Spine 1 Spine 2 Leaf 1 Leaf 2 Leaf 3 Border leaf Every leaf to every spine; routed links; ECMP across spines

Any leaf reaches any other leaf through any spine in two hops. The border leaf connects to WAN, internet and campus.

Scaling and oversubscription in a fabric

You scale a fabric out, not up. Need more server ports? Add a leaf. Need more bandwidth between leaves? Add a spine and one more uplink on every leaf. The limits are simple: the number of spines is capped by the uplink ports on each leaf, and the number of leaves is capped by the ports on each spine.

Worked example. Each leaf has 48 server ports at 25 Gbps (1,200 Gbps down) and uplinks at 100 Gbps. With 4 spines (4 x 100 = 400 Gbps up) the leaf is 3:1 oversubscribed. With 6 spines (600 Gbps up) it is 2:1. If each spine has 32 ports of 100 Gbps, the fabric can hold up to 32 leaves, or 32 x 48 = 1,536 server ports at 25 Gbps.

LEAF1# show ip route 10.20.3.0
Routing entry for 10.20.3.0/24
  Known via "ospf 1", distance 110, metric 3, type intra area
  Routing Descriptor Blocks:
  * 10.255.1.1, from 10.255.0.3, via HundredGigE1/0/49
    10.255.2.1, from 10.255.0.3, via HundredGigE1/0/50

Two spines, two equal-cost next hops towards LEAF3's subnet.

Underlay and overlay

A fabric has two layers. The underlay is the routed physical network (OSPF, IS-IS or eBGP between leaves and spines) that simply makes every switch's loopback reachable. The overlay builds tenant networks on top, typically VXLAN tunnels between leaves, so a VLAN or subnet can appear on any leaf without stretching spanning tree. In data centres the overlay control plane is usually BGP EVPN.

The same idea came to the campus as SD-Access: a routed underlay (often IS-IS built by LAN automation), LISP as the control plane that tracks where each endpoint is, VXLAN as the data plane, and Scalable Group Tags as the policy plane, managed by Catalyst Center. Its roles are edge nodes (where users attach), border nodes (exits to other networks) and control-plane nodes (the LISP map server). You get routed access for resiliency, yet a subnet can still exist on every edge node. The SD-WAN/SD-Access and LISP/VXLAN modules go deep; here you only need to recognise fabric as a design option.

WAN and branch designs

Branches range from a two-person sales office to a thousand-person regional hub, so there is no single branch design. Choose by size and by how much downtime the branch can tolerate:

Branch sizeTypical designWeak point
SmallOne router or SD-WAN edge, one or two WAN links (for example broadband plus 4G/5G), access switch with integrated servicesSingle device
MediumTwo routers or edges, two transports (MPLS plus internet), FHRP or routing towards a switch stackSingle switch stack
LargeDual edges, dual transports, a small collapsed core or full access-distribution blockCost and complexity

Transports include MPLS L3VPN (provider-managed, SLAs), dedicated or broadband internet with IPsec or DMVPN, and cellular backup. SD-WAN overlays all of them, measures loss, latency and jitter on each, and steers each application to the best path. The internet edge module usually has redundant firewalls, two ISPs and BGP for multihoming.

Common mistakes. Cabling a leaf to only some spines, which breaks the equal-path promise. Connecting two spines together "for redundancy". Giving a branch two routers that both use the same ISP and the same last-mile cable.

Exam trap. In spine-leaf, the answer to "how many hops between servers on different leaves" is always leaf-spine-leaf, and the answer to "how do you add bandwidth" is add spines. In SD-Access, LISP is the control plane, VXLAN the data plane and SGT the policy plane.

Backups at 2 AM

A company's new data centre used four leaves and two spines, each leaf with 48 x 25 Gbps server ports and two 100 Gbps uplinks: 6:1 oversubscription. Every night at 2 AM the backup job pulled data from all servers to the storage leaf, and application monitoring showed drops and retransmissions. Interface counters on the leaf uplinks showed output drops, while spine CPU and links looked healthy. The fix needed no redesign: two more spines were added and every leaf got two more uplinks, giving 3:1 and four ECMP paths. The backup window shrank by half.

Lesson: spine-leaf scales out; when leaves run hot on uplinks, add spines rather than bigger boxes.

"Why do data centres use spine-leaf instead of three-tier?"

Because traffic became mostly east-west. Spine-leaf gives every pair of leaves the same two-hop path, uses routed links with ECMP so no link is blocked, gives predictable latency and scales out by adding spines for bandwidth or leaves for ports. Mention the underlay/overlay split (routed underlay, VXLAN overlay with EVPN) and that SD-Access brings the same model to the campus with LISP and VXLAN.

Key takeaways

  • Spine-leaf: every leaf to every spine; no spine-spine or leaf-leaf links; always leaf-spine-leaf.
  • Routed links with ECMP; scale out by adding spines (bandwidth) or leaves (ports).
  • Fabrics use a routed underlay and an overlay (VXLAN with EVPN in the DC, LISP plus VXLAN in SD-Access).
  • Branch design depends on size and tolerance for downtime: single, dual-edge, or small campus.
  • SD-WAN overlays MPLS, internet and cellular and picks a path per application.
08

On-premises versus cloud infrastructure

Owning a car and using a ride-hailing app both get you to work. Owning costs a lot up front, but the car is always in your driveway, you can modify it, and daily trips are cheap once it is paid for. Ride-hailing costs nothing up front, scales instantly when your whole family needs to travel, and someone else handles servicing, but you pay every trip, depend on the app and the roads, and cannot choose the engine. On-premises infrastructure is the owned car; public cloud is ride-hailing. ENCOR expects you to explain the difference and know how a network engineer connects the two.

Deployment models

On-premises
The organisation owns and runs the hardware in its own data centre or a rented colocation space. Full control, full responsibility.
Private cloud
On-premises (or dedicated hosted) infrastructure run with cloud-style self-service, automation and virtualisation for one organisation.
Public cloud
A provider's shared infrastructure, rented on demand and billed by use, reached over the internet or private interconnects.
Hybrid cloud
On-premises or private cloud connected to public cloud, with workloads placed where they fit best. This is what most enterprises actually run.
Multicloud
Using two or more public cloud providers, for resilience, features or commercial reasons.

Service models describe how much the provider manages: IaaS (you get virtual machines, storage and virtual networks and manage the OS and everything above), PaaS (you deploy code or containers onto a managed platform) and SaaS (you just use the application, for example hosted email or CRM). Networking itself can also be consumed as a service, such as cloud-managed switches and access points whose management plane lives in the provider's cloud.

Comparing on-premises and cloud

FactorOn-premisesPublic cloud
Cost modelCapEx: buy hardware up front, depreciate over yearsOpEx: pay per hour, per GB and per request
ScalingSize for peak; adding capacity takes weeksElastic; scale in minutes, scale down to save money
ControlFull control of hardware, software versions, topologyLimited to what the provider exposes
Latency and localityClose to campus users and factory systemsDepends on WAN or internet path to the region
ComplianceData stays where you put itChoose regions carefully; shared-responsibility model
OperationsYour team patches, replaces and cools everythingProvider runs facilities, hardware and hypervisor

The shared-responsibility model is key. With IaaS, the provider secures the buildings, hardware and hypervisor; you still own the guest operating systems, applications, identities, data and your virtual network rules such as security groups and route tables. Moving to the cloud does not move security responsibility away from you.

Campus / DCEDGE1, EDGE2 Cloud regionVPC / VNet SaaS apps IPsec over internet Private interconnect Direct internet

Hybrid connectivity: VPN tunnels, dedicated interconnects and direct internet access to SaaS.

How the network connects to the cloud

  • Site-to-site IPsec VPN over the internet: quick and cheap; use two tunnels to two provider gateways and run BGP over them.
  • Private or dedicated interconnect: a physical circuit from your router, usually in a colocation facility, into the provider's network. Predictable latency and bandwidth; build two in different locations for resilience.
  • SD-WAN cloud on-ramp: virtual SD-WAN routers (for example Catalyst 8000V) inside the cloud region become part of the SD-WAN fabric, and branches can reach SaaS directly instead of hairpinning through the head office.
  • Virtual network functions in the cloud: virtual routers, firewalls and even wireless controllers (Catalyst 9800-CL) run as cloud instances.
EDGE1# show ip bgp summary
Neighbor        V           AS MsgRcvd MsgSent   TblVer  InQ OutQ Up/Down  State/PfxRcd
169.254.10.1    4        64512     220     218       14    0    0 03:12:41        2
169.254.11.1    4        64512     219     218       14    0    0 03:12:39        2

Two VPN tunnels to the cloud gateway, each with an eBGP session learning the two cloud prefixes. If one tunnel fails, BGP converges to the other. Link-local 169.254.x.x addresses inside the tunnels are common with cloud VPN gateways.

Worked example. An online shop needs 20 web servers all year and 100 during a three-week festival sale. On-premises, it must buy and power 100 servers that sit 80 percent idle for 49 weeks. In the cloud it runs 20 instances and scales to 100 for three weeks, paying only for what it uses. Its payroll database, however, is steady, small and bound by data-residency rules, so it stays on-premises. The result is hybrid, and the network must now provide resilient, low-latency connectivity between the two.

Common mistakes. Moving applications to the cloud but leaving one VPN tunnel from one site as the only path. Forgetting that data leaving the cloud (egress) is usually charged. Assuming the provider patches your virtual machines. Ignoring latency for chatty applications split between cloud and on-premises.

Exam trap. CapEx versus OpEx, elasticity, and shared responsibility are the most tested points. Questions often give a workload (steady, regulated, latency-critical versus spiky, global, fast-changing) and ask where it belongs. "Hybrid" is right only when the scenario really needs both.

The ERP that lived behind one tunnel

A distributor moved its ERP system to a public cloud region. All 30 branches reached it through the head office, which had one ISP and one IPsec tunnel to the cloud. When the head-office ISP failed for four hours, the ERP was down for every branch even though the cloud itself was healthy. The post-incident review found a single failure domain: one ISP, one router, one tunnel. The fix: a second ISP and edge router at head office with two tunnels and BGP to the cloud, and SD-WAN at the branches so they could reach the cloud directly over their own internet links if the head office was unreachable.

Lesson: moving to the cloud moves the critical path onto the WAN; design that path with the same redundancy as the data centre.

"How do you decide whether a workload should run on-premises or in the cloud?"

Ask about the load pattern (steady versus spiky), latency and locality needs, data residency and compliance, integration with other systems, team skills and total cost over several years including egress. Steady, regulated or latency-critical workloads often stay on-premises; spiky, global or fast-changing ones fit the cloud. Then explain how you would connect them: redundant VPN or interconnects with BGP, SD-WAN for branches, and the shared-responsibility model for security.

Key takeaways

  • Models: on-premises, private, public, hybrid and multicloud; services: IaaS, PaaS, SaaS.
  • On-premises is CapEx with full control; cloud is OpEx with elasticity and less control.
  • Shared responsibility: the provider secures the platform, you secure what you put on it.
  • Connect with redundant IPsec tunnels or private interconnects running BGP, or through SD-WAN.
  • Most enterprises are hybrid; the WAN path to the cloud becomes a critical failure domain.
09

WLAN deployment models: centralised, distributed, cloud and remote branch

Think of a restaurant chain. A tiny café can have its own cook who decides everything alone. A big mall food court sends every order to one central kitchen. A chain of small outlets might have a head chef who writes the menu centrally, while each outlet cooks locally so it keeps serving even when the phone line to head office is down. And some chains let a remote company run their ordering system from an app. Wireless networks are deployed in exactly these ways. ENCOR 1.2 asks you to compare them at a design level; the enterprise wireless module later covers RF, roaming and controller configuration.

Controller-less versus controller-based

The first split is whether there is a controller at all.

  • Controller-less (autonomous): each access point is a complete, independent device with its own configuration, SSIDs, security and RF settings. Fine for one or two APs; painful beyond that, because every change is repeated on every AP and features such as coordinated radio resource management and fast roaming are limited. Some material also calls an AP running an embedded wireless controller (one AP acting as controller for its neighbours) controller-less, because there is no separate controller box.
  • Controller-based (lightweight APs): APs are managed by a wireless LAN controller (WLC) using CAPWAP. This is the split-MAC model: real-time functions (beacons, acknowledgements, encryption) stay on the AP, while management, authentication, RF management and roaming decisions move to the controller. CAPWAP uses UDP 5246 for control and UDP 5247 for data.

Where the controller and the data path live

ModelControllerClient data pathBest fit
CentralisedWLC (appliance or pair) in the data centre or services blockTunnelled in CAPWAP from AP (local mode) to the WLC, then switched thereCampus with good LAN bandwidth; central policy and easy roaming
DistributedController function close to the edge: embedded on Catalyst 9000 switches, or central WLC with locally switched APsSwitched at the access layer (in SD-Access, VXLAN from AP to fabric edge)Large campuses and fabric designs; avoids hairpinning all traffic to one place
CloudManagement in the cloud (cloud-managed APs) or a virtual controller such as Catalyst 9800-CL in a public or private cloudLocal at the site; only management goes to the cloudMany small sites, lean IT teams, fast rollout
Remote branchCentral WLC over the WAN with FlexConnect APs, or an embedded controller on a branch APLocally switched onto the branch VLAN; can keep working if the WAN failsBranches and retail stores
WLC in DC Campus AP (local) Branch AP (Flex) control + data control only Branch VLAN (local) All client traffic to WLC

Local mode tunnels everything to the controller; FlexConnect keeps data at the branch.

FlexConnect for branches

A FlexConnect AP keeps its CAPWAP control connection to a central WLC but switches client traffic onto a local VLAN. If the WAN fails, it enters standalone mode: already-connected clients stay connected, and with local authentication configured new clients can still join. On a Catalyst 9800 controller, APs become FlexConnect through a site tag that is not a local site:

wireless profile flex BR1-FLEX
 native-vlan-id 30
wireless tag site BR1-SITE
 flex-profile BR1-FLEX
 no local-site
WLC1# show ap summary
Number of APs: 2
AP Name     Slots  AP Model     Ethernet MAC    Radio MAC       Location  Country  IP Address   State
AP-HQ-01    2      C9120AXI-D   00aa.bb11.0001  00aa.bb22.0001  HQ-FL1    IN       10.1.30.21   Registered
AP-BR1-01   2      C9120AXI-D   00aa.bb11.0002  00aa.bb22.0002  Branch1   IN       10.50.30.21  Registered

Design factors beyond the model

  • Client density: auditoriums, classrooms and stadiums are designed for capacity, not coverage. Use more APs at lower power, prefer 5 GHz and 6 GHz, use narrower channels to get more of them, and plan the number of clients per radio.
  • Location services: accurate location needs every point heard by at least three APs at about -75 dBm or better, with APs placed along the perimeter and staggered, not just down the corridor.
  • Controller redundancy: N+1 (one spare WLC for many), N+N (two sets sharing load), N+N+1, or an HA SSO pair in which a standby controller holds AP and client state, so a failure does not force APs to rejoin.
  • Latency and bandwidth between AP and controller: local mode depends on the LAN or WAN path for every packet, which is why remote sites use FlexConnect.

Worked example. A retail chain has 400 stores with 3 APs each and a 20 Mbps WAN link per store. In local mode, a video stream from a store tablet would cross the WAN to the WLC and back to the store's printer server: wasted bandwidth and a total Wi-Fi outage when the WAN drops. With FlexConnect and local switching, only CAPWAP control (a few kbps per AP) crosses the WAN, and point-of-sale devices keep working in standalone mode during a WAN outage.

Common mistakes. Putting branch APs in local mode across a slow WAN. Designing a lecture hall for coverage (a few high-power APs) instead of capacity. Expecting controller-less APs to give seamless roaming and central RF management.

Exam trap. Local mode = data tunnelled to the WLC (centralised). FlexConnect = data switched locally (remote branch), control still to the WLC. Cloud-managed = management in the cloud, data stays local. CAPWAP control is UDP 5246, data UDP 5247.

Card payments stopped when the WAN did

A pharmacy chain ran all store APs in local mode against a WLC pair in the head-office data centre. When a provider fault cut the WAN to 60 stores for two hours, every handheld card terminal on Wi-Fi went offline, even though each store's internet breakout for payments was working. The engineer confirmed with show ap summary that the store APs had dropped their CAPWAP sessions. The redesign moved stores to FlexConnect with local switching of the payment VLAN and local authentication, so terminals stay online in standalone mode.

Lesson: choose the wireless model per site type; branches need data to stay local.

"Compare centralised (local mode) and FlexConnect wireless deployments."

In centralised mode APs tunnel all client traffic in CAPWAP to the WLC, which gives central policy, simple roaming and one place to inspect traffic, but depends on the path to the controller for every packet. FlexConnect keeps CAPWAP control to the WLC but switches data locally, saves WAN bandwidth and survives WAN outages in standalone mode, at the cost of some features and local VLAN planning. Use local mode in the campus and FlexConnect in branches. Mention cloud-managed and embedded controllers as further options.

Key takeaways

  • Controller-less = autonomous APs; controller-based = lightweight APs with CAPWAP split-MAC.
  • Centralised: WLC in the DC, data tunnelled to it; best for campuses.
  • Distributed: controller or switching at the edge (embedded WLC, SD-Access fabric wireless).
  • Cloud: management in the cloud or a virtual WLC in the cloud; data local.
  • Remote branch: FlexConnect with local switching and standalone mode; plan for density, location and WLC redundancy.
10

Hardware and software switching: CEF, RIB, FIB, adjacency, CAM and TCAM

A busy post office can sort letters in three ways. A clerk can read every envelope and look the address up in a big book (slow, but always correct). A clerk can keep sticky notes for addresses seen recently, so only the first letter to a new address is slow. Or the office can print a complete sorting table for every possible destination before the morning mail arrives, with a pre-printed label for each outgoing van, and feed it to an automatic sorting machine. Routers and switches went through exactly these three stages: process switching, fast switching and Cisco Express Forwarding (CEF) in hardware. ENCOR 1.7 asks you to explain them and the tables behind them.

Control plane and data plane

The control plane is the device's brain: routing protocols, spanning tree, ARP, building tables. The data plane is its muscle: moving each packet from an input port to an output port as fast as possible. The management plane is how you reach the device (SSH, SNMP, NETCONF). Good switching design keeps the data plane in hardware and the control plane on the CPU.

From process switching to CEF

MethodHow it worksWeakness
Process switchingThe CPU handles every packet: routing table lookup, ARP lookup, rewriteSlow, CPU bound
Fast switchingFirst packet to a destination is process-switched; the result is cached and later packets use the cache ("route once, switch many")Demand-driven: first packet always slow; cache churn with many flows
CEFTables are built in advance from the routing table and ARP, before any packet arrivesTables use memory; hardware tables have finite size

CEF is topology-driven: whenever the routing table or ARP changes, CEF updates its tables immediately, so there is no first-packet penalty. It is the default on all modern Cisco platforms.

The four tables

RIB (routing information base)
The routing table built by the control plane from connected, static and dynamic routes. show ip route.
FIB (forwarding information base)
CEF's copy of the RIB, optimised for lookup, with recursive next hops already resolved to a directly connected next hop. show ip cef.
Adjacency table
For every directly connected next hop, the pre-built Layer 2 header (destination MAC, source MAC, EtherType) learned from ARP. show adjacency detail.
CAM / MAC address table
Layer 2 switching: VLAN plus MAC to port, an exact-match lookup. show mac address-table.

Special adjacencies handle exceptions: glean (the destination is on a connected subnet but its ARP entry is missing, so the packet is punted to trigger ARP), punt (send to the CPU, for example traffic addressed to the device), drop, discard and null (routes to Null0).

Control plane (CPU) OSPF, BGP, static RIB ARP Data plane (ASIC) FIB in TCAM Adjacency MAC in CAM CEF programs hardware

The CPU builds the RIB and ARP; CEF turns them into the FIB and adjacency table that the hardware uses for every packet.

CAM and TCAM: the hardware memories

CAM (content-addressable memory) is searched by content rather than by address, and answers in one lookup. It is binary: every bit must match exactly (0 or 1), which is perfect for the MAC address table, where a VLAN plus MAC either matches or does not.

TCAM (ternary CAM) adds a third state, X, "don't care". Each entry is a value, mask and result (VMR). That makes it ideal for lookups that are not exact: longest-prefix match for the FIB, ACLs, QoS classification and policy-based routing. All entries are compared in parallel, so a 1,000-line ACL is checked as fast as a 10-line one. TCAM is expensive and limited; on Catalyst switches SDM templates decide how it is shared between routes, MAC addresses, ACLs and QoS.

Worked example. An ACL line permit ip 10.1.10.0 0.0.0.255 any becomes a TCAM entry: value 10.1.10.0, mask "compare the first 24 bits, ignore the last 8", result permit. A packet from 10.1.10.11 matches in one clock cycle. For routing, R1's FIB entries 10.1.20.0/24 and 10.1.0.0/16 are both in TCAM ordered by prefix length, so a packet to 10.1.20.12 hits the /24 first: longest match wins without scanning.

Verification on the lab network

R1# show ip cef 10.1.20.0/24
10.1.20.0/24
  nexthop 10.0.1.2 GigabitEthernet0/0
  nexthop 10.0.2.2 GigabitEthernet0/1

R1# show ip cef exact-route 172.16.0.10 10.1.20.12
172.16.0.10 -> 10.1.20.12 =>IP adj out of GigabitEthernet0/1, addr 10.0.2.2

AS1# show mac address-table vlan 10
Vlan    Mac Address       Type        Ports
----    -----------       --------    -----
  10    0000.0c07.ac0a    DYNAMIC     Gi0/0
  10    0050.7966.6801    DYNAMIC     Gi0/2

CEF shares load per destination by default: a hash of source and destination address picks one of the equal-cost next hops, so one flow stays on one path and packets are not reordered. On AS1 the HSRP virtual MAC is learned on Gi0/0, the uplink towards DS1, which is exactly what an aligned design should show.

When hardware falls back to software

Some packets are always punted to the CPU: packets addressed to the device, TTL expiry, IP options, glean adjacencies, and traffic that needs features the ASIC cannot do. If TCAM runs out, new routes or ACL entries cannot be programmed in hardware and matching traffic may be software-switched or dropped, which shows up as high CPU. On Catalyst 9000 check resources with show platform hardware fed switch active fwd-asic resource tcam utilization and the template with show sdm prefer.

Common mistakes. Troubleshooting with show ip route only, when the FIB or adjacency is what forwards. Assuming a huge ACL is free because "it is in hardware" while TCAM is nearly full. Changing an SDM template and forgetting it needs a reload.

Exam trap. The RIB is control plane; the FIB and adjacency table are data plane. CAM is binary exact match (MAC table); TCAM is ternary with don't-care bits (routes, ACLs, QoS). Fast switching is demand-driven; CEF is topology-driven. A glean adjacency means ARP is still needed.

The ACL that pushed the CPU to 99 percent

A security team added a 3,000-line ACL to a distribution switch's user SVIs. Users complained about slow applications, and show processes cpu sorted showed the CPU near 99 percent, driven by packet-forwarding processes. The log reported that hardware resources for ACLs were exhausted, and the TCAM utilisation command confirmed the ACL region was full, so part of the traffic was being handled in software. The engineer summarised the ACL into 400 lines using object groups and wider masks, moved to an SDM template with more ACL space during a maintenance window, and CPU returned to normal.

Lesson: hardware switching is only as fast as the space you leave it in TCAM.

"Explain the difference between the RIB and the FIB, and between CAM and TCAM."

The RIB is the routing table built by routing protocols in the control plane; CEF copies it into the FIB, resolves recursive next hops and pairs each entry with an adjacency that holds the pre-built Layer 2 rewrite from ARP, so the data plane forwards without asking the CPU. CAM is binary exact-match memory used for the MAC table; TCAM adds a don't-care bit, stores value-mask-result entries and does longest-prefix match, ACL and QoS lookups in a single parallel search. Mention punts and TCAM exhaustion for a senior-level answer.

Key takeaways

  • Process switching uses the CPU per packet; fast switching caches after the first packet; CEF prebuilds tables.
  • RIB = routing table (control plane); FIB + adjacency table = CEF forwarding tables (data plane).
  • Glean, punt, drop, discard and null are special adjacencies.
  • CAM is binary exact match for MAC tables; TCAM is ternary (value, mask, result) for routes, ACLs and QoS.
  • TCAM is finite; SDM templates share it, and exhaustion pushes traffic to software.
11

Design review troubleshooting: when traffic takes the long way

A building inspector does not start by knocking down walls. She takes the architect's drawing, walks the building floor by floor, and marks every place where reality differs from the plan. A design review of a network works the same way: you start from the design intent, check each layer in a fixed order, and list every difference. Most "the network is slow" tickets in redundant campuses are not broken links; they are redundancy that no longer matches the design. This chapter gives you the workflow and walks through the second lab, Design review: traffic takes the long way.

The design-review workflow

1 Intent 2 Bundles 3 STP 4 FHRP 5 Routing 6 Verify Compare every layer with the design, on both distribution switches

Work bottom-up from the physical bundle to the gateway and routing, always on both peers.

  1. Design intent. Write down what should be true: which switch owns which VLAN (STP root and HSRP active), the inter-switch bundle (members and protocol), and which switches are in OSPF.
  2. Bundles. show etherchannel summary and show lacp neighbor on both ends. Healthy is Po(SU) with every member (P).
  3. Spanning tree. On the access switch, show spanning-tree vlan N shows the root port; show cdp neighbors maps that port to a switch. On distribution, show spanning-tree root.
  4. FHRP. show standby brief on both switches: active, standby, priority, preempt flag.
  5. Routing. show ip ospf neighbor on distribution, show ip route ospf on the WAN router for equal-cost paths.
  6. Verify forwarding. Ping and traceroute from hosts; on the access switch check which uplink learned the virtual MAC.

Symptom to cause

What you seeLikely causeNext check
Inter-switch link hot, one uplink per access switch idleSTP root and HSRP active on different switchesshow spanning-tree vlan N on access, show standby brief
Po1(SD) on one side, members (I) or (s)Channel mode mismatch (on versus LACP) or LACP not negotiatingshow etherchannel summary both sides, show lacp neighbor
Preferred switch stays HSRP standby after recoveryPreempt missingshow standby, look for "Preemption enabled"
WAN router has one path to a user subnetA distribution switch missing from OSPFshow ip ospf neighbor, show ip route ospf

Lab walkthrough: traffic takes the long way

Ticket: the DS1–DS2 port-channel is down and VLAN 10 users take a detour, going up to DS2 and hairpinning through AS1 to reach DS1. Design intent: DS1 owns VLAN 10, DS2 owns VLAN 20, the interconnect is a two-member LACP bundle.

Task 1: who is the root for VLAN 10?

AS1# show spanning-tree vlan 10
VLAN0010
  Root ID    Priority    24586
             Cost        4
             Port        2 (GigabitEthernet0/1)
AS1# show cdp neighbors
Device ID    Local Intrfce   Holdtme   Capability  Platform  Port ID
DS1          Gig 0/0         152       R S I                 Gig 0/0
DS2          Gig 0/1         147       R S I                 Gig 0/0

The root port is Gi0/1 and CDP says Gi0/1 connects to DS2, so the answer is DS2. The running configuration explains why: DS1 has spanning-tree vlan 10 priority 28672 and vlan 20 priority 24576, DS2 the opposite. The priorities are swapped, while HSRP (DS1 priority 110 for group 10, DS2 for group 20) follows the design.

Task 2: realign the roots

DS1(config)# spanning-tree vlan 10 root primary
DS1(config)# spanning-tree vlan 20 root secondary
DS2(config)# spanning-tree vlan 20 root primary
DS2(config)# spanning-tree vlan 10 root secondary

Because DS2 already had 24576 for VLAN 10, the root primary macro on DS1 writes 20480 (4096 lower). Check the result on AS1: VLAN 10 root port Gi0/0 (DS1), VLAN 20 root port Gi0/1 (DS2).

Task 3: repair the bundle

DS1# show etherchannel summary
Group  Port-channel  Protocol    Ports
1      Po1(SD)         LACP      Gi0/6(I)    Gi0/7(I)

DS2# show etherchannel summary
Group  Port-channel  Protocol    Ports
1      Po1(SU)          -        Gi0/6(P)    Gi0/7(P)

DS1 speaks LACP (mode active) and receives no LACP packets, so its ports stay stand-alone (I) and its Po1 is down. DS2 uses mode on: no protocol ("-"), so it bundles unconditionally and believes Po1 is up. The two sides disagree, the interconnect is unusable, and anything that must cross between DS1 and DS2 at Layer 2 is forced down through AS1 and back up: the hairpin in the ticket.

DS2(config)# interface range g0/6 - 7
DS2(config-if-range)# no channel-group 1
DS2(config-if-range)# channel-group 1 mode active

The Port-channel1 interface and its trunk settings stay in place, and the members inherit them when they rejoin. Verify Po1(SU) with both members (P) and protocol LACP on both switches.

Task 4: verify gateways and reachability

DS1# show standby brief
Interface   Grp  Pri P State   Active          Standby         Virtual IP
Vl10        10   110 P Active  local           10.1.10.3       10.1.10.1
Vl20        20   100 P Standby 10.1.20.3       local           10.1.20.1
PC1> ping 172.16.0.10
84 bytes from 172.16.0.10 icmp_seq=1 ttl=62 time=3.112 ms

TTL 62 means two routed hops (DS1 and R1), the direct path. Repeat from PC2 through DS2.

Worked example. Before the fix, a VLAN 10 frame from PC1 took AS1 → DS2 (STP forwarding uplink) → and, with the bundle broken, back down through AS1 → DS1 to reach the gateway: four link crossings for one hop of routing. After the fix it takes AS1 → DS1: one crossing. Multiply by every VLAN 10 packet in the building.

Common mistakes. Fixing only the first difference you find: the lab has two independent faults. Checking the bundle on one switch only; mode on looks perfectly healthy from its own side. Raising HSRP priority on DS2 to "follow" the wrong STP root, which just moves the problem.

Exam trap. mode on never exchanges LACP or PAgP packets, so it cannot detect a mismatch; pairing it with active leaves the LACP side stand-alone or suspended. on works only with on; active works with active or passive.

The RMA that undid the design

A distribution switch failed and was replaced under RMA. The engineer restored an old configuration backup taken before the VLAN load-sharing project. It had HSRP priorities but no spanning-tree priority lines, so the surviving switch became STP root for every VLAN while HSRP still split the VLANs. A week later users on half the floors reported slow file transfers. The design review found the root/active mismatch in minutes using show spanning-tree root and show standby brief side by side. The team restored the root macros and added a post-change check that compares both commands against the design table.

Lesson: after any hardware swap or restore, re-run the design review, not just a ping.

"Users say a redundant campus block is slow but nothing is down. How do you approach it?"

Start from the design intent, then check bottom-up on both distribution switches: EtherChannel state and mode, STP root per VLAN from the access switch, HSRP active per VLAN, routing adjacencies and equal-cost paths, then verify with traceroute and the MAC table on the access switch. Explain that the classic causes are a root/active mismatch, a broken or mismatched bundle, missing preempt, or a switch missing from routing, and that you document and fix every difference, not just the first.

Key takeaways

  • A design review compares reality with intent, layer by layer, on both peers.
  • Order: intent, bundles, STP, FHRP, routing, then verify forwarding.
  • Root port on the access switch plus CDP tells you which switch is root.
  • mode on against active gives Po(SD) with (I) members on the LACP side.
  • The lab had two faults: swapped STP priorities and a channel-mode mismatch.
12

Summary and exam checklist

You have walked the whole city: lanes, arterial roads and highways; the rule that every district has two roads in; the traffic police that must stand where the cars actually arrive; the hot-standby pilot; the metro-style fabric; the rented cars of the cloud; the restaurant kitchens of wireless; and the sorting machine inside every switch. This chapter packs it into a checklist you can revise from the night before the exam or an interview.

Can you do all of this?

  • Explain hierarchy, modularity, resiliency and flexibility, and give a network example of each.
  • Define a failure domain and describe three ways to keep it small.
  • Calculate an oversubscription ratio, counting only forwarding links, and compare it with 20:1 (access to distribution) and 4:1 (distribution to core).
  • Describe the roles of access, distribution and core, and say when a collapsed core is enough.
  • Compare L2 access, routed access and virtual switching with MEC.
  • Align STP root and HSRP active per VLAN, with preempt, and bundle the interconnect with LACP.
  • Build and verify the collapsed-core lab: Po1(SU), root ports on AS1, show standby brief, ECMP on R1, failover and the virtual MAC 0000.0c07.ac0a.
  • Explain RPR, RPR+, SSO, NSF, graceful restart, NSR, StackWise and StackWise Virtual.
  • Describe spine-leaf rules, how to scale a fabric, and the SD-Access planes.
  • Choose between on-premises, cloud and hybrid for a given workload and describe how to connect them.
  • Compare centralised, distributed, controller-less, controller-based, cloud and remote-branch WLAN models.
  • Explain process switching, fast switching and CEF, and the RIB, FIB, adjacency table, CAM and TCAM.
  • Run a design review and find a root/active mismatch and a channel-mode mismatch.
LinkEtherChannel, ECMP GatewayHSRP, VRRP, GLBP DeviceSSO, NSF, NSR Pair as oneStackWise, SVL, MEC PathOSPF equal-cost paths, BGP multihoming DesignPairs, small failure domains, alignment

Each HA tool covers one kind of failure; a resilient design layers them.

Mini glossary

Collapsed core
Two-tier design in which distribution and core run on one switch pair.
Failure domain
The part of the network affected by one failure.
Oversubscription
Downstream bandwidth divided by usable upstream bandwidth.
Routed access
Access switch is the default gateway and routes on its uplinks; no FHRP, no uplink blocking.
MEC
Multichassis EtherChannel: one bundle to two physical switches acting as one.
SSO / NSF / NSR
Stateful standby supervisor / keep forwarding on the last FIB / keep routing state on the standby.
Graceful restart
Protocol signalling that asks NSF-aware neighbours to keep routes during a restart.
Spine-leaf
Fabric in which every leaf connects to every spine; always leaf-spine-leaf.
FlexConnect
AP mode with control to a central WLC and data switched locally; survives WAN loss.
FIB / adjacency
CEF's forwarding table and its pre-built Layer 2 rewrites.
TCAM
Ternary memory with value, mask and result entries for routes, ACLs and QoS.

Most-tested facts

TopicRemember
Collapsed coreMerges distribution and core
Oversubscription guidanceAbout 20:1 access to distribution, 4:1 distribution to core
Routed accessNo FHRP, no STP blocking on uplinks, VLAN cannot span
AlignmentSame switch is STP root and HSRP active per VLAN, with preempt
Root macrosprimary = 24576 or 4096 below current root; secondary = 28672
HSRPv1 MAC0000.0c07.acXX (group in hex)
EtherChannel modeson only with on; active with active or passive
SSO vs NSF vs NSRstate sync; forward on last FIB with GR helpers; routing state on standby, no helper
Spine-leafAdd spines for bandwidth, leaves for ports
SD-Access planesLISP control, VXLAN data, SGT policy
CAPWAPUDP 5246 control, UDP 5247 data
CAM vs TCAMbinary exact match vs ternary with don't-care
CEFtopology-driven, per-destination load sharing by default

Command cheat-sheet

! Bundle
interface range g0/6 - 7
 channel-group 1 mode active
show etherchannel summary
show lacp neighbor
! STP and HSRP alignment
spanning-tree vlan 10 root primary
spanning-tree vlan 20 root secondary
interface vlan 10
 standby 10 priority 110
 standby 10 preempt
show spanning-tree vlan 10
show standby brief
! Routing and forwarding
show ip route ospf
show ip cef 10.1.20.0/24
show ip cef exact-route 172.16.0.10 10.1.20.12
show adjacency detail
show mac address-table vlan 10
! Device HA
redundancy
 mode sso
show redundancy states
show switch
show stackwise-virtual

Worked example: a two-minute design answer. "Medium office, one building, 600 users, voice and Wi-Fi." Answer: collapsed core with two Catalyst distribution/core switches (or a StackWise Virtual pair), access stacks with two uplinks each, routed or MEC uplinks if possible, otherwise L2 access with STP root and HSRP active aligned per VLAN and preempt; LACP interconnect; OSPF to two WAN edges with SD-WAN; centralised wireless with an HA SSO WLC pair; SSO on any modular chassis; oversubscription checked against 20:1.

Last-minute mistakes to avoid. Confusing collapsed core with merged access. Saying NSF needs no neighbour support. Forgetting that routed access removes FHRP. Mixing up CAPWAP ports. Calling TCAM "binary".

Exam trap round-up. When two answers both "add redundancy", pick the one that also makes failover fast and keeps protocols aligned. When a question mentions heavy east-west traffic, think spine-leaf; branches and WAN outages, think FlexConnect and SD-WAN; supervisor failure, think SSO plus NSF or NSR.

The interview that turned into a whiteboard

A candidate for a network engineer role was asked to "draw the network of a hospital with three buildings". Instead of drawing boxes first, she asked about users, critical systems and uptime, then drew a core pair, a distribution pair per building with routed links to the core, access stacks with two uplinks, and explained STP/HSRP alignment, SSO on the core chassis, FlexConnect for the two clinics and a hybrid cloud link for the imaging archive. The panel stopped her after ten minutes: she had covered their whole question bank.

Lesson: design answers score when every box comes with the failure it survives.

"If you could change only one thing in a campus that keeps having slow periods, what would you look at first?"

Say you would first compare the running network with the design intent: STP root and FHRP active per VLAN, EtherChannel health on the distribution interconnect, and oversubscription on uplinks, because misaligned redundancy is the most common silent cause. Then explain how you would measure it with show commands and interface counters before changing anything.

Key takeaways

  • Design principles: hierarchy, modularity, resiliency, flexibility; measure failure domains and oversubscription.
  • Collapsed core for one site, three-tier for many blocks, spine-leaf for heavy east-west traffic.
  • Redundancy needs alignment: STP root with HSRP active, LACP bundles, both switches in routing.
  • SSO, NSF/GR, NSR, StackWise and StackWise Virtual protect inside the device or pair.
  • Know the WLAN models, on-premises versus cloud trade-offs and CEF, CAM and TCAM.
🎓 For educational purposes only — all devices are simulationsTerms of UsePrivacy Policy© 2026 Network Kings
CONFIG by Network Kings — an educational IT simulation platform for learning purposes only. It is not Cisco IOS, Junos, FortiOS or PAN-OS and contains no Cisco, Juniper, Fortinet or Palo Alto Networks software. Cisco, IOS, CCNA, CCNP, Juniper, JNCIA, JNCIS, JNCIP, Fortinet, FortiGate, FortiOS, NSE, Palo Alto Networks, PAN-OS and PCNSE are trademarks of their respective owners. Network Kings is not affiliated with or endorsed by Cisco Systems, Inc., Juniper Networks, Inc., Fortinet, Inc. or Palo Alto Networks, Inc.