Jump to chapter (12)
The enterprise network: what you will learn and the big picture
What you will learn in this module. By the end of these chapters you will be able to look at any enterprise network drawing and explain why it is built the way it is: why there are two of every box, where Layer 2 ends and Layer 3 begins, why the spanning-tree root and the default gateway must live on the same switch, how a supervisor can fail without users noticing, which wireless deployment model fits a campus or a branch, when a workload belongs in the cloud, and how a switch actually forwards a frame in hardware. You will also build a redundant collapsed-core block in the lab and review a broken design until traffic stops taking the long way.
Prerequisites. You should be comfortable with CCNA-level VLANs and trunks, spanning tree (root bridge, root port, blocking), EtherChannel with LACP, HSRP basics and single-area OSPF. If any of these feel shaky, revisit the CCNA modules on STP, EtherChannel and FHRP first. This module does not re-teach them; it uses them as building blocks for design.
An analogy: a network is a city's road system
Picture a well-planned city. Every house sits on a quiet local lane. Lanes feed into arterial roads where the traffic lights, speed checks and turn restrictions live. Arterial roads join a fast highway with no traffic lights at all, because its only job is to move a lot of cars quickly between districts. Every district is planned as a copy of the same template, so a new district can be added without redesigning the city. If a lane is dug up for repairs, only that street is affected, and there is always a second road into every district.
An enterprise network works the same way. Local lanes are the access layer where users plug in. Arterial roads are the distribution layer where policy and routing decisions happen. The highway is the core. Districts are modules (a building, a data centre, a WAN edge). The second road is redundancy. The rule that a dug-up lane only affects one street is the idea of a failure domain.
An enterprise is a set of modules joined by a core. Each module is designed on its own and can fail on its own.
Where this module sits in the exam
The ENCOR 350-401 Architecture domain opens with design. This module covers the design-level items; later modules go deep on the technologies they mention.
| Blueprint item | What you must be able to do | Chapters |
|---|---|---|
| 1.1 Design principles | Explain 2-tier, 3-tier, fabric and cloud designs; high availability with redundancy, FHRP and SSO | 2 to 7 |
| 1.2 WLAN deployment models | Compare centralised, distributed, controller-less, controller-based, cloud and remote-branch designs | 9 |
| 1.3 On-premises vs cloud | Choose where infrastructure should live and how it connects | 8 |
| 1.7 Hardware vs software switching | Explain process switching, CEF, RIB, FIB, adjacency table, CAM and TCAM | 10 |
SD-WAN and SD-Access, LISP and VXLAN, HSRP object tracking and deep wireless configuration each get their own modules later. Here you learn why they exist and where they fit.
The lab network you will use
Both labs use the same small collapsed-core block. Two distribution/core switches, DS1 and DS2, are joined by two links (Gi0/6 and Gi0/7). One access switch, AS1, has an uplink to each: Gi0/0 to DS1 and Gi0/1 to DS2. PC1 sits in VLAN 10 (10.1.10.0/24) and PC2 in VLAN 20 (10.1.20.0/24). DS1 owns .2 and DS2 owns .3 in each VLAN, and HSRP provides the virtual gateway .1. A WAN router R1 connects to DS1 over 10.0.1.0/30 and to DS2 over 10.0.2.0/30, and reaches the data-centre server SRV (172.16.0.10). OSPF area 0 runs between R1 and the distribution switches.
Worked example. PC1 (10.1.10.11) opens a file on SRV (172.16.0.10). The frame leaves AS1 on the uplink towards the HSRP active gateway, DS1 routes it onto 10.0.1.0/30, R1 routes it to the data-centre LAN. If DS1 dies, HSRP moves 10.1.10.1 to DS2, spanning tree unblocks the DS2 uplink if needed, and OSPF on R1 already has a second equal-cost path through DS2. Three independent mechanisms, each covering a different failure. That is design.
Common mistake. Treating design as drawing boxes. Every design choice has a protocol consequence. "Two distribution switches" immediately means you must decide the STP root, the HSRP active router, the EtherChannel mode between them and how routing sees both paths. Most outages in well-funded networks are caused by redundancy that was bought but never aligned.
Exam trap. ENCOR design questions rarely ask for commands. They describe a situation (many buildings, a small branch, heavy east-west traffic, a supervisor failure) and ask which design or feature fits. Learn the reason behind each option, not just its name.
The campus that grew one switch at a time
A college started with one core switch and added access switches whenever a lab opened, daisy-chaining some of them. After five years a single faulty NIC in one hostel caused a broadcast storm that took down every building, because all 40 switches shared one flat VLAN and one spanning tree. The redesign split the campus into building modules, each with a distribution pair, routed links to a core pair and VLANs that never leave the building.
Lesson: hierarchy and modularity are not paperwork. They are what keeps one bad cable from becoming a campus-wide outage.
"Walk me through how you would design a new enterprise campus."
Start with requirements: number of users and buildings, applications, growth, availability target. Then describe the hierarchy (access, distribution, core, or collapsed core if small), the Layer 2/3 boundary, redundancy at every layer, how STP and FHRP are aligned, how the campus connects to the data centre, WAN and cloud, and how wireless is deployed. Mentioning failure domains and oversubscription shows you think like a designer, not just a configurator.
Key takeaways
- Enterprise networks are built from modules joined by a core, like districts joined by a highway.
- Access connects users, distribution applies policy and is the L2/L3 boundary, core moves traffic fast.
- Redundancy only works when protocols such as STP, FHRP and routing are aligned with it.
- This module covers ENCOR 1.1, 1.2, 1.3 and 1.7 at a design level; later modules go deep.
- The labs use DS1, DS2, AS1, R1, SRV, PC1 and PC2 in a collapsed-core block.
Design principles: hierarchy, modularity, resiliency, flexibility
Before you pick a single switch model, you need a way to judge whether a design is good. Architects of buildings use principles such as "load-bearing walls line up floor to floor" and "every floor has two exits". Network designers use four: hierarchy, modularity, resiliency and flexibility. Two measurements keep those principles honest: the size of each failure domain and the oversubscription ratio of each layer.
Hierarchy: every layer has one job
Hierarchy means splitting the network into layers with clearly separated roles: access, distribution and core (the next chapter looks at each). Because every device in a layer does the same job, the network behaves predictably. You know where to look for a policy (distribution), where users attach (access) and which devices must never be slowed down (core). Hierarchy also makes route summarisation possible: a distribution pair can advertise one summary such as 10.1.0.0/16 for a whole building, so a flapping user subnet does not ripple through the core.
Modularity: build with repeatable blocks
Modularity means building the enterprise from self-contained blocks: an access-distribution block per building, a data-centre block, a WAN edge block, an internet edge block. Each block has the same internal design and connects to the core in the same way. The benefits are practical. A new building is a copy of a tested template. A change in one block (say a new VLAN in building B) does not touch the others. Troubleshooting starts by finding the block, then working inside it. Modularity also lets different teams own different blocks.
Resiliency: survive failures, normal and abnormal
Resiliency is the ability to keep working when something goes wrong. That includes the obvious failures (a link, a line card, a whole switch, a power feed) and the abnormal ones: a broadcast storm, a loop created by a user's home switch, a traffic flood, a bad software upgrade. Redundancy is the raw material: two uplinks, two distribution switches, dual supervisors and power supplies. But redundancy only becomes resiliency when failover is fast and automatic, which is the job of EtherChannel, FHRP, routing protocols, SSO and NSF, all covered in later chapters.
Availability is measured in "nines". 99.9 % allows about 8.8 hours of downtime a year, 99.99 % about 53 minutes, 99.999 % about 5 minutes. Each extra nine means removing another single point of failure and shortening another convergence time.
Flexibility: ready for what comes next
Flexibility is the ability to adapt without a redesign: adding wireless, IoT, a new cloud region, a merger, a new security policy. Hierarchical and modular designs are naturally flexible because you change one block or add one block. Designs with fixed, hand-crafted exceptions are not. Scalability is closely related: can the design grow from 5 to 50 access switches without changing its shape?
Failure domains: how far does one fault spread?
A failure domain is the part of the network that is affected when one component fails or misbehaves. A Layer 2 VLAN is a failure domain for broadcasts and loops: a storm reaches every port in that VLAN. A spanning-tree instance is a failure domain for topology changes. The good design habit is to keep Layer 2 domains small, ideally inside one access-distribution block, and to use routed links between blocks, because a routing protocol does not forward broadcasts and contains a failure to one subnet.
Left: one VLAN across buildings is one failure domain. Right: VLANs end at each building's distribution pair.
Oversubscription: honest about bandwidth
Oversubscription is the ratio between the bandwidth that could arrive from below and the bandwidth available upward. Not every user transmits at line rate at the same time, so some oversubscription is normal and saves money. Common starting guidance for a campus is up to about 20:1 from access to distribution and about 4:1 from distribution to core. Data-centre fabrics aim much lower, often 3:1 or better, because servers really do send at high rates.
Worked example. An access switch has 48 user ports at 1 Gbps (48 Gbps down) and two 10 Gbps uplinks (20 Gbps up). If both uplinks forward, the ratio is 48 : 20 = 2.4:1. If spanning tree blocks one uplink, only 10 Gbps is usable and the ratio becomes 4.8:1. Now stack four such switches behind the same two uplinks: 192 Gbps down, 20 Gbps up, 9.6:1, still inside the 20:1 guideline. Count only links that actually forward.
Common mistakes. Counting an STP-blocked uplink as capacity. Stretching a VLAN between buildings "just for one application", which silently merges failure domains. Buying redundant boxes but feeding both from one power strip or one cable tray.
Exam trap. "Redundancy" and "resiliency" are not the same word. A question that says a network has redundant links but takes 50 seconds to recover is describing redundancy without resiliency; the fix is a faster mechanism (RSTP, EtherChannel, routed links, tuned FHRP), not more links.
Cameras everywhere, outage everywhere
A factory stretched VLAN 50 across three buildings so every security camera could sit in one subnet with the recording server. A contractor plugged a small unmanaged switch into two wall ports in building C, creating a loop. The broadcast storm ran through VLAN 50 into buildings A and B, CPU on the distribution switches hit 100 percent and users in all three buildings lost the network. The engineer found the storm with interface counters and show spanning-tree topology change counts, shut the looped ports, and later redesigned: a camera VLAN per building, routed links between buildings, BPDU guard on access ports and storm control.
Lesson: a VLAN stretched for convenience is a failure domain stretched for disaster.
"What is a failure domain, and how do you keep it small in a campus?"
Define it as the set of devices and users affected by one failure. Then list the tools: keep VLANs local to one access-distribution block, make links between blocks and to the core routed, use routed access where the design allows, summarise routes at distribution, and protect the edge with BPDU guard and storm control. A senior touch is mentioning that control-plane failures such as a flapping route are also contained by summarisation.
Key takeaways
- Hierarchy gives each layer one role; modularity builds the network from repeatable blocks.
- Resiliency is redundancy plus fast, automatic failover, against both failures and abnormal traffic.
- Flexibility and scalability come from changing or adding blocks, not redesigning.
- Keep Layer 2 failure domains small; route between blocks.
- Oversubscription guideline: about 20:1 access to distribution, 4:1 distribution to core; count only forwarding links.
Three-tier, two-tier and the collapsed core
Think of a big railway network. Local trains stop at every small station (access). Several local lines meet at a junction where passengers change trains and tickets are checked (distribution). Junctions are linked by express trains that never stop at small stations (core). A small town, though, does not need an express line: one junction that also acts as the main station is enough. That is the difference between a three-tier campus and a two-tier (collapsed-core) campus.
The three layers and their jobs
| Layer | Main job | Typical features |
|---|---|---|
| Access | Connect end devices: PCs, phones, APs, printers, cameras | VLANs, PoE, 802.1X, port security, DHCP snooping, dynamic ARP inspection, QoS trust boundary, PortFast and BPDU guard |
| Distribution | Aggregate access switches; the Layer 2/Layer 3 boundary | SVIs and FHRP gateways, STP root, route summarisation, ACLs and policy, redundant uplinks to core |
| Core | Fast, reliable transport between distribution blocks, data centre, WAN and internet edge | Layer 3 only, high-speed links, fast convergence, minimal policy, no users attached |
The core is deliberately simple. Anything that costs CPU or adds risk (complex ACLs, NAT, packet inspection) belongs at the distribution layer or in dedicated blocks such as the internet edge. A core switch should do one thing: move packets between blocks as fast and as reliably as possible.
Three-tier adds a core pair once there are several distribution blocks. Two-tier merges core and distribution into one pair, as in this module's lab.
Why a separate core? The full-mesh problem
Without a core, every distribution block must connect directly to every other block and to the WAN, data centre and internet edge. The number of block-to-block connections grows as n(n-1)/2.
Worked example. Six building distribution pairs fully meshed need 6 x 5 / 2 = 15 pair-to-pair connections. If each connection is four links (every switch to every switch of the other pair), that is 60 fibres, and building seven adds 24 more. With a core pair in the middle, each distribution pair needs only 4 links (two switches, each to both core switches): 6 x 4 = 24 links, and building seven adds just 4. Routing adjacencies drop the same way.
A common rule of thumb is to introduce a dedicated core once you have more than two or three distribution blocks, or when the campus spans several buildings, or when you need one clean meeting point for the data centre, WAN and internet edge.
Two-tier: the collapsed core
In a collapsed-core design the core and distribution functions run on the same pair of switches. Access switches uplink to the pair; the pair also connects to the WAN routers, firewalls or a small server room. This is the right design for a single building or a medium site: fewer boxes, less cabling, less cost, and still fully redundant if built in pairs. In the lab, DS1 and DS2 are a collapsed core: they are the default gateways for VLANs 10 and 20 (distribution role) and route towards R1 and the data centre (core role).
The trade-offs: the collapsed pair carries both user-facing policy and core transport, so a mistake there affects everything; scaling to many buildings eventually brings back the full-mesh problem; and upgrades must be planned more carefully because there is no separate core to carry traffic.
Summarise at the distribution layer
Hierarchy pays off in routing. If each building uses one block of addresses, its distribution pair can advertise a single summary to the core. In OSPF, with the building in its own area and the distribution switches as ABRs:
! On each distribution switch acting as ABR for area 1
router ospf 1
area 1 range 10.1.0.0 255.255.0.0
CORE1# show ip route ospf
O IA 10.1.0.0/16 [110/3] via 10.0.11.2, 00:05:12, TenGigabitEthernet1/0/1
[110/3] via 10.0.12.2, 00:05:12, TenGigabitEthernet1/0/2
A user subnet flapping inside the building no longer triggers SPF recalculation across the whole campus. The multi-area OSPF module covers the details.
Common mistakes. Attaching servers or users directly to core switches "because there were free ports". Putting heavy ACLs or NAT in the core. Building a collapsed core with a single switch, which creates a single point of failure for the whole site.
Exam trap. A collapsed core merges distribution and core, not access and distribution. When a question describes a small site where one pair of switches provides gateways and connects to the WAN, the answer is two-tier/collapsed core.
Building five broke the budget
A company grew from one office building to four, each with its own collapsed-core pair, and connected the pairs in a full mesh. When the fifth building was approved, the plan needed eight new fibre runs, four new routing adjacencies per switch, and changes on every existing pair. The architect proposed a core pair in the main data room instead. New buildings now connect with four links to the core, existing pairs were migrated one at a time, and OSPF summarisation at each building cut the routing table in the core from hundreds of routes to a handful.
Lesson: a collapsed core is perfect until you have several of them; then a dedicated core restores modularity.
"When would you choose a two-tier collapsed core instead of a three-tier design?"
For a single building or a small-to-medium site where one pair of switches can provide the gateways and connect to the WAN, data centre and internet edge. It saves cost and cabling while keeping redundancy. Move to three-tier when there are several distribution blocks or buildings, because a core avoids the full mesh, gives one meeting point for the other modules and keeps each block's failures local. Mention the n(n-1)/2 link growth.
Key takeaways
- Access connects devices, distribution is the L2/L3 and policy boundary, core is fast Layer 3 transport.
- Keep the core simple: no users, no heavy policy, fast convergence.
- A core avoids the n(n-1)/2 full mesh once there are several distribution blocks.
- Collapsed core (two-tier) merges distribution and core; ideal for one building or a medium site.
- Summarise each block at the distribution layer to contain routing churn.
Where Layer 2 ends: L2 access, routed access and virtual switching
Every campus design must answer one question early: where does Layer 2 stop and Layer 3 begin? Think of a housing society. Inside the gate, neighbours walk freely between houses (Layer 2, same VLAN). At the gate, a guard checks where you are going and sends you on the right road (Layer 3, routing). You can put the gate at the society entrance, or at each building's door. Both work, but they change how far a problem spreads and how far a resident can wander. This chapter compares the three common answers.
Option 1: Layer 2 access (the looped triangle)
Access switches are pure Layer 2. Every access switch has an uplink trunk to each distribution switch, and the two distribution switches are linked by a Layer 2 trunk (usually an EtherChannel). The VLAN's gateway lives on distribution SVIs protected by an FHRP. Each access switch plus the distribution pair forms a triangle, which is a physical loop, so spanning tree blocks one uplink per VLAN.
- Pros: a VLAN can span several access switches; simple access configuration; works with any access switch.
- Cons: STP blocks half the uplinks (unless you load-share per VLAN); failover depends on STP and FHRP timers; the VLAN is a larger failure domain; the STP root and FHRP active must be aligned (next chapter).
This is the design used in the lab: AS1 has Gi0/0 to DS1 and Gi0/1 to DS2, and DS1 and DS2 are linked by Gi0/6 and Gi0/7.
Two variants avoid the loop. In a loop-free U, the two access switches of a pair are joined by a Layer 2 link and the distribution link is routed, so VLANs span only that access pair and nothing blocks. In a loop-free inverted U, the distribution link is Layer 2 and each access switch has only one uplink, so nothing blocks but a single uplink failure isolates that switch. Both are niche; the looped triangle and routed access are what you will meet most.
Option 2: Routed access
The Layer 3 boundary moves down to the access switch. Each uplink is a routed point-to-point link, the access switch is the default gateway for its own VLANs, and it runs OSPF or EIGRP towards the distribution pair.
- Pros: no STP blocking on uplinks, both uplinks forward with equal-cost multipath (ECMP); no FHRP needed because the access switch is the gateway; routing convergence is fast and deterministic; broadcasts and loops stay on one switch, so failure domains are tiny.
- Cons: a VLAN cannot span access switches; more IP planning (a /30 or /31 per uplink, subnets per access switch); the access switches need Layer 3 features and licensing.
! Routed access switch ACC-R1 (example)
ip routing
interface GigabitEthernet1/0/49
description UPLINK-DIST1
no switchport
ip address 10.0.101.1 255.255.255.252
ip ospf network point-to-point
interface GigabitEthernet1/0/50
description UPLINK-DIST2
no switchport
ip address 10.0.102.1 255.255.255.252
ip ospf network point-to-point
interface Vlan110
ip address 10.1.110.1 255.255.255.0
router ospf 1
router-id 10.255.0.21
passive-interface default
no passive-interface GigabitEthernet1/0/49
no passive-interface GigabitEthernet1/0/50
network 10.0.101.0 0.0.0.3 area 1
network 10.0.102.0 0.0.0.3 area 1
network 10.1.110.0 0.0.0.255 area 1
ACC-R1# show ip route ospf O*IA 0.0.0.0/0 [110/2] via 10.0.101.2, 00:02:10, GigabitEthernet1/0/49 [110/2] via 10.0.102.2, 00:02:10, GigabitEthernet1/0/50
Two equal-cost default routes: both uplinks carry traffic. Making the access area totally stubby (or the access switch an EIGRP stub) keeps its routing table tiny and stops it from ever being used as a transit path.
Option 3: Virtual switching (StackWise Virtual)
The two distribution switches are combined into one logical switch with one control plane (StackWise Virtual on Catalyst 9000, covered in the high-availability chapter). The access switch bundles its two uplinks into a single multichassis EtherChannel (MEC) with LACP. To spanning tree that is one port, so nothing blocks. The gateway is a single SVI on the logical switch, so no FHRP is needed, and VLANs can still span access switches.
L2 access blocks one uplink; routed access and virtual switching use both.
| Question | L2 access | Routed access | Virtual switching |
|---|---|---|---|
| Uplinks forwarding | One per VLAN | Both (ECMP) | Both (MEC) |
| STP role | Active, blocking | Only on edge ports | Loop protection only |
| FHRP needed | Yes | No | No |
| VLAN spans access switches | Yes | No | Yes |
| Failure domain | Whole VLAN | One access switch | Whole VLAN |
Worked example. A floor has four access switches with two 10 Gbps uplinks each. With L2 access and no per-VLAN load sharing, each switch really has 10 Gbps upward (80 Gbps of uplinks installed, 40 Gbps used). With routed access or MEC, all 80 Gbps forward. Same cables, twice the usable bandwidth.
Common mistakes. Choosing routed access and then discovering a legacy application that needs one subnet on several floors. Mixing models inside one block without a plan. Forgetting that routed access still needs STP on edge ports with BPDU guard, because users can still plug in switches.
Exam trap. In routed access there is no FHRP and no STP blocking on uplinks; the trade-off is that VLANs cannot span access switches. A question asking how to use both uplinks without spanning VLANs across switches points to routed access; one that must keep VLANs spanning points to virtual switching with MEC.
The lab that needed one subnet
A university converted its science block to routed access. Convergence after an uplink failure dropped from several seconds to well under one second and the uplink utilisation doubled. Two weeks later the physics lab complained: an old instrument-control application discovered its devices with broadcasts and needed all benches on one subnet, but the benches were on two floors with different access switches. The engineer kept routed access for the building, placed the instrument benches on one dedicated access switch per lab, and documented the constraint for future moves.
Lesson: check application Layer 2 requirements before moving the L3 boundary to the access layer.
"Compare Layer 2 access with routed access. Which would you choose?"
Explain that L2 access keeps VLANs spanning access switches but relies on STP (one uplink blocked) and an FHRP aligned with the STP root. Routed access moves the gateway to the access switch, uses both uplinks with ECMP, converges quickly and keeps failure domains tiny, but a VLAN cannot leave its access switch. Choose routed access for new builds without VLAN-spanning needs; choose L2 access or virtual switching with MEC when VLANs must span. Mentioning SD-Access, which gives routed access plus stretched subnets through an overlay, is a strong finish.
Key takeaways
- L2 access: VLANs can span, but STP blocks an uplink and an FHRP is required.
- Routed access: the access switch is the gateway, both uplinks forward with ECMP, no FHRP, tiny failure domains.
- Virtual switching (StackWise Virtual + MEC): no blocking, no FHRP, VLANs can still span.
- Loop-free U and inverted U remove the loop but have scaling or isolation limits.
- The lab uses L2 access in a looped triangle, so STP and HSRP must be aligned.
Aligning STP, HSRP and EtherChannel: build the redundant block
Imagine an office with two reception desks, one at each end of a long corridor. Visitors always enter through the east door, but the only receptionist on duty sits at the west desk. Every visitor walks the whole corridor, then walks back. Nothing is broken, yet everything is slower and the corridor is crowded. That is what happens in an L2 access design when spanning tree sends frames up one uplink while the HSRP active gateway sits on the other distribution switch. This chapter shows how to align them and then walks through the first lab, Build a redundant collapsed-core block.
Why alignment matters
In the looped triangle, spanning tree picks one root bridge per VLAN and blocks the access uplink that leads away from it. HSRP independently picks one active router per VLAN. If DS2 is the STP root for VLAN 10 but DS1 is HSRP active, PC1's frames go AS1 → DS2 (the forwarding uplink), cross the DS1–DS2 link and only then reach the gateway on DS1. The inter-switch link carries all of VLAN 10's routed traffic and every packet takes an extra hop.
Make the same switch STP root and HSRP active for each VLAN.
The design rules
- Per VLAN, one owner. The same distribution switch is STP root and HSRP active; the other is secondary root and HSRP standby.
- Load-share by VLAN. DS1 owns VLAN 10, DS2 owns VLAN 20, so both uplinks carry traffic.
- Preempt on both switches. Without
standby preempt, a recovered owner stays standby and alignment is lost after the first failure. In production addstandby 10 preempt delay minimum 60so the switch waits for routing to converge before taking over. - Bundle the inter-switch link with LACP. Two links in one Port-channel look like one link to STP, so a member failure causes no topology change, and LACP detects mis-cabling that
mode onwould hide. - Both distribution switches in the routing protocol, so the WAN router has two equal-cost paths back to every user subnet.
Lab walkthrough: build a redundant collapsed-core block
Start state: addressing, SVIs and HSRP virtual IPs (10.1.10.1 and 10.1.20.1) exist; nothing is aligned, the DS1–DS2 links are separate, and DS2 is missing from OSPF.
Task 1: LACP bundle between DS1 and DS2
! On DS1 and on DS2
interface range g0/6 - 7
channel-group 1 mode active
interface port-channel 1
switchport mode trunk
switchport trunk allowed vlan 10,20
DS1# show etherchannel summary
Group Port-channel Protocol Ports
------+-------------+-----------+-----------------------------------------------
1 Po1(SU) LACP Gi0/6(P) Gi0/7(P)
SU means Layer 2 and in use; P means bundled. Anything else (I stand-alone, s suspended, D down) means the bundle is not healthy.
Task 2: STP roots per VLAN
DS1(config)# spanning-tree vlan 10 root primary DS1(config)# spanning-tree vlan 20 root secondary DS2(config)# spanning-tree vlan 20 root primary DS2(config)# spanning-tree vlan 10 root secondary
The root primary macro writes priority 24576, or 4096 less than the current root if that is already lower; root secondary writes 28672. It is a one-time calculation saved as a number in the configuration.
AS1# show spanning-tree vlan 10 Root ID Priority 24586 Cost 4 Port 1 (GigabitEthernet0/0) Interface Role Sts Cost Prio.Nbr Type Gi0/0 Root FWD 4 128.1 P2p Gi0/1 Altn BLK 4 128.2 P2p
24586 is 24576 plus the VLAN number (extended system ID). Gi0/0 towards DS1 is the root port; for VLAN 20 the roles reverse.
Task 3: HSRP owners with preempt
DS1(config)# interface vlan 10
DS1(config-if)# standby 10 priority 110
DS1(config-if)# standby 10 preempt
DS1(config)# interface vlan 20
DS1(config-if)# standby 20 preempt
! DS2 mirrors it: priority 110 + preempt for group 20, preempt for group 10
DS1# show standby brief
Interface Grp Pri P State Active Standby Virtual IP
Vl10 10 110 P Active local 10.1.10.3 10.1.10.1
Vl20 20 100 P Standby 10.1.20.3 local 10.1.20.1
Task 4: DS2 into OSPF
DS2(config)# router ospf 1 DS2(config-router)# router-id 1.1.1.3 DS2(config-router)# network 10.0.2.0 0.0.0.3 area 0 DS2(config-router)# network 10.1.0.0 0.0.255.255 area 0
R1# show ip route ospf O 10.1.20.0/24 [110/2] via 10.0.2.2, 00:00:21, GigabitEthernet0/1 [110/2] via 10.0.1.2, 00:00:21, GigabitEthernet0/0
Two equal-cost paths: return traffic can reach the users through either distribution switch.
Tasks 5 and 6: failover test and the virtual MAC
Shut interface vlan 10 on DS1. DS2 logs %HSRP-5-STATECHANGE: Vlan10 Grp 10 state Standby -> Active and PC1 still pings 172.16.0.10. Bring the SVI back with no shutdown; after the hold and preempt timers DS1 logs Speak -> Active or Standby -> Active and owns VLAN 10 again. Finally, show standby vlan 10 shows the virtual MAC.
DS1# show standby vlan 10
Vlan10 - Group 10
State is Active
Virtual IP address is 10.1.10.1
Active virtual MAC address is 0000.0c07.ac0a (MAC In Use)
Preemption enabled
Priority 110 (configured 110)
Worked example. HSRP version 1 builds its virtual MAC as 0000.0c07.acXX, where XX is the group number in hex. Group 10 is 0x0a, so the answer is 0000.0c07.ac0a. Group 20 would be 0000.0c07.ac14. Because the MAC moves with the active role, PC1's ARP entry for 10.1.10.1 never changes during failover.
Common mistakes. Configuring priority 110 but forgetting preempt, so alignment breaks after the first reboot. Using channel-group 1 mode on on one side and active on the other. Setting STP roots with the macro, then later adding a switch with a lower priority that silently steals the root.
Exam trap. STP root and FHRP active are chosen by different, independent protocols. Nothing aligns them automatically. If a question shows traffic crossing the inter-switch link to reach the gateway, the answer is to align root and active for that VLAN, not to change HSRP timers.
The inter-switch link that was always at 90 percent
A retail head office saw its 2 x 1 Gbps distribution interconnect run near 90 percent every afternoon, while one access uplink per switch sat idle. show spanning-tree root showed DS-B as root for every VLAN (default priorities, lowest MAC), while HSRP was active on DS-A for all VLANs. Every routed packet crossed the interconnect. The engineer set root primary/secondary per VLAN to match the HSRP design, split odd VLANs to DS-A and even VLANs to DS-B with matching HSRP priorities and preempt, and interconnect load fell below 20 percent.
Lesson: default STP elects the switch with the lowest MAC, which is almost never the switch you intended.
"Why should the STP root bridge and the HSRP active router be on the same switch?"
Explain that STP decides which uplink forwards and HSRP decides where the gateway is. If they differ, traffic climbs to the root, crosses the inter-switch link to the gateway, and the link becomes a bottleneck with an extra hop. Align them per VLAN, load-share VLANs across the pair, use preempt so alignment survives failures, and bundle the interconnect with LACP. Mention that routed access or StackWise Virtual removes the problem entirely.
Key takeaways
- Per VLAN, make one switch both STP root and HSRP active; load-share VLANs across the pair.
- Use preempt on both switches; add a preempt delay in production.
- Bundle the inter-switch link with LACP active on both sides; verify Po1(SU) and members (P).
- Put both distribution switches in OSPF so R1 has two equal-cost return paths.
- HSRPv1 virtual MAC = 0000.0c07.acXX; group 10 = 0000.0c07.ac0a.
Device-level high availability: SSO, NSF, NSR, graceful restart and StackWise
A plane has two pilots. If the captain falls ill, the first officer is already in the cockpit, already knows the altitude, speed and route, and simply takes the controls. The engines never stop. Air-traffic control keeps talking to the same flight number. Network devices have exactly this idea inside one chassis: a second supervisor (the pilot) that knows the full state, hardware that keeps forwarding (the engines), and neighbours that are told not to panic (air-traffic control). This chapter covers the tools that make a single box, or a pair acting as one box, survive its own failures.
Layers of redundancy
So far you have protected against link failures (EtherChannel, dual uplinks), gateway failures (HSRP, VRRP, GLBP) and path failures (OSPF ECMP). Device-level HA protects against failures inside a device: a supervisor engine crash, a software fault, a stack member dying. It matters most where a single box is a single point of failure, such as a modular core chassis or an access switch with users attached.
Supervisor redundancy modes: RPR, RPR+ and SSO
A modular switch (for example a Catalyst 9400 or 9600) can hold two supervisors: one active, one standby. How ready the standby is depends on the redundancy mode:
| Mode | Standby state | Switchover impact |
|---|---|---|
| RPR | Partially booted, not synchronised | Line cards reset; minutes of outage |
| RPR+ | Fully booted, configuration synchronised | Links stay up but state is rebuilt; tens of seconds |
| SSO | Fully booted, configuration and state synchronised (STANDBY HOT) | Layer 2 state and links kept; typically around a second or less |
SSO (stateful switchover) continuously copies the running configuration and protocol state (interface state, MAC tables, STP, LACP, and more) to the standby. On failure the standby takes over with no link flaps and no spanning-tree reconvergence.
redundancy mode sso
CORE1# show redundancy states my state = 13 -ACTIVE peer state = 8 -STANDBY HOT Mode = Duplex Redundancy Mode (Operational) = sso Redundancy Mode (Configured) = sso
NSF: keep forwarding while routing restarts
SSO synchronises Layer 2 state, but routing protocol adjacencies traditionally restart on the new supervisor. NSF (non-stop forwarding) lets the hardware keep forwarding packets using the last known FIB while the new active supervisor rebuilds its routing table. To stop neighbours from tearing down adjacencies and rerouting around the switch, the routing protocol uses graceful restart (GR): the restarting router signals that it is restarting, and NSF-aware (helper) neighbours keep its routes and adjacency for a grace period while databases resynchronise.
- NSF-capable: the device that can restart gracefully (dual supervisors with SSO).
- NSF-aware / helper: a neighbour that understands graceful restart and keeps forwarding towards the restarting device. Most modern IOS XE devices are NSF-aware by default.
router ospf 1
nsf ! Cisco NSF; "nsf ietf" for the RFC version
router bgp 65001
bgp graceful-restart
NSR: neighbours never notice
NSR (non-stop routing) goes one step further. The standby supervisor keeps a live copy of the routing protocol state itself: adjacencies, databases, BGP sessions and TCP state. After a switchover the new active continues the conversation as if nothing happened, so no helper is needed and neighbours never see a restart. NSR is useful when neighbours are not NSF-aware or belong to another organisation, such as a service provider.
router ospf 1 nsr
SSO keeps Layer 2 state, NSF keeps packets moving, graceful restart keeps neighbours from rerouting.
StackWise: many switches, one brain
StackWise (for example StackWise-480 on Catalyst 9300) joins up to eight access switches with stack cables into one logical switch with one management IP and one configuration. One member is active, one is standby (SSO between them), the rest are members. You can build a cross-stack EtherChannel with uplinks on different members, so losing a member does not isolate the stack. The member with the highest stack priority (1 to 15) becomes active.
ACC-STACK# show switch Switch# Role Mac Address Priority Version State ------------------------------------------------------------- *1 Active 00aa.bb00.1100 15 V01 Ready 2 Standby 00aa.bb00.2200 14 V01 Ready 3 Member 00aa.bb00.3300 1 V01 Ready
StackWise Virtual: two chassis, one logical switch
StackWise Virtual (SVL) joins two distribution or core switches (Catalyst 9400, 9500 or 9600) over a StackWise Virtual link of one or more Ethernet ports. The pair has one control plane (active plus hot standby with SSO) and two data planes, both forwarding. Access switches connect with a multichassis EtherChannel, so no STP blocking and no FHRP are needed. A separate dual-active detection (DAD) link prevents both switches from becoming active if the SVL fails.
stackwise-virtual
domain 10
interface range TenGigabitEthernet1/0/47 - 48
stackwise-virtual link 1
interface TenGigabitEthernet1/0/46
stackwise-virtual dual-active-detection
! save and reload both switches to form the pair; verify with show stackwise-virtual
Worked example. A core chassis handles 40 Gbps. Without SSO, a supervisor crash takes about three minutes to recover: 40 Gbps x 180 s of dropped traffic and every OSPF neighbour reroutes twice. With SSO and NSF, Layer 2 state survives, the ASICs keep forwarding with the last FIB, and NSF-aware neighbours keep their routes; the users see at most a sub-second blip. Add ISSU (in-service software upgrade, which relies on SSO) and even planned upgrades avoid an outage window.
Common mistakes. Buying dual supervisors but leaving the mode at RPR, or running different software versions that force a lower mode. Combining NSF with very aggressive hello timers or BFD, so neighbours declare the router dead before graceful restart can begin. Forgetting the DAD link on StackWise Virtual.
Exam trap. SSO keeps state on the standby; NSF keeps forwarding during the switchover and needs graceful restart with NSF-aware helpers; NSR keeps routing state on the standby so no helper is needed. HSRP/VRRP protect the gateway between two boxes; SSO protects inside one box.
The switchover that still dropped routes
A bank enabled SSO and NSF on its new core chassis and tested a supervisor failover in a change window. Layer 2 stayed up, but the two old WAN routers dropped their OSPF adjacencies and rerouted all branch traffic for about 40 seconds. show ip ospf neighbor history and the logs showed the WAN routers ran old software that was not NSF-aware, so they ignored the grace signal. The team upgraded the WAN routers to NSF-aware software and enabled NSR on the core for the BGP sessions with the service provider, whose routers they did not control. The next test showed no routing change at all.
Lesson: NSF works only when neighbours help; NSR works even when they cannot.
"Explain the difference between SSO, NSF and NSR."
SSO synchronises configuration and state to a hot standby supervisor so a switchover keeps links and Layer 2 state. NSF lets the data plane keep forwarding with the last FIB during that switchover, and uses graceful restart so NSF-aware neighbours keep adjacencies and routes while the new active rebuilds. NSR keeps the routing protocol state itself on the standby, so neighbours do not need to help and never see a restart. Add StackWise and StackWise Virtual as the switch-level equivalents that use SSO between members.
Key takeaways
- SSO: standby supervisor is STANDBY HOT with synchronised config and state; configure with redundancy / mode sso.
- NSF: forwarding continues on the last FIB; needs graceful restart and NSF-aware helper neighbours.
- NSR: routing state lives on the standby too; no helper required.
- StackWise joins up to eight access switches into one; StackWise Virtual joins two chassis with SVL and DAD.
- With StackWise Virtual and MEC, access uplinks do not block and no FHRP is needed.
Fabric designs, spine-leaf and the WAN and branch edge
A metro rail network with a ring line and radial lines has a nice property: from any station you can reach any other station with at most one change, and if one ring segment closes, trains simply use the other direction. Compare that with an old tree-shaped bus network where every trip goes to the central depot first. Data-centre and modern campus fabrics are the metro: every edge switch is exactly one hop from every other edge switch through a set of equal paths. This chapter covers fabric designs (ENCOR 1.1 lists "fabric" next to 2-tier and 3-tier) and the WAN and branch designs that connect an enterprise's sites.
Why the data centre moved to spine-leaf
Traditional three-tier data centres were built for north-south traffic: clients outside talking to servers inside. Modern applications are split into many services that talk to each other, so most traffic is now east-west, server to server. In a three-tier design with spanning tree, east-west traffic often climbs to the aggregation or core layer and back down, half the links are blocked, and latency differs depending on where two servers sit.
A spine-leaf (Clos) fabric fixes this with two strict rules:
- Every leaf (top-of-rack switch where servers, firewalls and routers attach) connects to every spine.
- Spines never connect to each other, and leaves never connect to each other (apart from special pairs such as vPC or MLAG peers).
So any server reaches any other server on another leaf in exactly leaf → spine → leaf: the same number of hops and predictable latency. The links are routed and every spine offers an equal-cost path, so traffic is spread with ECMP and nothing is blocked.
Any leaf reaches any other leaf through any spine in two hops. The border leaf connects to WAN, internet and campus.
Scaling and oversubscription in a fabric
You scale a fabric out, not up. Need more server ports? Add a leaf. Need more bandwidth between leaves? Add a spine and one more uplink on every leaf. The limits are simple: the number of spines is capped by the uplink ports on each leaf, and the number of leaves is capped by the ports on each spine.
Worked example. Each leaf has 48 server ports at 25 Gbps (1,200 Gbps down) and uplinks at 100 Gbps. With 4 spines (4 x 100 = 400 Gbps up) the leaf is 3:1 oversubscribed. With 6 spines (600 Gbps up) it is 2:1. If each spine has 32 ports of 100 Gbps, the fabric can hold up to 32 leaves, or 32 x 48 = 1,536 server ports at 25 Gbps.
LEAF1# show ip route 10.20.3.0 Routing entry for 10.20.3.0/24 Known via "ospf 1", distance 110, metric 3, type intra area Routing Descriptor Blocks: * 10.255.1.1, from 10.255.0.3, via HundredGigE1/0/49 10.255.2.1, from 10.255.0.3, via HundredGigE1/0/50
Two spines, two equal-cost next hops towards LEAF3's subnet.
Underlay and overlay
A fabric has two layers. The underlay is the routed physical network (OSPF, IS-IS or eBGP between leaves and spines) that simply makes every switch's loopback reachable. The overlay builds tenant networks on top, typically VXLAN tunnels between leaves, so a VLAN or subnet can appear on any leaf without stretching spanning tree. In data centres the overlay control plane is usually BGP EVPN.
The same idea came to the campus as SD-Access: a routed underlay (often IS-IS built by LAN automation), LISP as the control plane that tracks where each endpoint is, VXLAN as the data plane, and Scalable Group Tags as the policy plane, managed by Catalyst Center. Its roles are edge nodes (where users attach), border nodes (exits to other networks) and control-plane nodes (the LISP map server). You get routed access for resiliency, yet a subnet can still exist on every edge node. The SD-WAN/SD-Access and LISP/VXLAN modules go deep; here you only need to recognise fabric as a design option.
WAN and branch designs
Branches range from a two-person sales office to a thousand-person regional hub, so there is no single branch design. Choose by size and by how much downtime the branch can tolerate:
| Branch size | Typical design | Weak point |
|---|---|---|
| Small | One router or SD-WAN edge, one or two WAN links (for example broadband plus 4G/5G), access switch with integrated services | Single device |
| Medium | Two routers or edges, two transports (MPLS plus internet), FHRP or routing towards a switch stack | Single switch stack |
| Large | Dual edges, dual transports, a small collapsed core or full access-distribution block | Cost and complexity |
Transports include MPLS L3VPN (provider-managed, SLAs), dedicated or broadband internet with IPsec or DMVPN, and cellular backup. SD-WAN overlays all of them, measures loss, latency and jitter on each, and steers each application to the best path. The internet edge module usually has redundant firewalls, two ISPs and BGP for multihoming.
Common mistakes. Cabling a leaf to only some spines, which breaks the equal-path promise. Connecting two spines together "for redundancy". Giving a branch two routers that both use the same ISP and the same last-mile cable.
Exam trap. In spine-leaf, the answer to "how many hops between servers on different leaves" is always leaf-spine-leaf, and the answer to "how do you add bandwidth" is add spines. In SD-Access, LISP is the control plane, VXLAN the data plane and SGT the policy plane.
Backups at 2 AM
A company's new data centre used four leaves and two spines, each leaf with 48 x 25 Gbps server ports and two 100 Gbps uplinks: 6:1 oversubscription. Every night at 2 AM the backup job pulled data from all servers to the storage leaf, and application monitoring showed drops and retransmissions. Interface counters on the leaf uplinks showed output drops, while spine CPU and links looked healthy. The fix needed no redesign: two more spines were added and every leaf got two more uplinks, giving 3:1 and four ECMP paths. The backup window shrank by half.
Lesson: spine-leaf scales out; when leaves run hot on uplinks, add spines rather than bigger boxes.
"Why do data centres use spine-leaf instead of three-tier?"
Because traffic became mostly east-west. Spine-leaf gives every pair of leaves the same two-hop path, uses routed links with ECMP so no link is blocked, gives predictable latency and scales out by adding spines for bandwidth or leaves for ports. Mention the underlay/overlay split (routed underlay, VXLAN overlay with EVPN) and that SD-Access brings the same model to the campus with LISP and VXLAN.
Key takeaways
- Spine-leaf: every leaf to every spine; no spine-spine or leaf-leaf links; always leaf-spine-leaf.
- Routed links with ECMP; scale out by adding spines (bandwidth) or leaves (ports).
- Fabrics use a routed underlay and an overlay (VXLAN with EVPN in the DC, LISP plus VXLAN in SD-Access).
- Branch design depends on size and tolerance for downtime: single, dual-edge, or small campus.
- SD-WAN overlays MPLS, internet and cellular and picks a path per application.
On-premises versus cloud infrastructure
Owning a car and using a ride-hailing app both get you to work. Owning costs a lot up front, but the car is always in your driveway, you can modify it, and daily trips are cheap once it is paid for. Ride-hailing costs nothing up front, scales instantly when your whole family needs to travel, and someone else handles servicing, but you pay every trip, depend on the app and the roads, and cannot choose the engine. On-premises infrastructure is the owned car; public cloud is ride-hailing. ENCOR expects you to explain the difference and know how a network engineer connects the two.
Deployment models
- On-premises
- The organisation owns and runs the hardware in its own data centre or a rented colocation space. Full control, full responsibility.
- Private cloud
- On-premises (or dedicated hosted) infrastructure run with cloud-style self-service, automation and virtualisation for one organisation.
- Public cloud
- A provider's shared infrastructure, rented on demand and billed by use, reached over the internet or private interconnects.
- Hybrid cloud
- On-premises or private cloud connected to public cloud, with workloads placed where they fit best. This is what most enterprises actually run.
- Multicloud
- Using two or more public cloud providers, for resilience, features or commercial reasons.
Service models describe how much the provider manages: IaaS (you get virtual machines, storage and virtual networks and manage the OS and everything above), PaaS (you deploy code or containers onto a managed platform) and SaaS (you just use the application, for example hosted email or CRM). Networking itself can also be consumed as a service, such as cloud-managed switches and access points whose management plane lives in the provider's cloud.
Comparing on-premises and cloud
| Factor | On-premises | Public cloud |
|---|---|---|
| Cost model | CapEx: buy hardware up front, depreciate over years | OpEx: pay per hour, per GB and per request |
| Scaling | Size for peak; adding capacity takes weeks | Elastic; scale in minutes, scale down to save money |
| Control | Full control of hardware, software versions, topology | Limited to what the provider exposes |
| Latency and locality | Close to campus users and factory systems | Depends on WAN or internet path to the region |
| Compliance | Data stays where you put it | Choose regions carefully; shared-responsibility model |
| Operations | Your team patches, replaces and cools everything | Provider runs facilities, hardware and hypervisor |
The shared-responsibility model is key. With IaaS, the provider secures the buildings, hardware and hypervisor; you still own the guest operating systems, applications, identities, data and your virtual network rules such as security groups and route tables. Moving to the cloud does not move security responsibility away from you.
Hybrid connectivity: VPN tunnels, dedicated interconnects and direct internet access to SaaS.
How the network connects to the cloud
- Site-to-site IPsec VPN over the internet: quick and cheap; use two tunnels to two provider gateways and run BGP over them.
- Private or dedicated interconnect: a physical circuit from your router, usually in a colocation facility, into the provider's network. Predictable latency and bandwidth; build two in different locations for resilience.
- SD-WAN cloud on-ramp: virtual SD-WAN routers (for example Catalyst 8000V) inside the cloud region become part of the SD-WAN fabric, and branches can reach SaaS directly instead of hairpinning through the head office.
- Virtual network functions in the cloud: virtual routers, firewalls and even wireless controllers (Catalyst 9800-CL) run as cloud instances.
EDGE1# show ip bgp summary Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd 169.254.10.1 4 64512 220 218 14 0 0 03:12:41 2 169.254.11.1 4 64512 219 218 14 0 0 03:12:39 2
Two VPN tunnels to the cloud gateway, each with an eBGP session learning the two cloud prefixes. If one tunnel fails, BGP converges to the other. Link-local 169.254.x.x addresses inside the tunnels are common with cloud VPN gateways.
Worked example. An online shop needs 20 web servers all year and 100 during a three-week festival sale. On-premises, it must buy and power 100 servers that sit 80 percent idle for 49 weeks. In the cloud it runs 20 instances and scales to 100 for three weeks, paying only for what it uses. Its payroll database, however, is steady, small and bound by data-residency rules, so it stays on-premises. The result is hybrid, and the network must now provide resilient, low-latency connectivity between the two.
Common mistakes. Moving applications to the cloud but leaving one VPN tunnel from one site as the only path. Forgetting that data leaving the cloud (egress) is usually charged. Assuming the provider patches your virtual machines. Ignoring latency for chatty applications split between cloud and on-premises.
Exam trap. CapEx versus OpEx, elasticity, and shared responsibility are the most tested points. Questions often give a workload (steady, regulated, latency-critical versus spiky, global, fast-changing) and ask where it belongs. "Hybrid" is right only when the scenario really needs both.
The ERP that lived behind one tunnel
A distributor moved its ERP system to a public cloud region. All 30 branches reached it through the head office, which had one ISP and one IPsec tunnel to the cloud. When the head-office ISP failed for four hours, the ERP was down for every branch even though the cloud itself was healthy. The post-incident review found a single failure domain: one ISP, one router, one tunnel. The fix: a second ISP and edge router at head office with two tunnels and BGP to the cloud, and SD-WAN at the branches so they could reach the cloud directly over their own internet links if the head office was unreachable.
Lesson: moving to the cloud moves the critical path onto the WAN; design that path with the same redundancy as the data centre.
"How do you decide whether a workload should run on-premises or in the cloud?"
Ask about the load pattern (steady versus spiky), latency and locality needs, data residency and compliance, integration with other systems, team skills and total cost over several years including egress. Steady, regulated or latency-critical workloads often stay on-premises; spiky, global or fast-changing ones fit the cloud. Then explain how you would connect them: redundant VPN or interconnects with BGP, SD-WAN for branches, and the shared-responsibility model for security.
Key takeaways
- Models: on-premises, private, public, hybrid and multicloud; services: IaaS, PaaS, SaaS.
- On-premises is CapEx with full control; cloud is OpEx with elasticity and less control.
- Shared responsibility: the provider secures the platform, you secure what you put on it.
- Connect with redundant IPsec tunnels or private interconnects running BGP, or through SD-WAN.
- Most enterprises are hybrid; the WAN path to the cloud becomes a critical failure domain.
WLAN deployment models: centralised, distributed, cloud and remote branch
Think of a restaurant chain. A tiny café can have its own cook who decides everything alone. A big mall food court sends every order to one central kitchen. A chain of small outlets might have a head chef who writes the menu centrally, while each outlet cooks locally so it keeps serving even when the phone line to head office is down. And some chains let a remote company run their ordering system from an app. Wireless networks are deployed in exactly these ways. ENCOR 1.2 asks you to compare them at a design level; the enterprise wireless module later covers RF, roaming and controller configuration.
Controller-less versus controller-based
The first split is whether there is a controller at all.
- Controller-less (autonomous): each access point is a complete, independent device with its own configuration, SSIDs, security and RF settings. Fine for one or two APs; painful beyond that, because every change is repeated on every AP and features such as coordinated radio resource management and fast roaming are limited. Some material also calls an AP running an embedded wireless controller (one AP acting as controller for its neighbours) controller-less, because there is no separate controller box.
- Controller-based (lightweight APs): APs are managed by a wireless LAN controller (WLC) using CAPWAP. This is the split-MAC model: real-time functions (beacons, acknowledgements, encryption) stay on the AP, while management, authentication, RF management and roaming decisions move to the controller. CAPWAP uses UDP 5246 for control and UDP 5247 for data.
Where the controller and the data path live
| Model | Controller | Client data path | Best fit |
|---|---|---|---|
| Centralised | WLC (appliance or pair) in the data centre or services block | Tunnelled in CAPWAP from AP (local mode) to the WLC, then switched there | Campus with good LAN bandwidth; central policy and easy roaming |
| Distributed | Controller function close to the edge: embedded on Catalyst 9000 switches, or central WLC with locally switched APs | Switched at the access layer (in SD-Access, VXLAN from AP to fabric edge) | Large campuses and fabric designs; avoids hairpinning all traffic to one place |
| Cloud | Management in the cloud (cloud-managed APs) or a virtual controller such as Catalyst 9800-CL in a public or private cloud | Local at the site; only management goes to the cloud | Many small sites, lean IT teams, fast rollout |
| Remote branch | Central WLC over the WAN with FlexConnect APs, or an embedded controller on a branch AP | Locally switched onto the branch VLAN; can keep working if the WAN fails | Branches and retail stores |
Local mode tunnels everything to the controller; FlexConnect keeps data at the branch.
FlexConnect for branches
A FlexConnect AP keeps its CAPWAP control connection to a central WLC but switches client traffic onto a local VLAN. If the WAN fails, it enters standalone mode: already-connected clients stay connected, and with local authentication configured new clients can still join. On a Catalyst 9800 controller, APs become FlexConnect through a site tag that is not a local site:
wireless profile flex BR1-FLEX native-vlan-id 30 wireless tag site BR1-SITE flex-profile BR1-FLEX no local-site
WLC1# show ap summary Number of APs: 2 AP Name Slots AP Model Ethernet MAC Radio MAC Location Country IP Address State AP-HQ-01 2 C9120AXI-D 00aa.bb11.0001 00aa.bb22.0001 HQ-FL1 IN 10.1.30.21 Registered AP-BR1-01 2 C9120AXI-D 00aa.bb11.0002 00aa.bb22.0002 Branch1 IN 10.50.30.21 Registered
Design factors beyond the model
- Client density: auditoriums, classrooms and stadiums are designed for capacity, not coverage. Use more APs at lower power, prefer 5 GHz and 6 GHz, use narrower channels to get more of them, and plan the number of clients per radio.
- Location services: accurate location needs every point heard by at least three APs at about -75 dBm or better, with APs placed along the perimeter and staggered, not just down the corridor.
- Controller redundancy: N+1 (one spare WLC for many), N+N (two sets sharing load), N+N+1, or an HA SSO pair in which a standby controller holds AP and client state, so a failure does not force APs to rejoin.
- Latency and bandwidth between AP and controller: local mode depends on the LAN or WAN path for every packet, which is why remote sites use FlexConnect.
Worked example. A retail chain has 400 stores with 3 APs each and a 20 Mbps WAN link per store. In local mode, a video stream from a store tablet would cross the WAN to the WLC and back to the store's printer server: wasted bandwidth and a total Wi-Fi outage when the WAN drops. With FlexConnect and local switching, only CAPWAP control (a few kbps per AP) crosses the WAN, and point-of-sale devices keep working in standalone mode during a WAN outage.
Common mistakes. Putting branch APs in local mode across a slow WAN. Designing a lecture hall for coverage (a few high-power APs) instead of capacity. Expecting controller-less APs to give seamless roaming and central RF management.
Exam trap. Local mode = data tunnelled to the WLC (centralised). FlexConnect = data switched locally (remote branch), control still to the WLC. Cloud-managed = management in the cloud, data stays local. CAPWAP control is UDP 5246, data UDP 5247.
Card payments stopped when the WAN did
A pharmacy chain ran all store APs in local mode against a WLC pair in the head-office data centre. When a provider fault cut the WAN to 60 stores for two hours, every handheld card terminal on Wi-Fi went offline, even though each store's internet breakout for payments was working. The engineer confirmed with show ap summary that the store APs had dropped their CAPWAP sessions. The redesign moved stores to FlexConnect with local switching of the payment VLAN and local authentication, so terminals stay online in standalone mode.
Lesson: choose the wireless model per site type; branches need data to stay local.
"Compare centralised (local mode) and FlexConnect wireless deployments."
In centralised mode APs tunnel all client traffic in CAPWAP to the WLC, which gives central policy, simple roaming and one place to inspect traffic, but depends on the path to the controller for every packet. FlexConnect keeps CAPWAP control to the WLC but switches data locally, saves WAN bandwidth and survives WAN outages in standalone mode, at the cost of some features and local VLAN planning. Use local mode in the campus and FlexConnect in branches. Mention cloud-managed and embedded controllers as further options.
Key takeaways
- Controller-less = autonomous APs; controller-based = lightweight APs with CAPWAP split-MAC.
- Centralised: WLC in the DC, data tunnelled to it; best for campuses.
- Distributed: controller or switching at the edge (embedded WLC, SD-Access fabric wireless).
- Cloud: management in the cloud or a virtual WLC in the cloud; data local.
- Remote branch: FlexConnect with local switching and standalone mode; plan for density, location and WLC redundancy.
Hardware and software switching: CEF, RIB, FIB, adjacency, CAM and TCAM
A busy post office can sort letters in three ways. A clerk can read every envelope and look the address up in a big book (slow, but always correct). A clerk can keep sticky notes for addresses seen recently, so only the first letter to a new address is slow. Or the office can print a complete sorting table for every possible destination before the morning mail arrives, with a pre-printed label for each outgoing van, and feed it to an automatic sorting machine. Routers and switches went through exactly these three stages: process switching, fast switching and Cisco Express Forwarding (CEF) in hardware. ENCOR 1.7 asks you to explain them and the tables behind them.
Control plane and data plane
The control plane is the device's brain: routing protocols, spanning tree, ARP, building tables. The data plane is its muscle: moving each packet from an input port to an output port as fast as possible. The management plane is how you reach the device (SSH, SNMP, NETCONF). Good switching design keeps the data plane in hardware and the control plane on the CPU.
From process switching to CEF
| Method | How it works | Weakness |
|---|---|---|
| Process switching | The CPU handles every packet: routing table lookup, ARP lookup, rewrite | Slow, CPU bound |
| Fast switching | First packet to a destination is process-switched; the result is cached and later packets use the cache ("route once, switch many") | Demand-driven: first packet always slow; cache churn with many flows |
| CEF | Tables are built in advance from the routing table and ARP, before any packet arrives | Tables use memory; hardware tables have finite size |
CEF is topology-driven: whenever the routing table or ARP changes, CEF updates its tables immediately, so there is no first-packet penalty. It is the default on all modern Cisco platforms.
The four tables
- RIB (routing information base)
- The routing table built by the control plane from connected, static and dynamic routes.
show ip route. - FIB (forwarding information base)
- CEF's copy of the RIB, optimised for lookup, with recursive next hops already resolved to a directly connected next hop.
show ip cef. - Adjacency table
- For every directly connected next hop, the pre-built Layer 2 header (destination MAC, source MAC, EtherType) learned from ARP.
show adjacency detail. - CAM / MAC address table
- Layer 2 switching: VLAN plus MAC to port, an exact-match lookup.
show mac address-table.
Special adjacencies handle exceptions: glean (the destination is on a connected subnet but its ARP entry is missing, so the packet is punted to trigger ARP), punt (send to the CPU, for example traffic addressed to the device), drop, discard and null (routes to Null0).
The CPU builds the RIB and ARP; CEF turns them into the FIB and adjacency table that the hardware uses for every packet.
CAM and TCAM: the hardware memories
CAM (content-addressable memory) is searched by content rather than by address, and answers in one lookup. It is binary: every bit must match exactly (0 or 1), which is perfect for the MAC address table, where a VLAN plus MAC either matches or does not.
TCAM (ternary CAM) adds a third state, X, "don't care". Each entry is a value, mask and result (VMR). That makes it ideal for lookups that are not exact: longest-prefix match for the FIB, ACLs, QoS classification and policy-based routing. All entries are compared in parallel, so a 1,000-line ACL is checked as fast as a 10-line one. TCAM is expensive and limited; on Catalyst switches SDM templates decide how it is shared between routes, MAC addresses, ACLs and QoS.
Worked example. An ACL line permit ip 10.1.10.0 0.0.0.255 any becomes a TCAM entry: value 10.1.10.0, mask "compare the first 24 bits, ignore the last 8", result permit. A packet from 10.1.10.11 matches in one clock cycle. For routing, R1's FIB entries 10.1.20.0/24 and 10.1.0.0/16 are both in TCAM ordered by prefix length, so a packet to 10.1.20.12 hits the /24 first: longest match wins without scanning.
Verification on the lab network
R1# show ip cef 10.1.20.0/24 10.1.20.0/24 nexthop 10.0.1.2 GigabitEthernet0/0 nexthop 10.0.2.2 GigabitEthernet0/1 R1# show ip cef exact-route 172.16.0.10 10.1.20.12 172.16.0.10 -> 10.1.20.12 =>IP adj out of GigabitEthernet0/1, addr 10.0.2.2 AS1# show mac address-table vlan 10 Vlan Mac Address Type Ports ---- ----------- -------- ----- 10 0000.0c07.ac0a DYNAMIC Gi0/0 10 0050.7966.6801 DYNAMIC Gi0/2
CEF shares load per destination by default: a hash of source and destination address picks one of the equal-cost next hops, so one flow stays on one path and packets are not reordered. On AS1 the HSRP virtual MAC is learned on Gi0/0, the uplink towards DS1, which is exactly what an aligned design should show.
When hardware falls back to software
Some packets are always punted to the CPU: packets addressed to the device, TTL expiry, IP options, glean adjacencies, and traffic that needs features the ASIC cannot do. If TCAM runs out, new routes or ACL entries cannot be programmed in hardware and matching traffic may be software-switched or dropped, which shows up as high CPU. On Catalyst 9000 check resources with show platform hardware fed switch active fwd-asic resource tcam utilization and the template with show sdm prefer.
Common mistakes. Troubleshooting with show ip route only, when the FIB or adjacency is what forwards. Assuming a huge ACL is free because "it is in hardware" while TCAM is nearly full. Changing an SDM template and forgetting it needs a reload.
Exam trap. The RIB is control plane; the FIB and adjacency table are data plane. CAM is binary exact match (MAC table); TCAM is ternary with don't-care bits (routes, ACLs, QoS). Fast switching is demand-driven; CEF is topology-driven. A glean adjacency means ARP is still needed.
The ACL that pushed the CPU to 99 percent
A security team added a 3,000-line ACL to a distribution switch's user SVIs. Users complained about slow applications, and show processes cpu sorted showed the CPU near 99 percent, driven by packet-forwarding processes. The log reported that hardware resources for ACLs were exhausted, and the TCAM utilisation command confirmed the ACL region was full, so part of the traffic was being handled in software. The engineer summarised the ACL into 400 lines using object groups and wider masks, moved to an SDM template with more ACL space during a maintenance window, and CPU returned to normal.
Lesson: hardware switching is only as fast as the space you leave it in TCAM.
"Explain the difference between the RIB and the FIB, and between CAM and TCAM."
The RIB is the routing table built by routing protocols in the control plane; CEF copies it into the FIB, resolves recursive next hops and pairs each entry with an adjacency that holds the pre-built Layer 2 rewrite from ARP, so the data plane forwards without asking the CPU. CAM is binary exact-match memory used for the MAC table; TCAM adds a don't-care bit, stores value-mask-result entries and does longest-prefix match, ACL and QoS lookups in a single parallel search. Mention punts and TCAM exhaustion for a senior-level answer.
Key takeaways
- Process switching uses the CPU per packet; fast switching caches after the first packet; CEF prebuilds tables.
- RIB = routing table (control plane); FIB + adjacency table = CEF forwarding tables (data plane).
- Glean, punt, drop, discard and null are special adjacencies.
- CAM is binary exact match for MAC tables; TCAM is ternary (value, mask, result) for routes, ACLs and QoS.
- TCAM is finite; SDM templates share it, and exhaustion pushes traffic to software.
Design review troubleshooting: when traffic takes the long way
A building inspector does not start by knocking down walls. She takes the architect's drawing, walks the building floor by floor, and marks every place where reality differs from the plan. A design review of a network works the same way: you start from the design intent, check each layer in a fixed order, and list every difference. Most "the network is slow" tickets in redundant campuses are not broken links; they are redundancy that no longer matches the design. This chapter gives you the workflow and walks through the second lab, Design review: traffic takes the long way.
The design-review workflow
Work bottom-up from the physical bundle to the gateway and routing, always on both peers.
- Design intent. Write down what should be true: which switch owns which VLAN (STP root and HSRP active), the inter-switch bundle (members and protocol), and which switches are in OSPF.
- Bundles.
show etherchannel summaryandshow lacp neighboron both ends. Healthy is Po(SU) with every member (P). - Spanning tree. On the access switch,
show spanning-tree vlan Nshows the root port;show cdp neighborsmaps that port to a switch. On distribution,show spanning-tree root. - FHRP.
show standby briefon both switches: active, standby, priority, preempt flag. - Routing.
show ip ospf neighboron distribution,show ip route ospfon the WAN router for equal-cost paths. - Verify forwarding. Ping and traceroute from hosts; on the access switch check which uplink learned the virtual MAC.
Symptom to cause
| What you see | Likely cause | Next check |
|---|---|---|
| Inter-switch link hot, one uplink per access switch idle | STP root and HSRP active on different switches | show spanning-tree vlan N on access, show standby brief |
| Po1(SD) on one side, members (I) or (s) | Channel mode mismatch (on versus LACP) or LACP not negotiating | show etherchannel summary both sides, show lacp neighbor |
| Preferred switch stays HSRP standby after recovery | Preempt missing | show standby, look for "Preemption enabled" |
| WAN router has one path to a user subnet | A distribution switch missing from OSPF | show ip ospf neighbor, show ip route ospf |
Lab walkthrough: traffic takes the long way
Ticket: the DS1–DS2 port-channel is down and VLAN 10 users take a detour, going up to DS2 and hairpinning through AS1 to reach DS1. Design intent: DS1 owns VLAN 10, DS2 owns VLAN 20, the interconnect is a two-member LACP bundle.
Task 1: who is the root for VLAN 10?
AS1# show spanning-tree vlan 10 VLAN0010 Root ID Priority 24586 Cost 4 Port 2 (GigabitEthernet0/1) AS1# show cdp neighbors Device ID Local Intrfce Holdtme Capability Platform Port ID DS1 Gig 0/0 152 R S I Gig 0/0 DS2 Gig 0/1 147 R S I Gig 0/0
The root port is Gi0/1 and CDP says Gi0/1 connects to DS2, so the answer is DS2. The running configuration explains why: DS1 has spanning-tree vlan 10 priority 28672 and vlan 20 priority 24576, DS2 the opposite. The priorities are swapped, while HSRP (DS1 priority 110 for group 10, DS2 for group 20) follows the design.
Task 2: realign the roots
DS1(config)# spanning-tree vlan 10 root primary DS1(config)# spanning-tree vlan 20 root secondary DS2(config)# spanning-tree vlan 20 root primary DS2(config)# spanning-tree vlan 10 root secondary
Because DS2 already had 24576 for VLAN 10, the root primary macro on DS1 writes 20480 (4096 lower). Check the result on AS1: VLAN 10 root port Gi0/0 (DS1), VLAN 20 root port Gi0/1 (DS2).
Task 3: repair the bundle
DS1# show etherchannel summary Group Port-channel Protocol Ports 1 Po1(SD) LACP Gi0/6(I) Gi0/7(I) DS2# show etherchannel summary Group Port-channel Protocol Ports 1 Po1(SU) - Gi0/6(P) Gi0/7(P)
DS1 speaks LACP (mode active) and receives no LACP packets, so its ports stay stand-alone (I) and its Po1 is down. DS2 uses mode on: no protocol ("-"), so it bundles unconditionally and believes Po1 is up. The two sides disagree, the interconnect is unusable, and anything that must cross between DS1 and DS2 at Layer 2 is forced down through AS1 and back up: the hairpin in the ticket.
DS2(config)# interface range g0/6 - 7 DS2(config-if-range)# no channel-group 1 DS2(config-if-range)# channel-group 1 mode active
The Port-channel1 interface and its trunk settings stay in place, and the members inherit them when they rejoin. Verify Po1(SU) with both members (P) and protocol LACP on both switches.
Task 4: verify gateways and reachability
DS1# show standby brief Interface Grp Pri P State Active Standby Virtual IP Vl10 10 110 P Active local 10.1.10.3 10.1.10.1 Vl20 20 100 P Standby 10.1.20.3 local 10.1.20.1 PC1> ping 172.16.0.10 84 bytes from 172.16.0.10 icmp_seq=1 ttl=62 time=3.112 ms
TTL 62 means two routed hops (DS1 and R1), the direct path. Repeat from PC2 through DS2.
Worked example. Before the fix, a VLAN 10 frame from PC1 took AS1 → DS2 (STP forwarding uplink) → and, with the bundle broken, back down through AS1 → DS1 to reach the gateway: four link crossings for one hop of routing. After the fix it takes AS1 → DS1: one crossing. Multiply by every VLAN 10 packet in the building.
Common mistakes. Fixing only the first difference you find: the lab has two independent faults. Checking the bundle on one switch only; mode on looks perfectly healthy from its own side. Raising HSRP priority on DS2 to "follow" the wrong STP root, which just moves the problem.
Exam trap. mode on never exchanges LACP or PAgP packets, so it cannot detect a mismatch; pairing it with active leaves the LACP side stand-alone or suspended. on works only with on; active works with active or passive.
The RMA that undid the design
A distribution switch failed and was replaced under RMA. The engineer restored an old configuration backup taken before the VLAN load-sharing project. It had HSRP priorities but no spanning-tree priority lines, so the surviving switch became STP root for every VLAN while HSRP still split the VLANs. A week later users on half the floors reported slow file transfers. The design review found the root/active mismatch in minutes using show spanning-tree root and show standby brief side by side. The team restored the root macros and added a post-change check that compares both commands against the design table.
Lesson: after any hardware swap or restore, re-run the design review, not just a ping.
"Users say a redundant campus block is slow but nothing is down. How do you approach it?"
Start from the design intent, then check bottom-up on both distribution switches: EtherChannel state and mode, STP root per VLAN from the access switch, HSRP active per VLAN, routing adjacencies and equal-cost paths, then verify with traceroute and the MAC table on the access switch. Explain that the classic causes are a root/active mismatch, a broken or mismatched bundle, missing preempt, or a switch missing from routing, and that you document and fix every difference, not just the first.
Key takeaways
- A design review compares reality with intent, layer by layer, on both peers.
- Order: intent, bundles, STP, FHRP, routing, then verify forwarding.
- Root port on the access switch plus CDP tells you which switch is root.
mode onagainstactivegives Po(SD) with (I) members on the LACP side.- The lab had two faults: swapped STP priorities and a channel-mode mismatch.
Summary and exam checklist
You have walked the whole city: lanes, arterial roads and highways; the rule that every district has two roads in; the traffic police that must stand where the cars actually arrive; the hot-standby pilot; the metro-style fabric; the rented cars of the cloud; the restaurant kitchens of wireless; and the sorting machine inside every switch. This chapter packs it into a checklist you can revise from the night before the exam or an interview.
Can you do all of this?
- Explain hierarchy, modularity, resiliency and flexibility, and give a network example of each.
- Define a failure domain and describe three ways to keep it small.
- Calculate an oversubscription ratio, counting only forwarding links, and compare it with 20:1 (access to distribution) and 4:1 (distribution to core).
- Describe the roles of access, distribution and core, and say when a collapsed core is enough.
- Compare L2 access, routed access and virtual switching with MEC.
- Align STP root and HSRP active per VLAN, with preempt, and bundle the interconnect with LACP.
- Build and verify the collapsed-core lab: Po1(SU), root ports on AS1,
show standby brief, ECMP on R1, failover and the virtual MAC 0000.0c07.ac0a. - Explain RPR, RPR+, SSO, NSF, graceful restart, NSR, StackWise and StackWise Virtual.
- Describe spine-leaf rules, how to scale a fabric, and the SD-Access planes.
- Choose between on-premises, cloud and hybrid for a given workload and describe how to connect them.
- Compare centralised, distributed, controller-less, controller-based, cloud and remote-branch WLAN models.
- Explain process switching, fast switching and CEF, and the RIB, FIB, adjacency table, CAM and TCAM.
- Run a design review and find a root/active mismatch and a channel-mode mismatch.
Each HA tool covers one kind of failure; a resilient design layers them.
Mini glossary
- Collapsed core
- Two-tier design in which distribution and core run on one switch pair.
- Failure domain
- The part of the network affected by one failure.
- Oversubscription
- Downstream bandwidth divided by usable upstream bandwidth.
- Routed access
- Access switch is the default gateway and routes on its uplinks; no FHRP, no uplink blocking.
- MEC
- Multichassis EtherChannel: one bundle to two physical switches acting as one.
- SSO / NSF / NSR
- Stateful standby supervisor / keep forwarding on the last FIB / keep routing state on the standby.
- Graceful restart
- Protocol signalling that asks NSF-aware neighbours to keep routes during a restart.
- Spine-leaf
- Fabric in which every leaf connects to every spine; always leaf-spine-leaf.
- FlexConnect
- AP mode with control to a central WLC and data switched locally; survives WAN loss.
- FIB / adjacency
- CEF's forwarding table and its pre-built Layer 2 rewrites.
- TCAM
- Ternary memory with value, mask and result entries for routes, ACLs and QoS.
Most-tested facts
| Topic | Remember |
|---|---|
| Collapsed core | Merges distribution and core |
| Oversubscription guidance | About 20:1 access to distribution, 4:1 distribution to core |
| Routed access | No FHRP, no STP blocking on uplinks, VLAN cannot span |
| Alignment | Same switch is STP root and HSRP active per VLAN, with preempt |
| Root macros | primary = 24576 or 4096 below current root; secondary = 28672 |
| HSRPv1 MAC | 0000.0c07.acXX (group in hex) |
| EtherChannel modes | on only with on; active with active or passive |
| SSO vs NSF vs NSR | state sync; forward on last FIB with GR helpers; routing state on standby, no helper |
| Spine-leaf | Add spines for bandwidth, leaves for ports |
| SD-Access planes | LISP control, VXLAN data, SGT policy |
| CAPWAP | UDP 5246 control, UDP 5247 data |
| CAM vs TCAM | binary exact match vs ternary with don't-care |
| CEF | topology-driven, per-destination load sharing by default |
Command cheat-sheet
! Bundle interface range g0/6 - 7 channel-group 1 mode active show etherchannel summary show lacp neighbor ! STP and HSRP alignment spanning-tree vlan 10 root primary spanning-tree vlan 20 root secondary interface vlan 10 standby 10 priority 110 standby 10 preempt show spanning-tree vlan 10 show standby brief ! Routing and forwarding show ip route ospf show ip cef 10.1.20.0/24 show ip cef exact-route 172.16.0.10 10.1.20.12 show adjacency detail show mac address-table vlan 10 ! Device HA redundancy mode sso show redundancy states show switch show stackwise-virtual
Worked example: a two-minute design answer. "Medium office, one building, 600 users, voice and Wi-Fi." Answer: collapsed core with two Catalyst distribution/core switches (or a StackWise Virtual pair), access stacks with two uplinks each, routed or MEC uplinks if possible, otherwise L2 access with STP root and HSRP active aligned per VLAN and preempt; LACP interconnect; OSPF to two WAN edges with SD-WAN; centralised wireless with an HA SSO WLC pair; SSO on any modular chassis; oversubscription checked against 20:1.
Last-minute mistakes to avoid. Confusing collapsed core with merged access. Saying NSF needs no neighbour support. Forgetting that routed access removes FHRP. Mixing up CAPWAP ports. Calling TCAM "binary".
Exam trap round-up. When two answers both "add redundancy", pick the one that also makes failover fast and keeps protocols aligned. When a question mentions heavy east-west traffic, think spine-leaf; branches and WAN outages, think FlexConnect and SD-WAN; supervisor failure, think SSO plus NSF or NSR.
The interview that turned into a whiteboard
A candidate for a network engineer role was asked to "draw the network of a hospital with three buildings". Instead of drawing boxes first, she asked about users, critical systems and uptime, then drew a core pair, a distribution pair per building with routed links to the core, access stacks with two uplinks, and explained STP/HSRP alignment, SSO on the core chassis, FlexConnect for the two clinics and a hybrid cloud link for the imaging archive. The panel stopped her after ten minutes: she had covered their whole question bank.
Lesson: design answers score when every box comes with the failure it survives.
"If you could change only one thing in a campus that keeps having slow periods, what would you look at first?"
Say you would first compare the running network with the design intent: STP root and FHRP active per VLAN, EtherChannel health on the distribution interconnect, and oversubscription on uplinks, because misaligned redundancy is the most common silent cause. Then explain how you would measure it with show commands and interface counters before changing anything.
Key takeaways
- Design principles: hierarchy, modularity, resiliency, flexibility; measure failure domains and oversubscription.
- Collapsed core for one site, three-tier for many blocks, spine-leaf for heavy east-west traffic.
- Redundancy needs alignment: STP root with HSRP active, LACP bundles, both switches in routing.
- SSO, NSF/GR, NSR, StackWise and StackWise Virtual protect inside the device or pair.
- Know the WLAN models, on-premises versus cloud trade-offs and CEF, CAM and TCAM.