Jump to chapter (9)
The big picture: one fabric, three labs and a map of the components
What you will learn in this module. By the end you will be able to read a live Cisco Catalyst SD-WAN network from the command line and explain what you see. You will name each component and say what it does, describe how TLOCs, colours, OMP and BFD fit together, follow one packet from a PC in Delhi to a server in Hyderabad, and build a second segment (a guest VPN) on a working fabric. These are the skills the architecture domain of exam 300-415 ENSDWI v1.2 tests, and they are also the skills a NOC engineer uses on day one.
Prerequisites. The starter module SD-WAN from zero, or the same knowledge: the four planes, the four components, the idea of an overlay and an underlay, and basic IP routing. You do not need to have configured SD-WAN before. Each command is explained when it first appears.
Start with an analogy: a metro rail network
A metro system has stations (the WAN edges) where passengers get on and off, and trains that carry them between stations (the encrypted tunnels). There is a central control room that knows every station and publishes the timetable: which station can be reached from which platform (this is the Controller and its routing information). There is a head office where the rule book, staff rosters and monitoring screens live (the Manager). And there is a ticket and ID desk at the entrance of the network where every new employee is checked before being sent on to the right office (the Validator). The control room never carries a passenger; the trains do.
The fabric you will work on
The labs in this module use a small but realistic network for a retailer called Bharat Retail. The control components sit in a data centre behind a router, and the data centre reaches two transports: a simulated internet cloud and a simulated MPLS cloud. Branch routers called WAN edges (names start with E-) attach to these clouds.
The Bharat Retail lab fabric. Manager, Controller and Validator sit in the data centre; each store edge reaches them over the internet or MPLS.
| Device | Role | Identity |
|---|---|---|
| VMANAGE | SD-WAN Manager | system IP 10.255.255.1, address 198.51.100.10 |
| VSMART | SD-WAN Controller | system IP 10.255.255.2, address 198.51.100.20 |
| VBOND | SD-WAN Validator | system IP 10.255.255.3, address 198.51.100.30 |
| E-DEL-1, E-MUM-1 | WAN edges with internet and MPLS | system IPs 10.255.0.11, 10.255.0.12; sites 100, 200 |
| E-BLR-1 | WAN edge with internet only | system IP 10.255.0.13; site 300 |
| INET, MPLS | Simulated carrier clouds | underlay only |
Naming: new product names and the older ones
Cisco renamed the components. The current names are SD-WAN Manager, SD-WAN Controller and SD-WAN Validator. The lab devices are still labelled VMANAGE, VSMART and VBOND, and many commands, APIs and job interviews still use vManage, vSmart and vBond. WAN edge is the router; cEdge means an IOS XE router in controller mode and vEdge means a legacy Viptela router. Learn both names for every item once and you will never be confused.
How the three labs map to the chapters
- Read a live fabric (chapters 2 to 5): control connections, OMP peers, BFD sessions and an end-to-end test in VPN 10.
- Follow a packet (chapters 4, 5 and 6): routing table, OMP route, TLOC, BFD session and traceroute; then add a second transport.
- Guest segmentation (chapter 7): create VPN 20 on two sites and prove it is isolated from VPN 10.
Chapter 8 gives a troubleshooting routine that works on all three, and chapter 9 is the summary and checklist.
Worked example. In one glance at the table above you can already predict facts the labs will confirm. E-DEL-1 has two transports (internet and MPLS), so it has two TLOCs. E-BLR-1 has one transport, so one TLOC. The three store LANs are 10.1.10.0/24, 10.2.10.0/24 and 10.3.10.0/24, all in VPN 10, each behind a different site ID. If a PC in Delhi reaches a PC in Bengaluru, the packet will travel in an IPsec tunnel over the internet because E-BLR-1 has no MPLS.
Common mistake. Treating the simulator devices as ordinary routers. VMANAGE, VSMART and VBOND use the Viptela-style CLI (system, vpn 0, commit), while the edges use IOS XE in controller mode. The commands differ by device type, and the next chapters show which one to use where.
The new NOC joiner's first hour
On her first day a NOC engineer was told "the fabric is healthy, just have a look". She did not know where to start, so she opened the Manager and clicked every menu. Her lead showed her a better routine: check the control connections on the Validator and an edge; then OMP peers; then BFD sessions; finally a test between two LANs. In ten minutes she could tell the lead what was normal and what a failure would look like.
Lesson: a fixed reading order (control, OMP, BFD, data) gives you a map to hold the whole architecture in your head. This module teaches that order.
"Describe the architecture of a Catalyst SD-WAN deployment you have worked on or studied."
Strong answer: name the three control components and where they sit (for example the data centre or a cloud), say that each WAN edge keeps control connections to the Controller and the Manager and builds IPsec tunnels with BFD to other edges, and give numbers: how many sites, how many transports, which VPNs. Show that you know which parts carry user traffic (only the edges).
Key takeaways
- This module teaches you to read, trace and extend a working Catalyst SD-WAN fabric.
- The lab fabric has VMANAGE, VSMART and VBOND in a data centre and three edges: Delhi, Mumbai, Bengaluru.
- New names (Manager, Controller, Validator) and old names (vManage, vSmart, vBond) are both used in practice.
- Controllers use the Viptela-style CLI; WAN edges use IOS XE in controller mode.
- Read in this order: control connections, OMP, BFD, then the data path.
The four planes and the control connections between components
The previous chapter introduced the fabric. This chapter looks at how the control components relate to each other and to the edges: who connects to whom, over what protocol, and what you see when you ask the CLI. This is the first thing you check on any SD-WAN network, so it is worth learning well.
The four planes in one table
| Plane | Component | Main jobs |
|---|---|---|
| Orchestration | SD-WAN Validator (vBond) | First point of contact for every device. Authenticates it (certificate, organization name, authorised serial list), tells it which Controllers and Managers exist, and discovers the public address of devices behind NAT. Needs a public or statically NATed address. Its connection to an edge is transient. |
| Management | SD-WAN Manager (vManage) | Single pane of glass: GUI and REST API, templates and configuration groups, the authorised WAN edge list, certificates, software upgrades, monitoring and alarms. Keeps a permanent control connection to every device; NETCONF travels over it. |
| Control | SD-WAN Controller (vSmart) | OMP route reflector: receives routes, TLOCs and service routes from edges, applies control policy and reflects them to other edges; distributes data policy and IPsec keys. Does not forward user traffic. |
| Data | WAN edges | Build the IPsec tunnels, run BFD in each tunnel, forward user traffic, enforce data and application-aware routing policy, QoS, NAT and security. |
Who connects to whom
All control connections are secured with DTLS (UDP, base port 12346) by default; TLS (TCP) can be configured instead; if two peers are configured differently, TLS is used. The rules are easy to remember if you picture them as a star:
- Every Manager and Controller keeps a permanent connection to the Validator. This is how the Validator knows which Controllers and Managers to advertise to the edges.
- Every edge briefly connects to the Validator, learns the addresses, then drops that connection.
- Every edge keeps one connection to a Controller per TLOC (up to the configured
max-control-connections, default 2) and exactly one connection to the Manager. - Controllers also connect to the Manager. Controllers connect to each other in a full mesh so that they share the same view.
- Edges never form control connections with each other. Their relationship is the data-plane tunnels (chapter 6).
Control connections form a star around the controllers. The edges never connect to each other except with data-plane tunnels.
Reading the Validator
On the Validator the command is show orchestrator connections. Only Controllers and Managers stay registered there, because edge connections are transient.
VBOND# show orchestrator connections
PEER PEER PEER SITE DOMAIN PEER ORGANIZATION REMOTE COLOR PROXY STATE
INSTANCE TYPE PROT SYSTEM IP ID ID PRIVATE IP ...
------------------------------------------------------------------------------------------------
0 vmanage dtls 10.255.255.1 1 0 198.51.100.10 BHARAT-RETAIL default No up
0 vsmart dtls 10.255.255.2 1 1 198.51.100.20 BHARAT-RETAIL default No up
Check the type, the system IP, the organization name and that the STATE is up. The system IP of the Controller, 10.255.255.2, is the answer to the first lab question.
Reading an edge
On a WAN edge running IOS XE the command carries the sdwan keyword. E-DEL-1 has two TLOCs, so it holds two Controller connections and one Manager connection.
E-DEL-1# show sdwan control connections PEER PEER PEER SITE DOMAIN PEER PEER PUBLIC PUB ORGANIZATION LOCAL PROXY STATE UPTIME TYPE PROT SYSTEM IP ID ID PRIVATE IP IP PORT COLOR -------------------------------------------------------------------------------------------------------------------- vsmart dtls 10.255.255.2 1 1 198.51.100.20 198.51.100.20 12346 BHARAT-RETAIL biz-internet No up 0:00:02:31 vsmart dtls 10.255.255.2 1 1 198.51.100.20 198.51.100.20 12346 BHARAT-RETAIL mpls No up 0:00:02:30 vmanage dtls 10.255.255.1 1 0 198.51.100.10 198.51.100.10 12346 BHARAT-RETAIL biz-internet No up 0:00:02:33
There is no vbond row, and that is normal. Once the edge has a Controller connection it closes the Validator connection. A missing vbond row on a healthy edge is not a fault.
Reading the Controller
On the Controller the commands use the Viptela style, without the sdwan keyword. show control connections lists the same peers from the other side, and show omp peers lists one OMP session per edge:
VSMART# show omp peers
R -> routes received
I -> routes installed
S -> routes sent
DOMAIN OVERLAY SITE
PEER TYPE ID ID ID STATE UPTIME R/I/S
---------------------------------------------------------------------
10.255.0.11 vedge 1 1 100 up 0:02:11:30 3/0/4
10.255.0.12 vedge 1 1 200 up 0:02:11:28 3/0/4
10.255.0.13 vedge 1 1 300 up 0:02:11:25 2/0/5
"Type vedge" here is just the OMP peer type name for any WAN edge, including IOS XE edges. The R/I/S column shows routes received, installed and sent; a Controller installs nothing for itself, so the middle number is 0.
Worked example. Count the sessions expected for Bharat Retail. Edges: 3. Each edge has one Manager connection (3 total). Controller connections: E-DEL-1 two TLOCs gives 2, E-MUM-1 gives 2, E-BLR-1 gives 1, a total of 5 edge-to-Controller connections. Plus VSMART and VMANAGE each keep one connection to VBOND. If a row is missing in your output, you now know exactly which connection it should be.
Common mistakes. Expecting a vbond row on an edge; expecting one Controller row regardless of the number of TLOCs; typing the Viptela form show control connections on an IOS XE edge (it needs show sdwan control connections) or the IOS XE form on a Controller.
Exam trap. "Which component maintains a permanent connection with every other control component and a temporary connection with edges?" is the Validator. "Which has one connection per edge regardless of TLOC count?" is the Manager.
Missing row after a link change
After a carrier change at the Mumbai store, E-MUM-1 showed only one vsmart row where it used to show two. The NOC compared the output with what they expected from the TLOC count and saw that the MPLS-colour connection was gone. The cause was a carrier-side address change on the MPLS circuit. Because the NOC knew the expected rows by heart, the missing row was found in under a minute, before the store noticed anything.
Lesson: knowing the expected number of control connections turns the output into a checklist.
"How many control connections should a WAN edge with two transports have?"
Strong answer: one to the Manager over one transport, and one to a Controller per TLOC (limited by max-control-connections, default 2). The Validator connection is only temporary. So two TLOCs normally give two Controller connections and one Manager connection, all DTLS over UDP 12346 by default.
Key takeaways
- Validator = orchestration, Manager = management, Controller = control, WAN edges = data.
- Manager and Controller keep permanent connections to the Validator; edges use it only briefly.
- An edge has one Manager connection and one Controller connection per TLOC.
- Use
show orchestrator connectionson the Validator,show control connectionsandshow omp peerson the Controller, andshow sdwan control connectionson an edge. - Edges never form control connections with each other.
WAN edge platforms, controller mode and Cloud OnRamp
The WAN edge is where the overlay meets the real world. It has LAN ports towards users and transport ports towards carriers, and it is the only component that forwards user packets. This chapter explains the two edge families, how to talk to an edge at the command line, how an edge is identified and authorised, and how edges extend to the cloud.
Two families: cEdge and vEdge
| cEdge (IOS XE) | vEdge (legacy Viptela OS) | |
|---|---|---|
| Hardware / virtual | Catalyst 8000 series (8200, 8300, 8500), Catalyst 8000V, ISR 1000 and 4000, ASR 1000 | vEdge 100, 1000, 2000, 5000 and vEdge Cloud; end of sale, but still in many networks |
| CLI | IOS XE in controller mode: normal IOS commands for interfaces, VRFs and routing plus the system, sdwan and omp sections; changes with config-transaction and commit | Viptela CLI: vpn 0, vpn 512, vpn 10 blocks and commit |
| Service VPN | A VRF (vrf definition 10); VPN 0 is the global table; VPN 512 is Mgmt-intf | vpn N natively |
| Features | Full IOS XE feature set: UTD for IPS and URL filtering, AMP, voice (SRST, CUBE), advanced QoS, NAT, many interface types | Smaller feature set, no voice, limited embedded security |
The same Manager, Controller and Validator support both families in one network, which makes migration from vEdge to cEdge possible without a flag day.
Autonomous mode versus controller mode
An IOS XE router runs in one of two modes. In autonomous mode it is an ordinary router configured with configure terminal. In controller mode it is managed by the SD-WAN Manager. The mode is changed in exec mode with controller-mode enable or controller-mode disable, and the switch erases the configuration and reloads the device, so it is a deliberate one-time action when onboarding or retiring a router.
In controller mode the router refuses the usual configure terminal. You enter configuration with config-transaction, and the changes take effect only when you type commit:
E-DEL-1# configure terminal % This WAN Edge runs in controller mode: use "config-transaction" (changes take effect on "commit"). E-DEL-1# config-transaction E-DEL-1(config)# ... E-DEL-1(config)# commit Commit complete.
On a router that is managed by templates from the Manager, changes you type on the device can be overwritten the next time the template is pushed. The CLI is mainly for verification and, in the lab, for building a small piece of configuration directly.
How an edge is identified and authorised
Each edge carries a chassis number and a serial number. For hardware routers these come from the board and its built-in certificate (the secure unique device identifier, SUDI). For virtual routers the Manager issues them together with a one-time token. The Manager keeps a list of the edges that the company allows: the authorised WAN edge list. It sends the list to the Validator and the Controllers. A device that is not on the list fails authentication and cannot join. A device can also be present on the list but set to invalid or staging to prevent or limit its participation; the controller deployment module covers this in detail.
Besides identity, each edge needs three items of configuration so that it can find and prove itself to the controllers. In the lab they look like this:
system system-ip 10.255.0.11 site-id 100 organization-name BHARAT-RETAIL vbond 198.51.100.30
system-ip and site-id identify the device and its location; organization-name must be spelled exactly like the controllers' value; vbond is the address where the edge starts. Together with a transport interface (chapter 4) these four lines are what a router needs to find its way into the fabric.
Where edges sit: branch, hub and cloud
Edge roles. Whatever its size or location, an edge is identified, authenticated and configured the same way.
Cloud OnRamp: reaching SaaS and public cloud
The architecture extends to the cloud in three named ways. They are part of the blueprint, so learn the question each one answers:
| Feature | Question it answers | How it works |
|---|---|---|
| Cloud OnRamp for SaaS | Which exit gives this SaaS application the best quality right now? | Edges probe the application over every possible exit (local internet or a gateway site across the overlay), compute a quality score from loss and latency, and send traffic out of the best exit. It moves when quality degrades. |
| Cloud OnRamp for IaaS / Multicloud | How do branches reach workloads in the public cloud with the same VPNs and policy? | The Manager uses the cloud provider's API to build the transit (for example a transit gateway or virtual WAN hub), deploys Catalyst 8000V edges or integrates the provider's SD-WAN hub, and maps cloud networks to service VPNs. The cloud becomes another site. |
| Cloud OnRamp for Colocation | Where do I host shared regional security and routing services? | Virtualised edges and security functions run in a colocation facility close to the cloud and the branches, forming a regional hub. |
Worked example. Bharat Retail adds a Catalyst 8000V in a public cloud so that stores can reach the cloud-based inventory system. The 8000V is onboarded like any edge (system IP 10.255.0.101, site ID 900, same organization name) and advertises the cloud subnet 172.20.0.0/24 in VPN 10. Store edges learn that prefix through the Controller and build tunnels to the cloud edge. No store configuration changes.
Common mistakes. Typing configure terminal on a controller-mode router and wondering why it is refused; forgetting commit so the change never applies; mixing up controller-mode enable with a mode that can be toggled casually (it wipes the configuration).
Exam trap. The OMP and control-plane behaviour is the same for vEdge and cEdge. Differences are in the CLI and features: IOS XE has the full IOS feature set; vEdge has the Viptela CLI and fewer features.
The router that would not take a command
A contractor tried to add a loopback to a newly installed Catalyst 8200 with configure terminal and was told the device was in controller mode. He assumed the router was broken and requested a replacement. The NOC lead showed him config-transaction and commit; the loopback was added in seconds, and the replacement request was cancelled.
Lesson: controller mode changes how you configure the device, not whether you can.
"What is the difference between a cEdge and a vEdge, and what is controller mode?"
Strong answer: cEdge is an IOS XE router (Catalyst 8000, ISR, ASR) running SD-WAN software with the full IOS XE feature set; vEdge is the legacy Viptela OS platform. Both join the same overlay. An IOS XE router must be in controller mode to be managed by the SD-WAN Manager; in that mode configuration is done with config-transaction and commit, and switching the mode erases the configuration.
Key takeaways
- cEdge runs IOS XE in controller mode; vEdge runs the legacy Viptela OS. Both use the same controllers.
- In controller mode use
config-transactionandcommit;configure terminalis refused. - An edge needs system-ip, site-id, organization-name and vbond, plus a transport interface, to join.
- Only edges on the authorised WAN edge list can join; chassis and serial numbers identify them.
- Cloud OnRamp for SaaS, IaaS/Multicloud and Colocation extend the fabric to cloud applications and workloads.
TLOCs, system IP and colours: how an edge attaches to a transport
If you understand TLOCs, most of the rest of SD-WAN falls into place, because every route and every tunnel in the overlay is expressed in terms of TLOCs. This chapter starts with identities, builds up to the TLOC and its colour, explains how colours decide which address a tunnel uses, and finishes with the commands to read and to configure a TLOC.
System IP and site ID: who and where
Every device has a system IP: a /32 identifier written like an IPv4 address, similar to a router ID. It does not have to be reachable in the underlay; it is simply the device's name in the overlay. In the lab fabric the controllers use 10.255.255.1 to 3 and the edges use 10.255.0.11 and up. Every device also has a site ID, a number shared by all devices in one location. Two edges with the same site ID never build tunnels to each other, so duplicates must be avoided between different locations.
The TLOC
A TLOC (transport locator) is how an edge attaches to a transport. It is identified by three items:
- the system IP of the edge,
- the colour, a label for the type of transport, and
- the encapsulation, normally IPsec (GRE is an option mostly used towards service devices).
An edge with MPLS and broadband has two TLOCs: 10.255.0.11 / mpls / ipsec and 10.255.0.11 / biz-internet / ipsec. Each TLOC also carries the private and public IP address and the UDP port the edge uses for its tunnels. The interface IP address is an attribute, not part of the identity.
Colours and why they matter
The colour is a label chosen from a fixed list. Its job is practical: it tells the edge which address of the remote TLOC to aim at, and it lets policy refer to a transport by name.
| Private colours | Public colours |
|---|---|
| mpls, metro-ethernet, private1 to private6 | default, biz-internet, public-internet, lte, 3g, blue, green, red, gold, silver, bronze, custom1 to custom3 |
| Between two private colours the tunnel is built to the peer's private (pre-NAT) address | If either side uses a public colour, the tunnel is built to the peer's public (post-NAT) address, which the Validator and Controller learned |
Why? A private transport such as MPLS has no NAT: the address on the interface is the real address, so edges can talk to each other directly. The internet is different: an edge behind a NAT device is known to the outside only by its translated address and port, so the tunnel must be built to that. The Validator discovers this translated (public) address when the edge first connects.
Restrict: keeping colours apart
By default each TLOC tries to build a tunnel to every remote TLOC, whatever its colour. If the MPLS cloud has no way to reach internet addresses, those attempts are wasted and the edge may log failures. The keyword restrict after a colour limits that TLOC to peers with the same colour. In the lab, the MPLS TLOCs use color mpls restrict, so an MPLS TLOC only pairs with other MPLS TLOCs.
With restrict on the MPLS colour, only same-colour pairs form tunnels.
Reading the TLOCs
show sdwan omp tlocs lists the TLOCs known to an edge: its own and those learned from the Controller.
E-DEL-1# show sdwan omp tlocs ADDRESS PSEUDO PUBLIC PRIVATE BFD FAMILY TLOC IP COLOR ENCAP FROM PEER STATUS KEY PUBLIC IP PORT PRIVATE IP PORT STATUS ------------------------------------------------------------------------------------------------------------------- ipv4 10.255.0.11 biz-internet ipsec 0.0.0.0 C,Red,R 1 203.0.113.6 12346 203.0.113.6 12346 - ipv4 10.255.0.11 mpls ipsec 0.0.0.0 C,Red,R 1 172.16.1.2 12366 172.16.1.2 12366 - ipv4 10.255.0.12 biz-internet ipsec 10.255.255.2 C,I,R 1 203.0.113.10 12346 203.0.113.10 12346 up ipv4 10.255.0.12 mpls ipsec 10.255.255.2 C,I,R 1 172.16.2.2 12366 172.16.2.2 12366 up ipv4 10.255.0.13 biz-internet ipsec 10.255.255.2 C,I,R 1 203.0.113.14 12346 203.0.113.14 12346 up
The first two lines are E-DEL-1's own TLOCs (FROM PEER 0.0.0.0, status Red for redistributed); the rest were reflected by the Controller (10.255.255.2). The last column tells you whether a BFD session to that TLOC is up. Notice the ports: IPsec tunnels can use 12346, 12366, 12386 and so on (steps of 20), so that several TLOCs or several devices behind one NAT address do not collide. In this lab the second TLOC of each edge shows 12366.
Configuring a TLOC on an IOS XE edge
Two pieces are needed: a tunnel interface that is bound to the physical transport interface, and a tunnel-interface section under sdwan that sets the encapsulation and colour. The physical interface must already have an IP address and a route out.
! Make GigabitEthernet1 a TLOC with colour biz-internet
config-transaction
interface Tunnel1
ip unnumbered GigabitEthernet1
tunnel source GigabitEthernet1
tunnel mode sdwan
exit
sdwan
interface GigabitEthernet1
tunnel-interface
encapsulation ipsec
color biz-internet
exit
exit
exit
commit
end
The tunnel number is only a local name; pick one that is not already used (Tunnel2 may belong to the MPLS transport). tunnel mode sdwan tells IOS XE that this tunnel interface belongs to the SD-WAN overlay.
Worked example. E-HYD-1 in another fabric has only an MPLS TLOC (10.255.0.14 / mpls / ipsec, restricted). After the configuration above, it also advertises 10.255.0.14 / biz-internet / ipsec. The Controller reflects the new TLOC to every other edge; E-DEL-1, which has a biz-internet TLOC, now builds a second tunnel to E-HYD-1 over the internet. E-DEL-1 goes from one BFD session with E-HYD-1 to two.
Common mistakes. Forgetting tunnel mode sdwan; forgetting commit; choosing a colour that does not match the transport (a private colour on an internet link means tunnels are built to the private address, which is not reachable); using restrict on a colour that other sites use as public; assuming the system IP is the tunnel source.
Exam traps. A TLOC is system IP + colour + encapsulation, not an interface address. Private to private uses private addresses; anything with a public colour uses public addresses. restrict means same colour only. Only default (public) is the colour assigned when none is configured.
The internet link configured as MPLS
A new store was onboarded with its broadband link given the colour private1 because a template was copied from a store with a private circuit. Control connections came up, but data tunnels to other stores never did: the other edges tried to reach the store's private (pre-NAT) address, which sits behind the ISP's NAT and is unreachable from outside. The fix was changing the colour to biz-internet, which makes peers use the public address.
Lesson: the colour is not cosmetic; it decides which address peers aim at.
"What is the role of colours in SD-WAN?"
Strong answer: colours label transports on TLOCs. They split into private (mpls, metro-ethernet, private1 to 6) and public (biz-internet, public-internet, lte and others). Colour determines which remote address a tunnel uses: private to private uses the private (pre-NAT) address; any public colour uses the public (post-NAT) address learned through the Validator. The restrict keyword limits tunnels to the same colour, and policies refer to colours to select a path.
Key takeaways
- System IP names the device; site ID names the location; same site ID means no tunnels between those edges.
- A TLOC is system IP + colour + encapsulation; one edge has one TLOC per transport.
- Colour decides private versus public addressing for the tunnel; restrict limits tunnels to the same colour.
- Use
show sdwan omp tlocsto read TLOCs and their BFD status. - To add a TLOC: a Tunnel interface with tunnel mode sdwan, plus tunnel-interface with encapsulation and colour under sdwan, then commit.
OMP: the routing protocol of the overlay
Tunnels alone do not tell an edge which prefixes live behind which remote site. That is the job of the Overlay Management Protocol (OMP). This chapter explains what OMP carries, how an edge turns OMP information into a route, and how to read OMP output without getting lost in the status codes.
What OMP is
OMP runs inside the DTLS or TLS control connection between each edge and each Controller. Edges never peer with each other. It behaves like BGP with a route reflector: every edge tells the Controller what it knows, and the Controller reflects the (policy-filtered) result to the other edges. Unlike BGP, it was designed for this job, so it also carries encryption keys and policy and needs very little configuration. It is enabled by default on a WAN edge.
The three kinds of OMP route
- OMP routes (also called vRoutes): service-side prefixes learned in a VPN (connected, static, OSPF, BGP, EIGRP and so on). Attributes include the VPN and a label, the TLOC that is the next hop, the origin protocol and metric, preference, site ID and tag.
- TLOC routes: each TLOC with its private and public address and port, colour, encapsulation, weight and preference (chapter 4).
- Service routes: services such as a firewall or IPS reachable in a VPN, used for service chaining (policy modules).
Each edge sends its prefixes and TLOCs to the Controller; the Controller sends the others' back. Data never goes through the Controller.
From OMP route to forwarding
An edge installs an OMP route in a VPN's routing table only if its TLOC is resolved: that is, a BFD session to that TLOC is up. This is why a prefix can be present in OMP yet unusable. Installed routes appear with code m, administrative distance 251, and point to the remote system IP out of the logical Sdwan-system-intf interface.
E-DEL-1# show sdwan omp routes vpn 10
PATH ATTRIBUTE
VPN PREFIX FROM PEER ID LABEL STATUS TYPE TLOC IP COLOR ENCAP PREFERENCE
-----------------------------------------------------------------------------------------------------
10 10.1.10.0/24 0.0.0.0 64 1012 C,Red,R installed 10.255.0.11 biz-internet ipsec -
10 10.1.10.0/24 0.0.0.0 65 1012 C,Red,R installed 10.255.0.11 mpls ipsec -
10 10.3.10.0/24 10.255.255.2 1 1012 C,I,R installed 10.255.0.13 biz-internet ipsec -
E-DEL-1# show ip route vrf 10 | begin Gateway
10.0.0.0/8 is variably subnetted, 3 subnets, 2 masks
C 10.1.10.0/24 is directly connected, GigabitEthernet3
L 10.1.10.1/32 is directly connected, GigabitEthernet3
m 10.3.10.0/24 [251/0] via 10.255.0.13, 00:00:03, Sdwan-system-intf
Read it from the left: VPN 10, prefix 10.3.10.0/24, learned from the Controller (10.255.255.2), next hop is the TLOC of E-BLR-1 over biz-internet. In the routing table the same prefix appears with code m. The first two lines (FROM PEER 0.0.0.0) are the edge's own prefix, sent to the Controller.
The status codes
| Code | Meaning |
|---|---|
| C | Chosen: selected as best |
| I | Installed in the routing table |
| R | Resolved: the TLOC is reachable (BFD up) |
| Red | Redistributed: our own prefix, sent to the Controller |
| Inv, U | Invalid or unresolved: received but unusable, normally no BFD to that TLOC |
| S | Stale: kept during a graceful restart |
Looking inside one route
Add a prefix and the word detail to see all attributes of the route, including which TLOC is the next hop and which peer sent it. In the Zenith Logistics fabric (lab 2), E-DEL-1 learns 10.4.10.0/24 from E-HYD-1:
E-DEL-1# show sdwan omp routes vpn 10 10.4.10.0/24 detail
---------------------------------------------------
omp route entries for vpn 10 route 10.4.10.0/24
---------------------------------------------------
RECEIVED FROM:
peer 10.255.255.2
path-id 1
label 1002
status C,I,R
Attributes:
originator 10.255.0.14
type installed
tloc 10.255.0.14, mpls, ipsec
site-id 400
origin-proto connected
origin-metric 0
Two key answers: the originator (10.255.0.14, the system IP of E-HYD-1) and the TLOC (colour mpls, so today the traffic uses the MPLS tunnel). The path-ID and label are internal identifiers; the label is what marks the packet as belonging to VPN 10 inside the tunnel.
What an edge advertises, and how many paths
By default an edge advertises the connected and static routes of each VPN. Routes learned from OSPF, BGP or EIGRP are not advertised unless you ask, under sdwan, omp, address-family ipv4 vrf. The Controller sends at most send-path-limit paths per prefix (default 4) and the edge installs up to ecmp-limit (default 4) equal-cost paths. You can check the OMP state of an edge quickly with show sdwan omp summary.
! Reference configuration: also advertise BGP and OSPF routes of VPN 10
config-transaction
sdwan
omp
address-family ipv4 vrf 10
advertise connected
advertise static
advertise bgp
advertise ospf external
exit
exit
exit
commit
Worked example. E-DEL-1 has two TLOCs and E-HYD-1 (in lab 2, after the second transport is added) also has two. For prefix 10.4.10.0/24 the Controller sends two paths: via 10.255.0.14 mpls and via 10.255.0.14 biz-internet. E-DEL-1 installs both (ecmp-limit 4), so show sdwan omp routes shows two installed lines. Before the second transport was added there was one.
Common mistakes. Reading FROM PEER 0.0.0.0 as a fault (it is our own route); expecting OSPF routes to reach other sites without advertise ospf; forgetting that a received prefix with Inv,U is not in the routing table; confusing the TLOC's system IP with the interface used for the tunnel.
Exam traps. OMP administrative distance on IOS XE is 251. By default only connected and static routes are redistributed into OMP. A route is installed only if its TLOC is resolved. The Controller reflects, it does not forward. Default send-path-limit is 4 and ecmp-limit is 4.
The route that was in OMP but not in the table
After a firewall change at the Mumbai data centre, users in Delhi could not reach the 10.2.10.0/24 servers. The NOC saw the prefix in show sdwan omp routes but with status Inv,U, and no m route in the VRF. That pointed straight to the TLOC: no BFD session was up to it, because the new firewall rule blocked the IPsec UDP ports between the sites. After the rule was corrected, BFD came up, the route became C,I,R and traffic resumed.
Lesson: an unresolved route is a data-plane problem, not a control-plane one. Look at BFD.
"How does a prefix learned at one branch reach another branch in Catalyst SD-WAN?"
Strong answer: the source edge advertises the prefix to the Controller in OMP with its TLOC as next hop and a VPN label; the Controller applies control policy and reflects it to other edges; each receiving edge installs it only if its TLOC is resolved through a BFD session, with administrative distance 251; the traffic then goes straight between the edges inside an IPsec tunnel using the TLOC's addresses.
Key takeaways
- OMP runs edge to Controller only; it carries OMP routes, TLOC routes and service routes, plus keys and policy.
- A route is installed only when its TLOC is resolved (BFD up); it then appears with code m and AD 251.
- Status codes: C chosen, I installed, R resolved, Red redistributed, Inv,U unusable, S stale.
- Connected and static are advertised by default; OSPF, BGP and EIGRP need explicit advertise commands.
- Use show sdwan omp routes (with detail), show ip route vrf, and show sdwan omp tlocs together.
The data plane: IPsec tunnels, BFD, and following one packet
The control plane decides where traffic should go; the data plane actually moves it. In this chapter you will see how the tunnels are built and protected, what BFD does inside them, and how to follow a single packet from a PC in Delhi to a server in Hyderabad, hop by hop, using the commands from the earlier chapters. This is the second lab of the module.
IPsec tunnels without IKE
User traffic travels in IPsec ESP tunnels between TLOCs. The tunnel is carried inside UDP so it can cross NAT. Unlike a classic VPN there is no IKE negotiation between edges by default. Instead, each edge generates its own encryption key (AES-256-GCM) and advertises it to the Controller, which hands it out to the other edges inside OMP. Keys are refreshed on a timer (the rekey interval, default one day), and an edge keeps the previous key for a short time so traffic does not drop during the change. Newer software also offers pairwise keys, where each pair of edges uses its own key; the principle, that the Controller distributes the key material, stays the same. GRE is a possible encapsulation, mostly toward service devices, but IPsec is the normal case.
BFD inside every tunnel
Inside each IPsec tunnel the two edges run BFD, a very light hello exchange. It is the only way an edge learns that the remote TLOC is alive over that transport. Default values: hello interval 1000 ms, multiplier 7, so a dead tunnel is detected in about seven seconds. BFD also samples loss, latency and jitter for each tunnel; these samples feed application-aware routing (the poll interval and multiplier are set under bfd app-route). BFD up means the TLOC is resolved and OMP routes that point to it can be installed.
E-DEL-1# show sdwan bfd sessions
SOURCE TLOC REMOTE TLOC DST PUBLIC DST PUBLIC DETECT TX
SYSTEM IP SITE ID STATE COLOR COLOR SOURCE IP IP PORT ENCAP MULTIPLIER INTERVAL(msec) UPTIME TRANSITIONS
-----------------------------------------------------------------------------------------------------------------------------------------------
10.255.0.12 200 up biz-internet biz-internet 203.0.113.6 203.0.113.10 12346 ipsec 7 1000 0:00:00:03 0
10.255.0.12 200 up mpls mpls 172.16.1.2 172.16.2.2 12366 ipsec 7 1000 0:00:00:03 0
10.255.0.13 300 up biz-internet biz-internet 203.0.113.6 203.0.113.14 12346 ipsec 7 1000 0:00:00:03 0
E-DEL-1# show sdwan bfd summary
sessions-total 3
sessions-up 3
sessions-max 3
sessions-flap 0
Count the sessions. E-DEL-1 has two TLOCs (internet and MPLS), E-MUM-1 two, E-BLR-1 one. Because MPLS is restricted, E-DEL-1 forms internet-to-internet and MPLS-to-MPLS with Mumbai (2) and internet-to-internet with Bengaluru (1): 3 sessions. The TRANSITIONS column counts how often a session changed state; a large number means a flapping tunnel. Edges with the same site ID never build tunnels to each other.
Packet path in the Zenith Logistics lab. The underlay routers do not appear in a traceroute from a service VPN.
Follow the packet: PC-DEL to SRV-HYD (lab 2)
The developer asks how traffic from 10.1.10.10 (Delhi) reaches the scanner server 10.4.10.10 (Hyderabad). Answer it in four steps, one command each.
- VRF routing table on E-DEL-1. The PC's gateway is the edge, so the lookup happens in VRF 10.
E-DEL-1# show ip route vrf 10 | begin Gateway C 10.1.10.0/24 is directly connected, GigabitEthernet3 m 10.4.10.0/24 [251/0] via 10.255.0.14, 00:05:12, Sdwan-system-intfCodem, administrative distance 251, next hop the system IP 10.255.0.14 of E-HYD-1. - OMP route detail.
show sdwan omp routes vpn 10 10.4.10.0/24 detailshows that the Controller 10.255.255.2 sent the route and that the TLOC is10.255.0.14, mpls, ipsec(see chapter 5). - BFD.
show sdwan bfd sessionslists the session to 10.255.0.14. Because E-HYD-1 has only an MPLS TLOC, and MPLS is restricted, the only possible pair is MPLS to MPLS: one session, colourmpls. That is the colour the flow uses today. - Traceroute. From the PC (or with
traceroute vrf 10 10.4.10.10on the edge) you see the remote edge's LAN address and then the server. The WAN is one hop:E-DEL-1# traceroute vrf 10 10.4.10.10 Tracing the route to 10.4.10.10 1 10.4.10.1 1 msec 2 msec 1 msec 2 10.4.10.10 2 msec 3 msec 2 msec
The underlay routers in the MPLS cloud are invisible because the original packet is wrapped in ESP and travels edge to edge.
Add the second transport and check the result
The last task of lab 2 brings Hyderabad's new broadband line (GigabitEthernet1, already addressed) into the overlay using the TLOC configuration from chapter 4: interface Tunnel1 with tunnel mode sdwan, and tunnel-interface with encapsulation ipsec and color biz-internet under sdwan, then commit. Expected result on E-DEL-1:
E-DEL-1# show sdwan bfd sessions SYSTEM IP SITE ID STATE COLOR COLOR SOURCE IP IP PORT ENCAP ... 10.255.0.14 400 up biz-internet biz-internet 203.0.113.6 203.0.113.10 12346 ipsec ... 10.255.0.14 400 up mpls mpls 172.16.1.2 172.16.2.2 12366 ipsec ... E-DEL-1# show sdwan omp routes vpn 10 10.4.10.0/24 VPN PREFIX FROM PEER ID LABEL STATUS TYPE TLOC IP COLOR ENCAP 10 10.4.10.0/24 10.255.255.2 1 1002 C,I,R installed 10.255.0.14 biz-internet ipsec 10 10.4.10.0/24 10.255.255.2 2 1002 C,I,R installed 10.255.0.14 mpls ipsec
Two BFD sessions (one per colour pair) and two installed paths: the prefix is now reachable over both transports, and the PC can still ping the server.
Worked example. A flow is going over MPLS and the carrier has a failure. BFD over MPLS stops answering; after the detect time (7 x 1000 ms) the session goes down, the MPLS TLOC becomes unresolved, the MPLS path of the OMP route is no longer usable, and the edge keeps forwarding over the remaining internet path. No routing protocol reconvergence is needed, and the PC sees only a short pause.
Common mistakes. Looking for the underlay routers in a traceroute; assuming "BFD up" means "quality good" (it only says the TLOC is alive; quality comes from the measurements and policy); expecting tunnels between edges that share a site ID; testing from the edge's CLI without vrf 10 and then wondering why the ping uses the wrong table.
Exam traps. No IKE between edges by default: the Controller distributes keys inside OMP. BFD runs inside the IPsec tunnels, default hello 1000 ms and multiplier 7. BFD is what resolves a TLOC and feeds application-aware routing with loss, latency and jitter.
Flapping tunnel to Bengaluru
Delhi users reported that file transfers to Bengaluru kept stalling. show sdwan bfd sessions showed the internet session to E-BLR-1 with the state "up" but a TRANSITIONS counter of 37 in a day and an uptime of only a few minutes. The NOC concluded the underlay was dropping BFD packets intermittently. They opened a ticket with the Bengaluru ISP using the transition counts as evidence, and in the meantime application-aware routing moved the critical flows to MPLS-backed paths through Mumbai.
Lesson: state alone is not enough; look at uptime and transitions to catch instability.
"How does a failure of one transport get detected and handled?"
Strong answer: BFD runs inside each tunnel; after the detect time (hello 1 s, multiplier 7 by default) the session goes down, the TLOC is marked unresolved, OMP routes via that TLOC stop being used, and the edge uses the remaining tunnels. Add that BFD also measures loss, latency and jitter so application-aware policy can react to degradation, not just failure.
Key takeaways
- User data travels edge to edge in IPsec tunnels over UDP; the Controller distributes the keys through OMP.
- BFD runs inside each tunnel (hello 1000 ms, multiplier 7) and resolves the TLOC.
- To follow a packet: show ip route vrf, show sdwan omp routes detail, show sdwan bfd sessions, traceroute vrf.
- Session count depends on TLOC pairs, restrict and site IDs; same site ID means no tunnels.
- Adding a TLOC at one site adds a BFD session and an installed path at its peers.
Segmentation with VPNs: building a guest segment (lab 3)
A hotel, a hospital or a retailer carries traffic of very different trust levels over the same routers: staff systems, guest Wi-Fi, payment terminals, cameras. Putting all of it in one routing table and then filtering with access lists is fragile. SD-WAN solves this with segmentation: separate routing domains that stay separate across the whole overlay. This chapter explains the VPN types, how routes and labels keep segments apart, and walks through the third lab step by step.
VPN means segment
In Catalyst SD-WAN a VPN is a separate routing domain that exists end to end, in every edge that has it. It is not an encrypted tunnel; the encryption is done by IPsec between edges. On IOS XE edges each VPN is a VRF. Three kinds exist:
| VPN | Purpose | IOS XE form |
|---|---|---|
| 0 | Transport: WAN interfaces, TLOCs, control connections, underlay routes | Global routing table |
| 512 | Out-of-band management | VRF Mgmt-intf |
| 1 to 511 and 513 to 65527 | Service VPNs: users, servers, guest, IoT, payment and so on | vrf definition N and vrf forwarding N on the interface |
Each service VPN has its own OMP routes (carried with a VPN label), its own routing table on every edge, and its traffic is marked with that label inside the IPsec tunnel. The receiving edge uses the label to put the packet into the right VRF. A host in VPN 20 has no route to VPN 10 unless a policy deliberately leaks routes between them (for example with a control policy or through a firewall). Only edges that have a VPN configured receive routes for it.
Two segments across the same pair of edges. Guest hosts have no route to corporate prefixes.
Lab 3: guest Wi-Fi for Sahyadri Hotels
Sahyadri Hotels runs its corporate systems in VPN 10 at Pune (E-PNQ-1, 10.1.10.0/24) and Goa (E-GOI-1, 10.2.10.0/24). Guest Wi-Fi has to travel over the same edges in its own segment: VPN 20 on GigabitEthernet4 (Pune 10.1.20.1/24, Goa 10.2.20.1/24). Guests at both hotels may reach each other's guest portal, but must never reach anything in VPN 10.
Three things are needed on each edge: the VRF, the interface in the VRF, and permission for OMP to advertise that VRF's connected network.
! E-PNQ-1 (repeat on E-GOI-1 with 10.2.20.1)
config-transaction
vrf definition 20
address-family ipv4
exit-address-family
exit
interface GigabitEthernet4
description GUEST-WIFI
vrf forwarding 20
ip address 10.1.20.1 255.255.255.0
no shutdown
exit
sdwan
omp
address-family ipv4 vrf 20
advertise connected
exit
exit
exit
commit
end
Order matters: define the VRF before putting the interface in it (vrf forwarding removes any existing IP address from the interface, so re-enter the address after it). Production templates generate extra lines for a VRF, such as a route distinguisher; the lab accepts the short form shown here. advertise connected is what makes OMP carry 10.1.20.0/24 into the overlay. Without it the guest prefix is local only.
Verify, step by step
1. The Controller holds both guest prefixes in VPN 20.
VSMART# show omp routes vpn 20
PATH ATTRIBUTE
VPN PREFIX FROM PEER ID LABEL STATUS TYPE TLOC IP COLOR ENCAP PREFERENCE
-----------------------------------------------------------------------------------------------------------
20 10.1.20.0/24 10.255.0.21 66 1004 C,R installed 10.255.0.21 biz-internet ipsec -
20 10.2.20.0/24 10.255.0.22 66 1004 C,R installed 10.255.0.22 biz-internet ipsec -
2. The edge installs the remote guest prefix in VRF 20.
E-PNQ-1# show ip route vrf 20 | begin Gateway
10.0.0.0/8 is variably subnetted, 3 subnets, 2 masks
C 10.1.20.0/24 is directly connected, GigabitEthernet4
L 10.1.20.1/32 is directly connected, GigabitEthernet4
m 10.2.20.0/24 [251/0] via 10.255.0.22, 00:00:41, Sdwan-system-intf
3. Guests reach guests, and nothing else. From GST-PNQ, ping 10.2.20.10 succeeds. ping 10.2.10.10 (the Goa corporate host) fails, and the VRF 20 table above contains no 10.2.10.0/24 or any other VPN 10 prefix. VPN 10 itself is unchanged: CORP-PNQ still reaches 10.2.10.10. No access list was needed: the guest VRF simply has no route.
Worked example. Count the routes. Pune's VRF 20 has two connected entries (the network and the local address) and one OMP route to Goa, three lines. VRF 10 at Pune is separate and holds the corporate prefixes. If a third hotel, Nashik, adds VPN 20, Pune and Goa learn its guest prefix automatically; if Nashik only has VPN 10, its guest hosts do not exist and Pune never receives anything for VPN 20 from it.
Segmentation design tips
- Use the same VPN numbering at every site (VPN 10 corporate everywhere). The same number joins the pieces into one segment.
- Keep VPN 0 free of user traffic; keep VPN 512 for management only.
- When two segments must talk (for example guests reaching a public portal in a DMZ), do it on purpose with a control policy that leaks routes, or send the traffic through a firewall, and document it.
Common mistakes. Forgetting commit; forgetting advertise connected in the VRF (the prefix never reaches the Controller); typing the IP address before vrf forwarding so it is wiped; creating VPN 20 on one site only and expecting the other to learn it; thinking an ACL is needed to isolate VPNs.
Exam traps. VPN 0 is transport (global table on cEdge), VPN 512 is management (Mgmt-intf), service VPNs are everything else. A service VPN is a segment, not an encrypted tunnel. Routes of a VPN reach only edges that have that VPN. Each VPN has its own label inside the IPsec tunnel.
Guests that could see the finance printer
A hotel launched guest Wi-Fi by copying the staff LAN interface configuration, which put guests in VPN 10. Security noticed that a guest laptop could reach the finance printer. The fix was to create VPN 20 on the router, move the guest interface into it and advertise it in OMP. After the change the guest VRF had no route to any corporate prefix, and a penetration tester's scan from the guest network found nothing in VPN 10.
Lesson: isolation comes from separate routing tables, not from remembering to write the right access list.
"How does SD-WAN keep two segments separate across the WAN?"
Strong answer: each segment is a VPN (a VRF on IOS XE) with its own routing table on every edge; OMP carries routes per VPN with a label, and the label inside the IPsec tunnel tells the receiving edge which VRF to use. Prefixes of a VPN are advertised only to edges that have it. Between VPNs there is no route unless you configure leaking by policy or via a firewall.
Key takeaways
- A VPN is a routing segment (a VRF on IOS XE); VPN 0 transport, VPN 512 management, others service VPNs.
- OMP routes carry a VPN label; one IPsec tunnel carries many segments.
- To add a segment: vrf definition, vrf forwarding on the interface, advertise connected under sdwan omp, commit.
- Verify with show omp routes vpn on the Controller and show ip route vrf on the edge.
- Isolation needs no ACL; the segment simply has no route to the other.
A troubleshooting workflow: control, OMP, BFD, then the data path
SD-WAN is built in layers that depend on each other: the control connections must be up before OMP can run, OMP must deliver a route and its TLOC before BFD is needed, and a resolved TLOC is required before traffic flows. If you check the layers in that order, you find the first broken layer and you can stop. This chapter turns the commands from chapters 2 to 7 into one routine, then applies it to realistic faults. It is also the right routine for the three labs.
The routine in five steps
Layer by layer. A failure in an earlier layer explains symptoms in every later one.
| Step | Question | Command | Healthy sign |
|---|---|---|---|
| 1 | Is the device connected to the controllers? | show orchestrator connections (VBOND), show control connections (VSMART), show sdwan control connections (edge) | State up; expected number of rows; organization name equal everywhere |
| 2 | Do OMP sessions and routes exist? | show omp peers (VSMART), show sdwan omp summary and show sdwan omp routes vpn N (edge) | Peer up; the prefix is present with a TLOC |
| 3 | Does the TLOC look right? | show sdwan omp tlocs | Correct colour, public address and port |
| 4 | Is BFD up to the remote TLOC? | show sdwan bfd sessions, show sdwan bfd summary | Expected count, all up, few transitions |
| 5 | Is the route in the VRF, and does traffic flow? | show ip route vrf N, ping, traceroute vrf N | Code m route, successful ping, one WAN hop |
If control connections are down, show sdwan control connection-history on the edge explains why. It lists past attempts with error codes. The commonly seen codes include DCONFAIL (the DTLS connection could not be established, often blocked ports or no route), CRTVERFL (certificate could not be verified), CTORGNMMIS (organization name mismatch), and VB_TMO, VM_TMO, VS_TMO (timeouts towards Validator, Manager or Controller). The controller deployment module treats these in depth.
Symptom to likely cause
| What you see | Likely cause | What to check or fix |
|---|---|---|
Edge not in show omp peers, no vsmart row on the edge | Control connection down: wrong vbond, organization name, certificate, clock, blocked ports | show sdwan control connection-history; the system block; reachability of VBOND |
Prefix in OMP with Inv,U | TLOC unresolved: no BFD to that TLOC | BFD sessions; colours and restrict; firewall between sites |
| Prefix missing everywhere | Remote site is not advertising it: forgot advertise connected or commit, wrong VRF, or protocol not advertised | On the remote edge: show sdwan omp routes (Red lines); config under sdwan omp |
| Fewer BFD sessions than expected | Restrict, colour mismatch, same site ID, or one transport down | TLOC list; site IDs; the interface state |
| Ping fails but route is installed | Remote LAN, gateway or host issue; one-way asymmetry; policy | Ping from the remote edge in the VRF; check the LAN interface |
| A segment can reach another it should not | Both interfaces are in the same VRF | show ip vrf interfaces; vrf forwarding lines |
Applying the routine to the labs
Lab 1 (read a live fabric). This is the routine in healthy mode: confirm what normal looks like. Validator rows (2 controllers), edge rows (2 vsmart + 1 vmanage), OMP peers (3), BFD sessions (3), ping and traceroute.
Lab 2 (follow a packet). Steps 5 to 2 in reverse: VRF route (m, 251), OMP route detail (TLOC), BFD (colour), traceroute. After adding the broadband TLOC at Hyderabad, check step 1 (control connection for the new TLOC), step 3 (new TLOC in the list) and step 4 (a second session) and step 5 (two installed paths).
Lab 3 (guest segmentation). If the guest ping fails, check step 2 on the Controller first: show omp routes vpn 20. If a prefix is missing, the source edge did not advertise it; look at its config for the VRF, the interface in the VRF and advertise connected, and make sure you typed commit.
Worked ticket. "Delhi cannot reach Bengaluru." Step 1: E-DEL-1 and E-BLR-1 both show up control connections. Step 2: show sdwan omp routes vpn 10 on E-DEL-1 lists 10.3.10.0/24 with status Inv,U. Step 3: the TLOC is 10.255.0.13 biz-internet, looks right. Step 4: show sdwan bfd sessions has no session to 10.255.0.13. The prefix is received but unresolved, so the fault is in the tunnel between the edges. A firewall change between the internet breakout of the two sites is blocking the IPsec UDP ports. When BFD comes up, the route changes to C,I,R and the ping works. The five-step routine got there without guessing.
Common mistakes. Starting at the ping and assuming the routing is wrong; changing configuration before reading the current state; forgetting commit and then testing a change that was never applied; reading only one side (the Controller says everything is fine because it only sees control, not data).
Triage in five minutes
At 09:10 the Bengaluru store reported that the shared inventory application was unreachable. The on-call engineer ran the routine. Control connections: up. OMP peers: up. For the routes: 10.3.10.0/24 was fine, but the Delhi prefix was missing at Bengaluru. On E-DEL-1 the prefix was there with the line Red absent. A change at 08:50 had replaced the Delhi LAN interface configuration and the interface was now outside VRF 10, so the connected route was no longer advertised into VPN 10. The engineer restored vrf forwarding 10 and the IP address, committed, and the route appeared on the Controller within seconds.
Lesson: the fault was at the very first edge, and the routine pointed to it quickly because it checked the sender before the receiver.
"A user cannot reach a server at another site over SD-WAN. How do you troubleshoot?"
Strong answer: follow the layers. Check control connections of both edges; check OMP peers and whether the prefix is received on the local edge with a TLOC and status C,I,R; check the TLOC colour and address; check BFD to the remote TLOC; then the VRF route, ping from the edge in the VRF and traceroute. Say that the first layer that fails is the cause, and that you compare against the design (expected TLOC and session counts) before changing anything.
Key takeaways
- Check in this order: control connections, OMP, TLOC, BFD, then VRF route and data path.
- Compare outputs with the design: expected rows, TLOCs and sessions.
- Inv,U means received but unresolved: look at BFD and the firewall between sites.
- A prefix missing everywhere points at the originating edge: advertise commands, VRF membership, commit.
- Use show sdwan control connection-history for control failures and read the error codes.
Summary: checklist, glossary and the most-tested facts
This chapter is your revision sheet for the architecture module. Read it the evening before an exam or an interview, then go back to any chapter where a line does not feel automatic. Everything here was built up in chapters 1 to 8 and used in the three labs.
The whole architecture on one page
Four planes, four jobs. If you can redraw this from memory, you own the module.
Can-do checklist
- I can name the four planes and say which component belongs to each, using both new and old names.
- I can say who connects to whom: every device reaches the Validator first, then the Controller and the Manager.
- I can explain controller mode on a cEdge and why the authorised WAN edge list matters.
- I can read a TLOC (system IP, colour, encapsulation) and explain private and public colours and
restrict. - I can read
show sdwan omp routes, including the status codes C, I, R, Inv and U. - I can explain why IPsec runs without IKE and what BFD adds.
- I can explain VPN 0, VPN 512 and service VPNs, and build a guest segment.
- I can walk through the five-step troubleshooting routine.
Mini glossary
| Term | Meaning |
|---|---|
| Manager, Controller, Validator | Formerly vManage, vSmart, vBond |
| WAN edge | Router in the data plane: cEdge (IOS XE) or vEdge |
| TLOC | Transport locator: system IP + colour + encapsulation |
| Colour | Label of a transport; public colours are assumed to be behind NAT, private ones are not |
| OMP | Overlay Management Protocol, runs between each device and the Controller |
| BFD | Liveness and quality probe inside each tunnel |
| VPN / VRF | Segment; the number is carried as a label in the data plane |
| Site ID | Identifies a location; edges of one site share it |
| System IP | Permanent device identity, like a router ID |
Most-tested facts
| Fact | Value |
|---|---|
| OMP administrative distance | 251 |
| Control connection default | DTLS over UDP, base port 12346 |
| BFD default hello and multiplier | 1000 ms and 7 |
| Default OMP path limits | send-path-limit 4, ecmp-limit 4 |
| Management VPN | 512 |
| Transport VPN | 0 |
| Route not usable because its TLOC is unresolved | Inv,U |
| Normal healthy installed route | C,I,R |
Command cheat-sheet
# Validator show orchestrator connections # Controller show control connections show omp peers show omp routes vpn 10 # WAN edge show sdwan control connections show sdwan control connection-history show sdwan omp summary show sdwan omp routes vpn 10 show sdwan omp tlocs show sdwan bfd sessions show ip route vrf 10 traceroute vrf 10 10.3.10.1
Common mistakes. Mixing up the planes (the Manager is not in the control plane); believing the Validator carries user traffic; reading a service VPN as an encrypted tunnel; forgetting that OMP routes need a resolved TLOC; testing before commit.
Exam traps. The Validator is the first contact and orchestrates; the Controller runs OMP; the Manager configures and monitors. OMP has three route types: routes, TLOC routes and service routes. A TLOC is not an interface address. A colour is a label with a public or private meaning. BFD runs inside IPsec tunnels, not between edge and Controller.
Redraw test
A new colleague asks you to explain why Mumbai can reach Delhi even though the two edges never exchange any routing protocol packets. Your answer: both edges have OMP sessions to the Controller. Each advertises its prefixes with a TLOC; the Controller reflects them. The edges reach each other's TLOC address directly, build an IPsec tunnel, and BFD proves it works. The route becomes usable. Control goes through the centre; data does not.
"Describe the Catalyst SD-WAN architecture in one minute."
Strong answer: four planes. The Validator orchestrates the first contact and authenticates. The Manager is the management plane: templates, monitoring, certificates. The Controller is the control plane and runs OMP, distributing routes, TLOCs, keys and policy. WAN edges form the data plane with IPsec tunnels over any transport, measured by BFD. Segmentation uses VPNs, with VPN 0 for transport and VPN 512 for management.
Key takeaways
- Four planes: orchestration, management, control, data.
- TLOC = system IP + colour + encapsulation; OMP carries three route types.
- IPsec without IKE; BFD gives liveness and quality.
- VPN 0 transport, VPN 512 management, service VPNs for users.
- Troubleshoot layer by layer and compare to the design.