Fortinet Secure SD-WAN ยท Basic SD-WAN Setup

SD-WAN: members, SLAs & rules

FortiGate SD-WAN: members and zones, performance SLAs, health checks, steering rules and link failover.

29 min read9 chapters3 labs15 quiz7 scenarios15 interview Q&A

This first module is free: read the lesson and take the quiz. Create a free account to run up to 3 hands-on labs.

Log inStart free
Jump to chapter (9)
01

SD-WAN on the FortiGate: what you will learn and the big picture

What you will learn in this module. You will learn how FortiGate SD-WAN bundles several internet or private links into one logical interface, measures their quality, and steers each type of traffic over the best link. You will build members and zones, health checks and SLA targets, SD-WAN rules, and understand the implicit rule, load balancing and failover. These are the SD-WAN objectives of the Fortinet NSE 4 / FortiOS 7.6 Administrator exam.

Note. Honest note about the simulator. The SD-WAN core runs in the simulator labs of this module: zones, members, health checks, SLA targets, rules, failover and diagnose sys sdwan. Try Two ISPs become one SD-WAN zone, A performance SLA and failover when ISP-A degrades and the ticket lab Half the users cannot browse. Overlay IPsec, ADVPN, FortiManager, application-control steering, passive measurement and traffic shaping are reference-only: they are taught in the later fsd modules but are not modelled. The routing parts that SD-WAN builds on (routing table, floating routes, debug flow) are also real labs in nse4-routing. For hands-on policy routing and tracking concepts use enarsi-pbr-vrf-bfd; for Cisco's different SD-WAN architecture see encor-sdwan-sda.

Prerequisites. Finish nse4-routing first: you must be comfortable with the routing table, distance, floating routes and firewall policies per interface.

Analogy: a delivery company with a live traffic app

A floating static route is like telling every driver: "always take the highway; if it is closed, take the village road". It cannot tell that the highway is open but jammed, and it treats an urgent medicine parcel and a catalogue the same. SD-WAN is a dispatcher with a live traffic app. It measures every road all the time (loss, delay, jitter), knows what each parcel is, and says: "medicine over the fastest road, catalogues over the cheap road, and if a road gets bad, switch."

LAN users10.0.1.0/24SD-WAN zonevirtual-wan-linkmember 1: port1ISP-Amember 2: port4ISP-BRules + SLAsteer by applicationHealth checksloss, latency, jitter

One logical interface (the zone) hides several physical links. Rules decide which member carries which traffic, health checks supply the measurements.

The five building blocks

BlockWhat it is
MemberOne link (interface and gateway) that joins SD-WAN.
ZoneA group of members; the object that routes and policies point to.
Health checkA probe (ping, DNS and so on) that measures each member.
SLAThe thresholds that decide whether a member is good enough.
Rule (service)Which traffic goes over which members and by which strategy.

Traffic that matches no rule falls to the implicit rule, which uses the normal routing table and a load-balancing mode. The next chapters build each block in order.

Why SD-WAN and not floating routes

  • Failover on quality, not only on a dead interface.
  • Per-application steering: voice on the best link, backups on the cheap one.
  • Use of all links at once instead of leaving the backup idle.
  • One policy for the zone instead of one per link.

Common mistake. Believing SD-WAN needs a special licence or a controller. On a FortiGate it is a built-in feature of FortiOS configured on each unit; FortiManager is optional for managing many sites.

Exam trap. SD-WAN is a FortiOS feature configured under config system sdwan. Policies and routes refer to the SD-WAN zone, not to the member interfaces.

The backup link that was never used

A retailer paid for two 100 Mbit links per store but used one; the second sat idle until an outage. After moving to SD-WAN, store cards and cloud tills used the primary while software updates and guest Wi-Fi used the second link. The same money bought double the capacity.

Lesson: SD-WAN turns an idle backup into usable capacity.

"What problem does SD-WAN solve compared with floating static routes?"

A floating route reacts only to interface state and sends everything over one link. SD-WAN measures loss, latency and jitter per link with health checks, steers each application over the best link by rule, can use all links at once, and gives one zone object for routes and policies.

Key takeaways

  • SD-WAN bundles links (members) into a zone and steers traffic by measured quality.
  • Building blocks: members, zone, health check, SLA, rules, implicit rule.
  • This module has three simulator labs: build the zone, add an SLA with failover, and fix a ticket.
  • Practise the routing foundations in nse4-routing.
02

SD-WAN members and zones

A member is a single WAN link described by an interface and, normally, the gateway of that link. A physical port, a VLAN, an aggregate or an IPsec tunnel interface can be a member. A zone is a named group of members. Out of the box FortiOS has one zone called virtual-wan-link; you can create more, for example UNDERLAY for ISP links and OVERLAY for VPN tunnels, so that rules and policies can address them separately.

! Try it: lab "Two ISPs become one SD-WAN zone"
config system sdwan
    set status enable
    config zone
        edit "virtual-wan-link"
        next
    end
    config members
        edit 1
            set interface "port1"
            set zone "virtual-wan-link"
            set gateway 203.0.113.1
        next
        edit 2
            set interface "port4"
            set zone "virtual-wan-link"
            set gateway 198.51.100.1
            set cost 10
        next
    end
end

The addresses are the same as the Pune branch of nse4-routing: ISP-A on port1 (203.0.113.2/30, gateway 203.0.113.1) and ISP-B on port4 (198.51.100.2/30, gateway 198.51.100.1). The member number is only an ID; it is the number you use later in rules and health checks.

SD-WAN: config system sdwanzone: UNDERLAYISP linksport1ISP-Aport4ISP-Bzone: OVERLAYVPN tunnelsHUB1-T1IPsecHUB1-T2IPsec

Two zones: rules can target ISP links and VPN tunnels separately.

Member attributes that matter

AttributeMeaning
interfaceThe port, VLAN, aggregate or tunnel interface.
zoneThe zone this member belongs to; default virtual-wan-link.
gatewayNext hop of this link. Not needed for DHCP or PPPoE links, which learn it.
costUsed by the cost-based strategy (lowest cost link preferred).
priorityPreference of the member when the routing table must pick one.
weightShare of sessions in weight-based load balancing.
statusenable or disable the member without deleting it.

The reference rule: clean the interface first

An interface cannot join SD-WAN while other parts of the configuration still use it by name. Before adding port1 and port4 as members you must remove or move every reference: static routes that point to the port, firewall policies that use it as source or destination, DHCP servers, and similar objects. Typical order of work on an existing branch:

  1. Note existing routes, policies and NAT that use port1 and port4.
  2. Delete the old default routes and the policy entries that name those ports.
  3. Create the members and zone.
  4. Create one default route for the zone and one policy to the zone (next chapter).
  5. Test from a client, then add health checks and rules.

Common mistake. Doing this on a production firewall without a maintenance window or console access. The moment the old routes and policies are removed, traffic stops until the new ones exist. Prepare the whole change, apply it quickly, or use a staging unit.

Exam trap. Know why an interface cannot be added as a member while it is referenced by a route, policy or DHCP server, and that the default zone is virtual-wan-link.

The member that would not save

An engineer added port4 as a member and the unit refused with an error that the interface was in use. A forgotten guest policy still listed port4 as destination. After removing that policy entry and recreating it against the zone, the member saved.

Lesson: Search the configuration for every use of an interface before turning it into a member.

"What is an SD-WAN zone and why use more than one?"

A zone is a named group of members that routes and policies refer to. Separate zones, for example underlay ISP links and overlay VPN tunnels, let rules and policies treat them differently and make it easy to keep internet-bound and site-to-site traffic apart.

Key takeaways

  • A member is an interface plus usually its gateway; members live in zones.
  • The default zone is virtual-wan-link; extra zones separate underlay and overlay.
  • Cost, priority and weight feed strategies and load balancing.
  • Remove references to an interface before adding it as a member.
03

Wiring it together: the zone route and the firewall policy

Members and zones alone carry no traffic. You need two more things that you already know from nse4-routing: a route and a policy. In SD-WAN, both point at the zone. The route tells the FortiGate "internet traffic leaves through the zone"; the policy says "the LAN may go to the zone, with NAT". SD-WAN then chooses the member.

! Try it: lab "Two ISPs become one SD-WAN zone"
config router static
    edit 1
        set sdwan-zone "virtual-wan-link"
        set comment "Default route over the SD-WAN zone"
    next
end
config firewall policy
    edit 10
        set name "LAN-to-SDWAN"
        set srcintf "port2"
        set dstintf "virtual-wan-link"
        set srcaddr "all"
        set dstaddr "all"
        set action accept
        set schedule "always"
        set service "ALL"
        set nat enable
    next
end

Compare this with the floating-route design of nse4-routing. There you needed two default routes with different distances and one policy that listed both interfaces. Here there is one default route and one policy, and the choice between ISP-A and ISP-B moves into SD-WAN where it can use quality measurements. The route has no gateway of its own because each member carries its gateway.

port2 LAN10.0.1.0/24Policy 10port2 -> virtual-wan-linkRoute 10.0.0.0/0 -> zoneSD-WAN picks a memberport1 / port4

The route and the policy both point at the zone; SD-WAN picks the member.

What changes in the routing table

With SD-WAN enabled, the zone's routes appear in the routing table through the member interfaces, and the table behaves like the ECMP example in chapter 6 of nse4-routing: the member gateways are the next hops. SD-WAN rules and the implicit rule decide which of them a given session uses. When a member is considered dead by its health check, its next hop is not used for new sessions even though the interface itself may still be up. That is the big difference from a plain static route.

Other routes and policies

  • Routes to internal networks stay normal static routes (or OSPF/BGP) on the internal interfaces.
  • A static route can name a member interface directly when you need a specific path, but then that traffic skips SD-WAN decisions.
  • Policies to a zone apply to every member; you no longer need one per ISP.
  • NAT still uses the address of the egress interface (see the failover story in nse4-routing).

Common mistake. Leaving an old floating default route (distance 20 via port4) in place after SD-WAN is enabled. It overlaps with the zone route and confuses troubleshooting. Remove it, or its interface will not be allowed as a member.

Exam trap. Policies and the default route refer to the SD-WAN zone. The members' own interfaces are not used in policies. NAT still follows the egress interface.

One policy instead of four

A bank branch template had four policies, one per ISP direction and protocol, and every new link meant new copies. Moving the template to a zone gave one policy and one default route per branch; adding a third link later was a three-line change on each unit.

Lesson: Zones make templates smaller and growth cheaper.

"Which objects point to the SD-WAN zone instead of to physical interfaces?"

The default static route (set with sdwan-zone) and the firewall policies (srcintf or dstintf). Internal routes and policies still use physical interfaces. NAT uses the address of whichever member carries the session.

Key takeaways

  • One default route and one policy point at the zone.
  • Each member carries its own gateway.
  • A dead member is skipped for new sessions even if its interface is up.
  • Remove old floating routes and per-ISP policies after migrating.
04

Health checks (performance SLAs): measuring every link

A static route knows only whether the cable is up. SD-WAN wants to know whether the path works. A health check (called a performance SLA in the GUI) sends small probes through each member toward a reliable server, waits for answers and measures three things: packet loss, latency (delay) and jitter (variation of the delay). From those it derives a state for every member: alive or dead.

! Try it: lab "A performance SLA and failover when ISP-A degrades"
config system sdwan
    config health-check
        edit "INTERNET-PING"
            set server "192.0.2.53"
            set protocol ping
            set interval 1000
            set failtime 5
            set recoverytime 5
            set members 0
        next
    end
end

The probe target 192.0.2.53 is a documentation address standing for a stable, always-on internet server that you trust; in real life pick something reliable and not controlled by the ISP you are measuring, such as a public DNS service or a server you own.

FortiGatesends probesport1 ISP-Aaliveport4 ISP-BaliveProbe server192.0.2.53Each member is measured separately through its own link.

A health check sends probes through each member and records loss, latency and jitter per member.

The parameters

ParameterMeaning
serverThe probe target: an IP address or a name.
protocolping, tcp-echo, udp-echo, http, twamp, dns, tcp-connect, ftp.
intervalMilliseconds between probes (default 500).
failtimeConsecutive lost probes before a member is declared dead (default 5).
recoverytimeConsecutive good probes before a dead member is alive again (default 5).
membersWhich members to probe; 0 means all.
update-static-routeRemove static routes of a dead member from the table.

How long does a failure take to detect?

Roughly interval times failtime. With an interval of 1000 ms and a failtime of 5, a member is declared dead after about five seconds of lost probes. Recovery works the same way with recoverytime. Short values react fast but make links "flap" on a single bad moment; long values are stable but slow. For most branches a few seconds is a good compromise. There is also a probe timeout: a probe that gets no answer within it counts as lost.

Choosing the protocol

  • ping: simplest; proves the path to a host exists. Some servers or ISPs rate-limit ICMP.
  • dns: asks a real name server, which proves that name resolution over that link works.
  • http or tcp-connect: proves that a TCP service answers; useful for a specific SaaS application.
  • twamp: standards-based latency measurement between FortiGates or a TWAMP responder.

Reading the result

On a real FortiGate diagnose sys sdwan health-check prints one block per health check and one line per member. Each line names the member, shows whether it is alive or dead, and lists loss, latency, jitter and an SLA map that says which SLA targets the member currently meets. Run it in the lab A performance SLA and failover when ISP-A degrades: the impair tool on a provider router adds delay, jitter or loss to a link, and you can watch the member's numbers and sla_map change. Probe latency in the simulator is one-way, so expect smaller numbers than a real round-trip probe.

Common mistake. Probing a target that is itself unreliable, or that only one ISP can reach. The member then looks dead although the link is fine. Use more than one server or a robust public service.

Exam trap. A health check is attached to members, measures loss, latency and jitter, and marks a member dead after failtime lost probes. Detection time is about interval times failtime.

The ISP that answered pings but dropped web

A branch used a ping health check against the ISP's own router. During a peering problem the ISP router kept answering while internet traffic failed, so SD-WAN saw the link as alive. Changing the health check to probe a public DNS server through the link made the failure visible within seconds.

Lesson: Probe something that proves the service you care about, not only the first hop.

"How does an SD-WAN health check decide that a link is dead?"

It sends probes at the configured interval through each member; after failtime consecutive lost probes the member is marked dead and is no longer used for new sessions. After recoverytime consecutive successes it is alive again. Detection time is about interval times failtime.

Key takeaways

  • A health check probes each member and measures loss, latency and jitter.
  • failtime and recoverytime decide when a member is dead or alive again.
  • Detection time is about interval times failtime.
  • Choose a reliable target and a protocol that proves your service works.
05

SLA targets: when is a link good enough?

"Alive" is not the same as "good". A link can answer probes but lose 8 percent of packets or take 400 ms, which ruins a voice call. An SLA target is a set of thresholds inside a health check: a member meets the SLA when its measured latency, jitter and packet loss are all within the thresholds. A member that is alive but misses the SLA is "out of SLA"; it still exists, but rules that require the SLA stop choosing it.

! Try it: lab "A performance SLA and failover when ISP-A degrades"
config system sdwan
    config health-check
        edit "INTERNET-PING"
            config sla
                edit 1
                    set link-cost-factor latency jitter packet-loss
                    set latency-threshold 100
                    set jitter-threshold 30
                    set packetloss-threshold 2
                next
            end
        next
    end
end

This target, number 1 inside the health check, says: latency up to 100 ms, jitter up to 30 ms, loss up to 2 percent. A health check can hold several SLA targets with different numbers, for example a strict one for voice and a relaxed one for web. Rules refer to a target as "health check name, SLA number".

Measured values against SLA 1 (thresholds: 100 ms, 30 ms, 2 percent)port1: meets SLAall three within limitsport4: misses SLAloss above thresholdport3: deadfailtime reachedalive + in SLA - alive + out of SLA - deadOnly the first can carry traffic that requires the SLA.Values in this picture are an illustration, not output from a device.

Three states of a member: in SLA, out of SLA, and dead.

link-cost-factor: ranking links by quality

The attribute link-cost-factor tells the "best quality" strategy which measurement to compare when several members are eligible: latency, jitter, packet-loss, inbound-bandwidth, outbound-bandwidth, bibandwidth, custom-profile-1 and others. Listing more than one makes the FortiGate combine them by weight. For a voice rule people usually choose latency and jitter; for a bulk-transfer rule, bandwidth.

Dead versus out of SLA

StateCauseEffect
Deadfailtime probes lost in a rowRemoved from every decision for new sessions
Alive, in SLAAll thresholds metEligible for all rules
Alive, out of SLAA threshold exceededSkipped by rules that require this SLA; still usable by others

Two optional behaviours help in practice. Logging of SLA changes (sla-fail-log-period and sla-pass-log-period) writes an event when a member crosses a threshold, which is the evidence you need for an ISP complaint. A threshold that is too strict makes the decision flip back and forth; start from the application's real need (voice is comfortable under about 150 ms one way and a few percent loss) and tune with data.

Common mistake. Setting all thresholds to tiny values "to be safe". Every link then looks out of SLA, the rule has no eligible member, and traffic falls back to default behaviour, which is exactly what you were trying to avoid.

Exam trap. SLA thresholds are latency, jitter and packet loss inside a health check. A member can be alive and still fail the SLA. link-cost-factor ranks members for the best-quality strategy.

Video calls that stuttered on a "healthy" link

Users complained about choppy video calls although the monitoring showed all links up. There was no SLA, so only dead links were avoided. After adding an SLA of 100 ms latency, 30 ms jitter and 2 percent loss and a rule that required it, calls moved to the second link whenever the first became lossy, and the log of SLA failures backed the ISP complaint.

Lesson: Up is not good; define good with an SLA.

"What is the difference between a dead member and a member that fails its SLA?"

Dead means the health check lost failtime probes in a row, so the member is excluded for new sessions. Failing the SLA means the member is alive but latency, jitter or loss exceeds a threshold, so only rules that require that SLA skip it. Describe an example with voice.

Key takeaways

  • An SLA target sets thresholds for latency, jitter and packet loss.
  • Three states: in SLA, out of SLA, dead.
  • link-cost-factor ranks eligible members by a chosen measurement.
  • Set thresholds from the application's real need and tune with data.
06

SD-WAN rules: steering traffic with strategies

An SD-WAN rule (a "service" in the CLI) says: "traffic that looks like this should use these members, chosen this way". It has a match part (source, destination, protocol and port, an Internet Service Database application such as a SaaS, or user) and a strategy part. Rules are read top to bottom and the first rule that matches decides; this is the same idea as a policy route, evaluated before the routing table.

! Try it: lab "A performance SLA and failover when ISP-A degrades"
config system sdwan
    config service
        edit 1
            set name "VOICE-LOW-LATENCY"
            set mode sla
            set dst "all"
            set src "VOICE-NET"
            set protocol 17
            set start-port 5060
            set end-port 5061
            config sla
                edit "INTERNET-PING"
                    set id 1
                next
            end
            set priority-members 1 2
        next
        edit 2
            set name "BACKUPS-CHEAP"
            set mode manual
            set dst "BACKUP-SERVERS"
            set priority-members 2
        next
    end
end

Rule 1 sends the voice network's SIP traffic over member 1 as long as it meets SLA 1 of the health check, otherwise over member 2. Rule 2 pins traffic to the backup servers to member 2, the cheap link (manual mode). In a rule, priority-members lists members in order of preference.

The strategies (modes)

GUI nameCLI modeWhat it does
ManualmanualAlways use the listed member (if alive).
Best Qualitypriority (link-cost-factor)Use the member with the best measured value of the chosen factor, for example latency.
Lowest Cost (SLA)slaUse the first member in priority-members that meets the SLA.
Maximize Bandwidth (SLA)load-balanceShare sessions over all members that meet the SLA.

Note. Mode names have changed between FortiOS releases and an older auto mode still exists for the quality-based behaviour. Always check the choices with set mode ? on the firmware you are using, and read the GUI strategy name in the exam as the authoritative description.

Matching on applications

Instead of addresses a rule can match Internet Service Database (ISDB) objects, which FortiGuard keeps up to date with the addresses of well-known SaaS and cloud services. In the CLI that is set internet-service enable and set internet-service-name, choosing the name from the list shown by ? on your build. Application control signatures can be used as well, but they only work after the first packets have been inspected, so a session can start on one link and be steered later.

Tie-break

When two members are equally good, tie-break decides: zone (the default), cfg-order (order in the configuration), fib-best-match or input-device. Set it explicitly when the result matters.

Common mistake. Putting a broad rule (dst all, src all) above a specific one. The first match wins, so the specific rule below it never runs. Order rules from most specific to least specific.

Exam trap. Rules are matched top-down and are evaluated before the routing table. A rule needs a strategy (mode) and members; SLA-based modes need a health check and SLA id.

Voice never moved

The company created a voice rule with a strict SLA, but calls stayed on the lossy link. The log showed that an older catch-all rule placed first matched all traffic. Moving the voice rule to position 1 fixed it without touching the SLA.

Lesson: Check rule order before suspecting the measurements.

"Name the SD-WAN rule strategies and when you would use each."

Manual for a fixed path; Best Quality to pick the lowest latency, jitter or loss link for voice or video; Lowest Cost (SLA) to use the cheap link as long as it meets the SLA; Maximize Bandwidth (SLA) to share bulk traffic across all good links. Mention that rules match top-down and before the routing table.

Key takeaways

  • A rule has a match part and a strategy; rules are read top-down.
  • Strategies: manual, best quality, lowest cost (SLA), maximize bandwidth (SLA).
  • priority-members lists members in order of preference.
  • ISDB objects let rules match SaaS applications.
07

The implicit rule, load balancing and failover

What happens to traffic that matches none of your rules? It is handled by the implicit rule, the last, invisible rule of every SD-WAN configuration. The implicit rule does not look at quality. It uses the normal routing table and shares sessions among the members that have a route, using the load-balance mode of the SD-WAN settings. If you write no rules at all, everything uses the implicit rule, and SD-WAN behaves like ECMP with health-aware members.

! Try it: lab "A performance SLA and failover when ISP-A degrades"
config system sdwan
    set load-balance-mode weight-based
    config members
        edit 1
            set interface "port1"
            set weight 70
        next
        edit 2
            set interface "port4"
            set weight 30
        next
    end
end
load-balance-modeHow sessions are shared
source-ip-basedDefault. Hash on the source address, so one client keeps one link.
weight-basedIn proportion to member weights (70/30 above).
usage-basedFill one link up to a limit, then use the next.
source-dest-ip-basedHash on source and destination addresses.
measured-volume-basedShare by the volume each member has already carried.
New sessionsrc, dst, portRule 1 matches?strategy picks memberRule N matches?first match winsImplicit rulerouting table + load balanceEgress memberalive and in SLA when the rule needs it

Every new session passes the rules top-down; if none matches, the implicit rule decides.

Failover in three situations

  1. Interface down. The member's routes are removed, exactly as with a floating route (nse4-routing chapter 5). No health check is needed.
  2. Member dead. The health check lost failtime probes. The member is dropped from every decision for new sessions, although its interface stays up.
  3. Member out of SLA. Only rules that require that SLA stop using it. The rule moves to the next member in priority-members that meets the SLA.

Two details to plan for. First, what a rule does when no member meets the SLA depends on the strategy and the firmware version, so build a last-resort member into the design and test it. Second, sessions that are already established are not always moved the moment a decision changes; behaviour for existing sessions has options and differs between failure types, so read the release notes of your version and test with a long download and a voice call.

Overlay members and routing protocols (a first look)

A member does not have to be an ISP port. IPsec tunnels towards a hub can be members of an overlay zone, so that one SD-WAN rule chooses between, say, an MPLS-based tunnel and an internet tunnel. Dynamic routing (BGP or OSPF over the tunnels, see nse4-dynamic) then supplies the routes to the other sites while SD-WAN decides which tunnel carries each flow. On a hub, SD-WAN can also use the BGP neighbours' state; that is an advanced topic beyond this exam.

Common mistake. Expecting the implicit rule to avoid a lossy link. It does not measure quality; it only skips dead members. Write a rule with an SLA for anything quality-sensitive.

Exam trap. Implicit rule = no rule matched; it uses the routing table and the load-balance mode (default source-ip-based). It only removes dead members, not members that fail an SLA.

Sixty percent on the cheap link

A retailer wanted roughly 60 percent of unclassified traffic on the cheap 200 Mbit link and 40 percent on the premium 100 Mbit link. Using weight-based mode with weights 60 and 40 achieved that, while voice and card payments had explicit rules with SLAs that kept them on the premium link whenever it was good.

Lesson: Rules for what matters, load balancing for the rest.

"What is the implicit SD-WAN rule?"

The built-in last rule that handles traffic matching no explicit rule. It uses the routing table and the global load-balance mode (default source-IP based) over the members that are alive. It does not evaluate SLAs, so quality-sensitive traffic needs explicit rules.

Key takeaways

  • The implicit rule handles unmatched traffic with the routing table and a load-balance mode.
  • Modes: source-ip-based (default), weight-based, usage-based, source-dest-ip-based, measured-volume-based.
  • Failover happens on interface down, member dead or out of SLA.
  • Design and test a last-resort member and the behaviour of existing sessions.
08

Troubleshooting SD-WAN: a step-by-step workflow

SD-WAN adds three layers on top of ordinary routing: members, measurements and rules. A good workflow checks them in order and then falls back to the routing basics you practised in nse4-routing. The FortiOS commands for the SD-WAN layers all run in the simulator: practise them in the ticket lab Half the users cannot browse. The routing and flow outputs at the end were captured from the real nse4-routing labs.

LayerCommand (real FortiGate)Question
Membersdiagnose sys sdwan memberAre the right interfaces and gateways members of the right zone?
Zonesdiagnose sys sdwan zoneWhich members does each zone contain?
Healthdiagnose sys sdwan health-checkIs each member alive, and does it meet the SLA?
Rulesdiagnose sys sdwan serviceWhich members does each rule currently choose?
Routingget router info routing-table allIs the zone default route present?
1 Membersconfigured?2 Healthalive / SLA3 Rulesorder, match4 Routezone default5 Policyto zone6 Flowdebug flowMeasure first, then steer, then route, then permit.

Six checks in a fixed order.

Common causes and what you see

SymptomLikely cause
Member shows dead although the ISP is fineProbe target unreachable through that member, probe blocked, wrong protocol, or no route/NAT for the probe
Rule never chooses the expected memberAn earlier rule matches first, or the SLA is too strict so no member qualifies
All traffic on one linkOnly the first match or default source-ip hashing; check weights and rules
No internet after enabling SD-WANMissing zone default route or no policy to the zone
Interface refused as memberIt is still referenced by a route, policy or DHCP server
SaaS application not steeredISDB name wrong, or classification only happens after the first packets

The routing and flow checks (real output)

The two lab firewalls of nse4-routing have no SD-WAN, but the generic checks are identical. A dual-link default route with equal distance shows both next hops, which is what a zone with two members looks like in the table:

NK-PUNE-FGT # get router info routing-table details 0.0.0.0
Routing table for VRF=0
Routing entry for 0.0.0.0/0
  Known via "static", distance 10, metric 0, candidate default, best
  * vrf 0 203.0.113.1, via port1
  * vrf 0 198.51.100.1, via port4

And a missing policy shows up in the flow trace as a denied packet even when the route and the member are right. This is the same message you would see for a missing policy to the SD-WAN zone:

id=65308 trace_id=1 msg="find a route: flag=04000000 gw-198.51.100.1 via port4"
id=65308 trace_id=1 msg="in-[port2], out-[port4], skb_flags-02000000, vid-0, app_id: 0, url_cat_id: 0"
id=65308 trace_id=1 msg="Denied by forward policy check (policy 0)"

For an SD-WAN session the flow trace also shows a Match policy routing line naming the rule, and diagnose firewall proute list and diagnose sys session list (sdwan_mbr_seq, sdwan_service_id) show which rule and member a session used. The simulator prints these too.

Common mistake. Changing rules before looking at the health check. A "rule problem" is often a measurement problem: the SLA is failing or the probe is dead. Look at the measurements first.

Exam trap. Expect questions that give a symptom and ask which diagnose command to use: member for membership, health-check for alive/SLA, service for the rule decision.

Two links alive, one rule dead

After a change window nobody could stream video over the secondary link as planned. The health check showed both members alive and in SLA, the service view showed the rule choosing only member 1. A new catch-all rule inserted above it matched all traffic first. Reordering the rules fixed it in one minute.

Lesson: Read measurements, then rule choice, then rule order.

"Users cannot reach the internet after SD-WAN was enabled. How do you troubleshoot?"

I check that the members and zone exist, that the health checks show alive members, that a default route and a firewall policy point at the zone, then run a flow trace for one client to see whether the packet is routed and permitted. I look at the service view to see which member the rule picks, and fix one thing at a time.

Key takeaways

  • Check members, health, rules, route, policy, flow in that order.
  • diagnose sys sdwan member, health-check and service are the main commands on a real FortiGate.
  • Most "rule" problems are measurement or rule-order problems.
  • The routing and flow basics from nse4-routing apply unchanged.
09

Summary and exam checklist

This chapter gathers the SD-WAN module into a checklist, a glossary and the facts that appear most often in exam questions. The SD-WAN core in this module (zones, members, health checks, SLA, rules, failover and diagnose sys sdwan) runs in the three simulator labs; overlay IPsec, ADVPN, FortiManager and shaping are reference-only.

Can-do checklist

  • Explain why SD-WAN beats floating routes: quality-based failover, per-application steering, all links in use.
  • Create members and zones and list what must be removed from an interface first.
  • Point one default route and one firewall policy at the zone.
  • Configure a health check with server, protocol, interval, failtime and recoverytime.
  • Define an SLA with latency, jitter and packet-loss thresholds.
  • Write a rule with the right strategy and priority-members, and order rules correctly.
  • Explain the implicit rule and the load-balance modes.
  • Troubleshoot with diagnose sys sdwan commands and the routing basics.

Mini glossary

Member
A WAN link (interface plus gateway) in SD-WAN.
Zone
A named group of members; routes and policies point to it.
Health check
A probe that measures loss, latency and jitter per member.
SLA
Thresholds that define a good link.
Rule (service)
Match plus strategy, evaluated top-down.
Implicit rule
Handles unmatched traffic with routing and load balancing.
ISDB
Internet Service Database of SaaS and cloud addresses.

Most-tested facts

  • Default zone: virtual-wan-link. Policies and the default route use the zone.
  • Health check measures latency, jitter and packet loss; dead after failtime lost probes.
  • A member can be alive and still fail the SLA.
  • Strategies: manual, best quality, lowest cost (SLA), maximize bandwidth (SLA).
  • Rules are matched top-down; unmatched traffic uses the implicit rule (default load balancing is source-ip-based).

Command cheat-sheet (reference)

! Try it: lab "Half the users cannot browse" and the other two labs of this module
config system sdwan
diagnose sys sdwan member
diagnose sys sdwan zone
diagnose sys sdwan health-check
diagnose sys sdwan service
get router info routing-table all
diagnose debug flow filter addr 10.0.1.10
diagnose debug flow trace start 5
diagnose debug enable

Where to practise

Repeat the three labs of this module until you can finish them without hints, then strengthen the pieces SD-WAN builds on: repeat both labs of nse4-routing (floating route and reverse-path check), practise tracking and policy routing in enarsi-pbr-vrf-bfd, and learn the dynamic routing used between SD-WAN sites in nse4-dynamic, enarsi-ospf, enarsi-bgp and jncis-ospf.

Exam trap. If an answer mentions configuring members by IP address only, or adding member interfaces to the policy, it is probably wrong: policies use the zone.

A whiteboard answer that got the job

In an interview a candidate drew the zone, two members, a health check and a voice rule, then explained what happens when the primary link is alive but lossy: it fails the SLA, the voice rule moves to the second member, while web traffic stays on the implicit rule. The panel noted that he separated alive, out-of-SLA and dead.

Lesson: Clear states and clear decisions win interviews.

"Explain FortiGate SD-WAN in two minutes."

Links become members of a zone; one route and one policy point at the zone. Health checks measure loss, latency and jitter per member, SLA thresholds say what is good, and rules steer traffic by strategy, evaluated top-down before the routing table. Unmatched traffic uses the implicit rule with a load-balance mode. Dead members are skipped and out-of-SLA members are skipped by rules that need the SLA.

Key takeaways

  • Members, zone, health check, SLA, rules, implicit rule: know all six.
  • Alive, out of SLA and dead are different states.
  • Policies and routes use the zone.
  • Practise the three SD-WAN labs here and the foundations in nse4-routing.
๐ŸŽ“ For educational purposes only โ€” all devices are simulationsTerms of UsePrivacy Policyยฉ 2026 Network Kings
CONFIG by Network Kings โ€” an educational IT simulation platform for learning purposes only. It is not Cisco IOS, Junos, FortiOS or PAN-OS and contains no Cisco, Juniper, Fortinet or Palo Alto Networks software. Cisco, IOS, CCNA, CCNP, Juniper, JNCIA, JNCIS, JNCIP, Fortinet, FortiGate, FortiOS, NSE, Palo Alto Networks, PAN-OS and PCNSE are trademarks of their respective owners. Network Kings is not affiliated with or endorsed by Cisco Systems, Inc., Juniper Networks, Inc., Fortinet, Inc. or Palo Alto Networks, Inc.