CCNA Automation ยท Start here: automation from zero

Turning CLI output into data: split, regex & parsers

A script that logs in over SSH gets back plain text made for humans. Learn three ways to turn it into data you can test and store: cutting lines with split, matching patterns with regular expressions, and ready-made parsers that return dictionaries.

49 min read9 chapters2 labs15 quiz7 scenarios15 interview Q&A

This first module is free: read the lesson and take the quiz. Create a free account to run up to 3 hands-on labs.

Log inStart free
Jump to chapter (9)
01

Why parse? From CLI text to data you can use

What you will learn in this module. Network devices answer show commands with text for humans. A script needs data: lists, dictionaries, numbers. Turning the first into the second is called parsing. You will learn three ways to parse, from the simplest to the most robust: cutting text with split(), finding patterns with regular expressions (the re module), and using a parser that returns ready-made dictionaries. In the labs you build an "interfaces that are up" list and a CSV inventory of routers. This is part of the CCNP Automation (350-901 AUTOCOR) track.

Prerequisites. Python 1 to 3 of this track: strings and f-strings, lists and dictionaries, loops and if, functions, files and try/except. The previous module showed how data is stored in JSON and YAML; this module is about getting data out of screen text.

Analogy: reading a parcel label

A courier label says "Mr Rao, Flat 12, Tower B, Pune 411001, Ph 98xxxxxx". A person reads it at a glance. A sorting machine needs fields: name, flat, city, pincode. To get fields out of the label, you can (1) cut it at the commas, (2) search it for a pattern such as "six digits is a pincode", or (3) use a barcode that already carries the fields. These are exactly the three methods of this module: split, regex and a parser.

What the raw output looks like

This is show ip interface brief from router MUM-R1 of the first lab, captured from the simulator.

MUM-R1# show ip interface brief
Interface              IP-Address      OK? Method Status                Protocol
GigabitEthernet0/0     10.99.0.1       YES manual up                    up
GigabitEthernet0/1     10.12.0.1       YES manual up                    up
GigabitEthernet0/2     10.50.0.1       YES manual down                  down
GigabitEthernet0/3     unassigned      YES unset  administratively down down
Loopback0              1.1.1.1         YES manual up                    up

A human sees a table. A program sees one long string with new-line characters. Column 1 is the interface name, column 2 the IP, and the last two columns the Status and Protocol. The question for the NOC: which interfaces are really up (status up and protocol up)?

The screen-scraping idea

Treating output as text and picking pieces out of it is called screen scraping. It is how automation started, and it is still needed because many devices and commands have no API. Its weakness is that it depends on the exact shape of the text:

  • a different software version may add a column,
  • an interface that is administratively down prints administratively down (two words) in one column,
  • an empty line, a banner or a warning can appear in the middle.
split()cut text at spacesRegexfind patternsParserready dictionariesDatalists, dicts, numbers

Three ways to parse

So the rule of this module is: start with the simplest method that is safe, and move up when the output gets harder.

The pipeline of every collection script

  1. Connect (SSH) and run the command: you get a string.
  2. Parse the string into data (list or dictionary).
  3. Use the data: filter, count, compare, save as JSON or CSV, build a report.

The labs use nkmiko for step 1 and focus on steps 2 and 3. Here is the whole idea in a few lines of plain Python, using a short piece of the output:

text = """GigabitEthernet0/0     10.99.0.1       YES manual up                    up
GigabitEthernet0/2     10.50.0.1       YES manual down                  down"""

for line in text.splitlines():
    parts = line.split()
    print(parts[0], "is", parts[-1])

Real output, checked in the lab terminal.

GigabitEthernet0/0 is up
GigabitEthernet0/2 is down

splitlines() cuts the text into lines, split() cuts a line into words, and parts[0] and parts[-1] pick the first and last word. You will practice every piece in the coming chapters.

What the labs ask you to do

Lab 1 (up.py, up2.py) collects show ip interface brief from MUM-R1 and MUM-R2 and lists only the interfaces that are up/up, first with split(), then with a parser. Lab 2 (inv.py) reads show version from three Bengaluru routers with regular expressions, fixes one broken pattern, and writes inventory.csv with hostname, version, uptime and serial.

Common beginner mistake. Believing parsing is a one-time job. Output changes with software versions, so a good script checks what it found (for example, that a pattern matched) and reports clearly when it did not.

Exam trap. The exam contrasts unstructured CLI text with structured data from APIs, and asks which approach is more reliable. Structured data (RESTCONF, NETCONF) is more reliable than screen scraping; parsers such as TextFSM and Genie sit in between.

The report that broke after an upgrade

A team wrote a script that read the fifth word of every line of show ip interface brief. It worked for two years. After a software upgrade, one device printed an extra column and the script began reporting interfaces as "up" when they were down. Nobody noticed for a week because the script never failed; it just returned wrong answers. The team rewrote it with a parser keyed by field names and added a check that the number of interfaces matched the device's own count.

Lesson: a parser that silently returns wrong data is worse than one that crashes. Verify what you parse.

"Why is parsing CLI output fragile, and how do you reduce the risk?"

CLI output is meant for humans, so its columns, spacing and wording can change between versions and platforms. I reduce the risk by preferring structured sources (RESTCONF or NETCONF) when available, otherwise using a tested parser keyed by field names, anchoring regular expressions to labels instead of positions, and always checking that the parse found what I expected.

Key takeaways

  • Parsing turns CLI text for humans into lists and dictionaries for programs.
  • Three methods: split(), regular expressions, and ready-made parsers.
  • Screen scraping depends on the exact shape of the text and can break after upgrades.
  • Pipeline: run the command, parse the string, use the data, verify.
  • Always check that the parse succeeded; silent wrong data is worse than a crash.
02

Cutting text: splitlines, split, strip and indexes

The simplest parsing tools are string methods you may already know. This chapter uses them on the real output of show ip interface brief from lab 1. One small idea per section.

Step 1: the output is one string

When a script receives a command's output, it is a single string with \n (new-line) characters inside. To make it easy to practice, we keep a copy of the output of router MUM-R1 in a variable. The triple quotes """ let a string span many lines.

Step 2: splitlines cuts text into lines

text.splitlines() returns a list with one item per line, without the new-line characters.

output = """Interface              IP-Address      OK? Method Status                Protocol
GigabitEthernet0/0     10.99.0.1       YES manual up                    up
GigabitEthernet0/1     10.12.0.1       YES manual up                    up
GigabitEthernet0/2     10.50.0.1       YES manual down                  down
GigabitEthernet0/3     unassigned      YES unset  administratively down down
Loopback0              1.1.1.1         YES manual up                    up"""

lines = output.splitlines()
print(len(lines))
print(lines[0])
print(lines[1])

Real output, checked in the lab terminal.

6
Interface              IP-Address      OK? Method Status                Protocol
GigabitEthernet0/0     10.99.0.1       YES manual up                    up

Six lines: one heading and five interfaces. lines[0] is the heading, because counting starts at 0. We do not want the heading in our data, so we skip it with a slice: lines[1:] means "from item 1 to the end".

Step 3: split cuts a line into words

line.split() with no argument cuts at any run of spaces and returns a list of words. The spaces, however many, vanish. That is the magic: columns lined up with many spaces become tidy list items.

line = "GigabitEthernet0/1     10.12.0.1       YES manual up                    up"
parts = line.split()
print(parts)
print(len(parts))
print(parts[0])
print(parts[1])
print(parts[-1])

Real output, checked in the lab terminal.

['GigabitEthernet0/1', '10.12.0.1', 'YES', 'manual', 'up', 'up']
6
GigabitEthernet0/1
10.12.0.1
up

The line became six words. Each can be reached by its position: parts[0] is the interface name, parts[1] the IP address, and parts[-1] the last word, the protocol. A negative index counts from the end, so -1 is the last item and -2 the one before it. Using the last word is handy because it stays at the end even if columns are added in the middle.

[[flow From string to fields | output/one long string | splitlines()/list of lines | split()/list of words | parts[0], parts[-1]/pick fields]]

Step 4: loop and decide

Combine the pieces with the loop and if you know. This is the exact TODO of lab 1: print the interface name only when the last word is up.

output = """Interface              IP-Address      OK? Method Status                Protocol
GigabitEthernet0/0     10.99.0.1       YES manual up                    up
GigabitEthernet0/1     10.12.0.1       YES manual up                    up
GigabitEthernet0/2     10.50.0.1       YES manual down                  down
GigabitEthernet0/3     unassigned      YES unset  administratively down down
Loopback0              1.1.1.1         YES manual up                    up"""

for line in output.splitlines()[1:]:
    parts = line.split()
    if parts[-1] == "up":
        print(parts[0])

Real output, checked in the lab terminal.

GigabitEthernet0/0
GigabitEthernet0/1
Loopback0

Three interfaces are up/up. GigabitEthernet0/2 is down/down and GigabitEthernet0/3 is administratively down, so neither is printed.

What the lab script prints before you finish it

In the lab, up.py logs in to both routers and prints each row as a list of words, before the TODO is solved. This is the real output from the lab terminal. Notice the two lists that are longer than the others.

from nkmiko import ConnectHandler

ROUTERS = {"MUM-R1": "10.99.0.1", "MUM-R2": "10.99.0.2"}

for name, ip in ROUTERS.items():
    conn = ConnectHandler(device_type="nk_ios", host=ip, username="netops", password="NK@2026")
    output = conn.send_command("show ip interface brief")
    conn.disconnect()
    print(f"=== {name} ===")
    for line in output.splitlines()[1:]:
        parts = line.split()
        print(parts)

Run in the lab terminal (it uses the lab routers). Output captured from the simulator.

=== MUM-R1 ===
['GigabitEthernet0/0', '10.99.0.1', 'YES', 'manual', 'up', 'up']
['GigabitEthernet0/1', '10.12.0.1', 'YES', 'manual', 'up', 'up']
['GigabitEthernet0/2', '10.50.0.1', 'YES', 'manual', 'down', 'down']
['GigabitEthernet0/3', 'unassigned', 'YES', 'unset', 'administratively', 'down', 'down']
['Loopback0', '1.1.1.1', 'YES', 'manual', 'up', 'up']
=== MUM-R2 ===
['GigabitEthernet0/0', '10.99.0.2', 'YES', 'manual', 'up', 'up']
['GigabitEthernet0/1', '10.12.0.2', 'YES', 'manual', 'up', 'up']
['GigabitEthernet0/2', '10.60.0.1', 'YES', 'manual', 'administratively', 'down', 'down']
['GigabitEthernet0/3', 'unassigned', 'YES', 'unset', 'administratively', 'down', 'down']
['Loopback0', '2.2.2.2', 'YES', 'manual', 'up', 'up']
['Loopback1', '22.22.22.22', 'YES', 'manual', 'up', 'up']

The rows for the administratively down interfaces have seven words, not six, because administratively down is two words. The next chapter explains why that matters.

Useful companions

MethodWhat it doesExample result
s.strip()removes spaces and new-lines at both ends" up " becomes "up"
s.lower()lowercase"UP" becomes "up"
s.startswith("Gig")True if the text starts with thatfor choosing rows
"up" in sTrue if the text is insidea quick filter
s.replace("a", "b")replaces textclean-up
", ".join(list)glues a list into textfor reports
line = "  GigabitEthernet0/1  "
print(repr(line.strip()))
print(line.strip().startswith("Gig"))
print("Loopback" in "Loopback0 1.1.1.1")
print(", ".join(["R1", "R2", "R3"]))

Real output, checked in the lab terminal.

'GigabitEthernet0/1'
True
True
R1, R2, R3

Common mistakes. (1) Forgetting [1:] and treating the heading line as an interface, which gives wrong rows or an index error. (2) Using split(" ") with one space: it creates empty items when there are many spaces. Use split() with no argument. (3) Using a fixed index such as parts[4] for the status. (4) Calling split on an empty line: "".split() gives an empty list, and parts[0] then raises IndexError.

Exam trap. "a b".split() gives ['a', 'b'] but "a b".split(" ") gives ['a', '', 'b']. Negative index -1 is the last item. splitlines() returns a list.

The empty line that crashed the report

A script looped over every line of a command output and read parts[0]. It ran fine on lab devices, then crashed in production with IndexError. A trailing blank line in the output produced an empty list, and index 0 of an empty list does not exist. Adding if not parts: continue (skip blank lines) fixed it.

Lesson: real output contains blank lines; always guard against empty rows.

"How do you split a multi-line command output into fields in Python?"

I call splitlines() to get the lines, skip the header with a slice such as [1:], and call split() with no argument on each line to get the words. I use negative indexes for fields near the end and I skip empty lines. For anything less regular than a simple table I switch to regular expressions.

Key takeaways

  • splitlines() makes a list of lines; [1:] skips the heading.
  • split() with no argument cuts at any run of spaces into words.
  • parts[0] is the first word, parts[-1] the last; counting starts at 0.
  • Combine with for and if to filter rows, for example last word equal to up.
  • Guard against blank lines and always use split() without an argument.
03

Where split() breaks: ragged columns and a safer way

split() is wonderful for tidy tables. This chapter shows exactly where it breaks, so you can recognise the danger early and choose a safer method.

The trap: a column with two words

Look again at the rows of show ip interface brief. Most rows have six words, but the row of an interface that is administratively down has seven, because the status is the two-word phrase administratively down. If your script counts positions from the left, the columns shift.

rows = [
    "GigabitEthernet0/1     10.12.0.1       YES manual up                    up",
    "GigabitEthernet0/3     unassigned      YES unset  administratively down down",
]
for row in rows:
    parts = row.split()
    print(len(parts), "words | status =", parts[4], "| protocol =", parts[5])

Real output, checked in the lab terminal.

6 words | status = up | protocol = up
7 words | status = administratively | protocol = down

For the first row everything looks right. For the second, parts[4] is administratively and parts[5] is down, so a script that reads the fixed positions 4 and 5 reports a nonsense status, and it does so without any error. Silent wrong data is the dangerous kind.

Row A6 words: status at index 4Row B7 words: status at index 4 and 5Resultcolumns shift, no error

Why position counting breaks

Fix 1: count from the end for the last column

The protocol is always the last word, so parts[-1] stays correct for both rows. That is exactly why the lab filter parts[-1] == "up" works. Counting from the end is safe for the last column, but not for the middle ones.

Fix 2: split in a limited number of places

line.split(None, 4) cuts at the first 4 gaps only and leaves the remainder as one piece. Then rsplit(None, 1) cuts one word from the right. Together they treat the two-word status as a single field.

rows = [
    "GigabitEthernet0/1     10.12.0.1       YES manual up                    up",
    "GigabitEthernet0/3     unassigned      YES unset  administratively down down",
]
for row in rows:
    name, ip, ok, method, rest = row.split(None, 4)
    status, protocol = rest.rsplit(None, 1)
    print(name, "|", ip, "|", status, "|", protocol)

Real output, checked in the lab terminal.

GigabitEthernet0/1 | 10.12.0.1 | up | up
GigabitEthernet0/3 | unassigned | administratively down | down

Now both rows give the right status, including administratively down. This works, but you can see the price: the code is harder to read, and it still depends on the exact number of columns.

Fix 3: cut at a label instead of a position

For lines such as BLR-CORE uptime is 0 minutes, the better anchor is the words that never change. split("separator") accepts a multi-character separator.

line = "BLR-CORE uptime is 0 minutes"
host, uptime = line.split(" uptime is ")
print(host)
print(uptime)

board = "Processor board ID NK3C15U4X00"
print(board.split()[-1])

image = 'System image file is "flash:nkos-router-universal.2.0.bin"'
print(image.split('"')[1])

Real output, checked in the lab terminal.

BLR-CORE
0 minutes
NK3C15U4X00
flash:nkos-router-universal.2.0.bin

Three useful moves are shown: cut at a phrase, take the last word, and take the text between quotes (split('"')[1] is the second piece, the one inside the quotes). Anchoring on a label survives column changes better than counting positions.

Choosing the right line first

Most of the output is not data you want. Select the line before splitting it, using in, startswith or a loop with if.

show_version = """Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE
ROM: NK-OS Bootstrap, Version 2.0, RELEASE SOFTWARE

BLR-CORE uptime is 0 minutes
System returned to ROM by power-on
Processor board ID NK3C15U4X00
4 Gigabit Ethernet interfaces"""

for line in show_version.splitlines():
    if "uptime is" in line:
        print("uptime line:", line)
    if line.startswith("Processor board ID"):
        print("serial:", line.split()[-1])

Real output, checked in the lab terminal.

uptime line: BLR-CORE uptime is 0 minutes
serial: NK3C15U4X00

When to stop using split

Use split() while all these are true: the output is a regular table; every row has the same number of fields; you only need the first or last column; and the format will not change soon. Switch to a stronger tool when any of these fails:

SituationBetter tool
Value is inside a sentence (Version 2.0,)Regular expression
Fields contain spacesRegular expression or a parser
Many commands, many platformsA ready-made parser
You must be sure of the resultA parser plus checks

The next two chapters teach regular expressions, the first stronger tool.

Common mistakes. (1) Reading parts[4] and assuming that position never moves. (2) Not checking len(parts) before indexing, which gives IndexError on short rows. (3) Splitting a line that you have not first selected, so a banner line accidentally produces data. (4) Forgetting that trailing punctuation sticks to words, such as 2.0, with a comma.

Exam trap. Screen scraping with positions is fragile because output formats change; the exam's recommended order is structured API first, then a template-based parser, then regex or split as a last resort.

The outage that was only a column

A dashboard counted a port healthy when parts[4] == "up". After a software upgrade added one new column before the status, parts[4] held a different word on every row, so the dashboard showed zero healthy ports and the NOC was paged for an outage that did not exist. The network was fine; the parser was reading the wrong column. Switching the check to the last word, and adding a test that compares the count of up ports with the device's own summary, ended the false alarms.

Lesson: test your parser on the ugly rows and after every upgrade, not just on the clean lab output.

"A split-based parser works on your lab but misreads a few production rows. Why?"

Production output has irregular rows: status phrases with two words, wrapped lines, blank lines, or extra columns after an upgrade. Position-based splitting shifts the columns. I would anchor on labels or the last column, use limited splits where the shape is known, add length checks, and move to regular expressions or a tested parser when the table is not regular.

Key takeaways

  • A row with a two-word value has more words and shifts the columns.
  • The last word is safe to read with parts[-1]; middle positions are not.
  • split(None, n) and rsplit(None, 1) limit where the cuts happen.
  • Cut at a label such as " uptime is " and select the right line first.
  • When values sit inside sentences or fields hold spaces, use regular expressions or a parser.
04

Regular expressions 1: search for a pattern with re

A regular expression (regex) is a small language for describing text patterns, such as "the word Version, a space, then some non-space characters". Instead of counting words or positions, you describe what the value looks like, and the re module finds it. This is the tool for values hidden inside sentences, and it is what lab 2 uses.

The idea: describe the shape

Think of a shape-sorting toy. Each hole accepts only one shape. A regex is a set of holes; text passes through only where its shape fits. A pattern such as \d+ means "one or more digits", so it fits 2026 and 10 but not abc.

Your first search

re.search(pattern, text) looks anywhere in the text for the first place where the pattern fits. It returns a match object if found, or None if not.

import re

line = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE"
m = re.search(r"Version", line)
print(m)
print(m.group())
print(re.search(r"Cisco", line))

Real output, checked in the lab terminal.

<re.Match object; span=(51, 58), match='Version'>
Version
None

The first search found the word and printed a match object, which says where the match sits (span) and what it matched. m.group() returns the matched text. The last search found nothing and returned None. Always remember: no match means None, and calling a method on None is an error you will meet in the lab.

The raw string prefix

Notice r"Version": a string with an r in front is a raw string. Python leaves backslashes alone inside it. Regex uses backslashes a lot (\d, \s), so write patterns as raw strings every time.

The building blocks

PieceMeaningExample
abcthe literal textVersion finds Version
\done digit\d\d finds 20
\sone whitespace (space, tab)
\Sone non-whitespace
\wletter, digit or underscore
.any one charactera.c finds abc
\.a real dot10\.99
+one or more of the previous\d+ finds 2026
*zero or more
?zero or one
^ and $start and end of the line
Literal textVersionSpace\sGroup of characters\S+Comma,

A pattern is a row of holes

Combine them: Version \S+, means "the word Version, a space, one or more non-space characters, and a comma". Try it on the line above.

import re

line = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE"
m = re.search(r"Version \S+,", line)
print(m.group())
print(re.search(r"\d+\.\d+", line).group())
print(re.findall(r"[A-Z]{2,}", line))

Real output, checked in the lab terminal.

Version 2.0,
2.0
['OS', 'NK', 'OS', 'RELEASE', 'SOFTWARE']

The first pattern matched Version 2.0,, including the comma, which is usually not what we want. The second, \d+\.\d+, matched the number 2.0 (digits, a real dot, digits). The third uses [A-Z]{2,}: a character class [A-Z] (any capital letter) repeated 2 or more times; it found the runs of capital letters in the line. In the next chapter, groups let you keep only the part you need.

search, match and findall

  • re.search(p, text): the pattern may be anywhere.
  • re.match(p, text): the pattern must fit at the start of the text.
  • re.findall(p, text): returns a list of all non-overlapping matches.
import re

text = "R1 10.99.0.1 up\nR2 10.99.0.2 down"
print(bool(re.match(r"R1", text)))
print(bool(re.match(r"10", text)))
print(bool(re.search(r"10", text)))
print(re.findall(r"\d+\.\d+\.\d+\.\d+", text))

Real output, checked in the lab terminal.

True
False
True
['10.99.0.1', '10.99.0.2']

match fails for 10 because the text starts with R1; search finds it anywhere. findall collected both IP addresses. Remember the escaped dots \. in the IP pattern: an unescaped dot would match any character.

Case matters

A regex is case-sensitive by default. Version does not match version. You can switch case-insensitive on with re.IGNORECASE (or re.I). This small fact is behind a lab bug.

import re

line = "System Version 2.0"
print(re.search(r"version", line))
print(re.search(r"version", line, re.I).group())

Real output, checked in the lab terminal.

None
Version

Test a pattern before you trust it

Write the pattern, run it on a sample of real output, and print the result. If the result is None, adjust. Most regex work is this loop of small experiments.

Common mistakes. (1) Forgetting the r before the quotes. (2) An unescaped dot, so 10.99 also matches 10x99. (3) Using match when you meant search. (4) Using the result without checking for None. (5) Writing one giant pattern at once instead of building it piece by piece.

Exam trap. re.match anchors at the start; re.search does not. \d is a digit, \s whitespace, \S non-whitespace, . any character, \. a real dot. A failed search returns None.

The dot that matched too much

A script validated management IPs with the pattern 10.99.0.1. A typo address 10x99y0z1 passed the check, because each dot matched any character. The fix was escaping the dots (10\.99\.0\.1) and anchoring the pattern with ^ and $.

Lesson: escape dots, and anchor patterns when you need an exact match.

"What is the difference between re.match, re.search and re.findall?"

match succeeds only if the pattern fits at the start of the string, search finds the first fit anywhere, and findall returns a list of all matches. match and search return a match object or None, so I check the result before using it. I write patterns as raw strings and escape literal dots.

Key takeaways

  • A regex describes the shape of text; re.search returns a match object or None.
  • Write patterns as raw strings: r"...".
  • Key pieces: \d, \s, \S, ., \., +, *, ?, ^, $, [A-Z], {2,}.
  • match anchors at the start; search looks anywhere; findall returns every match.
  • Regex is case-sensitive unless you pass re.I.
05

Regular expressions 2: groups, named groups and show version

The previous chapter found text. This one extracts it. The tool is the group: round brackets ( ) around the part of the pattern whose text you want to keep. With groups you can pull the version, hostname, uptime and serial number out of show version, which is what lab 2 does.

Why groups?

re.search(r"Version \S+,", line) matched Version 2.0,, but we only want 2.0. Put the interesting part in brackets and ask for it with .group(1):

import re

line = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE"
m = re.search(r"Version (\S+),", line)
print(m.group())
print(m.group(1))

Real output, checked in the lab terminal.

Version 2.0,
2.0

m.group() (or group(0)) is the whole match, including the label. m.group(1) is only what the first pair of brackets captured: 2.0. The label Version and the comma act as anchors: they tell the engine where the value begins and ends, and they are not kept. This is anchoring on text that does not change, the idea from the split chapter.

LabelVersionGroup(\S+) the valueEnd anchor,Resultm.group(1)

Anchors and a group

Parsing show version with four patterns

This is the output of show version of router BLR-CORE in lab 2 (shortened to the useful lines), and four patterns, one per fact. Each pattern anchors on a stable label.

import re

out = """Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE
Technical Support: ns.example/support

BLR-CORE uptime is 0 minutes
System returned to ROM by power-on
Processor board ID NK3C15U4X00
4 Gigabit Ethernet interfaces"""

host = re.search(r"(\S+) uptime is", out).group(1)
version = re.search(r"Version (\S+),", out).group(1)
uptime = re.search(r"uptime is (.+)", out).group(1)
serial = re.search(r"Processor board ID (\S+)", out).group(1)
print(host, version, uptime, serial, sep=" | ")

Real output, checked in the lab terminal.

BLR-CORE | 2.0 | 0 minutes | NK3C15U4X00

Read each pattern aloud. (\S+) uptime is: a group of non-space characters, then a space and the words uptime is; the group is the hostname just before them. uptime is (.+): after those words, .+ means "one or more of any character", which takes the rest of the line. Processor board ID (\S+): the word after the label is the serial.

The case-sensitivity bug of lab 2

In the lab, the version pattern is written with a lowercase v: r"version (\S+),". The output spells it Version. Regex is case-sensitive, so search returns None, and None.group(1) crashes.

import re

out = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE"
version = re.search(r"version (\S+),", out).group(1)
print(version)

Real output, checked in the lab terminal.

Traceback (most recent call last):
  File "casebug.py", line 4, in <module>
    version = re.search(r"version (\S+),", out).group(1)
AttributeError: 'NoneType' object has no attribute 'group'

The last line, AttributeError: 'NoneType' object has no attribute 'group', is the classic regex failure. It does not mean group is broken. It means the pattern found nothing, so there is no match object. The fix in the lab is one letter, Version, or the flag re.I. A defensive script checks first:

import re

out = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE"
m = re.search(r"version (\S+),", out)
if m:
    print("version:", m.group(1))
else:
    print("version line not found - check the pattern")
print(re.search(r"version (\S+),", out, re.I).group(1))

Real output, checked in the lab terminal.

version line not found - check the pattern
2.0

Several groups at once and named groups

A pattern may have many groups: group(1), group(2) and so on, counted by the position of the opening bracket. For readability you can name a group with (?P<name>...) and read it by name.

import re

line = "GigabitEthernet0/3     unassigned      YES unset  administratively down down"
pat = r"^(?P<name>\S+)\s+(?P<ip>\S+)\s+YES\s+\S+\s+(?P<status>administratively down|up|down)\s+(?P<proto>up|down)$"
m = re.match(pat, line)
print(m.group("name"), "|", m.group("status"), "|", m.group("proto"))
print(m.groupdict())

Real output, checked in the lab terminal.

GigabitEthernet0/3 | administratively down | down
{'name': 'GigabitEthernet0/3', 'ip': 'unassigned', 'status': 'administratively down', 'proto': 'down'}

The pattern solves the two-word status problem from the previous chapter. (administratively down|up|down) is an alternation: one of three choices, tried left to right, so the two-word phrase is tried first. ^ and $ pin the pattern to the start and end of the line. groupdict() returns the named groups as a dictionary, which is already half way to structured data.

One pattern over a whole table

re.finditer walks through every match in a multi-line text, and the flag re.M (multi-line) lets ^ and $ match at every line, not only at the ends of the whole text.

import re

output = """Interface              IP-Address      OK? Method Status                Protocol
GigabitEthernet0/0     10.99.0.1       YES manual up                    up
GigabitEthernet0/2     10.50.0.1       YES manual down                  down
GigabitEthernet0/3     unassigned      YES unset  administratively down down
Loopback0              1.1.1.1         YES manual up                    up"""

pat = r"^(\S+)\s+(\S+)\s+YES\s+\S+\s+(administratively down|up|down)\s+(up|down)$"
rows = [m.groups() for m in re.finditer(pat, output, re.M)]
for row in rows:
    print(row)
print("up/up:", [r[0] for r in rows if r[2] == "up" and r[3] == "up"])

Real output, checked in the lab terminal.

('GigabitEthernet0/0', '10.99.0.1', 'up', 'up')
('GigabitEthernet0/2', '10.50.0.1', 'down', 'down')
('GigabitEthernet0/3', 'unassigned', 'administratively down', 'down')
('Loopback0', '1.1.1.1', 'up', 'up')
up/up: ['GigabitEthernet0/0', 'Loopback0']

The heading line was skipped automatically because it does not match the pattern (it has no YES). This is a strength of regex: non-matching lines simply do not count. The groups() method returns all groups as a tuple.

Greedy and lazy

.+ is greedy: it grabs as much as it can. .+? is lazy: as little as it can. For <b>x</b><b>y</b> style text the difference matters; for line-based network output you rarely need lazy, but you must know why uptime is (.+) takes the rest of the line.

Common mistakes. (1) Using .group(1) without checking for None. (2) Forgetting that spelling and case must match, such as Version versus version. (3) Using .* where \S+ is meant, so the match swallows a whole line. (4) Anchoring on position instead of on a stable label. (5) Forgetting re.M, so ^ and $ match only at the very start and end of the whole text.

Exam trap. group(0) is the whole match and group(1) the first bracket. A failed search returns None, so .group on it raises AttributeError. re.I makes a pattern case-insensitive.

One capital letter, three routers

Finance asked for an inventory of three routers. The engineer's script crashed on the first router with 'NoneType' object has no attribute 'group'. He suspected SSH, then the regex engine. A colleague printed the raw output next to the pattern and noticed Version with a capital V against version in the pattern. One letter later, all three routers printed with their versions.

Lesson: when a pattern returns None, compare it character by character with the real output.

"How do you extract the software version from show version using Python?"

I use re.search(r"Version (\S+),", output) and read group(1), anchoring on the stable word Version and the comma. I check the result for None before using it, remember that regex is case-sensitive, and test the pattern on real output from more than one device. For many devices and platforms I would use a tested parser instead.

Key takeaways

  • Round brackets create a group; group(1) returns only the captured text.
  • Anchor patterns on stable labels such as Version, uptime is and Processor board ID.
  • None from search means no match; .group on it raises AttributeError.
  • Named groups (?P<name>...) and groupdict() give dictionaries.
  • re.finditer with re.M runs one pattern over every line of a table.
06

Structured parsers: dictionaries without counting columns

Splitting and regex put the work on you: you design every pattern, test it on every platform, and repair it after every upgrade. A parser moves that work to a maintained library. You give it the command name and the raw text; it returns a dictionary with named fields. This is the third and most robust method of the module, and the second half of lab 1 (up2.py).

The idea: a library of patterns, indexed by command

A parser library contains, for each command, patterns that were written and tested by many people on many software versions. You call one function:

data = parse("show ip interface brief", output)

and get a dictionary in a known shape. Your code then reads fields by name (data["interface"]["Loopback0"]["ip_address"]) instead of by position.

Commandshow ip interface briefRaw textone stringparse()command name plus textDictionarynamed fields

Parser pipeline

The real tools you will meet

ToolStyleTypical use
Genie parsers (pyATS)Python classes, returns nested dictionariesTest frameworks, large Cisco coverage
TextFSM with ntc-templatesTemplate files, returns a list of dictionariesNetmiko use_textfsm=True
TTPYour own simple templatesCustom outputs

In the lab you use the simulator's library nkparse, which behaves like a Genie-style parser: the function is parse(command, output). The real tools above are reference code, not run in the simulator:

Reference code (not run in the simulator).

# Genie / pyATS (Genie-style dictionary)
# parsed = device.parse("show ip interface brief")

# Netmiko with TextFSM templates (list of dictionaries)
# rows = conn.send_command("show ip interface brief", use_textfsm=True)

# ntc-templates directly
# from ntc_templates.parse import parse_output
# rows = parse_output(platform="cisco_ios", command="show ip interface brief", data=raw)

The lab library in action

Here is the lab call on router MUM-R1. The script logs in over SSH, gets the output string, and passes the command name and the text to parse. Compare it with the dozens of lines you would need for split or regex.

from nkmiko import ConnectHandler
from nkparse import parse
import json

conn = ConnectHandler(device_type="nk_ios", host="10.99.0.1", username="netops", password="NK@2026")
output = conn.send_command("show ip interface brief")
conn.disconnect()

data = parse("show ip interface brief", output)
print(list(data.keys()))
print(list(data["interface"].keys()))
print(json.dumps(data["interface"]["GigabitEthernet0/3"], indent=2))

Run in the lab terminal (it uses the lab routers). Output captured from the simulator.

['interface']
['GigabitEthernet0/0', 'GigabitEthernet0/1', 'GigabitEthernet0/2', 'GigabitEthernet0/3', 'Loopback0']
{
  "ip_address": "unassigned",
  "interface_is_ok": "YES",
  "method": "unset",
  "status": "administratively down",
  "protocol": "down"
}

Read the shape. The top of the dictionary has one key, interface. Inside it is one entry per interface, keyed by the interface name. Each entry is a small dictionary with the same field names every time: ip_address, interface_is_ok, method, status and protocol. Notice that status for the administratively-down port is administratively down, one value. The parser solved the two-word problem for you.

datadictionaryinterfaceone keyGigabitEthernet03/one entry per interfacestatus, protocolnamed fields

The shape of the parsed data

Walking the dictionary

Everything from the earlier chapters applies: .items() gives key and value pairs, if filters. This example uses a small dictionary in the same shape, so you can run it anywhere.

data = {"interface": {
    "GigabitEthernet0/0": {"ip_address": "10.99.0.1", "status": "up", "protocol": "up"},
    "GigabitEthernet0/2": {"ip_address": "10.50.0.1", "status": "down", "protocol": "down"},
    "GigabitEthernet0/3": {"ip_address": "unassigned", "status": "administratively down", "protocol": "down"},
    "Loopback0": {"ip_address": "1.1.1.1", "status": "up", "protocol": "up"},
}}

for intf, info in data["interface"].items():
    if info["status"] == "up" and info["protocol"] == "up":
        print(f"{intf:<22}{info['ip_address']}")

shut = [n for n, i in data["interface"].items() if i["status"] == "administratively down"]
print("Shut by an engineer:", shut)
print("Total interfaces:", len(data["interface"]))

Real output, checked in the lab terminal.

GigabitEthernet0/0    10.99.0.1
Loopback0             1.1.1.1
Shut by an engineer: ['GigabitEthernet0/3']
Total interfaces: 4

The first loop is the heart of up2.py. The list comprehension shows how easy the admin-down question becomes with named fields.

Finishing up2.py

The lab file up2.py starts with data = {} and a TODO. With an empty dictionary the loop asks for data["interface"] and crashes with KeyError: 'interface'. Replacing the line with data = parse("show ip interface brief", output) fixes it, and the script prints each up/up interface with its IP. This is the real lab output of the finished script:

from nkmiko import ConnectHandler
from nkparse import parse

ROUTERS = {"MUM-R1": "10.99.0.1", "MUM-R2": "10.99.0.2"}

for name, ip in ROUTERS.items():
    conn = ConnectHandler(device_type="nk_ios", host=ip, username="netops", password="NK@2026")
    output = conn.send_command("show ip interface brief")
    conn.disconnect()
    data = parse("show ip interface brief", output)
    print(f"=== {name} ===")
    for intf, info in data["interface"].items():
        if info["status"] == "up" and info["protocol"] == "up":
            print(f"{intf:<22}{info['ip_address']}")

Run in the lab terminal (it uses the lab routers). Output captured from the simulator.

=== MUM-R1 ===
GigabitEthernet0/0    10.99.0.1
GigabitEthernet0/1    10.12.0.1
Loopback0             1.1.1.1
=== MUM-R2 ===
GigabitEthernet0/0    10.99.0.2
GigabitEthernet0/1    10.12.0.2
Loopback0             2.2.2.2
Loopback1             22.22.22.22

Count the answer to the lab question: on MUM-R2 the up/up interfaces are GigabitEthernet0/0, GigabitEthernet0/1, Loopback0 and Loopback1, which is 4. GigabitEthernet0/2 is administratively down and GigabitEthernet0/3 is unassigned and down.

Other commands, same idea

The same call parses other commands. For show version it returns hostname, version, uptime and serial under the key version:

from nkmiko import ConnectHandler
from nkparse import parse

conn = ConnectHandler(device_type="nk_ios", host="10.99.0.1", username="netops", password="NK@2026")
output = conn.send_command("show version")
conn.disconnect()
print(parse("show version", output))

Run in the lab terminal (it uses the lab routers). Output captured from the simulator.

{'version': {'version': '2.0', 'hostname': 'BLR-CORE', 'uptime': '0 minutes', 'chassis_sn': 'NK3C15U4X00', 'chassis': None}}

A command without a parser stops with a clear error, so you know to fall back to regex:

from nkparse import parse

try:
    parse("show clock", "10:00:00.000 UTC Fri Oct 2 2026")
except ValueError as e:
    print("ValueError:", e)

Run in the lab terminal (it uses the lab routers). Output captured from the simulator.

ValueError: no parser for 'show clock' (available: show ip interface brief, show ip ospf neighbor, show ip bgp summary, show ip route, show version, show vlan brief)

Parser or regex?

Prefer a parser when one exists for your command and platform: it is tested, and your code reads fields by name. Use regex for one-off outputs or custom commands with no template. Use split only for the simplest tables. Whatever you choose, still check that the result contains what you expect.

Common mistakes. (1) Assuming a parser exists for every command; it raises an error or returns nothing otherwise. (2) Forgetting that a parser is tied to a platform and version family. (3) Reading fields that are missing in some outputs; use .get() for optional fields. (4) Not checking an empty result, which looks like "no interfaces" instead of "parse failed".

Exam trap. Genie returns nested dictionaries; TextFSM returns a list of dictionaries (one per row). A structured API such as RESTCONF is still preferred over parsing when it is available.

Replacing 400 lines of regex

A team maintained a 400-line file of regex patterns for ten show commands. Each software upgrade broke two or three. They moved to a parser library for the common commands, kept regex only for two custom outputs, and added a small test that parses a saved sample of every command after each upgrade. Maintenance dropped from days to an hour.

Lesson: let a maintained library carry the parsing, and test it on saved real output.

"What is the advantage of using a parser such as Genie or TextFSM over your own regex?"

A parser library holds tested patterns for many commands, platforms and versions, and returns data in a consistent structure with named fields, so my code does not depend on column positions. It is easier to maintain after upgrades. I still validate that the result is not empty, and I fall back to regex only for commands the library does not cover.

Key takeaways

  • A parser takes the command name and raw text and returns a dictionary or list of dictionaries.
  • Lab parse("show ip interface brief", output) returns data["interface"][name][field].
  • Fields are read by name, so the two-word status is no longer a problem.
  • Genie returns nested dictionaries; TextFSM returns a list of row dictionaries.
  • Check that the parse worked and fall back to regex for commands without a parser.
07

From parsed data to a CSV report

Parsing is only half of a collection script. The other half is delivering the result: a file that a manager, an auditor or a spreadsheet can use. The most common delivery format is CSV (comma-separated values): plain text, one row per line, fields separated by commas. Excel and Google Sheets open it directly. This chapter finishes lab 2, where inv.py writes inventory.csv.

What a CSV file looks like

A CSV is just text. The first line is the header with the column names. Every other line is one record, with the same number of fields.

Raw textshow versionValueshost, version, uptime, serialRowsone string per routerFileinventory.csv

From parsing to a report

Method 1: build the lines yourself

For a small job you can build each line with an f-string and join the lines with a new-line character. This is what the lab does. The example uses three fixed rows so that it runs anywhere.

rows = ["hostname,version,uptime,serial"]
rows.append("BLR-CORE,2.0,0 minutes,NK3C15U4X00")
rows.append("BLR-WAN,2.0,0 minutes,NK45ZYX1X00")

with open("inventory.csv", "w") as f:
    f.write("\n".join(rows) + "\n")

with open("inventory.csv") as f:
    text = f.read()
print(text)
print("Lines in file:", len(text.splitlines()))

Real output, checked in the lab terminal.

hostname,version,uptime,serial
BLR-CORE,2.0,0 minutes,NK3C15U4X00
BLR-WAN,2.0,0 minutes,NK45ZYX1X00

Lines in file: 3

Three details matter. "w" opens the file for writing and replaces an old file of the same name. "\n".join(rows) puts a new line between rows, so you add one more "\n" at the end to finish the last line. The with block closes the file for you, which is what makes sure that the data is really saved.

The comma problem

Method 1 has one weakness: a value that contains a comma splits into two columns. A location such as Andheri, Mumbai would shift every column after it. The standard csv module solves this by putting quotes around such a value.

import csv
import io

buffer = io.StringIO()
writer = csv.writer(buffer, lineterminator="\n")
writer.writerow(["hostname", "location"])
writer.writerow(["MUM-R1", "Andheri, Mumbai"])
writer.writerow(["BLR-CORE", "Bengaluru"])
print(buffer.getvalue())

rows = list(csv.reader(io.StringIO(buffer.getvalue())))
print(rows[1])
print(len(rows[1]), "fields")

Reference code (not run in the simulator). Output below is from standard Python 3.

hostname,location
MUM-R1,"Andheri, Mumbai"
BLR-CORE,Bengaluru

['MUM-R1', 'Andheri, Mumbai']
2 fields

The second row is written with quotes around Andheri, Mumbai, and reading it back gives two fields, not three. Use the csv module whenever data comes from people or from free text. For serial numbers and versions Method 1 is enough. For the same job with dictionaries, csv.DictWriter writes a list of dictionaries and picks the columns from the keys.

Finishing lab 2

In lab 2, inv.py already has patterns for the hostname, the version and the uptime. The two open tasks are the serial number and the file. After fixing the Version spelling, this is the finished script and its real output in the lab:

import re
from nkmiko import ConnectHandler

ROUTERS = ["10.99.0.1", "10.99.0.2", "10.99.0.3"]
rows = ["hostname,version,uptime,serial"]

for ip in ROUTERS:
    conn = ConnectHandler(device_type="nk_ios", host=ip, username="netops", password="NK@2026")
    out = conn.send_command("show version")
    conn.disconnect()
    host = re.search(r"(\S+) uptime is", out).group(1)
    version = re.search(r"Version (\S+),", out).group(1)
    uptime = re.search(r"uptime is (.+)", out).group(1)
    serial = re.search(r"Processor board ID (\S+)", out).group(1)
    rows.append(f"{host},{version},{uptime},{serial}")

with open("inventory.csv", "w") as f:
    f.write("\n".join(rows) + "\n")

print("Wrote", len(rows) - 1, "routers to inventory.csv")
print(open("inventory.csv").read())

Run in the lab terminal (it uses the lab routers). Output captured from the simulator.

Wrote 3 routers to inventory.csv
hostname,version,uptime,serial
BLR-CORE,2.0,0 minutes,NK3C15U4X00
BLR-WAN,2.0,0 minutes,NK45ZYX1X00
BLR-LAB,2.0,0 minutes,NK3W0D82X00

The serial number of BLR-WAN, the lab answer, is NK45ZYX1X00. Check the file with the same discipline as any report: header present, one row per router, no row with TODO left in it.

Other report shapes

A CSV is flat: one row per device. When a record has lists inside it, such as several VLANs per switch, JSON from the previous module is the better file. In both cases the idea is the same: collect, parse, store the result, then report from the stored data. A small script that reads inventory.csv back can answer questions such as "which routers run an older version?" without logging in to a single device again.

Common mistakes. (1) Opening the file with "w" when you meant to add rows; "w" erases the old content, "a" appends. (2) Forgetting the final new line, so the next tool glues two records together. (3) Leaving the placeholder TODO in a column and shipping the report. (4) Putting values with commas into Method 1 lines. (5) Writing the report before checking that every pattern matched.

Exam trap. json.dump writes to a file object; json.dumps returns a string. For CSV, the csv module handles quoting; joining with commas does not. Opening a file with "w" truncates it.

The inventory with a shifted column

A team exported a switch inventory with joined strings. One location field held Floor 3, Annexe. In Excel every row below it shifted one column to the right, and the serial numbers appeared under the uptime heading. Nobody noticed for a week, until an audit compared serials with purchase records. Switching to the csv module added the quotes and the audit passed.

Lesson: a report is code too; test it with the ugly values, not only the neat ones.

"How would you produce an inventory report from several routers using Python?"

I log in to each router, run show version, and extract hostname, version, uptime and serial with regular expressions or a parser. I append one row per router to a list, with a header first, and write it to a CSV file with the csv module so that commas and quotes are handled. Then I check the row count and that no value is empty before sharing it.

Key takeaways

  • A CSV file is text: a header line, then one record per line.
  • "w" replaces the file, "a" appends, and with closes it safely.
  • Joined strings are fine for simple values; the csv module handles commas and quotes.
  • Lab 2 adds the serial pattern Processor board ID (\S+) and writes inventory.csv.
  • Check header, row count and empty values before you share a report.
08

Troubleshooting workflow: when parsing returns nothing or the wrong thing

Parsing failures come in two kinds. The loud kind stops the script with a traceback. The quiet kind is worse: the script finishes, but the data is wrong or incomplete. This chapter gives one workflow for both kinds, using the lab mistakes as examples. It works for split(), regular expressions and parsers.

The five-step workflow

  1. Read the last line of the traceback. The exception class tells you what kind of mismatch happened.
  2. Look at the raw text exactly as the script sees it, not as the terminal shows it.
  3. Test one line and one pattern on their own, in a few lines of code.
  4. Fix one thing and run again on all devices, because the second router may differ from the first.
  5. Add a check so that the next failure is loud and names the device.
Exceptionlast lineRaw textprint reprOne linetest aloneFix and re-runall devices

From traceback to fix

Step 1: what each exception is telling you

ExceptionTypical cause in parsingWhere to look
AttributeError: 'NoneType' ... 'group're.search found nothing, so .group was called on NoneThe pattern against the real text; case
KeyError: 'interface'The dictionary has no such key; it is empty or has another shapeWhat parse returned; was it assigned?
IndexError: list index out of rangeA line has fewer words than the position you asked forBlank lines, headings, short rows
ValueErrorText where a number or a fixed count of words was expectedThe row that carries a two-word value

Here is each one, produced on purpose, with the real message:

import re

line = "GigabitEthernet0/3 unassigned YES unset administratively down down"
parts = line.split()

try:
    re.search(r"version (\S+),", "Version 2.0, RELEASE SOFTWARE").group(1)
except AttributeError as e:
    print("AttributeError:", e)

try:
    print(parts[9])
except IndexError as e:
    print("IndexError:", e)

try:
    data = {"interface": {}}
    print(data["interfaces"])
except KeyError as e:
    print("KeyError:", e)

try:
    int(parts[1])
except ValueError as e:
    print("ValueError:", e)

Real output, checked in the lab terminal.

AttributeError: 'NoneType' object has no attribute 'group'
IndexError: list index out of range
KeyError: 'interfaces'
ValueError: invalid literal for int() with base 10: 'unassigned'

Step 2: see the real characters

The terminal hides what a script sees: trailing spaces, tabs and a stray carriage return (\r) at the end of a line. repr() prints a value with all of them visible. Notice the trailing spaces and the \r in the first line, and how split() and strip() remove them:

line = "Loopback0   1.1.1.1   YES manual up   up   \r"
print(repr(line))
print(line.split())
print(repr(line.strip()))

Real output, checked in the lab terminal.

'Loopback0   1.1.1.1   YES manual up   up   \r'
['Loopback0', '1.1.1.1', 'YES', 'manual', 'up', 'up']
'Loopback0   1.1.1.1   YES manual up   up'

split() with no argument ignores the extra spaces and the \r, which is another reason to never write split(" "). Whenever a pattern "should match but does not", print(repr(text)) is the first debugging step.

Step 3: test in isolation

Do not debug a pattern inside a loop over three routers. Copy one line of real output into a variable and try the pattern on it. Print what each group gives. Change one thing at a time.

In lab 1, the fault was not in the loop at all. up2.py started with data = {}, so the dictionary was empty and the first access stopped with this real traceback:

from nkmiko import ConnectHandler

conn = ConnectHandler(device_type="nk_ios", host="10.99.0.1", username="netops", password="NK@2026")
output = conn.send_command("show ip interface brief")
conn.disconnect()

data = {}
for intf, info in data["interface"].items():
    print(intf)

Run in the lab terminal (it uses the lab routers). Output captured from the simulator.

Traceback (most recent call last):
  File "keyerr.py", line 8, in <module>
    for intf, info in data["interface"].items():
KeyError: 'interface'

The last line, KeyError: 'interface', says: the key is missing. The fix was one line: data = parse("show ip interface brief", output). In lab 2 the last line was AttributeError, and one letter, v against V, was the cause.

Step 4: quiet failures need checks

A script that prints a report with an empty column does not crash, and nobody knows. Make the script check its own result. This helper reports empty values and placeholders instead of hiding them:

def check_rows(rows, expected_min):
    problems = []
    if len(rows) < expected_min:
        problems.append(f"only {len(rows)} rows, expected at least {expected_min}")
    for r in rows:
        for key, value in r.items():
            if value == "" or value == "TODO":
                problems.append(f"{r['hostname']}: empty or TODO in {key}")
    return problems

rows = [
    {"hostname": "BLR-CORE", "version": "2.0", "serial": "NK3C15U4X00"},
    {"hostname": "BLR-WAN", "version": "", "serial": "TODO"},
]
for p in check_rows(rows, 3):
    print(p)

Real output, checked in the lab terminal.

only 2 rows, expected at least 3
BLR-WAN: empty or TODO in version
BLR-WAN: empty or TODO in serial

Notice that the message names the device and the column, so the engineer who reads it knows where to look. The same idea applies to parser output: check that data["interface"] is not empty before you trust a report that says "no interfaces are up".

A safe helper, used carefully

You can wrap re.search so that a missing match returns a default instead of crashing. This is convenient, but it hides bugs, because a wrong pattern now produces a quiet not found. Count the defaults and report them.

import re

def grab(pattern, text, default="not found"):
    m = re.search(pattern, text)
    return m.group(1) if m else default

text = "BLR-CORE uptime is 0 minutes\nProcessor board ID NK3C15U4X00"
found = [grab(r"Version (\S+),", text), grab(r"Processor board ID (\S+)", text)]
print(found)
print("Missing values:", found.count("not found"))

Real output, checked in the lab terminal.

['not found', 'NK3C15U4X00']
Missing values: 1

Common mistakes. (1) Debugging by changing three things at once. (2) Testing on one router and assuming all are the same. (3) Wrapping everything in try with pass, which turns a loud failure into silent wrong data. (4) Trusting the terminal view instead of repr(). (5) Ending the fix without re-running the whole script.

Exam trap. A script that exits without an error is not proof that the result is correct. Questions often ask for the first check: read the exception, then look at the raw input.

The report that said nothing was wrong

A nightly script listed interfaces that were down and mailed the list. For a month the mail was empty, which everybody read as good news. After a software upgrade the heading text had changed, the pattern matched nothing, and the loop quietly ran zero times. A new check made the script fail if it parsed fewer than one interface per router.

Lesson: an empty result needs a reason; check that the parse produced data before you trust a negative.

"A parsing script crashes with an AttributeError on NoneType. How do you troubleshoot it?"

That means re.search returned None, so the pattern did not match the real text. I print repr of the raw output, copy one line into a small test, and compare it with the pattern, checking case, spaces and labels. After fixing it I run the script against every device, and I add a check so that a missing match reports the device name instead of crashing or going quiet.

Key takeaways

  • The exception class shows the mismatch: None, missing key, short row, wrong type.
  • repr() shows the characters a script really receives.
  • Test one line and one pattern alone, then re-run on every device.
  • Quiet failures are worse than crashes; add checks for empty and placeholder values.
  • A default value hides bugs unless you count and report it.
09

Summary and exam checklist

You can now turn the text of a show command into data, with three different methods, and write the result to a file. This chapter collects the facts for revision and for the exam.

Can-do checklist

Tick an item only if you can do it without looking:

  • Explain screen scraping and why it breaks when the text changes.
  • Cut a string into lines with splitlines() and a line into words with split().
  • Read the first word with parts[0] and the last with parts[-1], and explain why counting starts at 0.
  • Explain why administratively down shifts the columns and how the last word, split(None, n) or a label avoids it.
  • Write a regular expression with \d, \S, ., \., +, *, ? and use re.search, re.match and re.findall.
  • Use groups ( ) and group(1), named groups and groupdict().
  • Explain why search can return None and what the AttributeError on NoneType means.
  • Call parse(command, output) and walk data["interface"] by name.
  • Write rows to a CSV file with open(..., "w") and say when the csv module is needed.
  • Troubleshoot a parsing script: exception, raw text, one line alone, fix, re-run on all devices, add a check.

Which method, when

MethodStrong pointWeak pointUse it when
split()Very simple, no importBreaks on two-word values and new columnsQuick job, fixed simple table
Regular expressionOne precise value from any textCase, spacing, hard to readSingle values such as version or serial
ParserNamed fields, maintained, testedOnly for commands that have a parserTables and anything you will rely on
API returning JSONNo parsing at allDevice must support itWhenever it exists

Python quick reference

TaskCode
Lines of the outputoutput.splitlines()
Words of a lineline.split()
Last wordparts[-1]
Limited splitline.split(None, 1)
First match or Nonere.search(r"Version (\S+),", text)
Captured textm.group(1)
Named groups to dictionarym.groupdict()
All matchesre.findall(r"\d+", text)
Parser callparse("show ip interface brief", output)
Write a filewith open("inventory.csv", "w") as f:
Show hidden charactersrepr(text)

Mini glossary

parsing
Turning text for humans into data for programs.
screen scraping
Pulling values out of command output as plain text.
split
A string method that cuts text into a list of pieces.
regular expression
A pattern that describes the shape of text to find or extract.
raw string
A string written with r in front, so backslashes are kept as they are.
group
Round brackets in a pattern that capture part of the match.
None
The value meaning nothing was found.
parser
A library that turns the output of a known command into a dictionary or list.
CSV
Comma-separated values: a text file of rows that a spreadsheet opens.
repr
A function that shows a value with every hidden character visible.

Most tested facts

  • splitlines() gives lines; split() with no argument gives words and ignores repeated spaces.
  • administratively down is two words; the last word parts[-1] is still safe.
  • re.search returns a match object or None; .group on None raises AttributeError.
  • Regex is case-sensitive: version does not match Version.
  • group(1) is the first bracket; group(0) is the whole match.
  • A parser returns a dictionary (Genie style) or a list of dictionaries (TextFSM); lab parse gives data["interface"][name][field].
  • open(..., "w") replaces a file; "a" appends.
  • A script that exits without error can still hold wrong data: check the result.

Both labs in one paragraph

Lab 1, interfaces. The raw row of an administratively-down interface has 7 words. up.py prints only names whose last word is up; up2.py with data = parse("show ip interface brief", output) prints the up/up interfaces with IPs, and on MUM-R2 there are 4 of them.

Lab 2, inventory. inv.py crashes with AttributeError because the pattern says version while the output says Version. After the fix, the serial pattern Processor board ID (\S+) and the file write produce inventory.csv; the serial of BLR-WAN is NK45ZYX1X00.

Putting it together

A complete collection script has four stages: collect the text over SSH, parse it into data, store the data as JSON, YAML or CSV, and verify that the result is complete. Each stage can fail, so each stage gets a check. You now have the Python and data-format skills for that pipeline.

The script that found the missing router

A small team parsed show version from every branch router into a CSV each night. One morning the report had 41 rows instead of 42. The row-count check failed the job and named the router that returned no match. A firmware label on that device had a different spelling, a one-line fix to the pattern. Without the check, the missing router would have been noticed at the next audit.

Lesson: collect, parse, store and verify; the check is part of the script.

"When do you use split, regular expressions or a parser to process command output?"

I use split for a quick job on a simple fixed table, being careful with values that contain spaces. I use a regular expression when I need one precise value from free text, such as a version or serial, anchored on a stable label. I prefer a parser for tables and for anything I will rely on, because it returns named fields and is maintained. Whichever I choose, I check that the result is complete.

Key takeaways

  • Parsing turns CLI text into lists and dictionaries; three methods: split, regex, parser.
  • Split is simple but breaks on two-word values; regex needs care with case and labels; parsers give named fields.
  • search returns None on no match, and .group on None is an AttributeError.
  • Write reports with open(..., "w") or the csv module and check them before sharing.
  • Collect, parse, store, verify: always check the result.
๐ŸŽ“ For educational purposes only โ€” all devices are simulationsTerms of UsePrivacy Policyยฉ 2026 Network Kings
CONFIG by Network Kings โ€” an educational IT simulation platform for learning purposes only. It is not Cisco IOS, Junos, FortiOS or PAN-OS and contains no Cisco, Juniper, Fortinet or Palo Alto Networks software. Cisco, IOS, CCNA, CCNP, Juniper, JNCIA, JNCIS, JNCIP, Fortinet, FortiGate, FortiOS, NSE, Palo Alto Networks, PAN-OS and PCNSE are trademarks of their respective owners. Network Kings is not affiliated with or endorsed by Cisco Systems, Inc., Juniper Networks, Inc., Fortinet, Inc. or Palo Alto Networks, Inc.