Jump to chapter (9)
Why parse? From CLI text to data you can use
What you will learn in this module. Network devices answer show commands with text for humans. A script needs data: lists, dictionaries, numbers. Turning the first into the second is called parsing. You will learn three ways to parse, from the simplest to the most robust: cutting text with split(), finding patterns with regular expressions (the re module), and using a parser that returns ready-made dictionaries. In the labs you build an "interfaces that are up" list and a CSV inventory of routers. This is part of the CCNP Automation (350-901 AUTOCOR) track.
Prerequisites. Python 1 to 3 of this track: strings and f-strings, lists and dictionaries, loops and if, functions, files and try/except. The previous module showed how data is stored in JSON and YAML; this module is about getting data out of screen text.
Analogy: reading a parcel label
A courier label says "Mr Rao, Flat 12, Tower B, Pune 411001, Ph 98xxxxxx". A person reads it at a glance. A sorting machine needs fields: name, flat, city, pincode. To get fields out of the label, you can (1) cut it at the commas, (2) search it for a pattern such as "six digits is a pincode", or (3) use a barcode that already carries the fields. These are exactly the three methods of this module: split, regex and a parser.
What the raw output looks like
This is show ip interface brief from router MUM-R1 of the first lab, captured from the simulator.
MUM-R1# show ip interface brief Interface IP-Address OK? Method Status Protocol GigabitEthernet0/0 10.99.0.1 YES manual up up GigabitEthernet0/1 10.12.0.1 YES manual up up GigabitEthernet0/2 10.50.0.1 YES manual down down GigabitEthernet0/3 unassigned YES unset administratively down down Loopback0 1.1.1.1 YES manual up up
A human sees a table. A program sees one long string with new-line characters. Column 1 is the interface name, column 2 the IP, and the last two columns the Status and Protocol. The question for the NOC: which interfaces are really up (status up and protocol up)?
The screen-scraping idea
Treating output as text and picking pieces out of it is called screen scraping. It is how automation started, and it is still needed because many devices and commands have no API. Its weakness is that it depends on the exact shape of the text:
- a different software version may add a column,
- an interface that is administratively down prints
administratively down(two words) in one column, - an empty line, a banner or a warning can appear in the middle.
Three ways to parse
So the rule of this module is: start with the simplest method that is safe, and move up when the output gets harder.
The pipeline of every collection script
- Connect (SSH) and run the command: you get a string.
- Parse the string into data (list or dictionary).
- Use the data: filter, count, compare, save as JSON or CSV, build a report.
The labs use nkmiko for step 1 and focus on steps 2 and 3. Here is the whole idea in a few lines of plain Python, using a short piece of the output:
text = """GigabitEthernet0/0 10.99.0.1 YES manual up up
GigabitEthernet0/2 10.50.0.1 YES manual down down"""
for line in text.splitlines():
parts = line.split()
print(parts[0], "is", parts[-1])Real output, checked in the lab terminal.
GigabitEthernet0/0 is up GigabitEthernet0/2 is down
splitlines() cuts the text into lines, split() cuts a line into words, and parts[0] and parts[-1] pick the first and last word. You will practice every piece in the coming chapters.
What the labs ask you to do
Lab 1 (up.py, up2.py) collects show ip interface brief from MUM-R1 and MUM-R2 and lists only the interfaces that are up/up, first with split(), then with a parser. Lab 2 (inv.py) reads show version from three Bengaluru routers with regular expressions, fixes one broken pattern, and writes inventory.csv with hostname, version, uptime and serial.
Common beginner mistake. Believing parsing is a one-time job. Output changes with software versions, so a good script checks what it found (for example, that a pattern matched) and reports clearly when it did not.
Exam trap. The exam contrasts unstructured CLI text with structured data from APIs, and asks which approach is more reliable. Structured data (RESTCONF, NETCONF) is more reliable than screen scraping; parsers such as TextFSM and Genie sit in between.
The report that broke after an upgrade
A team wrote a script that read the fifth word of every line of show ip interface brief. It worked for two years. After a software upgrade, one device printed an extra column and the script began reporting interfaces as "up" when they were down. Nobody noticed for a week because the script never failed; it just returned wrong answers. The team rewrote it with a parser keyed by field names and added a check that the number of interfaces matched the device's own count.
Lesson: a parser that silently returns wrong data is worse than one that crashes. Verify what you parse.
"Why is parsing CLI output fragile, and how do you reduce the risk?"
CLI output is meant for humans, so its columns, spacing and wording can change between versions and platforms. I reduce the risk by preferring structured sources (RESTCONF or NETCONF) when available, otherwise using a tested parser keyed by field names, anchoring regular expressions to labels instead of positions, and always checking that the parse found what I expected.
Key takeaways
- Parsing turns CLI text for humans into lists and dictionaries for programs.
- Three methods:
split(), regular expressions, and ready-made parsers. - Screen scraping depends on the exact shape of the text and can break after upgrades.
- Pipeline: run the command, parse the string, use the data, verify.
- Always check that the parse succeeded; silent wrong data is worse than a crash.
Cutting text: splitlines, split, strip and indexes
The simplest parsing tools are string methods you may already know. This chapter uses them on the real output of show ip interface brief from lab 1. One small idea per section.
Step 1: the output is one string
When a script receives a command's output, it is a single string with \n (new-line) characters inside. To make it easy to practice, we keep a copy of the output of router MUM-R1 in a variable. The triple quotes """ let a string span many lines.
Step 2: splitlines cuts text into lines
text.splitlines() returns a list with one item per line, without the new-line characters.
output = """Interface IP-Address OK? Method Status Protocol GigabitEthernet0/0 10.99.0.1 YES manual up up GigabitEthernet0/1 10.12.0.1 YES manual up up GigabitEthernet0/2 10.50.0.1 YES manual down down GigabitEthernet0/3 unassigned YES unset administratively down down Loopback0 1.1.1.1 YES manual up up""" lines = output.splitlines() print(len(lines)) print(lines[0]) print(lines[1])
Real output, checked in the lab terminal.
6 Interface IP-Address OK? Method Status Protocol GigabitEthernet0/0 10.99.0.1 YES manual up up
Six lines: one heading and five interfaces. lines[0] is the heading, because counting starts at 0. We do not want the heading in our data, so we skip it with a slice: lines[1:] means "from item 1 to the end".
Step 3: split cuts a line into words
line.split() with no argument cuts at any run of spaces and returns a list of words. The spaces, however many, vanish. That is the magic: columns lined up with many spaces become tidy list items.
line = "GigabitEthernet0/1 10.12.0.1 YES manual up up" parts = line.split() print(parts) print(len(parts)) print(parts[0]) print(parts[1]) print(parts[-1])
Real output, checked in the lab terminal.
['GigabitEthernet0/1', '10.12.0.1', 'YES', 'manual', 'up', 'up'] 6 GigabitEthernet0/1 10.12.0.1 up
The line became six words. Each can be reached by its position: parts[0] is the interface name, parts[1] the IP address, and parts[-1] the last word, the protocol. A negative index counts from the end, so -1 is the last item and -2 the one before it. Using the last word is handy because it stays at the end even if columns are added in the middle.
[[flow From string to fields | output/one long string | splitlines()/list of lines | split()/list of words | parts[0], parts[-1]/pick fields]]
Step 4: loop and decide
Combine the pieces with the loop and if you know. This is the exact TODO of lab 1: print the interface name only when the last word is up.
output = """Interface IP-Address OK? Method Status Protocol
GigabitEthernet0/0 10.99.0.1 YES manual up up
GigabitEthernet0/1 10.12.0.1 YES manual up up
GigabitEthernet0/2 10.50.0.1 YES manual down down
GigabitEthernet0/3 unassigned YES unset administratively down down
Loopback0 1.1.1.1 YES manual up up"""
for line in output.splitlines()[1:]:
parts = line.split()
if parts[-1] == "up":
print(parts[0])Real output, checked in the lab terminal.
GigabitEthernet0/0 GigabitEthernet0/1 Loopback0
Three interfaces are up/up. GigabitEthernet0/2 is down/down and GigabitEthernet0/3 is administratively down, so neither is printed.
What the lab script prints before you finish it
In the lab, up.py logs in to both routers and prints each row as a list of words, before the TODO is solved. This is the real output from the lab terminal. Notice the two lists that are longer than the others.
from nkmiko import ConnectHandler
ROUTERS = {"MUM-R1": "10.99.0.1", "MUM-R2": "10.99.0.2"}
for name, ip in ROUTERS.items():
conn = ConnectHandler(device_type="nk_ios", host=ip, username="netops", password="NK@2026")
output = conn.send_command("show ip interface brief")
conn.disconnect()
print(f"=== {name} ===")
for line in output.splitlines()[1:]:
parts = line.split()
print(parts)Run in the lab terminal (it uses the lab routers). Output captured from the simulator.
=== MUM-R1 === ['GigabitEthernet0/0', '10.99.0.1', 'YES', 'manual', 'up', 'up'] ['GigabitEthernet0/1', '10.12.0.1', 'YES', 'manual', 'up', 'up'] ['GigabitEthernet0/2', '10.50.0.1', 'YES', 'manual', 'down', 'down'] ['GigabitEthernet0/3', 'unassigned', 'YES', 'unset', 'administratively', 'down', 'down'] ['Loopback0', '1.1.1.1', 'YES', 'manual', 'up', 'up'] === MUM-R2 === ['GigabitEthernet0/0', '10.99.0.2', 'YES', 'manual', 'up', 'up'] ['GigabitEthernet0/1', '10.12.0.2', 'YES', 'manual', 'up', 'up'] ['GigabitEthernet0/2', '10.60.0.1', 'YES', 'manual', 'administratively', 'down', 'down'] ['GigabitEthernet0/3', 'unassigned', 'YES', 'unset', 'administratively', 'down', 'down'] ['Loopback0', '2.2.2.2', 'YES', 'manual', 'up', 'up'] ['Loopback1', '22.22.22.22', 'YES', 'manual', 'up', 'up']
The rows for the administratively down interfaces have seven words, not six, because administratively down is two words. The next chapter explains why that matters.
Useful companions
| Method | What it does | Example result |
|---|---|---|
s.strip() | removes spaces and new-lines at both ends | " up " becomes "up" |
s.lower() | lowercase | "UP" becomes "up" |
s.startswith("Gig") | True if the text starts with that | for choosing rows |
"up" in s | True if the text is inside | a quick filter |
s.replace("a", "b") | replaces text | clean-up |
", ".join(list) | glues a list into text | for reports |
line = " GigabitEthernet0/1 "
print(repr(line.strip()))
print(line.strip().startswith("Gig"))
print("Loopback" in "Loopback0 1.1.1.1")
print(", ".join(["R1", "R2", "R3"]))Real output, checked in the lab terminal.
'GigabitEthernet0/1' True True R1, R2, R3
Common mistakes. (1) Forgetting [1:] and treating the heading line as an interface, which gives wrong rows or an index error. (2) Using split(" ") with one space: it creates empty items when there are many spaces. Use split() with no argument. (3) Using a fixed index such as parts[4] for the status. (4) Calling split on an empty line: "".split() gives an empty list, and parts[0] then raises IndexError.
Exam trap. "a b".split() gives ['a', 'b'] but "a b".split(" ") gives ['a', '', 'b']. Negative index -1 is the last item. splitlines() returns a list.
The empty line that crashed the report
A script looped over every line of a command output and read parts[0]. It ran fine on lab devices, then crashed in production with IndexError. A trailing blank line in the output produced an empty list, and index 0 of an empty list does not exist. Adding if not parts: continue (skip blank lines) fixed it.
Lesson: real output contains blank lines; always guard against empty rows.
"How do you split a multi-line command output into fields in Python?"
I call splitlines() to get the lines, skip the header with a slice such as [1:], and call split() with no argument on each line to get the words. I use negative indexes for fields near the end and I skip empty lines. For anything less regular than a simple table I switch to regular expressions.
Key takeaways
splitlines()makes a list of lines;[1:]skips the heading.split()with no argument cuts at any run of spaces into words.parts[0]is the first word,parts[-1]the last; counting starts at 0.- Combine with
forandifto filter rows, for example last word equal toup. - Guard against blank lines and always use
split()without an argument.
Where split() breaks: ragged columns and a safer way
split() is wonderful for tidy tables. This chapter shows exactly where it breaks, so you can recognise the danger early and choose a safer method.
The trap: a column with two words
Look again at the rows of show ip interface brief. Most rows have six words, but the row of an interface that is administratively down has seven, because the status is the two-word phrase administratively down. If your script counts positions from the left, the columns shift.
rows = [
"GigabitEthernet0/1 10.12.0.1 YES manual up up",
"GigabitEthernet0/3 unassigned YES unset administratively down down",
]
for row in rows:
parts = row.split()
print(len(parts), "words | status =", parts[4], "| protocol =", parts[5])Real output, checked in the lab terminal.
6 words | status = up | protocol = up 7 words | status = administratively | protocol = down
For the first row everything looks right. For the second, parts[4] is administratively and parts[5] is down, so a script that reads the fixed positions 4 and 5 reports a nonsense status, and it does so without any error. Silent wrong data is the dangerous kind.
Why position counting breaks
Fix 1: count from the end for the last column
The protocol is always the last word, so parts[-1] stays correct for both rows. That is exactly why the lab filter parts[-1] == "up" works. Counting from the end is safe for the last column, but not for the middle ones.
Fix 2: split in a limited number of places
line.split(None, 4) cuts at the first 4 gaps only and leaves the remainder as one piece. Then rsplit(None, 1) cuts one word from the right. Together they treat the two-word status as a single field.
rows = [
"GigabitEthernet0/1 10.12.0.1 YES manual up up",
"GigabitEthernet0/3 unassigned YES unset administratively down down",
]
for row in rows:
name, ip, ok, method, rest = row.split(None, 4)
status, protocol = rest.rsplit(None, 1)
print(name, "|", ip, "|", status, "|", protocol)Real output, checked in the lab terminal.
GigabitEthernet0/1 | 10.12.0.1 | up | up GigabitEthernet0/3 | unassigned | administratively down | down
Now both rows give the right status, including administratively down. This works, but you can see the price: the code is harder to read, and it still depends on the exact number of columns.
Fix 3: cut at a label instead of a position
For lines such as BLR-CORE uptime is 0 minutes, the better anchor is the words that never change. split("separator") accepts a multi-character separator.
line = "BLR-CORE uptime is 0 minutes"
host, uptime = line.split(" uptime is ")
print(host)
print(uptime)
board = "Processor board ID NK3C15U4X00"
print(board.split()[-1])
image = 'System image file is "flash:nkos-router-universal.2.0.bin"'
print(image.split('"')[1])Real output, checked in the lab terminal.
BLR-CORE 0 minutes NK3C15U4X00 flash:nkos-router-universal.2.0.bin
Three useful moves are shown: cut at a phrase, take the last word, and take the text between quotes (split('"')[1] is the second piece, the one inside the quotes). Anchoring on a label survives column changes better than counting positions.
Choosing the right line first
Most of the output is not data you want. Select the line before splitting it, using in, startswith or a loop with if.
show_version = """Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE
ROM: NK-OS Bootstrap, Version 2.0, RELEASE SOFTWARE
BLR-CORE uptime is 0 minutes
System returned to ROM by power-on
Processor board ID NK3C15U4X00
4 Gigabit Ethernet interfaces"""
for line in show_version.splitlines():
if "uptime is" in line:
print("uptime line:", line)
if line.startswith("Processor board ID"):
print("serial:", line.split()[-1])Real output, checked in the lab terminal.
uptime line: BLR-CORE uptime is 0 minutes serial: NK3C15U4X00
When to stop using split
Use split() while all these are true: the output is a regular table; every row has the same number of fields; you only need the first or last column; and the format will not change soon. Switch to a stronger tool when any of these fails:
| Situation | Better tool |
|---|---|
Value is inside a sentence (Version 2.0,) | Regular expression |
| Fields contain spaces | Regular expression or a parser |
| Many commands, many platforms | A ready-made parser |
| You must be sure of the result | A parser plus checks |
The next two chapters teach regular expressions, the first stronger tool.
Common mistakes. (1) Reading parts[4] and assuming that position never moves. (2) Not checking len(parts) before indexing, which gives IndexError on short rows. (3) Splitting a line that you have not first selected, so a banner line accidentally produces data. (4) Forgetting that trailing punctuation sticks to words, such as 2.0, with a comma.
Exam trap. Screen scraping with positions is fragile because output formats change; the exam's recommended order is structured API first, then a template-based parser, then regex or split as a last resort.
The outage that was only a column
A dashboard counted a port healthy when parts[4] == "up". After a software upgrade added one new column before the status, parts[4] held a different word on every row, so the dashboard showed zero healthy ports and the NOC was paged for an outage that did not exist. The network was fine; the parser was reading the wrong column. Switching the check to the last word, and adding a test that compares the count of up ports with the device's own summary, ended the false alarms.
Lesson: test your parser on the ugly rows and after every upgrade, not just on the clean lab output.
"A split-based parser works on your lab but misreads a few production rows. Why?"
Production output has irregular rows: status phrases with two words, wrapped lines, blank lines, or extra columns after an upgrade. Position-based splitting shifts the columns. I would anchor on labels or the last column, use limited splits where the shape is known, add length checks, and move to regular expressions or a tested parser when the table is not regular.
Key takeaways
- A row with a two-word value has more words and shifts the columns.
- The last word is safe to read with
parts[-1]; middle positions are not. split(None, n)andrsplit(None, 1)limit where the cuts happen.- Cut at a label such as
" uptime is "and select the right line first. - When values sit inside sentences or fields hold spaces, use regular expressions or a parser.
Regular expressions 1: search for a pattern with re
A regular expression (regex) is a small language for describing text patterns, such as "the word Version, a space, then some non-space characters". Instead of counting words or positions, you describe what the value looks like, and the re module finds it. This is the tool for values hidden inside sentences, and it is what lab 2 uses.
The idea: describe the shape
Think of a shape-sorting toy. Each hole accepts only one shape. A regex is a set of holes; text passes through only where its shape fits. A pattern such as \d+ means "one or more digits", so it fits 2026 and 10 but not abc.
Your first search
re.search(pattern, text) looks anywhere in the text for the first place where the pattern fits. It returns a match object if found, or None if not.
import re line = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE" m = re.search(r"Version", line) print(m) print(m.group()) print(re.search(r"Cisco", line))
Real output, checked in the lab terminal.
<re.Match object; span=(51, 58), match='Version'> Version None
The first search found the word and printed a match object, which says where the match sits (span) and what it matched. m.group() returns the matched text. The last search found nothing and returned None. Always remember: no match means None, and calling a method on None is an error you will meet in the lab.
The raw string prefix
Notice r"Version": a string with an r in front is a raw string. Python leaves backslashes alone inside it. Regex uses backslashes a lot (\d, \s), so write patterns as raw strings every time.
The building blocks
| Piece | Meaning | Example |
|---|---|---|
abc | the literal text | Version finds Version |
\d | one digit | \d\d finds 20 |
\s | one whitespace (space, tab) | |
\S | one non-whitespace | |
\w | letter, digit or underscore | |
. | any one character | a.c finds abc |
\. | a real dot | 10\.99 |
+ | one or more of the previous | \d+ finds 2026 |
* | zero or more | |
? | zero or one | |
^ and $ | start and end of the line |
A pattern is a row of holes
Combine them: Version \S+, means "the word Version, a space, one or more non-space characters, and a comma". Try it on the line above.
import re
line = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE"
m = re.search(r"Version \S+,", line)
print(m.group())
print(re.search(r"\d+\.\d+", line).group())
print(re.findall(r"[A-Z]{2,}", line))Real output, checked in the lab terminal.
Version 2.0, 2.0 ['OS', 'NK', 'OS', 'RELEASE', 'SOFTWARE']
The first pattern matched Version 2.0,, including the comma, which is usually not what we want. The second, \d+\.\d+, matched the number 2.0 (digits, a real dot, digits). The third uses [A-Z]{2,}: a character class [A-Z] (any capital letter) repeated 2 or more times; it found the runs of capital letters in the line. In the next chapter, groups let you keep only the part you need.
search, match and findall
re.search(p, text): the pattern may be anywhere.re.match(p, text): the pattern must fit at the start of the text.re.findall(p, text): returns a list of all non-overlapping matches.
import re text = "R1 10.99.0.1 up\nR2 10.99.0.2 down" print(bool(re.match(r"R1", text))) print(bool(re.match(r"10", text))) print(bool(re.search(r"10", text))) print(re.findall(r"\d+\.\d+\.\d+\.\d+", text))
Real output, checked in the lab terminal.
True False True ['10.99.0.1', '10.99.0.2']
match fails for 10 because the text starts with R1; search finds it anywhere. findall collected both IP addresses. Remember the escaped dots \. in the IP pattern: an unescaped dot would match any character.
Case matters
A regex is case-sensitive by default. Version does not match version. You can switch case-insensitive on with re.IGNORECASE (or re.I). This small fact is behind a lab bug.
import re line = "System Version 2.0" print(re.search(r"version", line)) print(re.search(r"version", line, re.I).group())
Real output, checked in the lab terminal.
None Version
Test a pattern before you trust it
Write the pattern, run it on a sample of real output, and print the result. If the result is None, adjust. Most regex work is this loop of small experiments.
Common mistakes. (1) Forgetting the r before the quotes. (2) An unescaped dot, so 10.99 also matches 10x99. (3) Using match when you meant search. (4) Using the result without checking for None. (5) Writing one giant pattern at once instead of building it piece by piece.
Exam trap. re.match anchors at the start; re.search does not. \d is a digit, \s whitespace, \S non-whitespace, . any character, \. a real dot. A failed search returns None.
The dot that matched too much
A script validated management IPs with the pattern 10.99.0.1. A typo address 10x99y0z1 passed the check, because each dot matched any character. The fix was escaping the dots (10\.99\.0\.1) and anchoring the pattern with ^ and $.
Lesson: escape dots, and anchor patterns when you need an exact match.
"What is the difference between re.match, re.search and re.findall?"
match succeeds only if the pattern fits at the start of the string, search finds the first fit anywhere, and findall returns a list of all matches. match and search return a match object or None, so I check the result before using it. I write patterns as raw strings and escape literal dots.
Key takeaways
- A regex describes the shape of text;
re.searchreturns a match object or None. - Write patterns as raw strings:
r"...". - Key pieces:
\d,\s,\S,.,\.,+,*,?,^,$,[A-Z],{2,}. matchanchors at the start;searchlooks anywhere;findallreturns every match.- Regex is case-sensitive unless you pass
re.I.
Regular expressions 2: groups, named groups and show version
The previous chapter found text. This one extracts it. The tool is the group: round brackets ( ) around the part of the pattern whose text you want to keep. With groups you can pull the version, hostname, uptime and serial number out of show version, which is what lab 2 does.
Why groups?
re.search(r"Version \S+,", line) matched Version 2.0,, but we only want 2.0. Put the interesting part in brackets and ask for it with .group(1):
import re line = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE" m = re.search(r"Version (\S+),", line) print(m.group()) print(m.group(1))
Real output, checked in the lab terminal.
Version 2.0, 2.0
m.group() (or group(0)) is the whole match, including the label. m.group(1) is only what the first pair of brackets captured: 2.0. The label Version and the comma act as anchors: they tell the engine where the value begins and ends, and they are not kept. This is anchoring on text that does not change, the idea from the split chapter.
Anchors and a group
Parsing show version with four patterns
This is the output of show version of router BLR-CORE in lab 2 (shortened to the useful lines), and four patterns, one per fact. Each pattern anchors on a stable label.
import re out = """Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE Technical Support: ns.example/support BLR-CORE uptime is 0 minutes System returned to ROM by power-on Processor board ID NK3C15U4X00 4 Gigabit Ethernet interfaces""" host = re.search(r"(\S+) uptime is", out).group(1) version = re.search(r"Version (\S+),", out).group(1) uptime = re.search(r"uptime is (.+)", out).group(1) serial = re.search(r"Processor board ID (\S+)", out).group(1) print(host, version, uptime, serial, sep=" | ")
Real output, checked in the lab terminal.
BLR-CORE | 2.0 | 0 minutes | NK3C15U4X00
Read each pattern aloud. (\S+) uptime is: a group of non-space characters, then a space and the words uptime is; the group is the hostname just before them. uptime is (.+): after those words, .+ means "one or more of any character", which takes the rest of the line. Processor board ID (\S+): the word after the label is the serial.
The case-sensitivity bug of lab 2
In the lab, the version pattern is written with a lowercase v: r"version (\S+),". The output spells it Version. Regex is case-sensitive, so search returns None, and None.group(1) crashes.
import re out = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE" version = re.search(r"version (\S+),", out).group(1) print(version)
Real output, checked in the lab terminal.
Traceback (most recent call last):
File "casebug.py", line 4, in <module>
version = re.search(r"version (\S+),", out).group(1)
AttributeError: 'NoneType' object has no attribute 'group'
The last line, AttributeError: 'NoneType' object has no attribute 'group', is the classic regex failure. It does not mean group is broken. It means the pattern found nothing, so there is no match object. The fix in the lab is one letter, Version, or the flag re.I. A defensive script checks first:
import re
out = "Network Kings Network OS Software, NK-OS Software, Version 2.0, RELEASE SOFTWARE"
m = re.search(r"version (\S+),", out)
if m:
print("version:", m.group(1))
else:
print("version line not found - check the pattern")
print(re.search(r"version (\S+),", out, re.I).group(1))Real output, checked in the lab terminal.
version line not found - check the pattern 2.0
Several groups at once and named groups
A pattern may have many groups: group(1), group(2) and so on, counted by the position of the opening bracket. For readability you can name a group with (?P<name>...) and read it by name.
import re
line = "GigabitEthernet0/3 unassigned YES unset administratively down down"
pat = r"^(?P<name>\S+)\s+(?P<ip>\S+)\s+YES\s+\S+\s+(?P<status>administratively down|up|down)\s+(?P<proto>up|down)$"
m = re.match(pat, line)
print(m.group("name"), "|", m.group("status"), "|", m.group("proto"))
print(m.groupdict())Real output, checked in the lab terminal.
GigabitEthernet0/3 | administratively down | down
{'name': 'GigabitEthernet0/3', 'ip': 'unassigned', 'status': 'administratively down', 'proto': 'down'}
The pattern solves the two-word status problem from the previous chapter. (administratively down|up|down) is an alternation: one of three choices, tried left to right, so the two-word phrase is tried first. ^ and $ pin the pattern to the start and end of the line. groupdict() returns the named groups as a dictionary, which is already half way to structured data.
One pattern over a whole table
re.finditer walks through every match in a multi-line text, and the flag re.M (multi-line) lets ^ and $ match at every line, not only at the ends of the whole text.
import re
output = """Interface IP-Address OK? Method Status Protocol
GigabitEthernet0/0 10.99.0.1 YES manual up up
GigabitEthernet0/2 10.50.0.1 YES manual down down
GigabitEthernet0/3 unassigned YES unset administratively down down
Loopback0 1.1.1.1 YES manual up up"""
pat = r"^(\S+)\s+(\S+)\s+YES\s+\S+\s+(administratively down|up|down)\s+(up|down)$"
rows = [m.groups() for m in re.finditer(pat, output, re.M)]
for row in rows:
print(row)
print("up/up:", [r[0] for r in rows if r[2] == "up" and r[3] == "up"])Real output, checked in the lab terminal.
('GigabitEthernet0/0', '10.99.0.1', 'up', 'up')
('GigabitEthernet0/2', '10.50.0.1', 'down', 'down')
('GigabitEthernet0/3', 'unassigned', 'administratively down', 'down')
('Loopback0', '1.1.1.1', 'up', 'up')
up/up: ['GigabitEthernet0/0', 'Loopback0']
The heading line was skipped automatically because it does not match the pattern (it has no YES). This is a strength of regex: non-matching lines simply do not count. The groups() method returns all groups as a tuple.
Greedy and lazy
.+ is greedy: it grabs as much as it can. .+? is lazy: as little as it can. For <b>x</b><b>y</b> style text the difference matters; for line-based network output you rarely need lazy, but you must know why uptime is (.+) takes the rest of the line.
Common mistakes. (1) Using .group(1) without checking for None. (2) Forgetting that spelling and case must match, such as Version versus version. (3) Using .* where \S+ is meant, so the match swallows a whole line. (4) Anchoring on position instead of on a stable label. (5) Forgetting re.M, so ^ and $ match only at the very start and end of the whole text.
Exam trap. group(0) is the whole match and group(1) the first bracket. A failed search returns None, so .group on it raises AttributeError. re.I makes a pattern case-insensitive.
One capital letter, three routers
Finance asked for an inventory of three routers. The engineer's script crashed on the first router with 'NoneType' object has no attribute 'group'. He suspected SSH, then the regex engine. A colleague printed the raw output next to the pattern and noticed Version with a capital V against version in the pattern. One letter later, all three routers printed with their versions.
Lesson: when a pattern returns None, compare it character by character with the real output.
"How do you extract the software version from show version using Python?"
I use re.search(r"Version (\S+),", output) and read group(1), anchoring on the stable word Version and the comma. I check the result for None before using it, remember that regex is case-sensitive, and test the pattern on real output from more than one device. For many devices and platforms I would use a tested parser instead.
Key takeaways
- Round brackets create a group;
group(1)returns only the captured text. - Anchor patterns on stable labels such as
Version,uptime isandProcessor board ID. Nonefromsearchmeans no match;.groupon it raises AttributeError.- Named groups
(?P<name>...)andgroupdict()give dictionaries. re.finditerwithre.Mruns one pattern over every line of a table.
Structured parsers: dictionaries without counting columns
Splitting and regex put the work on you: you design every pattern, test it on every platform, and repair it after every upgrade. A parser moves that work to a maintained library. You give it the command name and the raw text; it returns a dictionary with named fields. This is the third and most robust method of the module, and the second half of lab 1 (up2.py).
The idea: a library of patterns, indexed by command
A parser library contains, for each command, patterns that were written and tested by many people on many software versions. You call one function:
data = parse("show ip interface brief", output)
and get a dictionary in a known shape. Your code then reads fields by name (data["interface"]["Loopback0"]["ip_address"]) instead of by position.
Parser pipeline
The real tools you will meet
| Tool | Style | Typical use |
|---|---|---|
| Genie parsers (pyATS) | Python classes, returns nested dictionaries | Test frameworks, large Cisco coverage |
| TextFSM with ntc-templates | Template files, returns a list of dictionaries | Netmiko use_textfsm=True |
| TTP | Your own simple templates | Custom outputs |
In the lab you use the simulator's library nkparse, which behaves like a Genie-style parser: the function is parse(command, output). The real tools above are reference code, not run in the simulator:
Reference code (not run in the simulator).
# Genie / pyATS (Genie-style dictionary)
# parsed = device.parse("show ip interface brief")
# Netmiko with TextFSM templates (list of dictionaries)
# rows = conn.send_command("show ip interface brief", use_textfsm=True)
# ntc-templates directly
# from ntc_templates.parse import parse_output
# rows = parse_output(platform="cisco_ios", command="show ip interface brief", data=raw)
The lab library in action
Here is the lab call on router MUM-R1. The script logs in over SSH, gets the output string, and passes the command name and the text to parse. Compare it with the dozens of lines you would need for split or regex.
from nkmiko import ConnectHandler
from nkparse import parse
import json
conn = ConnectHandler(device_type="nk_ios", host="10.99.0.1", username="netops", password="NK@2026")
output = conn.send_command("show ip interface brief")
conn.disconnect()
data = parse("show ip interface brief", output)
print(list(data.keys()))
print(list(data["interface"].keys()))
print(json.dumps(data["interface"]["GigabitEthernet0/3"], indent=2))Run in the lab terminal (it uses the lab routers). Output captured from the simulator.
['interface']
['GigabitEthernet0/0', 'GigabitEthernet0/1', 'GigabitEthernet0/2', 'GigabitEthernet0/3', 'Loopback0']
{
"ip_address": "unassigned",
"interface_is_ok": "YES",
"method": "unset",
"status": "administratively down",
"protocol": "down"
}
Read the shape. The top of the dictionary has one key, interface. Inside it is one entry per interface, keyed by the interface name. Each entry is a small dictionary with the same field names every time: ip_address, interface_is_ok, method, status and protocol. Notice that status for the administratively-down port is administratively down, one value. The parser solved the two-word problem for you.
The shape of the parsed data
Walking the dictionary
Everything from the earlier chapters applies: .items() gives key and value pairs, if filters. This example uses a small dictionary in the same shape, so you can run it anywhere.
data = {"interface": {
"GigabitEthernet0/0": {"ip_address": "10.99.0.1", "status": "up", "protocol": "up"},
"GigabitEthernet0/2": {"ip_address": "10.50.0.1", "status": "down", "protocol": "down"},
"GigabitEthernet0/3": {"ip_address": "unassigned", "status": "administratively down", "protocol": "down"},
"Loopback0": {"ip_address": "1.1.1.1", "status": "up", "protocol": "up"},
}}
for intf, info in data["interface"].items():
if info["status"] == "up" and info["protocol"] == "up":
print(f"{intf:<22}{info['ip_address']}")
shut = [n for n, i in data["interface"].items() if i["status"] == "administratively down"]
print("Shut by an engineer:", shut)
print("Total interfaces:", len(data["interface"]))Real output, checked in the lab terminal.
GigabitEthernet0/0 10.99.0.1 Loopback0 1.1.1.1 Shut by an engineer: ['GigabitEthernet0/3'] Total interfaces: 4
The first loop is the heart of up2.py. The list comprehension shows how easy the admin-down question becomes with named fields.
Finishing up2.py
The lab file up2.py starts with data = {} and a TODO. With an empty dictionary the loop asks for data["interface"] and crashes with KeyError: 'interface'. Replacing the line with data = parse("show ip interface brief", output) fixes it, and the script prints each up/up interface with its IP. This is the real lab output of the finished script:
from nkmiko import ConnectHandler
from nkparse import parse
ROUTERS = {"MUM-R1": "10.99.0.1", "MUM-R2": "10.99.0.2"}
for name, ip in ROUTERS.items():
conn = ConnectHandler(device_type="nk_ios", host=ip, username="netops", password="NK@2026")
output = conn.send_command("show ip interface brief")
conn.disconnect()
data = parse("show ip interface brief", output)
print(f"=== {name} ===")
for intf, info in data["interface"].items():
if info["status"] == "up" and info["protocol"] == "up":
print(f"{intf:<22}{info['ip_address']}")Run in the lab terminal (it uses the lab routers). Output captured from the simulator.
=== MUM-R1 === GigabitEthernet0/0 10.99.0.1 GigabitEthernet0/1 10.12.0.1 Loopback0 1.1.1.1 === MUM-R2 === GigabitEthernet0/0 10.99.0.2 GigabitEthernet0/1 10.12.0.2 Loopback0 2.2.2.2 Loopback1 22.22.22.22
Count the answer to the lab question: on MUM-R2 the up/up interfaces are GigabitEthernet0/0, GigabitEthernet0/1, Loopback0 and Loopback1, which is 4. GigabitEthernet0/2 is administratively down and GigabitEthernet0/3 is unassigned and down.
Other commands, same idea
The same call parses other commands. For show version it returns hostname, version, uptime and serial under the key version:
from nkmiko import ConnectHandler
from nkparse import parse
conn = ConnectHandler(device_type="nk_ios", host="10.99.0.1", username="netops", password="NK@2026")
output = conn.send_command("show version")
conn.disconnect()
print(parse("show version", output))Run in the lab terminal (it uses the lab routers). Output captured from the simulator.
{'version': {'version': '2.0', 'hostname': 'BLR-CORE', 'uptime': '0 minutes', 'chassis_sn': 'NK3C15U4X00', 'chassis': None}}
A command without a parser stops with a clear error, so you know to fall back to regex:
from nkparse import parse
try:
parse("show clock", "10:00:00.000 UTC Fri Oct 2 2026")
except ValueError as e:
print("ValueError:", e)Run in the lab terminal (it uses the lab routers). Output captured from the simulator.
ValueError: no parser for 'show clock' (available: show ip interface brief, show ip ospf neighbor, show ip bgp summary, show ip route, show version, show vlan brief)
Parser or regex?
Prefer a parser when one exists for your command and platform: it is tested, and your code reads fields by name. Use regex for one-off outputs or custom commands with no template. Use split only for the simplest tables. Whatever you choose, still check that the result contains what you expect.
Common mistakes. (1) Assuming a parser exists for every command; it raises an error or returns nothing otherwise. (2) Forgetting that a parser is tied to a platform and version family. (3) Reading fields that are missing in some outputs; use .get() for optional fields. (4) Not checking an empty result, which looks like "no interfaces" instead of "parse failed".
Exam trap. Genie returns nested dictionaries; TextFSM returns a list of dictionaries (one per row). A structured API such as RESTCONF is still preferred over parsing when it is available.
Replacing 400 lines of regex
A team maintained a 400-line file of regex patterns for ten show commands. Each software upgrade broke two or three. They moved to a parser library for the common commands, kept regex only for two custom outputs, and added a small test that parses a saved sample of every command after each upgrade. Maintenance dropped from days to an hour.
Lesson: let a maintained library carry the parsing, and test it on saved real output.
"What is the advantage of using a parser such as Genie or TextFSM over your own regex?"
A parser library holds tested patterns for many commands, platforms and versions, and returns data in a consistent structure with named fields, so my code does not depend on column positions. It is easier to maintain after upgrades. I still validate that the result is not empty, and I fall back to regex only for commands the library does not cover.
Key takeaways
- A parser takes the command name and raw text and returns a dictionary or list of dictionaries.
- Lab
parse("show ip interface brief", output)returnsdata["interface"][name][field]. - Fields are read by name, so the two-word status is no longer a problem.
- Genie returns nested dictionaries; TextFSM returns a list of row dictionaries.
- Check that the parse worked and fall back to regex for commands without a parser.
From parsed data to a CSV report
Parsing is only half of a collection script. The other half is delivering the result: a file that a manager, an auditor or a spreadsheet can use. The most common delivery format is CSV (comma-separated values): plain text, one row per line, fields separated by commas. Excel and Google Sheets open it directly. This chapter finishes lab 2, where inv.py writes inventory.csv.
What a CSV file looks like
A CSV is just text. The first line is the header with the column names. Every other line is one record, with the same number of fields.
From parsing to a report
Method 1: build the lines yourself
For a small job you can build each line with an f-string and join the lines with a new-line character. This is what the lab does. The example uses three fixed rows so that it runs anywhere.
rows = ["hostname,version,uptime,serial"]
rows.append("BLR-CORE,2.0,0 minutes,NK3C15U4X00")
rows.append("BLR-WAN,2.0,0 minutes,NK45ZYX1X00")
with open("inventory.csv", "w") as f:
f.write("\n".join(rows) + "\n")
with open("inventory.csv") as f:
text = f.read()
print(text)
print("Lines in file:", len(text.splitlines()))Real output, checked in the lab terminal.
hostname,version,uptime,serial BLR-CORE,2.0,0 minutes,NK3C15U4X00 BLR-WAN,2.0,0 minutes,NK45ZYX1X00 Lines in file: 3
Three details matter. "w" opens the file for writing and replaces an old file of the same name. "\n".join(rows) puts a new line between rows, so you add one more "\n" at the end to finish the last line. The with block closes the file for you, which is what makes sure that the data is really saved.
The comma problem
Method 1 has one weakness: a value that contains a comma splits into two columns. A location such as Andheri, Mumbai would shift every column after it. The standard csv module solves this by putting quotes around such a value.
import csv import io buffer = io.StringIO() writer = csv.writer(buffer, lineterminator="\n") writer.writerow(["hostname", "location"]) writer.writerow(["MUM-R1", "Andheri, Mumbai"]) writer.writerow(["BLR-CORE", "Bengaluru"]) print(buffer.getvalue()) rows = list(csv.reader(io.StringIO(buffer.getvalue()))) print(rows[1]) print(len(rows[1]), "fields")
Reference code (not run in the simulator). Output below is from standard Python 3.
hostname,location MUM-R1,"Andheri, Mumbai" BLR-CORE,Bengaluru ['MUM-R1', 'Andheri, Mumbai'] 2 fields
The second row is written with quotes around Andheri, Mumbai, and reading it back gives two fields, not three. Use the csv module whenever data comes from people or from free text. For serial numbers and versions Method 1 is enough. For the same job with dictionaries, csv.DictWriter writes a list of dictionaries and picks the columns from the keys.
Finishing lab 2
In lab 2, inv.py already has patterns for the hostname, the version and the uptime. The two open tasks are the serial number and the file. After fixing the Version spelling, this is the finished script and its real output in the lab:
import re
from nkmiko import ConnectHandler
ROUTERS = ["10.99.0.1", "10.99.0.2", "10.99.0.3"]
rows = ["hostname,version,uptime,serial"]
for ip in ROUTERS:
conn = ConnectHandler(device_type="nk_ios", host=ip, username="netops", password="NK@2026")
out = conn.send_command("show version")
conn.disconnect()
host = re.search(r"(\S+) uptime is", out).group(1)
version = re.search(r"Version (\S+),", out).group(1)
uptime = re.search(r"uptime is (.+)", out).group(1)
serial = re.search(r"Processor board ID (\S+)", out).group(1)
rows.append(f"{host},{version},{uptime},{serial}")
with open("inventory.csv", "w") as f:
f.write("\n".join(rows) + "\n")
print("Wrote", len(rows) - 1, "routers to inventory.csv")
print(open("inventory.csv").read())Run in the lab terminal (it uses the lab routers). Output captured from the simulator.
Wrote 3 routers to inventory.csv hostname,version,uptime,serial BLR-CORE,2.0,0 minutes,NK3C15U4X00 BLR-WAN,2.0,0 minutes,NK45ZYX1X00 BLR-LAB,2.0,0 minutes,NK3W0D82X00
The serial number of BLR-WAN, the lab answer, is NK45ZYX1X00. Check the file with the same discipline as any report: header present, one row per router, no row with TODO left in it.
Other report shapes
A CSV is flat: one row per device. When a record has lists inside it, such as several VLANs per switch, JSON from the previous module is the better file. In both cases the idea is the same: collect, parse, store the result, then report from the stored data. A small script that reads inventory.csv back can answer questions such as "which routers run an older version?" without logging in to a single device again.
Common mistakes. (1) Opening the file with "w" when you meant to add rows; "w" erases the old content, "a" appends. (2) Forgetting the final new line, so the next tool glues two records together. (3) Leaving the placeholder TODO in a column and shipping the report. (4) Putting values with commas into Method 1 lines. (5) Writing the report before checking that every pattern matched.
Exam trap. json.dump writes to a file object; json.dumps returns a string. For CSV, the csv module handles quoting; joining with commas does not. Opening a file with "w" truncates it.
The inventory with a shifted column
A team exported a switch inventory with joined strings. One location field held Floor 3, Annexe. In Excel every row below it shifted one column to the right, and the serial numbers appeared under the uptime heading. Nobody noticed for a week, until an audit compared serials with purchase records. Switching to the csv module added the quotes and the audit passed.
Lesson: a report is code too; test it with the ugly values, not only the neat ones.
"How would you produce an inventory report from several routers using Python?"
I log in to each router, run show version, and extract hostname, version, uptime and serial with regular expressions or a parser. I append one row per router to a list, with a header first, and write it to a CSV file with the csv module so that commas and quotes are handled. Then I check the row count and that no value is empty before sharing it.
Key takeaways
- A CSV file is text: a header line, then one record per line.
"w"replaces the file,"a"appends, andwithcloses it safely.- Joined strings are fine for simple values; the
csvmodule handles commas and quotes. - Lab 2 adds the serial pattern
Processor board ID (\S+)and writesinventory.csv. - Check header, row count and empty values before you share a report.
Troubleshooting workflow: when parsing returns nothing or the wrong thing
Parsing failures come in two kinds. The loud kind stops the script with a traceback. The quiet kind is worse: the script finishes, but the data is wrong or incomplete. This chapter gives one workflow for both kinds, using the lab mistakes as examples. It works for split(), regular expressions and parsers.
The five-step workflow
- Read the last line of the traceback. The exception class tells you what kind of mismatch happened.
- Look at the raw text exactly as the script sees it, not as the terminal shows it.
- Test one line and one pattern on their own, in a few lines of code.
- Fix one thing and run again on all devices, because the second router may differ from the first.
- Add a check so that the next failure is loud and names the device.
From traceback to fix
Step 1: what each exception is telling you
| Exception | Typical cause in parsing | Where to look |
|---|---|---|
AttributeError: 'NoneType' ... 'group' | re.search found nothing, so .group was called on None | The pattern against the real text; case |
KeyError: 'interface' | The dictionary has no such key; it is empty or has another shape | What parse returned; was it assigned? |
IndexError: list index out of range | A line has fewer words than the position you asked for | Blank lines, headings, short rows |
ValueError | Text where a number or a fixed count of words was expected | The row that carries a two-word value |
Here is each one, produced on purpose, with the real message:
import re
line = "GigabitEthernet0/3 unassigned YES unset administratively down down"
parts = line.split()
try:
re.search(r"version (\S+),", "Version 2.0, RELEASE SOFTWARE").group(1)
except AttributeError as e:
print("AttributeError:", e)
try:
print(parts[9])
except IndexError as e:
print("IndexError:", e)
try:
data = {"interface": {}}
print(data["interfaces"])
except KeyError as e:
print("KeyError:", e)
try:
int(parts[1])
except ValueError as e:
print("ValueError:", e)Real output, checked in the lab terminal.
AttributeError: 'NoneType' object has no attribute 'group' IndexError: list index out of range KeyError: 'interfaces' ValueError: invalid literal for int() with base 10: 'unassigned'
Step 2: see the real characters
The terminal hides what a script sees: trailing spaces, tabs and a stray carriage return (\r) at the end of a line. repr() prints a value with all of them visible. Notice the trailing spaces and the \r in the first line, and how split() and strip() remove them:
line = "Loopback0 1.1.1.1 YES manual up up \r" print(repr(line)) print(line.split()) print(repr(line.strip()))
Real output, checked in the lab terminal.
'Loopback0 1.1.1.1 YES manual up up \r' ['Loopback0', '1.1.1.1', 'YES', 'manual', 'up', 'up'] 'Loopback0 1.1.1.1 YES manual up up'
split() with no argument ignores the extra spaces and the \r, which is another reason to never write split(" "). Whenever a pattern "should match but does not", print(repr(text)) is the first debugging step.
Step 3: test in isolation
Do not debug a pattern inside a loop over three routers. Copy one line of real output into a variable and try the pattern on it. Print what each group gives. Change one thing at a time.
In lab 1, the fault was not in the loop at all. up2.py started with data = {}, so the dictionary was empty and the first access stopped with this real traceback:
from nkmiko import ConnectHandler
conn = ConnectHandler(device_type="nk_ios", host="10.99.0.1", username="netops", password="NK@2026")
output = conn.send_command("show ip interface brief")
conn.disconnect()
data = {}
for intf, info in data["interface"].items():
print(intf)Run in the lab terminal (it uses the lab routers). Output captured from the simulator.
Traceback (most recent call last):
File "keyerr.py", line 8, in <module>
for intf, info in data["interface"].items():
KeyError: 'interface'
The last line, KeyError: 'interface', says: the key is missing. The fix was one line: data = parse("show ip interface brief", output). In lab 2 the last line was AttributeError, and one letter, v against V, was the cause.
Step 4: quiet failures need checks
A script that prints a report with an empty column does not crash, and nobody knows. Make the script check its own result. This helper reports empty values and placeholders instead of hiding them:
def check_rows(rows, expected_min):
problems = []
if len(rows) < expected_min:
problems.append(f"only {len(rows)} rows, expected at least {expected_min}")
for r in rows:
for key, value in r.items():
if value == "" or value == "TODO":
problems.append(f"{r['hostname']}: empty or TODO in {key}")
return problems
rows = [
{"hostname": "BLR-CORE", "version": "2.0", "serial": "NK3C15U4X00"},
{"hostname": "BLR-WAN", "version": "", "serial": "TODO"},
]
for p in check_rows(rows, 3):
print(p)Real output, checked in the lab terminal.
only 2 rows, expected at least 3 BLR-WAN: empty or TODO in version BLR-WAN: empty or TODO in serial
Notice that the message names the device and the column, so the engineer who reads it knows where to look. The same idea applies to parser output: check that data["interface"] is not empty before you trust a report that says "no interfaces are up".
A safe helper, used carefully
You can wrap re.search so that a missing match returns a default instead of crashing. This is convenient, but it hides bugs, because a wrong pattern now produces a quiet not found. Count the defaults and report them.
import re
def grab(pattern, text, default="not found"):
m = re.search(pattern, text)
return m.group(1) if m else default
text = "BLR-CORE uptime is 0 minutes\nProcessor board ID NK3C15U4X00"
found = [grab(r"Version (\S+),", text), grab(r"Processor board ID (\S+)", text)]
print(found)
print("Missing values:", found.count("not found"))Real output, checked in the lab terminal.
['not found', 'NK3C15U4X00'] Missing values: 1
Common mistakes. (1) Debugging by changing three things at once. (2) Testing on one router and assuming all are the same. (3) Wrapping everything in try with pass, which turns a loud failure into silent wrong data. (4) Trusting the terminal view instead of repr(). (5) Ending the fix without re-running the whole script.
Exam trap. A script that exits without an error is not proof that the result is correct. Questions often ask for the first check: read the exception, then look at the raw input.
The report that said nothing was wrong
A nightly script listed interfaces that were down and mailed the list. For a month the mail was empty, which everybody read as good news. After a software upgrade the heading text had changed, the pattern matched nothing, and the loop quietly ran zero times. A new check made the script fail if it parsed fewer than one interface per router.
Lesson: an empty result needs a reason; check that the parse produced data before you trust a negative.
"A parsing script crashes with an AttributeError on NoneType. How do you troubleshoot it?"
That means re.search returned None, so the pattern did not match the real text. I print repr of the raw output, copy one line into a small test, and compare it with the pattern, checking case, spaces and labels. After fixing it I run the script against every device, and I add a check so that a missing match reports the device name instead of crashing or going quiet.
Key takeaways
- The exception class shows the mismatch: None, missing key, short row, wrong type.
repr()shows the characters a script really receives.- Test one line and one pattern alone, then re-run on every device.
- Quiet failures are worse than crashes; add checks for empty and placeholder values.
- A default value hides bugs unless you count and report it.
Summary and exam checklist
You can now turn the text of a show command into data, with three different methods, and write the result to a file. This chapter collects the facts for revision and for the exam.
Can-do checklist
Tick an item only if you can do it without looking:
- Explain screen scraping and why it breaks when the text changes.
- Cut a string into lines with
splitlines()and a line into words withsplit(). - Read the first word with
parts[0]and the last withparts[-1], and explain why counting starts at 0. - Explain why
administratively downshifts the columns and how the last word,split(None, n)or a label avoids it. - Write a regular expression with
\d,\S,.,\.,+,*,?and usere.search,re.matchandre.findall. - Use groups
( )andgroup(1), named groups andgroupdict(). - Explain why
searchcan return None and what theAttributeErroron NoneType means. - Call
parse(command, output)and walkdata["interface"]by name. - Write rows to a CSV file with
open(..., "w")and say when thecsvmodule is needed. - Troubleshoot a parsing script: exception, raw text, one line alone, fix, re-run on all devices, add a check.
Which method, when
| Method | Strong point | Weak point | Use it when |
|---|---|---|---|
split() | Very simple, no import | Breaks on two-word values and new columns | Quick job, fixed simple table |
| Regular expression | One precise value from any text | Case, spacing, hard to read | Single values such as version or serial |
| Parser | Named fields, maintained, tested | Only for commands that have a parser | Tables and anything you will rely on |
| API returning JSON | No parsing at all | Device must support it | Whenever it exists |
Python quick reference
| Task | Code |
|---|---|
| Lines of the output | output.splitlines() |
| Words of a line | line.split() |
| Last word | parts[-1] |
| Limited split | line.split(None, 1) |
| First match or None | re.search(r"Version (\S+),", text) |
| Captured text | m.group(1) |
| Named groups to dictionary | m.groupdict() |
| All matches | re.findall(r"\d+", text) |
| Parser call | parse("show ip interface brief", output) |
| Write a file | with open("inventory.csv", "w") as f: |
| Show hidden characters | repr(text) |
Mini glossary
- parsing
- Turning text for humans into data for programs.
- screen scraping
- Pulling values out of command output as plain text.
- split
- A string method that cuts text into a list of pieces.
- regular expression
- A pattern that describes the shape of text to find or extract.
- raw string
- A string written with r in front, so backslashes are kept as they are.
- group
- Round brackets in a pattern that capture part of the match.
- None
- The value meaning nothing was found.
- parser
- A library that turns the output of a known command into a dictionary or list.
- CSV
- Comma-separated values: a text file of rows that a spreadsheet opens.
- repr
- A function that shows a value with every hidden character visible.
Most tested facts
splitlines()gives lines;split()with no argument gives words and ignores repeated spaces.administratively downis two words; the last wordparts[-1]is still safe.re.searchreturns a match object or None;.groupon None raises AttributeError.- Regex is case-sensitive:
versiondoes not matchVersion. group(1)is the first bracket;group(0)is the whole match.- A parser returns a dictionary (Genie style) or a list of dictionaries (TextFSM); lab
parsegivesdata["interface"][name][field]. open(..., "w")replaces a file;"a"appends.- A script that exits without error can still hold wrong data: check the result.
Both labs in one paragraph
Lab 1, interfaces. The raw row of an administratively-down interface has 7 words. up.py prints only names whose last word is up; up2.py with data = parse("show ip interface brief", output) prints the up/up interfaces with IPs, and on MUM-R2 there are 4 of them.
Lab 2, inventory. inv.py crashes with AttributeError because the pattern says version while the output says Version. After the fix, the serial pattern Processor board ID (\S+) and the file write produce inventory.csv; the serial of BLR-WAN is NK45ZYX1X00.
Putting it together
A complete collection script has four stages: collect the text over SSH, parse it into data, store the data as JSON, YAML or CSV, and verify that the result is complete. Each stage can fail, so each stage gets a check. You now have the Python and data-format skills for that pipeline.
The script that found the missing router
A small team parsed show version from every branch router into a CSV each night. One morning the report had 41 rows instead of 42. The row-count check failed the job and named the router that returned no match. A firmware label on that device had a different spelling, a one-line fix to the pattern. Without the check, the missing router would have been noticed at the next audit.
Lesson: collect, parse, store and verify; the check is part of the script.
"When do you use split, regular expressions or a parser to process command output?"
I use split for a quick job on a simple fixed table, being careful with values that contain spaces. I use a regular expression when I need one precise value from free text, such as a version or serial, anchored on a stable label. I prefer a parser for tables and for anything I will rely on, because it returns named fields and is maintained. Whichever I choose, I check that the result is complete.
Key takeaways
- Parsing turns CLI text into lists and dictionaries; three methods: split, regex, parser.
- Split is simple but breaks on two-word values; regex needs care with case and labels; parsers give named fields.
searchreturns None on no match, and.groupon None is an AttributeError.- Write reports with
open(..., "w")or thecsvmodule and check them before sharing. - Collect, parse, store, verify: always check the result.