Jump to chapter (9)
Why structured data? JSON, YAML and XML in one picture
What you will learn in this module. You will learn to read, write and repair the three data formats that network automation runs on: JSON, YAML and XML. You will load them into Python, change them, convert one into another, and find a syntax mistake quickly when a file will not load. In the labs you repair an inventory.json that nobody can read and a vlans.yml that makes a config generator crash, then grow a network by editing only data. This is part of the CCNP Automation (350-901 AUTOCOR) track.
Prerequisites. The previous modules: Python 1 (variables, strings), Python 2 (lists, dictionaries, loops) and Python 3 (functions, files and exceptions). You do not need to know any of the three formats yet.
Analogy: a letter versus a form
A letter says "Please send the parcel to Mr Rao, flat 12, Tower B, Pune 411001". A human reads it easily. A machine cannot reliably tell where the name ends and the address begins. Now look at a form: boxes labelled *Name*, *Flat*, *Tower*, *City*, *Pincode*, each box with its value. A machine can read a form without guessing, because every value has a label and a fixed place. Structured data is a form that a computer can read. JSON, YAML and XML are three ways of drawing the form.
The problem with CLI text
The output of show ip interface brief is a letter for humans. To use it in a program you must cut it with split() or regular expressions, and the cuts break the moment a column changes. The next module teaches that skill. But APIs, automation tools and source-of-truth files do not use letters; they exchange forms. So you must be fluent in them.
Same information, three ways to write it
One device, three formats
Here is the same small fact: router R1 has the management IP 10.99.0.1 and two VLANs, 10 and 20.
{"name": "R1", "mgmt_ip": "10.99.0.1", "vlans": [10, 20]}
That was JSON (JavaScript Object Notation): braces, quotes, commas.
name: R1 mgmt_ip: 10.99.0.1 vlans: - 10 - 20
That was YAML ("YAML Ain't Markup Language"): no braces, no quotes, and the structure comes from indentation. Easy for humans to write.
<device>
<name>R1</name>
<mgmt_ip>10.99.0.1</mgmt_ip>
<vlans>
<vlan>10</vlan>
<vlan>20</vlan>
</vlans>
</device>
That was XML (eXtensible Markup Language): every value sits between an opening tag and a closing tag. It is wordy but very strict, and NETCONF still uses it.
They all mean the same thing
Different spelling, same data. Python can prove it: it loads both texts and compares the results.
import json
import yaml
json_text = '{"name": "R1", "mgmt_ip": "10.99.0.1", "vlans": [10, 20]}'
yaml_text = """
name: R1
mgmt_ip: 10.99.0.1
vlans:
- 10
- 20
"""
a = json.loads(json_text)
b = yaml.safe_load(yaml_text)
print(a)
print(b)
print("Same data:", a == b)Real output, checked in the lab terminal.
{'name': 'R1', 'mgmt_ip': '10.99.0.1', 'vlans': [10, 20]}
{'name': 'R1', 'mgmt_ip': '10.99.0.1', 'vlans': [10, 20]}
Same data: True
Both texts became the same Python dictionary with a list inside. That is the key idea of this module: once the file is loaded, the format no longer matters. What matters is the shape: dictionaries (labelled values) and lists (ordered items), which you already know from Python 2.
Which format where
| Format | Typical home in networking | Strength |
|---|---|---|
| JSON | REST and RESTCONF APIs, controller replies, config files | Simple, universal, strict |
| YAML | Ansible playbooks and variables, CI pipelines, inventories | Easy for humans to edit, allows comments |
| XML | NETCONF, older vendor APIs | Strict, supports schemas and attributes |
What the labs ask you to do
Lab 1 gives you inventory.json that Python cannot read, plus a show_inv.py that prints a table from it. You validate, fix two syntax mistakes, correct a wrong key name in the script and query the file with jq. Lab 2 gives you vlans.yml and build.py: fix an indentation error, add a VLAN and a switch by editing only the data, and convert the plan to JSON with to_json.py.
Your path
Common beginner mistake. Treating the three formats as three different languages to learn. They are three notations for the same two building blocks, dictionaries and lists. Learn the blocks once and the notations follow.
Exam trap. Know which format goes with which technology: JSON with REST and RESTCONF, XML with NETCONF, YAML with Ansible. Also know that YAML uses indentation and JSON uses braces and brackets.
The inventory nobody could read
A company stored its device list as text pasted from an email. Each engineer parsed it with their own scripts and the three scripts disagreed about which devices existed. After moving the list to one JSON file, every script loaded it the same way, and a new device was added in one place. The data became the single source of truth.
Lesson: structured data gives every tool the same view of the network.
"What is the difference between JSON, YAML and XML?"
All three describe structured data. JSON uses braces, brackets and quotes and is the standard for REST APIs. YAML uses indentation, is easy for humans and allows comments, so Ansible uses it. XML uses opening and closing tags, is strict, and is used by NETCONF. Once loaded in Python, JSON and YAML become dictionaries and lists.
Key takeaways
- Structured data is a form a program can read without guessing.
- JSON, YAML and XML are three notations for dictionaries and lists.
- JSON is used by REST APIs, YAML by Ansible, XML by NETCONF.
- Once loaded into Python, the original format no longer matters.
- Two labs: repair an inventory.json and drive config from vlans.yml.
JSON syntax: objects, arrays and six value types
JSON is the most common format you will meet. Almost every REST API answers in JSON. It has only a handful of rules, and they are strict. Learn them once and you can spot every error.
The two containers
JSON has two containers. Everything else is a simple value.
- An object is written with curly braces
{ }. It holds key: value pairs separated by commas. The key is always text in double quotes. It is like a form with labelled boxes, and it becomes a Python dictionary. - An array is written with square brackets
[ ]. It holds values in order, separated by commas. It becomes a Python list.
{
"hostname": "HYD-CORE1",
"vlans": [10, 20, 30]
}
Read it aloud: "An object with two keys. The key hostname has the text HYD-CORE1. The key vlans has an array of three numbers."
[[flow Containers hold values | Object/{ "key": value }/ | Array/[ value, value ] | Value/text, number, true, false, null]]
The value types
A value can be one of exactly six things:
| Type | Example | Python gets |
|---|---|---|
| string | "GigabitEthernet0/1" | str |
| number | 10, 3.5, -1 | int or float |
| boolean | true, false | True, False |
| null | null | None |
| object | {"a": 1} | dict |
| array | [1, 2] | list |
Notice that booleans and null are lowercase, and not in quotes. The text "true" in quotes is a string, not a boolean. Also notice that 10 (a number) and "10" (text) are different: you can do arithmetic with the first.
Nesting: boxes inside boxes
Containers can hold other containers. This is how a whole inventory is described: an object with a key devices whose value is an array, and each item of the array is itself an object.
import json
text = """
{
"company": "Deccan Net",
"site": "Hyderabad",
"devices": [
{"name": "HYD-CORE1", "role": "core", "vlans": [10, 20, 30], "managed": true},
{"name": "HYD-ACC1", "role": "access", "vlans": [10, 20], "managed": false}
]
}
"""
data = json.loads(text)
print(data["site"])
print(data["devices"][0]["name"])
print(data["devices"][1]["vlans"][0])
print(data["devices"][0]["managed"])Real output, checked in the lab terminal.
Hyderabad HYD-CORE1 10 True
Follow each path like directions in a building: data["devices"] is the array, [0] is the first device (counting starts at 0), ["name"] is its name. The last line prints True, not true: Python spells the boolean its own way after loading.
The rules that cause most errors
- Keys and strings use double quotes only. Single quotes are not allowed.
- No trailing comma. A comma goes between items, never after the last one.
- A comma is required between items. Forgetting it is the other common mistake.
- Keys must be strings.
{name: "R1"}is invalid; write{"name": "R1"}. - No comments.
//and#are not part of JSON. - Every opening bracket must close. Count your
{ }and[ ]. - **
true,false,nullare lowercase** and unquoted.TrueandNoneare Python, not JSON.
Python can show you what happens with a broken file. Notice that the message names the line and the column where it got confused:
import json
text = """{
"name": "HYD-EDGE1",
"mgmt_ip": '10.99.0.2'
}"""
data = json.loads(text)Real output, checked in the lab terminal.
Traceback (most recent call last):
File "badjson.py", line 7, in <module>
data = json.loads(text)
json.decoder.JSONDecodeError: Expecting value: line 3 column 14 (char 38)
JSONDecodeError: Expecting value: line 3 column 14 says: at line 3, column 15 Python expected the start of a value but found something else, here the single quote. The place where Python complains is sometimes just after the real mistake, so also look at the line before it.
Pretty and compact
Machines do not care about spaces. JSON can be one long line (compact) or indented (pretty). Both are identical in meaning. Pretty is for people, and tools such as jq and json.dumps(data, indent=2) produce it.
Common mistakes. (1) Copying a Python dictionary into a .json file: it contains single quotes, True and None, all invalid in JSON. (2) A trailing comma after the last item, which many programmers add out of habit. (3) Writing numbers with leading zeros such as 010. (4) Putting IP addresses without quotes: 10.99.0.1 is not a valid number, it must be the string "10.99.0.1".
Exam trap. Which of these is valid JSON? The answer always has double quotes, lowercase true, no trailing comma and no comments. Also remember that JSON object keys are unordered in meaning, while array items are ordered.
The controller reply that "was not JSON"
A script that called a controller crashed with a decode error. The engineer opened the saved reply and saw a Python-style dictionary with True and single quotes. Someone had printed the Python object and saved the printout as the test file. The real controller reply was valid JSON. Replacing the saved sample with a proper json.dump of the data fixed the unit tests.
Lesson: printing a Python object is not the same as writing JSON; use json.dump to write it.
"What are the data types in JSON?"
There are six: string, number, boolean, null, object and array. Objects are key-value pairs in braces with double-quoted keys, and arrays are ordered lists in brackets. Booleans and null are lowercase. JSON has no comments, no single quotes and no trailing commas.
Key takeaways
- JSON has two containers: objects
{ }and arrays[ ]. - Six value types: string, number, true/false, null, object, array.
- Strings and keys use double quotes only; commas go between items, never after the last.
- JSON has no comments;
true,falseandnullare lowercase. - Nested paths read like directions:
data["devices"][0]["name"].
JSON in Python: load, dump and the four functions
Python ships with a module called json for reading and writing JSON. It has four main functions. They look alike, so this chapter gives you a simple memory trick.
The "s" means string
json.loads(text): load from a string. Text in, Python data out.json.load(file): load from a file object. File in, Python data out.json.dumps(data): dump to a string. Python data in, text out.json.dump(data, file): dump to a file object. Python data in, file written.
The trick: the function with an s works with a string. The one without works with a file. And load means "into Python", dump means "out of Python".
Four functions
Loading text into data
import json
text = '{"name": "HYD-CORE1", "mgmt_ip": "10.99.0.1", "vlans": [10, 20, 30], "managed": true, "owner": null}'
device = json.loads(text)
print(type(device))
print(device["name"])
print(device["vlans"][1])
print(device["managed"], type(device["managed"]))
print(device["owner"])Real output, checked in the lab terminal.
<class 'dict'> HYD-CORE1 20 True <class 'bool'> None
The text became a dictionary. Look at the last two lines: true became True, and null became None. After loading, you use ordinary Python: keys, indexes, loops.
Loading from a file
In lab 1 the data lives in inventory.json and show_inv.py reads it with json.load. Here is a self-contained version: it first writes a small file, then reads it back.
import json
with open("inventory.json", "w") as f:
f.write('{"site": "Hyderabad", "devices": [{"name": "HYD-CORE1", "vlans": [10, 20, 30]}, {"name": "HYD-ACC1", "vlans": [10, 20]}]}')
with open("inventory.json") as f:
inv = json.load(f)
print("Site:", inv["site"])
for dev in inv["devices"]:
print(dev["name"], dev["vlans"])Real output, checked in the lab terminal.
Site: Hyderabad HYD-CORE1 [10, 20, 30] HYD-ACC1 [10, 20]
Changing data and writing it back
Because the data is now a dictionary, you edit it like any dictionary. Then json.dump writes it out. The indent=2 argument makes the file pretty. Without it everything lands on one line.
import json
inv = {"site": "Hyderabad", "devices": [{"name": "HYD-CORE1", "vlans": [10, 20]}]}
inv["devices"][0]["vlans"].append(30)
inv["devices"].append({"name": "HYD-ACC2", "vlans": [10]})
with open("out.json", "w") as f:
json.dump(inv, f, indent=2)
with open("out.json") as f:
print(f.read())Real output, checked in the lab terminal.
{
"site": "Hyderabad",
"devices": [
{
"name": "HYD-CORE1",
"vlans": [
10,
20,
30
]
},
{
"name": "HYD-ACC2",
"vlans": [
10
]
}
]
}
Notice that Python's True and None would be written back as true and null, so a program can write valid JSON without you thinking about it. This is why you write JSON with json.dump and never by pasting Python's printout.
dumps: JSON as text
json.dumps is handy to show data or to send it in an API request. Useful options: indent=2 (pretty) and sort_keys=True (alphabetical keys, good for comparing two files).
import json
data = {"name": "R1", "up": True, "owner": None, "vlans": [10, 20]}
print(json.dumps(data))
print(json.dumps(data, indent=2, sort_keys=True))
print(type(json.dumps(data)))Real output, checked in the lab terminal.
{"name": "R1", "up": true, "owner": null, "vlans": [10, 20]}
{
"name": "R1",
"owner": null,
"up": true,
"vlans": [
10,
20
]
}
<class 'str'>
type shows str: dumps gives you text, not a dictionary. A very common beginner confusion is to call dumps and then try to use the result as a dictionary.
Walking through a loaded inventory
Loops and dictionary reads from the previous module work unchanged:
import json
inv = json.loads('{"devices": [{"name": "A", "managed": true}, {"name": "B", "managed": false}, {"name": "C", "managed": true}]}')
managed = [d["name"] for d in inv["devices"] if d["managed"]]
print("Managed devices:", managed)
print("Count:", len(inv["devices"]))Real output, checked in the lab terminal.
Managed devices: ['A', 'C'] Count: 3
Common mistakes. (1) Using json.load on a string or json.loads on a file object: you get a TypeError or AttributeError. (2) Forgetting to open the file with with open(...). (3) Calling json.dumps(data) and expecting a file to be written. (4) Reading a missing key and getting KeyError; use data.get("key") when the key may be absent.
Exam trap. Which function turns a string into a Python dictionary? loads. Which writes a dictionary to a file? dump. The one with the s is for strings.
One byte too many
A script saved device data with str(data) instead of json.dump. The saved file looked fine, but the next script crashed trying to load it, because str wrote True and single quotes. The fix was a one-line change to json.dump(data, f, indent=2). From then on both scripts read and wrote the same valid JSON.
Lesson: always write JSON with the json module, never with str() or print().
"What is the difference between json.load and json.loads?"
load reads JSON from a file object, loads from a string. The same pairing exists for dump and dumps. After loading, JSON objects become dictionaries and arrays become lists, with true, false and null becoming True, False and None.
Key takeaways
loadsanddumpswork with strings;loadanddumpwork with files.- load means into Python, dump means out of Python.
- JSON objects become dicts, arrays become lists, null becomes None.
json.dump(data, f, indent=2)writes a readable file.- Never write JSON by printing a Python object; use the json module.
Repairing broken JSON: validate, fix, query with jq
This chapter follows lab 1 step by step. The story: Deccan Net keeps its Hyderabad device list in inventory.json, and show_inv.py prints it as a table for the NOC. Someone edited the file by hand and now nothing can read it. You will find the mistakes with a validator, fix them, repair a wrong key name in the script, and then ask the file questions with jq.
The file and the symptom
The file describes three devices. These are the lines that matter in the broken version (line numbers on the left come from the lab command nl):
13 "name": "HYD-EDGE1",
14 "role": "edge",
15 "mgmt_ip": '10.99.0.2',
16 "vlans": [99],
...
23 "vlans": [10, 20],
24 "managed": false,
25 }
Running the script gives the Python view of the problem:
AUTO> python3 show_inv.py
Traceback (most recent call last):
File "show_inv.py", line 6, in <module>
inv = json.load(f)
json.decoder.JSONDecodeError: Expecting value: line 15 column 18 (char 282)
The last line is exactly what you learned to read: the class is JSONDecodeError, and Python points to line 15, column 18.
Use a validator
A validator checks a file against the rules and reports the first problem. The lab has a command json validate. On a real computer you can use python3 -m json.tool inventory.json (it prints the file if valid, otherwise an error with the line) or an editor with JSON support.
AUTO> json validate inventory.json
inventory.json: INVALID JSON
line 15, column 18: strings must use double quotes ("), not single quotes (')
15 | "mgmt_ip": '10.99.0.2',
| ^
The caret points at the single quote. Fix line 15 so the value uses double quotes, keeping the same indentation (6 spaces). In the lab you use edit inventory.json 15 ... to replace a line. Then validate again.
AUTO> json validate inventory.json
inventory.json: INVALID JSON
line 25, column 5: trailing comma is not allowed before '}'
25 | }
| ^
The first error is gone; the validator now finds the next one. A tool reports one mistake at a time, so you repeat until the file is clean. The message says the comma is "before }", meaning the comma is on the line above, line 24 ("managed": false,). Remove it.
AUTO> json validate inventory.json inventory.json: valid JSON (object with 3 top-level key(s): company, site, devices)
Valid JSON now. Always keep all three devices: the lab check expects the third one (HYD-ACC1, with "managed": false) to still be present.
The repair loop
Valid JSON is not the same as working code
Run the script again. JSON now loads, but a different error appears, this time in the script:
AUTO> python3 show_inv.py
Company: Deccan Net | Site: Hyderabad
NAME ROLE MGMT IP VLANS
Traceback (most recent call last):
File "show_inv.py", line 12, in <module>
print(f"{dev['name']:<12}{dev['role']:<8}{dev['mgmt']:<14}{dev['vlans']}")
KeyError: 'mgmt'
The data is correct. The script asks for a key mgmt but the file calls it mgmt_ip. KeyError: 'mgmt' is the dictionary telling you it has no such key. The cure is to make the script use the real key name. In the lab: sed -i "s/'mgmt'/'mgmt_ip'/" show_inv.py, a command that replaces text in a file.
AUTO> python3 show_inv.py Company: Deccan Net | Site: Hyderabad NAME ROLE MGMT IP VLANS HYD-CORE1 core 10.99.0.1 [10, 20, 30] HYD-EDGE1 edge 10.99.0.2 [99] HYD-ACC1 access 10.99.0.11 [10, 20]
Ask the file questions with jq
jq is a command-line tool for reading JSON. You give it a path and it prints the value. The dot . means "the whole document", and you follow it with keys and indexes, just like Python.
AUTO> jq '.devices[0].vlans[1]' inventory.json 20
Read the path: .devices is the array, [0] is the first device (HYD-CORE1), .vlans is its list [10, 20, 30], and [1] is the second item because counting starts at 0. The answer is 20.
Common patterns you can use on the same file: jq '.site' inventory.json prints "Hyderabad", jq '.devices[].name' inventory.json prints the name of every device, and jq '.devices | length' inventory.json prints 3.
The same job in Python needs one more line, and gives you the same answer:
import json
inv = json.loads('{"devices": [{"name": "HYD-CORE1", "vlans": [10, 20, 30]}, {"name": "HYD-EDGE1", "vlans": [99]}]}')
print(inv["devices"][0]["vlans"][1])
print([d["name"] for d in inv["devices"]])
print(len(inv["devices"]))Real output, checked in the lab terminal.
20 ['HYD-CORE1', 'HYD-EDGE1'] 2
A checklist for a file that will not load
- Run the validator and read the line and column.
- Look at the reported line and the line above.
- Check quotes, then commas (missing between items, extra after the last), then brackets.
- Fix one thing, validate again.
- When the file loads, check names: does the key in the code match the key in the data?
Common mistakes. (1) Fixing the wrong line because the error is reported after the real mistake, for example a missing comma at the end of the previous line. (2) Deleting a whole device to make the file valid, which silently loses data. (3) Assuming valid JSON means correct data. (4) Editing the file in a word processor that turns straight quotes into curly quotes, which JSON rejects.
Exam trap. A JSON file with a trailing comma is invalid, even though Python lists allow it. The same is true for single quotes and comments.
The curly quotes
An engineer pasted a snippet from a chat window into a JSON file. The chat app had changed " into curly quotes. The validator reported an error at the first quote of the first key, which looked perfectly fine on screen. Retyping the quotes in a code editor fixed it, and the team added a validation step before every commit.
Lesson: look at the exact character the validator points to, not only at what looks right.
"A teammate's JSON file fails to load. How do you debug it?"
I run a validator or python3 -m json.tool to get the line and column, check the line and the one above it for single quotes, missing or trailing commas and unclosed brackets, fix one error at a time, and re-run. Once it loads I verify that the keys my code uses match the keys in the data, using jq to inspect the structure.
Key takeaways
- A validator reports the first mistake; repeat until the file is valid.
- The reported line may be just after the real mistake; also check the line above.
- Valid JSON can still be wrong data or mismatch the code's key names.
KeyErrormeans the code asked for a key that the data does not have.jq '.devices[0].vlans[1]'follows keys and indexes; index counting starts at 0.
YAML syntax: indentation is the structure
YAML was designed to be easy for people to write. It drops braces and quotes and uses indentation and a few symbols instead. Ansible, many CI tools and most variable files use it. The price of the clean look: you must be exact with spaces.
The three building blocks
- Mapping:
key: value. A colon and a space. This is a dictionary. - Sequence: items that start with
-(dash and space), one per line. This is a list. - Scalar: a plain value: text, number, true/false, null.
import yaml text = """ name: HYD-CORE1 mgmt_ip: 10.99.0.1 managed: true vlans: - 10 - 20 - 30 """ data = yaml.safe_load(text) print(data)
Real output, checked in the lab terminal.
{'name': 'HYD-CORE1', 'mgmt_ip': '10.99.0.1', 'managed': True, 'vlans': [10, 20, 30]}
Compare with JSON from the previous chapters: no braces, no quotes, no commas. The two keys with simple values sit on one line each. The key vlans has nothing after the colon, so its value is the indented sequence below it. Python received a dictionary with a list inside, exactly as for JSON.
YAML building blocks
Indentation rules
- Indentation means belonging. Lines indented under a key belong to that key.
- Use spaces only, never tabs. Two spaces per level is the convention.
- Items at the same level must line up exactly. One extra space makes them a different level, and then YAML breaks.
A list of mappings
The most common structure in network files is a list where each item has several fields. Each item starts with a dash. The fields of that item line up under the first field.
import yaml
text = """
switches:
- name: WH-SW1
mgmt: 10.99.0.11
- name: WH-SW2
mgmt: 10.99.0.12
"""
data = yaml.safe_load(text)
for sw in data["switches"]:
print(sw["name"], sw["mgmt"])
print(data["switches"][1]["mgmt"])Real output, checked in the lab terminal.
WH-SW1 10.99.0.11 WH-SW2 10.99.0.12 10.99.0.12
Read the shape: the key switches holds a list. The first item begins with - name: WH-SW1, and mgmt: is lined up exactly under name: (4 spaces), so it belongs to the same item. The second dash starts a new item.
Comments and extra tools
A # starts a comment, which JSON cannot do. This makes YAML ideal for files that people maintain.
import yaml
text = """
# VLAN plan for the warehouse
site: Nagpur-WH # the site name
vlans: [10, 20, 30] # a short list can be written like JSON
ids: {office: 10, cctv: 30}
note: "text with: a colon needs quotes"
description: |
Line one
Line two
"""
data = yaml.safe_load(text)
for key, value in data.items():
print(key, "->", repr(value))Real output, checked in the lab terminal.
site -> 'Nagpur-WH'
vlans -> [10, 20, 30]
ids -> {'office': 10, 'cctv': 30}
note -> 'text with: a colon needs quotes'
description -> 'Line one\nLine two\n'
What you see: short lists and mappings may use JSON-like brackets (called flow style); a value that contains a colon followed by a space needs quotes; and | starts a block of text that keeps its line breaks.
YAML guesses types
Without quotes, YAML decides the type for you. That is convenient, and sometimes surprising.
import yaml
text = """
vlan: 10
version: 1.10
enabled: true
speed: 1G
code: 0123
answer: no
empty:
"""
data = yaml.safe_load(text)
for key, value in data.items():
print(key, repr(value), type(value).__name__)Reference code (not run in the simulator). Output below is from standard Python 3.
vlan 10 int version 1.1 float enabled True bool speed '1G' str code 83 int answer False bool empty None NoneType
Look at three traps. 1.10 became the number 1.1, so a version string lost its zero. 0123 was read as the number 83, because a leading zero meant octal in older YAML. And no became False, the so-called "Norway problem": a country code NO, or the word on or yes, turns into a boolean. The cure for all of them is the same: put quotes around values that must stay text (version: "1.10", country: "NO").
Syntax errors in YAML
A wrong indentation produces an error that names the line and what it expected. The lab tool yamllint reports it like this:
AUTO> yamllint vlans.yml vlans.yml:11:4: error: bad indentation of a sequence entry (this line is indented more than the list item above it)
It says line 11, column 4: the dash is one space too far to the right compared with the other items. On a real computer the Python yaml module and the yamllint program give similar messages. Tabs trigger their own message.
Common mistakes. (1) Using a tab instead of spaces. (2) Misaligning one item by one space. (3) Forgetting the space after the colon (key:value is a plain string, not a mapping). (4) Leaving unquoted values that YAML converts (no, on, 1.10, 010). (5) Copying JSON-style quotes unnecessarily; they are allowed but not needed.
Exam trap. YAML allows comments with #; JSON does not. YAML structure comes from indentation with spaces; tabs are forbidden. A dash starts a list item and key: value a mapping.
The country that became False
An inventory file had a field country: NO for a branch in Norway. After loading it, a report said the country was False. YAML 1.1 loaders read NO as the boolean false. The team quoted the value as "NO" and added a test that checks the type of every field after loading.
Lesson: when a value must stay text, quote it.
"How does YAML represent a list of dictionaries?"
Each list item starts with a dash and a space, and the fields of that item are written as key: value lines aligned under the first key. The indentation shows which keys belong to which item. For example, - name: SW1 then mgmt: 10.0.0.1 under it.
Key takeaways
- YAML has mappings (
key: value), sequences (- item) and scalars. - Indentation with spaces is the structure; tabs are not allowed.
#starts a comment; quotes protect values from being converted.no,yes,on,1.10and0123are guessed as other types unless quoted.- Both JSON and YAML load into the same Python dictionaries and lists.
YAML in Python: data in one file, code in another
This chapter follows lab 2. The story: Kaveri Foods describes its warehouse switches and VLANs in vlans.yml, and build.py turns that data into config for every switch. The big idea is to keep the code and the data apart. When the network grows, you edit only the YAML file and never touch the script.
Analogy: the recipe and the shopping list
A recipe (the code) stays the same all year. The shopping list (the data) changes every week. If every shopping change meant rewriting the recipe, cooking would be a nightmare. Network automation is the same: the build script is the recipe, and the YAML file is the list of switches, VLANs and addresses. The file with the data is often called the source of truth.
Loading YAML
You need the yaml module (the PyYAML package, already available in the lab). Two functions matter:
yaml.safe_load(text_or_file): YAML in, Python dictionaries and lists out.yaml.safe_dump(data): Python data in, YAML text out.
Use safe_load, not the plain load. The safe version builds only ordinary data (text, numbers, lists, dictionaries). The unsafe version can build arbitrary Python objects from a file, which is a security risk when the file comes from somewhere you do not control. Defensive habit: always safe_load.
The data file
This is the repaired vlans.yml from the lab. It has one site, two switches and three VLANs.
The build script
Read it slowly. The outer loop visits every switch; the inner loop visits every VLAN, so each switch receives all VLANs. The f-strings drop values into the config text.
File vlans.yml
# vlans.yml - VLAN plan for the Kaveri Foods warehouse (only data, no code)
site: Nagpur-WH
switches:
- name: WH-SW1
mgmt: 10.99.0.11
- name: WH-SW2
mgmt: 10.99.0.12
vlans:
- id: 10
name: OFFICE
- id: 20
name: SCANNERS
- id: 30
name: CCTVimport yaml
with open("vlans.yml") as f:
plan = yaml.safe_load(f)
print("! site:", plan["site"])
for sw in plan["switches"]:
print(f"! ---- {sw['name']} ({sw['mgmt']}) ----")
for v in plan["vlans"]:
print(f"vlan {v['id']}")
print(f" name {v['name']}")Real output, checked in the lab terminal.
! site: Nagpur-WH ! ---- WH-SW1 (10.99.0.11) ---- vlan 10 name OFFICE vlan 20 name SCANNERS vlan 30 name CCTV ! ---- WH-SW2 (10.99.0.12) ---- vlan 10 name OFFICE vlan 20 name SCANNERS vlan 30 name CCTV
Count what happened: 2 switches times 3 VLANs gave 6 vlan blocks. Every switch got the same VLAN list because the list lives once in the data.
Data and code are separate
Growing the network by editing only the data
Management asks for VLAN 40 (GUEST) and a third switch, WH-SW3 at 10.99.0.13. In the lab you add four lines to vlans.yml and run build.py again. Here is the extended data, and a smaller report script that counts what each switch now receives.
File vlans.yml
site: Nagpur-WH
switches:
- name: WH-SW1
mgmt: 10.99.0.11
- name: WH-SW2
mgmt: 10.99.0.12
- name: WH-SW3
mgmt: 10.99.0.13
vlans:
- id: 10
name: OFFICE
- id: 20
name: SCANNERS
- id: 30
name: CCTV
- id: 40
name: GUESTimport yaml
with open("vlans.yml") as f:
plan = yaml.safe_load(f)
names = [v["name"] for v in plan["vlans"]]
for sw in plan["switches"]:
print(sw["name"], "gets", len(names), "VLANs:", ", ".join(names))Real output, checked in the lab terminal.
WH-SW1 gets 4 VLANs: OFFICE, SCANNERS, CCTV, GUEST WH-SW2 gets 4 VLANs: OFFICE, SCANNERS, CCTV, GUEST WH-SW3 gets 4 VLANs: OFFICE, SCANNERS, CCTV, GUEST
The script did not change by a single character, yet the output grew from 2 switches and 3 VLANs to 3 switches and 4 VLANs. That is the point of data-driven automation: the change is a data edit that a person can review in a pull request.
YAML to JSON: same data, different notation
The monitoring team wants the plan as JSON. Loading with yaml and dumping with json is a two-line conversion. The lab version writes vlans.json with json.dump(plan, f, indent=2). Here is the same idea shown in memory:
import json
import yaml
plan = yaml.safe_load("""
site: Nagpur-WH
vlans:
- id: 10
name: OFFICE
- id: 20
name: SCANNERS
""")
print(json.dumps(plan, indent=2))
print(yaml.safe_dump(plan, default_flow_style=False))Real output, checked in the lab terminal.
{
"site": "Nagpur-WH",
"vlans": [
{
"id": 10,
"name": "OFFICE"
},
{
"id": 20,
"name": "SCANNERS"
}
]
}
site: Nagpur-WH
vlans:
- id: 10
name: OFFICE
- id: 20
name: SCANNERS
Both directions work because both formats describe the same dictionaries and lists. In the lab you then ask the file a question with jq:
AUTO> json validate vlans.json vlans.json: valid JSON (object with 3 top-level key(s): site, switches, vlans) AUTO> jq '.vlans[1].name' vlans.json "SCANNERS"
Index 1 is the second VLAN, so the answer is SCANNERS (with quotes, because it is a JSON string).
If the YAML is broken
When build.py met the original misaligned file, it did not produce config at all. The lab traceback ends like this:
AUTO> python3 build.py
Traceback (most recent call last):
File "build.py", line 6, in <module>
plan = yaml.safe_load(f)
yaml.YAMLError: bad indentation of a sequence entry (this line is indented more than the list item above it)
Handle it like any traceback: read the last line, then run yamllint vlans.yml to find line 11 and fix it. Fixing means giving the dash the same two-space indentation as the other items.
Common mistakes. (1) Using yaml.load instead of yaml.safe_load. (2) Hard-coding switch names in the script, which defeats the purpose of the data file. (3) Adding a VLAN with the wrong indentation so it silently attaches to the previous item. (4) Forgetting that YAML needs both id and name for every VLAN item, so a missing key later causes KeyError.
Exam trap. The safe loader is yaml.safe_load. Converting YAML to JSON means load with yaml, dump with json; there is no direct "convert" function. Also: Ansible variable files and inventories are YAML.
Forty switches, one pull request
A campus team generated VLAN config for forty switches from one YAML file and a 30-line script. When the security team requested a new guest VLAN, the change was four lines of YAML in a pull request. Reviewers read exactly what would change, approved it, and the script generated forty configs in seconds. Nobody edited forty files by hand.
Lesson: data-driven automation makes a change small, reviewable and repeatable.
"Why separate data from code in automation?"
Because the logic changes rarely and the data changes often. A YAML file of devices and VLANs can be edited and reviewed by anyone, version-controlled and reused by several tools, while the script stays untouched. It also acts as a source of truth for what the network should look like.
Key takeaways
- Keep data (YAML) and logic (Python) in separate files.
- Always use
yaml.safe_load;yaml.safe_dumpwrites YAML back. - Growing the network means editing the data, not the script.
- YAML to JSON: load with
yaml, write withjson.dump. - A broken YAML file stops
yaml.safe_load; useyamllintto find the line.
XML: tags, attributes and reading it in Python
XML is the oldest of the three formats and the wordiest. You will still meet it, mainly in NETCONF, the protocol that older and many current network devices use for model-driven configuration. You rarely write XML by hand in this track, but you must be able to read it and pick values out of it.
Anatomy of an XML document
- An element is a pair of tags:
<name>R1</name>. The first is the opening tag, the last the closing tag, and the text in between is the value. - Elements nest. An element that contains others is a parent; the ones inside are children. The single outermost element is the root.
- An attribute is extra information inside the opening tag:
<interface type="ethernet">. The value is always in quotes. - An element with no content can close itself:
<shutdown/>.
<device>
<name>R1</name>
<interfaces>
<interface type="ethernet">
<name>GigabitEthernet0/1</name>
<ip>10.10.1.1</ip>
<enabled>true</enabled>
</interface>
<interface type="loopback">
<name>Loopback0</name>
<ip>1.1.1.1</ip>
<enabled>true</enabled>
</interface>
</interfaces>
</device>
Read it like a tree: device is the root, it has one child interfaces, which has two interface children, and each of those has name, ip and enabled. The attribute type describes the interface itself.
XML is a tree
The strict rules
- Every opening tag needs a matching closing tag (or a self-closing tag).
- Tags must be properly nested: close the inner tag before the outer one.
- There is exactly one root element.
- Attribute values are in quotes.
- Tag names are case-sensitive:
<Name>and<name>differ. - The characters
<,>and&inside text must be written as<,>and&.
Reading XML with Python
The standard library module xml.etree.ElementTree loads XML into a tree you can search. This is reference code (not run in the simulator), but it runs on any normal Python 3 install, and the output below is from a real Python 3 run.
import xml.etree.ElementTree as ET
text = """
<device>
<name>R1</name>
<interfaces>
<interface type="ethernet"><name>GigabitEthernet0/1</name><ip>10.10.1.1</ip></interface>
<interface type="loopback"><name>Loopback0</name><ip>1.1.1.1</ip></interface>
</interfaces>
</device>
"""
root = ET.fromstring(text)
print(root.tag)
print(root.find("name").text)
for intf in root.findall("./interfaces/interface"):
print(intf.attrib["type"], intf.find("name").text, intf.find("ip").text)Reference code (not run in the simulator). Output below is from standard Python 3.
device R1 ethernet GigabitEthernet0/1 10.10.1.1 loopback Loopback0 1.1.1.1
Line by line: ET.fromstring(text) parses the text and returns the root element. root.tag is its tag name. root.find("name") finds the first child called name, and .text gives its value. findall("./interfaces/interface") is a small path expression (a limited form of XPath): start at the root, go into interfaces, then collect every interface. .attrib is a dictionary of the attributes.
Everything you read is text. If you need a number, convert it yourself, for example int(vlan.text).
Writing XML
You can build a tree with code and turn it into text:
import xml.etree.ElementTree as ET
root = ET.Element("interface", {"type": "ethernet"})
ET.SubElement(root, "name").text = "GigabitEthernet0/2"
ET.SubElement(root, "ip").text = "10.10.2.1"
print(ET.tostring(root, encoding="unicode"))Reference code (not run in the simulator). Output below is from standard Python 3.
<interface type="ethernet"><name>GigabitEthernet0/2</name><ip>10.10.2.1</ip></interface>
Namespaces in one minute
Large XML systems mix vocabularies, so tags may carry a namespace, a label that says which vocabulary a tag belongs to. In NETCONF replies you often see an attribute such as xmlns="urn:example:interfaces" on the root. In Python the namespace is glued in front of the tag name inside curly brackets.
import xml.etree.ElementTree as ET
text = '<interfaces xmlns="urn:example:interfaces"><interface><name>Gi0/1</name></interface></interfaces>'
root = ET.fromstring(text)
print(root.tag)
ns = {"i": "urn:example:interfaces"}
print(root.find("i:interface/i:name", ns).text)Reference code (not run in the simulator). Output below is from standard Python 3.
{urn:example:interfaces}interfaces
Gi0/1
If find returns None on a document that clearly has the element, a namespace is the usual reason. You then get AttributeError: 'NoneType' object has no attribute 'text', the same shape of error you saw with missing return values.
Broken XML
A wrong closing tag gives a ParseError that names the position. Catch it and print the message:
import xml.etree.ElementTree as ET
bad = "<device><name>R1</nme></device>"
try:
ET.fromstring(bad)
except ET.ParseError as e:
print("XML error:", e)Reference code (not run in the simulator). Output below is from standard Python 3.
XML error: mismatched tag: line 1, column 18
Where you will meet XML: NETCONF
A NETCONF request is an XML message. This one asks a device for its running configuration. It is shown only to help you recognise the shape; the NETCONF module of this course goes deeper.
<rpc message-id="101">
<get-config>
<source><running/></source>
</get-config>
</rpc>
JSON, YAML or XML? A quick comparison
| Question | JSON | YAML | XML |
|---|---|---|---|
| Comments | no | yes | yes (between comment markers) |
| Structure shown by | braces and brackets | indentation | opening and closing tags |
| Attributes | no | no | yes |
| Typical use | REST, RESTCONF | Ansible, variables | NETCONF |
| Python module | json | yaml (PyYAML) | xml.etree.ElementTree |
Common mistakes. (1) A mismatched or misspelled closing tag, which is case-sensitive. (2) Forgetting that every value read from XML is text. (3) Searching without the namespace and getting None. (4) Parsing untrusted XML from the internet with unsafe settings; in production use a hardened parser for external input.
Exam trap. XML elements have opening and closing tags; attributes live inside the opening tag. NETCONF uses XML; RESTCONF can use XML or JSON. XML can carry attributes, JSON cannot.
The None that was a namespace
An engineer parsed a NETCONF reply and root.find("interface") returned None even though the text clearly contained interface elements. The reply declared a namespace on the root, so the real tag was {urn:example:interfaces}interface. After passing the namespace mapping to find, the script read all interfaces.
Lesson: when an element that is obviously there cannot be found in XML, suspect a namespace.
"How would you read values from an XML reply in Python?"
I parse it with xml.etree.ElementTree.fromstring, then use find and findall with a path such as ./interfaces/interface and read .text for values and .attrib for attributes. I convert text to numbers myself, and if the document has a namespace I pass a namespace dictionary to find.
Key takeaways
- XML is a tree of elements with opening and closing tags; attributes sit in the opening tag.
- Tags must be properly nested, with one root, and are case-sensitive.
ElementTree.fromstring,find,findall,.textand.attribread it.- All values from XML are text; convert numbers yourself.
- A namespace is the usual reason
findreturns None; NETCONF uses XML.
Troubleshooting workflow: a data file that will not load
When automation fails, a surprising share of the time the cause is not the network and not the script logic. It is a data file with a small syntax mistake, or a data file that is valid but does not match what the code expects. This chapter gives you one workflow that works for JSON, YAML and XML, and uses the two labs as worked examples.
The five-step workflow
- Read the last line of the traceback. Is it a *parse* error (
JSONDecodeError,YAMLError,ParseError) or a *use* error (KeyError,TypeError)? - Parse errors: run a validator and read the line and column. Fix one thing, validate again.
- Use errors: the file loaded. Print what you loaded and compare the real keys with the keys your code asks for.
- Check the types: a number that became text, a version that lost a zero, a
nothat becameFalse. - Re-run the whole script and check the output, not only the validator.
Two kinds of failure
Step 1: parse error or use error?
The difference decides where to look. A parse error means the file is broken. A use error means the file is fine and the code is wrong. In lab 1 you met both in one ticket:
| Where it fails | Last line of the traceback | Meaning |
|---|---|---|
json.load(f) | JSONDecodeError: Expecting value: line 15 column 18 | The file is not valid JSON |
dev['mgmt'] | KeyError: 'mgmt' | The file is valid; the key is spelled differently |
yaml.safe_load(f) | YAMLError: bad indentation of a sequence entry | The file is not valid YAML |
Step 2: build a small checker
You can wrap the loaders so that every error becomes a readable line with a position. This is a handy script to keep. Parse errors in the JSON module carry the position as lineno and colno; YAML errors carry a problem_mark that counts lines from zero.
import json
import yaml
def check_json(text):
try:
json.loads(text)
return "JSON OK"
except json.JSONDecodeError as e:
return f"JSON error at line {e.lineno}, column {e.colno}: {e.msg}"
def check_yaml(text):
try:
yaml.safe_load(text)
return "YAML OK"
except yaml.YAMLError as e:
mark = getattr(e, "problem_mark", None)
where = f"line {mark.line + 1}, column {mark.column + 1}" if mark else "unknown position"
return f"YAML error at {where}"
print(check_json('{"name": "R1", "vlans": [10, 20]}'))
print(check_json('{"name": "R1" "vlans": [10, 20]}'))
print(check_json("{'name': 'R1'}"))
print(check_yaml("vlans:\n - 10\n - 20\n"))
print(check_yaml("vlans:\n - id: 10\n name: OFFICE\n - id: 20\n"))Reference code (not run in the simulator). Output below is from standard Python 3.
JSON OK JSON error at line 1, column 15: Expecting ',' delimiter JSON error at line 1, column 2: Expecting property name enclosed in double quotes YAML OK YAML error at line 4, column 4
The checker names the line and the column for each broken input. Notice that the missing comma between the two keys and the single-quoted keys are reported with a position, exactly like the lab validator.
Step 3: valid, but wrong shape
When the file loads and the code still fails, print the real structure. A pretty dump is the fastest way to see it:
import json
data = json.loads('{"devices": [{"name": "HYD-CORE1", "mgmt_ip": "10.99.0.1"}]}')
print(json.dumps(data, indent=2))
dev = data["devices"][0]
print(list(dev.keys()))
print(dev.get("mgmt"))
print(dev.get("mgmt_ip"))Real output, checked in the lab terminal.
{
"devices": [
{
"name": "HYD-CORE1",
"mgmt_ip": "10.99.0.1"
}
]
}
['name', 'mgmt_ip']
None
10.99.0.1
list(dev.keys()) shows the real key names, and dev.get("mgmt") returns None instead of crashing. Compare the names with the ones your code uses. This is exactly how the lab's mgmt versus mgmt_ip mismatch is found.
Step 4: check the types after loading
If a value looks right in the file but behaves wrongly in the script, print its type.
import json
import yaml
j = json.loads('{"vlan": "10", "enabled": "true"}')
y = yaml.safe_load('vlan: 10\nenabled: true\n')
print(j["vlan"] + j["vlan"])
print(y["vlan"] + y["vlan"])
print(type(j["enabled"]).__name__, type(y["enabled"]).__name__)Real output, checked in the lab terminal.
1010 20 str bool
In JSON the quoted "10" is text, so adding it to itself glues the text together (1010). In YAML the unquoted 10 is a number, so the sum is 20. And "true" in quotes is a string, while the unquoted true is a real boolean. When the numbers are off, look for quotes.
Lab 1 and lab 2 as one checklist
Lab 1 (JSON): json validate showed a single-quote error on line 15, then a trailing comma before } on line 25 (the comma on line 24). After both were fixed the script still failed with KeyError: 'mgmt' because the key in the file is mgmt_ip. The second VLAN of the first device, jq '.devices[0].vlans[1]' inventory.json, is 20.
Lab 2 (YAML): yamllint vlans.yml reported line 11, column 4: bad indentation of a sequence entry. After aligning the dash with the other items, build.py printed VLANs 10, 20 and 30 for both switches. Adding - id: 40 / name: GUEST and a third switch changed only the data. to_json.py then wrote vlans.json, and jq '.vlans[1].name' vlans.json returned "SCANNERS".
Converting safely
Conversion between formats is load in one module, dump in another. A good check is a round trip: dump to JSON, load it again, and compare with the original.
import json
import yaml
plan = yaml.safe_load("site: Nagpur-WH\nvlans:\n - id: 10\n name: OFFICE\n")
text = json.dumps(plan)
again = json.loads(text)
print(text)
print("Round trip identical:", plan == again)Real output, checked in the lab terminal.
{"site": "Nagpur-WH", "vlans": [{"id": 10, "name": "OFFICE"}]}
Round trip identical: True
Differences you may notice: JSON has no comments, so YAML comments are lost; and a YAML value that was converted (no to False) is already converted before it reaches JSON.
Tools to keep in your pocket
| Tool | Use |
|---|---|
python3 -m json.tool file.json | Validate and pretty-print JSON |
jq '.path' file.json | Pull a value out of JSON |
yamllint file.yml | Check YAML syntax and style |
print(json.dumps(data, indent=2)) | See the real structure inside Python |
type(x) | See what type a value really has |
Common mistakes. (1) Fixing the validator's complaint and assuming the job is done, without re-running the script. (2) Editing the data to match the code when the code should change (or the other way around) without agreeing on which side is correct. (3) Ignoring the line above the reported line. (4) Trusting a value's appearance in the file rather than checking its type after loading.
Exam trap. A parse error and a KeyError are different problems. Parse errors point to the file; KeyError points to a mismatch between data and code. The exam may also ask which format forbids comments (JSON) and which forbids tabs (YAML).
Two errors, one ticket
A monitoring team reported that the device table script "broke after a file edit". The first error was a JSON decode error caused by a trailing comma. The engineer fixed it and closed the ticket. An hour later the script failed again with a KeyError, because the same edit also renamed a key. Only a full re-run after the first fix would have caught it. The team added a step: validate, run, and compare the output to a known-good sample.
Lesson: after every fix, run the whole script again; the first error often hides the second.
"A data file that worked yesterday now breaks your script. What do you do?"
I read the last line of the traceback to decide whether it is a parse error or a use error. For parse errors I run a validator to get the line and column and check the line above as well. For use errors I print the loaded structure and compare key names and types. Then I re-run the script end to end, because fixing one error can reveal another, and I add a validation step to the pipeline.
Key takeaways
- Decide first: parse error (the file) or use error (the code or the key names).
- Validators report one mistake at a time; fix and repeat, and check the line above.
- Print the structure with
json.dumps(data, indent=2)and checktype(x)when values misbehave. - Quotes change types in JSON and YAML; unquoted
noand1.10surprise in YAML. - Re-run the full script after each fix.
Summary and exam checklist
You can now read, write, load and repair JSON, YAML and XML, and you know why automation keeps its data in files. This chapter collects the facts for revision and for the exam.
Can-do checklist
Tick an item only if you can do it without looking:
- Explain what structured data is and why a program prefers it to CLI text.
- Write a valid JSON object with strings, numbers, booleans, null, an array and a nested object.
- List the JSON rules: double quotes, no trailing comma, no comments, lowercase
true,false,null. - Load JSON with
json.load/json.loadsand write it withjson.dump/json.dumps. - Follow a path such as
data["devices"][0]["vlans"][1]and the same path injq. - Write YAML mappings, sequences and nested lists of mappings with correct indentation.
- Explain why YAML needs quotes around
no,1.10and010, and why tabs are forbidden. - Load YAML safely with
yaml.safe_loadand explain whysafe_loadis the habit. - Read an XML tree: elements, attributes, root, and use
find,findall,.text,.attrib. - Diagnose a file that will not load: parse error versus use error, validator, one fix at a time, re-run.
Side-by-side cheat-sheet
| JSON | YAML | XML | |
|---|---|---|---|
| Container for a dictionary | {"k": "v"} | k: v | element with child elements |
| Container for a list | [1, 2] | - 1 on each line | repeated child elements |
| Strings | always double quotes | quotes optional | element text |
| Comments | not allowed | # | allowed between comment markers |
| Booleans | true / false | true / false (and yes / no in older loaders) | text such as true |
| Python module | json | yaml | xml.etree.ElementTree |
| Network home | REST, RESTCONF | Ansible, variables | NETCONF |
Python quick reference
| Task | Code |
|---|---|
| JSON text to data | json.loads(text) |
| JSON file to data | json.load(f) |
| Data to JSON text | json.dumps(data, indent=2) |
| Data to JSON file | json.dump(data, f, indent=2) |
| YAML to data | yaml.safe_load(f) |
| Data to YAML text | yaml.safe_dump(data) |
| XML text to tree | ET.fromstring(text) |
| Read XML value | root.find("name").text |
Mini glossary
- structured data
- Information with labels and a fixed shape, readable by programs.
- object / mapping
- Labelled values: a JSON object, a YAML mapping, a Python dictionary.
- array / sequence
- Ordered items: a JSON array, a YAML sequence, a Python list.
- scalar
- A single plain value such as text, a number, a boolean or null.
- validator
- A tool that checks a file against the format rules and names the first mistake.
- jq
- A command-line tool that extracts values from JSON using a path.
- yamllint
- A tool that checks YAML syntax and style.
- namespace
- A label in XML saying which vocabulary a tag belongs to.
- source of truth
- The one place where the intended state of the network is recorded.
- round trip
- Dump data to a format, load it back and check it is identical.
Most tested facts
- JSON has no comments, no trailing commas, no single quotes; keys are double-quoted strings.
- JSON
true,false,nullbecome PythonTrue,False,None. loadsanddumpswork with strings;loadanddumpwork with files.- YAML structure comes from space indentation; tabs are not allowed; comments start with
#. - Always load YAML with
yaml.safe_load. - XML has opening and closing tags, one root, attributes in the opening tag, and is used by NETCONF.
KeyErrorpoints to a key mismatch;JSONDecodeErrorandYAMLErrorpoint to a broken file.- List indexes start at 0, so
[1]is the second item.
The two labs in one paragraph each
Lab 1, JSON. inventory.json fails with a single-quote error on line 15 and then a trailing comma on line 24 before the closing brace on line 25; after both fixes show_inv.py fails on KeyError: 'mgmt', which you cure by using mgmt_ip. Then jq '.devices[0].vlans[1]' gives 20.
Lab 2, YAML. vlans.yml has a misaligned dash on line 11; after the fix build.py prints VLANs 10, 20 and 30 for both switches, you add VLAN 40 (GUEST) and switch WH-SW3 in the data only, convert to vlans.json with json.dump(plan, f, indent=2), and jq '.vlans[1].name' returns "SCANNERS".
What comes next
The next module, parsing, handles the opposite problem: devices that only give text. You will turn show output into the dictionaries and lists you now know how to save as JSON and YAML.
The pipeline that rejected a pull request
A team added a validation step to its repository: every JSON and YAML file must pass a validator before a merge. The first week it blocked three pull requests, one for a trailing comma and two for YAML indentation. None of the three would have reached production anyway, but each would have broken a nightly job and cost an hour. After a month nobody missed the manual checking.
Lesson: validate data files automatically, the same way you test code.
"Why does network automation use YAML for variables and JSON for APIs?"
YAML is easy for humans to write and review and allows comments, so it suits files that engineers maintain, such as Ansible variables and inventories. JSON is strict and simple for machines to produce and parse, so it suits API traffic. Both load into the same dictionaries and lists, so the code does not care which one it received.
Key takeaways
- JSON, YAML and XML are three notations for dictionaries and lists.
- JSON is strict; YAML is human-friendly and indentation-based; XML is tag-based and used by NETCONF.
- Parse errors mean a broken file; KeyError means a mismatch between data and code.
- Keep data in files and logic in code; load YAML with
safe_load. - Validate, fix one error at a time, and re-run the whole script.