CCNP Automation 350-901 ยท Start here ยท Data

Data formats: JSON, YAML & XML

Scripts and APIs do not read "show" output the way you do โ€” they exchange structured data. Learn the rules of JSON, YAML and XML, spot the tiny mistakes that break them, and convert between them in Python.

47 min read9 chapters2 labs15 quiz7 scenarios15 interview Q&A

This first module is free: read the lesson and take the quiz. Create a free account to run up to 3 hands-on labs.

Log inStart free
Jump to chapter (9)
01

Why structured data? JSON, YAML and XML in one picture

What you will learn in this module. You will learn to read, write and repair the three data formats that network automation runs on: JSON, YAML and XML. You will load them into Python, change them, convert one into another, and find a syntax mistake quickly when a file will not load. In the labs you repair an inventory.json that nobody can read and a vlans.yml that makes a config generator crash, then grow a network by editing only data. This is part of the CCNP Automation (350-901 AUTOCOR) track.

Prerequisites. The previous modules: Python 1 (variables, strings), Python 2 (lists, dictionaries, loops) and Python 3 (functions, files and exceptions). You do not need to know any of the three formats yet.

Analogy: a letter versus a form

A letter says "Please send the parcel to Mr Rao, flat 12, Tower B, Pune 411001". A human reads it easily. A machine cannot reliably tell where the name ends and the address begins. Now look at a form: boxes labelled *Name*, *Flat*, *Tower*, *City*, *Pincode*, each box with its value. A machine can read a form without guessing, because every value has a label and a fixed place. Structured data is a form that a computer can read. JSON, YAML and XML are three ways of drawing the form.

The problem with CLI text

The output of show ip interface brief is a letter for humans. To use it in a program you must cut it with split() or regular expressions, and the cuts break the moment a column changes. The next module teaches that skill. But APIs, automation tools and source-of-truth files do not use letters; they exchange forms. So you must be fluent in them.

CLI textfor humansJSONAPIs, configsYAMLAnsible, variablesXMLNETCONF, older APIs

Same information, three ways to write it

One device, three formats

Here is the same small fact: router R1 has the management IP 10.99.0.1 and two VLANs, 10 and 20.

{"name": "R1", "mgmt_ip": "10.99.0.1", "vlans": [10, 20]}

That was JSON (JavaScript Object Notation): braces, quotes, commas.

name: R1
mgmt_ip: 10.99.0.1
vlans:
  - 10
  - 20

That was YAML ("YAML Ain't Markup Language"): no braces, no quotes, and the structure comes from indentation. Easy for humans to write.

<device>
  <name>R1</name>
  <mgmt_ip>10.99.0.1</mgmt_ip>
  <vlans>
    <vlan>10</vlan>
    <vlan>20</vlan>
  </vlans>
</device>

That was XML (eXtensible Markup Language): every value sits between an opening tag and a closing tag. It is wordy but very strict, and NETCONF still uses it.

They all mean the same thing

Different spelling, same data. Python can prove it: it loads both texts and compares the results.

import json
import yaml

json_text = '{"name": "R1", "mgmt_ip": "10.99.0.1", "vlans": [10, 20]}'
yaml_text = """
name: R1
mgmt_ip: 10.99.0.1
vlans:
  - 10
  - 20
"""

a = json.loads(json_text)
b = yaml.safe_load(yaml_text)
print(a)
print(b)
print("Same data:", a == b)

Real output, checked in the lab terminal.

{'name': 'R1', 'mgmt_ip': '10.99.0.1', 'vlans': [10, 20]}
{'name': 'R1', 'mgmt_ip': '10.99.0.1', 'vlans': [10, 20]}
Same data: True

Both texts became the same Python dictionary with a list inside. That is the key idea of this module: once the file is loaded, the format no longer matters. What matters is the shape: dictionaries (labelled values) and lists (ordered items), which you already know from Python 2.

Which format where

FormatTypical home in networkingStrength
JSONREST and RESTCONF APIs, controller replies, config filesSimple, universal, strict
YAMLAnsible playbooks and variables, CI pipelines, inventoriesEasy for humans to edit, allows comments
XMLNETCONF, older vendor APIsStrict, supports schemas and attributes

What the labs ask you to do

Lab 1 gives you inventory.json that Python cannot read, plus a show_inv.py that prints a table from it. You validate, fix two syntax mistakes, correct a wrong key name in the script and query the file with jq. Lab 2 gives you vlans.yml and build.py: fix an indentation error, add a VLAN and a switch by editing only the data, and convert the plan to JSON with to_json.py.

Readwhat each format looks likeWritethe syntax rulesLoadin PythonRepairvalidate and fix

Your path

Common beginner mistake. Treating the three formats as three different languages to learn. They are three notations for the same two building blocks, dictionaries and lists. Learn the blocks once and the notations follow.

Exam trap. Know which format goes with which technology: JSON with REST and RESTCONF, XML with NETCONF, YAML with Ansible. Also know that YAML uses indentation and JSON uses braces and brackets.

The inventory nobody could read

A company stored its device list as text pasted from an email. Each engineer parsed it with their own scripts and the three scripts disagreed about which devices existed. After moving the list to one JSON file, every script loaded it the same way, and a new device was added in one place. The data became the single source of truth.

Lesson: structured data gives every tool the same view of the network.

"What is the difference between JSON, YAML and XML?"

All three describe structured data. JSON uses braces, brackets and quotes and is the standard for REST APIs. YAML uses indentation, is easy for humans and allows comments, so Ansible uses it. XML uses opening and closing tags, is strict, and is used by NETCONF. Once loaded in Python, JSON and YAML become dictionaries and lists.

Key takeaways

  • Structured data is a form a program can read without guessing.
  • JSON, YAML and XML are three notations for dictionaries and lists.
  • JSON is used by REST APIs, YAML by Ansible, XML by NETCONF.
  • Once loaded into Python, the original format no longer matters.
  • Two labs: repair an inventory.json and drive config from vlans.yml.
02

JSON syntax: objects, arrays and six value types

JSON is the most common format you will meet. Almost every REST API answers in JSON. It has only a handful of rules, and they are strict. Learn them once and you can spot every error.

The two containers

JSON has two containers. Everything else is a simple value.

  • An object is written with curly braces { }. It holds key: value pairs separated by commas. The key is always text in double quotes. It is like a form with labelled boxes, and it becomes a Python dictionary.
  • An array is written with square brackets [ ]. It holds values in order, separated by commas. It becomes a Python list.
{
  "hostname": "HYD-CORE1",
  "vlans": [10, 20, 30]
}

Read it aloud: "An object with two keys. The key hostname has the text HYD-CORE1. The key vlans has an array of three numbers."

[[flow Containers hold values | Object/{ "key": value }/ | Array/[ value, value ] | Value/text, number, true, false, null]]

The value types

A value can be one of exactly six things:

TypeExamplePython gets
string"GigabitEthernet0/1"str
number10, 3.5, -1int or float
booleantrue, falseTrue, False
nullnullNone
object{"a": 1}dict
array[1, 2]list

Notice that booleans and null are lowercase, and not in quotes. The text "true" in quotes is a string, not a boolean. Also notice that 10 (a number) and "10" (text) are different: you can do arithmetic with the first.

Nesting: boxes inside boxes

Containers can hold other containers. This is how a whole inventory is described: an object with a key devices whose value is an array, and each item of the array is itself an object.

import json

text = """
{
  "company": "Deccan Net",
  "site": "Hyderabad",
  "devices": [
    {"name": "HYD-CORE1", "role": "core", "vlans": [10, 20, 30], "managed": true},
    {"name": "HYD-ACC1", "role": "access", "vlans": [10, 20], "managed": false}
  ]
}
"""
data = json.loads(text)
print(data["site"])
print(data["devices"][0]["name"])
print(data["devices"][1]["vlans"][0])
print(data["devices"][0]["managed"])

Real output, checked in the lab terminal.

Hyderabad
HYD-CORE1
10
True

Follow each path like directions in a building: data["devices"] is the array, [0] is the first device (counting starts at 0), ["name"] is its name. The last line prints True, not true: Python spells the boolean its own way after loading.

The rules that cause most errors

  1. Keys and strings use double quotes only. Single quotes are not allowed.
  2. No trailing comma. A comma goes between items, never after the last one.
  3. A comma is required between items. Forgetting it is the other common mistake.
  4. Keys must be strings. {name: "R1"} is invalid; write {"name": "R1"}.
  5. No comments. // and # are not part of JSON.
  6. Every opening bracket must close. Count your { } and [ ].
  7. **true, false, null are lowercase** and unquoted. True and None are Python, not JSON.

Python can show you what happens with a broken file. Notice that the message names the line and the column where it got confused:

import json

text = """{
  "name": "HYD-EDGE1",
  "mgmt_ip": '10.99.0.2'
}"""
data = json.loads(text)

Real output, checked in the lab terminal.

Traceback (most recent call last):
  File "badjson.py", line 7, in <module>
    data = json.loads(text)
json.decoder.JSONDecodeError: Expecting value: line 3 column 14 (char 38)

JSONDecodeError: Expecting value: line 3 column 14 says: at line 3, column 15 Python expected the start of a value but found something else, here the single quote. The place where Python complains is sometimes just after the real mistake, so also look at the line before it.

Pretty and compact

Machines do not care about spaces. JSON can be one long line (compact) or indented (pretty). Both are identical in meaning. Pretty is for people, and tools such as jq and json.dumps(data, indent=2) produce it.

Common mistakes. (1) Copying a Python dictionary into a .json file: it contains single quotes, True and None, all invalid in JSON. (2) A trailing comma after the last item, which many programmers add out of habit. (3) Writing numbers with leading zeros such as 010. (4) Putting IP addresses without quotes: 10.99.0.1 is not a valid number, it must be the string "10.99.0.1".

Exam trap. Which of these is valid JSON? The answer always has double quotes, lowercase true, no trailing comma and no comments. Also remember that JSON object keys are unordered in meaning, while array items are ordered.

The controller reply that "was not JSON"

A script that called a controller crashed with a decode error. The engineer opened the saved reply and saw a Python-style dictionary with True and single quotes. Someone had printed the Python object and saved the printout as the test file. The real controller reply was valid JSON. Replacing the saved sample with a proper json.dump of the data fixed the unit tests.

Lesson: printing a Python object is not the same as writing JSON; use json.dump to write it.

"What are the data types in JSON?"

There are six: string, number, boolean, null, object and array. Objects are key-value pairs in braces with double-quoted keys, and arrays are ordered lists in brackets. Booleans and null are lowercase. JSON has no comments, no single quotes and no trailing commas.

Key takeaways

  • JSON has two containers: objects { } and arrays [ ].
  • Six value types: string, number, true/false, null, object, array.
  • Strings and keys use double quotes only; commas go between items, never after the last.
  • JSON has no comments; true, false and null are lowercase.
  • Nested paths read like directions: data["devices"][0]["name"].
03

JSON in Python: load, dump and the four functions

Python ships with a module called json for reading and writing JSON. It has four main functions. They look alike, so this chapter gives you a simple memory trick.

The "s" means string

  • json.loads(text): load from a string. Text in, Python data out.
  • json.load(file): load from a file object. File in, Python data out.
  • json.dumps(data): dump to a string. Python data in, text out.
  • json.dump(data, file): dump to a file object. Python data in, file written.

The trick: the function with an s works with a string. The one without works with a file. And load means "into Python", dump means "out of Python".

Text or fileJSON on diskloadloads/text to PythonPython datadict and listdumpdumps/Python to text

Four functions

Loading text into data

import json

text = '{"name": "HYD-CORE1", "mgmt_ip": "10.99.0.1", "vlans": [10, 20, 30], "managed": true, "owner": null}'
device = json.loads(text)

print(type(device))
print(device["name"])
print(device["vlans"][1])
print(device["managed"], type(device["managed"]))
print(device["owner"])

Real output, checked in the lab terminal.

<class 'dict'>
HYD-CORE1
20
True <class 'bool'>
None

The text became a dictionary. Look at the last two lines: true became True, and null became None. After loading, you use ordinary Python: keys, indexes, loops.

Loading from a file

In lab 1 the data lives in inventory.json and show_inv.py reads it with json.load. Here is a self-contained version: it first writes a small file, then reads it back.

import json

with open("inventory.json", "w") as f:
    f.write('{"site": "Hyderabad", "devices": [{"name": "HYD-CORE1", "vlans": [10, 20, 30]}, {"name": "HYD-ACC1", "vlans": [10, 20]}]}')

with open("inventory.json") as f:
    inv = json.load(f)

print("Site:", inv["site"])
for dev in inv["devices"]:
    print(dev["name"], dev["vlans"])

Real output, checked in the lab terminal.

Site: Hyderabad
HYD-CORE1 [10, 20, 30]
HYD-ACC1 [10, 20]

Changing data and writing it back

Because the data is now a dictionary, you edit it like any dictionary. Then json.dump writes it out. The indent=2 argument makes the file pretty. Without it everything lands on one line.

import json

inv = {"site": "Hyderabad", "devices": [{"name": "HYD-CORE1", "vlans": [10, 20]}]}
inv["devices"][0]["vlans"].append(30)
inv["devices"].append({"name": "HYD-ACC2", "vlans": [10]})

with open("out.json", "w") as f:
    json.dump(inv, f, indent=2)

with open("out.json") as f:
    print(f.read())

Real output, checked in the lab terminal.

{
  "site": "Hyderabad",
  "devices": [
    {
      "name": "HYD-CORE1",
      "vlans": [
        10,
        20,
        30
      ]
    },
    {
      "name": "HYD-ACC2",
      "vlans": [
        10
      ]
    }
  ]
}

Notice that Python's True and None would be written back as true and null, so a program can write valid JSON without you thinking about it. This is why you write JSON with json.dump and never by pasting Python's printout.

dumps: JSON as text

json.dumps is handy to show data or to send it in an API request. Useful options: indent=2 (pretty) and sort_keys=True (alphabetical keys, good for comparing two files).

import json

data = {"name": "R1", "up": True, "owner": None, "vlans": [10, 20]}
print(json.dumps(data))
print(json.dumps(data, indent=2, sort_keys=True))
print(type(json.dumps(data)))

Real output, checked in the lab terminal.

{"name": "R1", "up": true, "owner": null, "vlans": [10, 20]}
{
  "name": "R1",
  "owner": null,
  "up": true,
  "vlans": [
    10,
    20
  ]
}
<class 'str'>

type shows str: dumps gives you text, not a dictionary. A very common beginner confusion is to call dumps and then try to use the result as a dictionary.

Walking through a loaded inventory

Loops and dictionary reads from the previous module work unchanged:

import json

inv = json.loads('{"devices": [{"name": "A", "managed": true}, {"name": "B", "managed": false}, {"name": "C", "managed": true}]}')
managed = [d["name"] for d in inv["devices"] if d["managed"]]
print("Managed devices:", managed)
print("Count:", len(inv["devices"]))

Real output, checked in the lab terminal.

Managed devices: ['A', 'C']
Count: 3

Common mistakes. (1) Using json.load on a string or json.loads on a file object: you get a TypeError or AttributeError. (2) Forgetting to open the file with with open(...). (3) Calling json.dumps(data) and expecting a file to be written. (4) Reading a missing key and getting KeyError; use data.get("key") when the key may be absent.

Exam trap. Which function turns a string into a Python dictionary? loads. Which writes a dictionary to a file? dump. The one with the s is for strings.

One byte too many

A script saved device data with str(data) instead of json.dump. The saved file looked fine, but the next script crashed trying to load it, because str wrote True and single quotes. The fix was a one-line change to json.dump(data, f, indent=2). From then on both scripts read and wrote the same valid JSON.

Lesson: always write JSON with the json module, never with str() or print().

"What is the difference between json.load and json.loads?"

load reads JSON from a file object, loads from a string. The same pairing exists for dump and dumps. After loading, JSON objects become dictionaries and arrays become lists, with true, false and null becoming True, False and None.

Key takeaways

  • loads and dumps work with strings; load and dump work with files.
  • load means into Python, dump means out of Python.
  • JSON objects become dicts, arrays become lists, null becomes None.
  • json.dump(data, f, indent=2) writes a readable file.
  • Never write JSON by printing a Python object; use the json module.
04

Repairing broken JSON: validate, fix, query with jq

This chapter follows lab 1 step by step. The story: Deccan Net keeps its Hyderabad device list in inventory.json, and show_inv.py prints it as a table for the NOC. Someone edited the file by hand and now nothing can read it. You will find the mistakes with a validator, fix them, repair a wrong key name in the script, and then ask the file questions with jq.

The file and the symptom

The file describes three devices. These are the lines that matter in the broken version (line numbers on the left come from the lab command nl):

    13        "name": "HYD-EDGE1",
    14        "role": "edge",
    15        "mgmt_ip": '10.99.0.2',
    16        "vlans": [99],
    ...
    23        "vlans": [10, 20],
    24        "managed": false,
    25      }

Running the script gives the Python view of the problem:

AUTO> python3 show_inv.py
Traceback (most recent call last):
  File "show_inv.py", line 6, in <module>
    inv = json.load(f)
json.decoder.JSONDecodeError: Expecting value: line 15 column 18 (char 282)

The last line is exactly what you learned to read: the class is JSONDecodeError, and Python points to line 15, column 18.

Use a validator

A validator checks a file against the rules and reports the first problem. The lab has a command json validate. On a real computer you can use python3 -m json.tool inventory.json (it prints the file if valid, otherwise an error with the line) or an editor with JSON support.

AUTO> json validate inventory.json
inventory.json: INVALID JSON
  line 15, column 18: strings must use double quotes ("), not single quotes (')
   15 |       "mgmt_ip": '10.99.0.2',
      |                  ^

The caret points at the single quote. Fix line 15 so the value uses double quotes, keeping the same indentation (6 spaces). In the lab you use edit inventory.json 15 ... to replace a line. Then validate again.

AUTO> json validate inventory.json
inventory.json: INVALID JSON
  line 25, column 5: trailing comma is not allowed before '}'
   25 |     }
      |     ^

The first error is gone; the validator now finds the next one. A tool reports one mistake at a time, so you repeat until the file is clean. The message says the comma is "before }", meaning the comma is on the line above, line 24 ("managed": false,). Remove it.

AUTO> json validate inventory.json
inventory.json: valid JSON (object with 3 top-level key(s): company, site, devices)

Valid JSON now. Always keep all three devices: the lab check expects the third one (HYD-ACC1, with "managed": false) to still be present.

Validaterun the checkerReadline and messageFixone mistakeRepeatuntil valid

The repair loop

Valid JSON is not the same as working code

Run the script again. JSON now loads, but a different error appears, this time in the script:

AUTO> python3 show_inv.py
Company: Deccan Net | Site: Hyderabad
NAME        ROLE    MGMT IP       VLANS
Traceback (most recent call last):
  File "show_inv.py", line 12, in <module>
    print(f"{dev['name']:<12}{dev['role']:<8}{dev['mgmt']:<14}{dev['vlans']}")
KeyError: 'mgmt'

The data is correct. The script asks for a key mgmt but the file calls it mgmt_ip. KeyError: 'mgmt' is the dictionary telling you it has no such key. The cure is to make the script use the real key name. In the lab: sed -i "s/'mgmt'/'mgmt_ip'/" show_inv.py, a command that replaces text in a file.

AUTO> python3 show_inv.py
Company: Deccan Net | Site: Hyderabad
NAME        ROLE    MGMT IP       VLANS
HYD-CORE1   core    10.99.0.1     [10, 20, 30]
HYD-EDGE1   edge    10.99.0.2     [99]
HYD-ACC1    access  10.99.0.11    [10, 20]

Ask the file questions with jq

jq is a command-line tool for reading JSON. You give it a path and it prints the value. The dot . means "the whole document", and you follow it with keys and indexes, just like Python.

AUTO> jq '.devices[0].vlans[1]' inventory.json
20

Read the path: .devices is the array, [0] is the first device (HYD-CORE1), .vlans is its list [10, 20, 30], and [1] is the second item because counting starts at 0. The answer is 20.

Common patterns you can use on the same file: jq '.site' inventory.json prints "Hyderabad", jq '.devices[].name' inventory.json prints the name of every device, and jq '.devices | length' inventory.json prints 3.

The same job in Python needs one more line, and gives you the same answer:

import json

inv = json.loads('{"devices": [{"name": "HYD-CORE1", "vlans": [10, 20, 30]}, {"name": "HYD-EDGE1", "vlans": [99]}]}')
print(inv["devices"][0]["vlans"][1])
print([d["name"] for d in inv["devices"]])
print(len(inv["devices"]))

Real output, checked in the lab terminal.

20
['HYD-CORE1', 'HYD-EDGE1']
2

A checklist for a file that will not load

  1. Run the validator and read the line and column.
  2. Look at the reported line and the line above.
  3. Check quotes, then commas (missing between items, extra after the last), then brackets.
  4. Fix one thing, validate again.
  5. When the file loads, check names: does the key in the code match the key in the data?

Common mistakes. (1) Fixing the wrong line because the error is reported after the real mistake, for example a missing comma at the end of the previous line. (2) Deleting a whole device to make the file valid, which silently loses data. (3) Assuming valid JSON means correct data. (4) Editing the file in a word processor that turns straight quotes into curly quotes, which JSON rejects.

Exam trap. A JSON file with a trailing comma is invalid, even though Python lists allow it. The same is true for single quotes and comments.

The curly quotes

An engineer pasted a snippet from a chat window into a JSON file. The chat app had changed " into curly quotes. The validator reported an error at the first quote of the first key, which looked perfectly fine on screen. Retyping the quotes in a code editor fixed it, and the team added a validation step before every commit.

Lesson: look at the exact character the validator points to, not only at what looks right.

"A teammate's JSON file fails to load. How do you debug it?"

I run a validator or python3 -m json.tool to get the line and column, check the line and the one above it for single quotes, missing or trailing commas and unclosed brackets, fix one error at a time, and re-run. Once it loads I verify that the keys my code uses match the keys in the data, using jq to inspect the structure.

Key takeaways

  • A validator reports the first mistake; repeat until the file is valid.
  • The reported line may be just after the real mistake; also check the line above.
  • Valid JSON can still be wrong data or mismatch the code's key names.
  • KeyError means the code asked for a key that the data does not have.
  • jq '.devices[0].vlans[1]' follows keys and indexes; index counting starts at 0.
05

YAML syntax: indentation is the structure

YAML was designed to be easy for people to write. It drops braces and quotes and uses indentation and a few symbols instead. Ansible, many CI tools and most variable files use it. The price of the clean look: you must be exact with spaces.

The three building blocks

  1. Mapping: key: value. A colon and a space. This is a dictionary.
  2. Sequence: items that start with - (dash and space), one per line. This is a list.
  3. Scalar: a plain value: text, number, true/false, null.
import yaml

text = """
name: HYD-CORE1
mgmt_ip: 10.99.0.1
managed: true
vlans:
  - 10
  - 20
  - 30
"""
data = yaml.safe_load(text)
print(data)

Real output, checked in the lab terminal.

{'name': 'HYD-CORE1', 'mgmt_ip': '10.99.0.1', 'managed': True, 'vlans': [10, 20, 30]}

Compare with JSON from the previous chapters: no braces, no quotes, no commas. The two keys with simple values sit on one line each. The key vlans has nothing after the colon, so its value is the indented sequence below it. Python received a dictionary with a list inside, exactly as for JSON.

Mappingkey: valueSequence- itemScalartext, number, true, null

YAML building blocks

Indentation rules

  • Indentation means belonging. Lines indented under a key belong to that key.
  • Use spaces only, never tabs. Two spaces per level is the convention.
  • Items at the same level must line up exactly. One extra space makes them a different level, and then YAML breaks.

A list of mappings

The most common structure in network files is a list where each item has several fields. Each item starts with a dash. The fields of that item line up under the first field.

import yaml

text = """
switches:
  - name: WH-SW1
    mgmt: 10.99.0.11
  - name: WH-SW2
    mgmt: 10.99.0.12
"""
data = yaml.safe_load(text)
for sw in data["switches"]:
    print(sw["name"], sw["mgmt"])
print(data["switches"][1]["mgmt"])

Real output, checked in the lab terminal.

WH-SW1 10.99.0.11
WH-SW2 10.99.0.12
10.99.0.12

Read the shape: the key switches holds a list. The first item begins with - name: WH-SW1, and mgmt: is lined up exactly under name: (4 spaces), so it belongs to the same item. The second dash starts a new item.

Comments and extra tools

A # starts a comment, which JSON cannot do. This makes YAML ideal for files that people maintain.

import yaml

text = """
# VLAN plan for the warehouse
site: Nagpur-WH      # the site name
vlans: [10, 20, 30]  # a short list can be written like JSON
ids: {office: 10, cctv: 30}
note: "text with: a colon needs quotes"
description: |
  Line one
  Line two
"""
data = yaml.safe_load(text)
for key, value in data.items():
    print(key, "->", repr(value))

Real output, checked in the lab terminal.

site -> 'Nagpur-WH'
vlans -> [10, 20, 30]
ids -> {'office': 10, 'cctv': 30}
note -> 'text with: a colon needs quotes'
description -> 'Line one\nLine two\n'

What you see: short lists and mappings may use JSON-like brackets (called flow style); a value that contains a colon followed by a space needs quotes; and | starts a block of text that keeps its line breaks.

YAML guesses types

Without quotes, YAML decides the type for you. That is convenient, and sometimes surprising.

import yaml

text = """
vlan: 10
version: 1.10
enabled: true
speed: 1G
code: 0123
answer: no
empty:
"""
data = yaml.safe_load(text)
for key, value in data.items():
    print(key, repr(value), type(value).__name__)

Reference code (not run in the simulator). Output below is from standard Python 3.

vlan 10 int
version 1.1 float
enabled True bool
speed '1G' str
code 83 int
answer False bool
empty None NoneType

Look at three traps. 1.10 became the number 1.1, so a version string lost its zero. 0123 was read as the number 83, because a leading zero meant octal in older YAML. And no became False, the so-called "Norway problem": a country code NO, or the word on or yes, turns into a boolean. The cure for all of them is the same: put quotes around values that must stay text (version: "1.10", country: "NO").

Syntax errors in YAML

A wrong indentation produces an error that names the line and what it expected. The lab tool yamllint reports it like this:

AUTO> yamllint vlans.yml
vlans.yml:11:4: error: bad indentation of a sequence entry (this line is indented more than the list item above it)

It says line 11, column 4: the dash is one space too far to the right compared with the other items. On a real computer the Python yaml module and the yamllint program give similar messages. Tabs trigger their own message.

Common mistakes. (1) Using a tab instead of spaces. (2) Misaligning one item by one space. (3) Forgetting the space after the colon (key:value is a plain string, not a mapping). (4) Leaving unquoted values that YAML converts (no, on, 1.10, 010). (5) Copying JSON-style quotes unnecessarily; they are allowed but not needed.

Exam trap. YAML allows comments with #; JSON does not. YAML structure comes from indentation with spaces; tabs are forbidden. A dash starts a list item and key: value a mapping.

The country that became False

An inventory file had a field country: NO for a branch in Norway. After loading it, a report said the country was False. YAML 1.1 loaders read NO as the boolean false. The team quoted the value as "NO" and added a test that checks the type of every field after loading.

Lesson: when a value must stay text, quote it.

"How does YAML represent a list of dictionaries?"

Each list item starts with a dash and a space, and the fields of that item are written as key: value lines aligned under the first key. The indentation shows which keys belong to which item. For example, - name: SW1 then mgmt: 10.0.0.1 under it.

Key takeaways

  • YAML has mappings (key: value), sequences (- item) and scalars.
  • Indentation with spaces is the structure; tabs are not allowed.
  • # starts a comment; quotes protect values from being converted.
  • no, yes, on, 1.10 and 0123 are guessed as other types unless quoted.
  • Both JSON and YAML load into the same Python dictionaries and lists.
06

YAML in Python: data in one file, code in another

This chapter follows lab 2. The story: Kaveri Foods describes its warehouse switches and VLANs in vlans.yml, and build.py turns that data into config for every switch. The big idea is to keep the code and the data apart. When the network grows, you edit only the YAML file and never touch the script.

Analogy: the recipe and the shopping list

A recipe (the code) stays the same all year. The shopping list (the data) changes every week. If every shopping change meant rewriting the recipe, cooking would be a nightmare. Network automation is the same: the build script is the recipe, and the YAML file is the list of switches, VLANs and addresses. The file with the data is often called the source of truth.

Loading YAML

You need the yaml module (the PyYAML package, already available in the lab). Two functions matter:

  • yaml.safe_load(text_or_file): YAML in, Python dictionaries and lists out.
  • yaml.safe_dump(data): Python data in, YAML text out.

Use safe_load, not the plain load. The safe version builds only ordinary data (text, numbers, lists, dictionaries). The unsafe version can build arbitrary Python objects from a file, which is a security risk when the file comes from somewhere you do not control. Defensive habit: always safe_load.

The data file

This is the repaired vlans.yml from the lab. It has one site, two switches and three VLANs.

The build script

Read it slowly. The outer loop visits every switch; the inner loop visits every VLAN, so each switch receives all VLANs. The f-strings drop values into the config text.

File vlans.yml

# vlans.yml - VLAN plan for the Kaveri Foods warehouse (only data, no code)
site: Nagpur-WH
switches:
  - name: WH-SW1
    mgmt: 10.99.0.11
  - name: WH-SW2
    mgmt: 10.99.0.12
vlans:
  - id: 10
    name: OFFICE
  - id: 20
    name: SCANNERS
  - id: 30
    name: CCTV
import yaml

with open("vlans.yml") as f:
    plan = yaml.safe_load(f)

print("! site:", plan["site"])
for sw in plan["switches"]:
    print(f"! ---- {sw['name']} ({sw['mgmt']}) ----")
    for v in plan["vlans"]:
        print(f"vlan {v['id']}")
        print(f" name {v['name']}")

Real output, checked in the lab terminal.

! site: Nagpur-WH
! ---- WH-SW1 (10.99.0.11) ----
vlan 10
 name OFFICE
vlan 20
 name SCANNERS
vlan 30
 name CCTV
! ---- WH-SW2 (10.99.0.12) ----
vlan 10
 name OFFICE
vlan 20
 name SCANNERS
vlan 30
 name CCTV

Count what happened: 2 switches times 3 VLANs gave 6 vlan blocks. Every switch got the same VLAN list because the list lives once in the data.

vlans.ymlthe databuild.pythe codeConfigtext for every switch

Data and code are separate

Growing the network by editing only the data

Management asks for VLAN 40 (GUEST) and a third switch, WH-SW3 at 10.99.0.13. In the lab you add four lines to vlans.yml and run build.py again. Here is the extended data, and a smaller report script that counts what each switch now receives.

File vlans.yml

site: Nagpur-WH
switches:
  - name: WH-SW1
    mgmt: 10.99.0.11
  - name: WH-SW2
    mgmt: 10.99.0.12
  - name: WH-SW3
    mgmt: 10.99.0.13
vlans:
  - id: 10
    name: OFFICE
  - id: 20
    name: SCANNERS
  - id: 30
    name: CCTV
  - id: 40
    name: GUEST
import yaml

with open("vlans.yml") as f:
    plan = yaml.safe_load(f)

names = [v["name"] for v in plan["vlans"]]
for sw in plan["switches"]:
    print(sw["name"], "gets", len(names), "VLANs:", ", ".join(names))

Real output, checked in the lab terminal.

WH-SW1 gets 4 VLANs: OFFICE, SCANNERS, CCTV, GUEST
WH-SW2 gets 4 VLANs: OFFICE, SCANNERS, CCTV, GUEST
WH-SW3 gets 4 VLANs: OFFICE, SCANNERS, CCTV, GUEST

The script did not change by a single character, yet the output grew from 2 switches and 3 VLANs to 3 switches and 4 VLANs. That is the point of data-driven automation: the change is a data edit that a person can review in a pull request.

YAML to JSON: same data, different notation

The monitoring team wants the plan as JSON. Loading with yaml and dumping with json is a two-line conversion. The lab version writes vlans.json with json.dump(plan, f, indent=2). Here is the same idea shown in memory:

import json
import yaml

plan = yaml.safe_load("""
site: Nagpur-WH
vlans:
  - id: 10
    name: OFFICE
  - id: 20
    name: SCANNERS
""")
print(json.dumps(plan, indent=2))
print(yaml.safe_dump(plan, default_flow_style=False))

Real output, checked in the lab terminal.

{
  "site": "Nagpur-WH",
  "vlans": [
    {
      "id": 10,
      "name": "OFFICE"
    },
    {
      "id": 20,
      "name": "SCANNERS"
    }
  ]
}
site: Nagpur-WH
vlans:
- id: 10
  name: OFFICE
- id: 20
  name: SCANNERS

Both directions work because both formats describe the same dictionaries and lists. In the lab you then ask the file a question with jq:

AUTO> json validate vlans.json
vlans.json: valid JSON (object with 3 top-level key(s): site, switches, vlans)

AUTO> jq '.vlans[1].name' vlans.json
"SCANNERS"

Index 1 is the second VLAN, so the answer is SCANNERS (with quotes, because it is a JSON string).

If the YAML is broken

When build.py met the original misaligned file, it did not produce config at all. The lab traceback ends like this:

AUTO> python3 build.py
Traceback (most recent call last):
  File "build.py", line 6, in <module>
    plan = yaml.safe_load(f)
yaml.YAMLError: bad indentation of a sequence entry (this line is indented more than the list item above it)

Handle it like any traceback: read the last line, then run yamllint vlans.yml to find line 11 and fix it. Fixing means giving the dash the same two-space indentation as the other items.

Common mistakes. (1) Using yaml.load instead of yaml.safe_load. (2) Hard-coding switch names in the script, which defeats the purpose of the data file. (3) Adding a VLAN with the wrong indentation so it silently attaches to the previous item. (4) Forgetting that YAML needs both id and name for every VLAN item, so a missing key later causes KeyError.

Exam trap. The safe loader is yaml.safe_load. Converting YAML to JSON means load with yaml, dump with json; there is no direct "convert" function. Also: Ansible variable files and inventories are YAML.

Forty switches, one pull request

A campus team generated VLAN config for forty switches from one YAML file and a 30-line script. When the security team requested a new guest VLAN, the change was four lines of YAML in a pull request. Reviewers read exactly what would change, approved it, and the script generated forty configs in seconds. Nobody edited forty files by hand.

Lesson: data-driven automation makes a change small, reviewable and repeatable.

"Why separate data from code in automation?"

Because the logic changes rarely and the data changes often. A YAML file of devices and VLANs can be edited and reviewed by anyone, version-controlled and reused by several tools, while the script stays untouched. It also acts as a source of truth for what the network should look like.

Key takeaways

  • Keep data (YAML) and logic (Python) in separate files.
  • Always use yaml.safe_load; yaml.safe_dump writes YAML back.
  • Growing the network means editing the data, not the script.
  • YAML to JSON: load with yaml, write with json.dump.
  • A broken YAML file stops yaml.safe_load; use yamllint to find the line.
07

XML: tags, attributes and reading it in Python

XML is the oldest of the three formats and the wordiest. You will still meet it, mainly in NETCONF, the protocol that older and many current network devices use for model-driven configuration. You rarely write XML by hand in this track, but you must be able to read it and pick values out of it.

Anatomy of an XML document

  • An element is a pair of tags: <name>R1</name>. The first is the opening tag, the last the closing tag, and the text in between is the value.
  • Elements nest. An element that contains others is a parent; the ones inside are children. The single outermost element is the root.
  • An attribute is extra information inside the opening tag: <interface type="ethernet">. The value is always in quotes.
  • An element with no content can close itself: <shutdown/>.
<device>
  <name>R1</name>
  <interfaces>
    <interface type="ethernet">
      <name>GigabitEthernet0/1</name>
      <ip>10.10.1.1</ip>
      <enabled>true</enabled>
    </interface>
    <interface type="loopback">
      <name>Loopback0</name>
      <ip>1.1.1.1</ip>
      <enabled>true</enabled>
    </interface>
  </interfaces>
</device>

Read it like a tree: device is the root, it has one child interfaces, which has two interface children, and each of those has name, ip and enabled. The attribute type describes the interface itself.

RootdeviceParentinterfacesChildinterface (type)Leavesname, ip, enabled

XML is a tree

The strict rules

  1. Every opening tag needs a matching closing tag (or a self-closing tag).
  2. Tags must be properly nested: close the inner tag before the outer one.
  3. There is exactly one root element.
  4. Attribute values are in quotes.
  5. Tag names are case-sensitive: <Name> and <name> differ.
  6. The characters <, > and & inside text must be written as &lt;, &gt; and &amp;.

Reading XML with Python

The standard library module xml.etree.ElementTree loads XML into a tree you can search. This is reference code (not run in the simulator), but it runs on any normal Python 3 install, and the output below is from a real Python 3 run.

import xml.etree.ElementTree as ET

text = """
<device>
  <name>R1</name>
  <interfaces>
    <interface type="ethernet"><name>GigabitEthernet0/1</name><ip>10.10.1.1</ip></interface>
    <interface type="loopback"><name>Loopback0</name><ip>1.1.1.1</ip></interface>
  </interfaces>
</device>
"""
root = ET.fromstring(text)
print(root.tag)
print(root.find("name").text)
for intf in root.findall("./interfaces/interface"):
    print(intf.attrib["type"], intf.find("name").text, intf.find("ip").text)

Reference code (not run in the simulator). Output below is from standard Python 3.

device
R1
ethernet GigabitEthernet0/1 10.10.1.1
loopback Loopback0 1.1.1.1

Line by line: ET.fromstring(text) parses the text and returns the root element. root.tag is its tag name. root.find("name") finds the first child called name, and .text gives its value. findall("./interfaces/interface") is a small path expression (a limited form of XPath): start at the root, go into interfaces, then collect every interface. .attrib is a dictionary of the attributes.

Everything you read is text. If you need a number, convert it yourself, for example int(vlan.text).

Writing XML

You can build a tree with code and turn it into text:

import xml.etree.ElementTree as ET

root = ET.Element("interface", {"type": "ethernet"})
ET.SubElement(root, "name").text = "GigabitEthernet0/2"
ET.SubElement(root, "ip").text = "10.10.2.1"
print(ET.tostring(root, encoding="unicode"))

Reference code (not run in the simulator). Output below is from standard Python 3.

<interface type="ethernet"><name>GigabitEthernet0/2</name><ip>10.10.2.1</ip></interface>

Namespaces in one minute

Large XML systems mix vocabularies, so tags may carry a namespace, a label that says which vocabulary a tag belongs to. In NETCONF replies you often see an attribute such as xmlns="urn:example:interfaces" on the root. In Python the namespace is glued in front of the tag name inside curly brackets.

import xml.etree.ElementTree as ET

text = '<interfaces xmlns="urn:example:interfaces"><interface><name>Gi0/1</name></interface></interfaces>'
root = ET.fromstring(text)
print(root.tag)
ns = {"i": "urn:example:interfaces"}
print(root.find("i:interface/i:name", ns).text)

Reference code (not run in the simulator). Output below is from standard Python 3.

{urn:example:interfaces}interfaces
Gi0/1

If find returns None on a document that clearly has the element, a namespace is the usual reason. You then get AttributeError: 'NoneType' object has no attribute 'text', the same shape of error you saw with missing return values.

Broken XML

A wrong closing tag gives a ParseError that names the position. Catch it and print the message:

import xml.etree.ElementTree as ET

bad = "<device><name>R1</nme></device>"
try:
    ET.fromstring(bad)
except ET.ParseError as e:
    print("XML error:", e)

Reference code (not run in the simulator). Output below is from standard Python 3.

XML error: mismatched tag: line 1, column 18

Where you will meet XML: NETCONF

A NETCONF request is an XML message. This one asks a device for its running configuration. It is shown only to help you recognise the shape; the NETCONF module of this course goes deeper.

<rpc message-id="101">
  <get-config>
    <source><running/></source>
  </get-config>
</rpc>

JSON, YAML or XML? A quick comparison

QuestionJSONYAMLXML
Commentsnoyesyes (between comment markers)
Structure shown bybraces and bracketsindentationopening and closing tags
Attributesnonoyes
Typical useREST, RESTCONFAnsible, variablesNETCONF
Python modulejsonyaml (PyYAML)xml.etree.ElementTree

Common mistakes. (1) A mismatched or misspelled closing tag, which is case-sensitive. (2) Forgetting that every value read from XML is text. (3) Searching without the namespace and getting None. (4) Parsing untrusted XML from the internet with unsafe settings; in production use a hardened parser for external input.

Exam trap. XML elements have opening and closing tags; attributes live inside the opening tag. NETCONF uses XML; RESTCONF can use XML or JSON. XML can carry attributes, JSON cannot.

The None that was a namespace

An engineer parsed a NETCONF reply and root.find("interface") returned None even though the text clearly contained interface elements. The reply declared a namespace on the root, so the real tag was {urn:example:interfaces}interface. After passing the namespace mapping to find, the script read all interfaces.

Lesson: when an element that is obviously there cannot be found in XML, suspect a namespace.

"How would you read values from an XML reply in Python?"

I parse it with xml.etree.ElementTree.fromstring, then use find and findall with a path such as ./interfaces/interface and read .text for values and .attrib for attributes. I convert text to numbers myself, and if the document has a namespace I pass a namespace dictionary to find.

Key takeaways

  • XML is a tree of elements with opening and closing tags; attributes sit in the opening tag.
  • Tags must be properly nested, with one root, and are case-sensitive.
  • ElementTree.fromstring, find, findall, .text and .attrib read it.
  • All values from XML are text; convert numbers yourself.
  • A namespace is the usual reason find returns None; NETCONF uses XML.
08

Troubleshooting workflow: a data file that will not load

When automation fails, a surprising share of the time the cause is not the network and not the script logic. It is a data file with a small syntax mistake, or a data file that is valid but does not match what the code expects. This chapter gives you one workflow that works for JSON, YAML and XML, and uses the two labs as worked examples.

The five-step workflow

  1. Read the last line of the traceback. Is it a *parse* error (JSONDecodeError, YAMLError, ParseError) or a *use* error (KeyError, TypeError)?
  2. Parse errors: run a validator and read the line and column. Fix one thing, validate again.
  3. Use errors: the file loaded. Print what you loaded and compare the real keys with the keys your code asks for.
  4. Check the types: a number that became text, a version that lost a zero, a no that became False.
  5. Re-run the whole script and check the output, not only the validator.
Parse errorfile is not validUse errorvalid, wrong shapeTypesloaded as another typeVerifyrun the script again

Two kinds of failure

Step 1: parse error or use error?

The difference decides where to look. A parse error means the file is broken. A use error means the file is fine and the code is wrong. In lab 1 you met both in one ticket:

Where it failsLast line of the tracebackMeaning
json.load(f)JSONDecodeError: Expecting value: line 15 column 18The file is not valid JSON
dev['mgmt']KeyError: 'mgmt'The file is valid; the key is spelled differently
yaml.safe_load(f)YAMLError: bad indentation of a sequence entryThe file is not valid YAML

Step 2: build a small checker

You can wrap the loaders so that every error becomes a readable line with a position. This is a handy script to keep. Parse errors in the JSON module carry the position as lineno and colno; YAML errors carry a problem_mark that counts lines from zero.

import json
import yaml

def check_json(text):
    try:
        json.loads(text)
        return "JSON OK"
    except json.JSONDecodeError as e:
        return f"JSON error at line {e.lineno}, column {e.colno}: {e.msg}"

def check_yaml(text):
    try:
        yaml.safe_load(text)
        return "YAML OK"
    except yaml.YAMLError as e:
        mark = getattr(e, "problem_mark", None)
        where = f"line {mark.line + 1}, column {mark.column + 1}" if mark else "unknown position"
        return f"YAML error at {where}"

print(check_json('{"name": "R1", "vlans": [10, 20]}'))
print(check_json('{"name": "R1" "vlans": [10, 20]}'))
print(check_json("{'name': 'R1'}"))
print(check_yaml("vlans:\n  - 10\n  - 20\n"))
print(check_yaml("vlans:\n  - id: 10\n    name: OFFICE\n   - id: 20\n"))

Reference code (not run in the simulator). Output below is from standard Python 3.

JSON OK
JSON error at line 1, column 15: Expecting ',' delimiter
JSON error at line 1, column 2: Expecting property name enclosed in double quotes
YAML OK
YAML error at line 4, column 4

The checker names the line and the column for each broken input. Notice that the missing comma between the two keys and the single-quoted keys are reported with a position, exactly like the lab validator.

Step 3: valid, but wrong shape

When the file loads and the code still fails, print the real structure. A pretty dump is the fastest way to see it:

import json

data = json.loads('{"devices": [{"name": "HYD-CORE1", "mgmt_ip": "10.99.0.1"}]}')
print(json.dumps(data, indent=2))
dev = data["devices"][0]
print(list(dev.keys()))
print(dev.get("mgmt"))
print(dev.get("mgmt_ip"))

Real output, checked in the lab terminal.

{
  "devices": [
    {
      "name": "HYD-CORE1",
      "mgmt_ip": "10.99.0.1"
    }
  ]
}
['name', 'mgmt_ip']
None
10.99.0.1

list(dev.keys()) shows the real key names, and dev.get("mgmt") returns None instead of crashing. Compare the names with the ones your code uses. This is exactly how the lab's mgmt versus mgmt_ip mismatch is found.

Step 4: check the types after loading

If a value looks right in the file but behaves wrongly in the script, print its type.

import json
import yaml

j = json.loads('{"vlan": "10", "enabled": "true"}')
y = yaml.safe_load('vlan: 10\nenabled: true\n')
print(j["vlan"] + j["vlan"])
print(y["vlan"] + y["vlan"])
print(type(j["enabled"]).__name__, type(y["enabled"]).__name__)

Real output, checked in the lab terminal.

1010
20
str bool

In JSON the quoted "10" is text, so adding it to itself glues the text together (1010). In YAML the unquoted 10 is a number, so the sum is 20. And "true" in quotes is a string, while the unquoted true is a real boolean. When the numbers are off, look for quotes.

Lab 1 and lab 2 as one checklist

Lab 1 (JSON): json validate showed a single-quote error on line 15, then a trailing comma before } on line 25 (the comma on line 24). After both were fixed the script still failed with KeyError: 'mgmt' because the key in the file is mgmt_ip. The second VLAN of the first device, jq '.devices[0].vlans[1]' inventory.json, is 20.

Lab 2 (YAML): yamllint vlans.yml reported line 11, column 4: bad indentation of a sequence entry. After aligning the dash with the other items, build.py printed VLANs 10, 20 and 30 for both switches. Adding - id: 40 / name: GUEST and a third switch changed only the data. to_json.py then wrote vlans.json, and jq '.vlans[1].name' vlans.json returned "SCANNERS".

Converting safely

Conversion between formats is load in one module, dump in another. A good check is a round trip: dump to JSON, load it again, and compare with the original.

import json
import yaml

plan = yaml.safe_load("site: Nagpur-WH\nvlans:\n  - id: 10\n    name: OFFICE\n")
text = json.dumps(plan)
again = json.loads(text)
print(text)
print("Round trip identical:", plan == again)

Real output, checked in the lab terminal.

{"site": "Nagpur-WH", "vlans": [{"id": 10, "name": "OFFICE"}]}
Round trip identical: True

Differences you may notice: JSON has no comments, so YAML comments are lost; and a YAML value that was converted (no to False) is already converted before it reaches JSON.

Tools to keep in your pocket

ToolUse
python3 -m json.tool file.jsonValidate and pretty-print JSON
jq '.path' file.jsonPull a value out of JSON
yamllint file.ymlCheck YAML syntax and style
print(json.dumps(data, indent=2))See the real structure inside Python
type(x)See what type a value really has

Common mistakes. (1) Fixing the validator's complaint and assuming the job is done, without re-running the script. (2) Editing the data to match the code when the code should change (or the other way around) without agreeing on which side is correct. (3) Ignoring the line above the reported line. (4) Trusting a value's appearance in the file rather than checking its type after loading.

Exam trap. A parse error and a KeyError are different problems. Parse errors point to the file; KeyError points to a mismatch between data and code. The exam may also ask which format forbids comments (JSON) and which forbids tabs (YAML).

Two errors, one ticket

A monitoring team reported that the device table script "broke after a file edit". The first error was a JSON decode error caused by a trailing comma. The engineer fixed it and closed the ticket. An hour later the script failed again with a KeyError, because the same edit also renamed a key. Only a full re-run after the first fix would have caught it. The team added a step: validate, run, and compare the output to a known-good sample.

Lesson: after every fix, run the whole script again; the first error often hides the second.

"A data file that worked yesterday now breaks your script. What do you do?"

I read the last line of the traceback to decide whether it is a parse error or a use error. For parse errors I run a validator to get the line and column and check the line above as well. For use errors I print the loaded structure and compare key names and types. Then I re-run the script end to end, because fixing one error can reveal another, and I add a validation step to the pipeline.

Key takeaways

  • Decide first: parse error (the file) or use error (the code or the key names).
  • Validators report one mistake at a time; fix and repeat, and check the line above.
  • Print the structure with json.dumps(data, indent=2) and check type(x) when values misbehave.
  • Quotes change types in JSON and YAML; unquoted no and 1.10 surprise in YAML.
  • Re-run the full script after each fix.
09

Summary and exam checklist

You can now read, write, load and repair JSON, YAML and XML, and you know why automation keeps its data in files. This chapter collects the facts for revision and for the exam.

Can-do checklist

Tick an item only if you can do it without looking:

  • Explain what structured data is and why a program prefers it to CLI text.
  • Write a valid JSON object with strings, numbers, booleans, null, an array and a nested object.
  • List the JSON rules: double quotes, no trailing comma, no comments, lowercase true, false, null.
  • Load JSON with json.load / json.loads and write it with json.dump / json.dumps.
  • Follow a path such as data["devices"][0]["vlans"][1] and the same path in jq.
  • Write YAML mappings, sequences and nested lists of mappings with correct indentation.
  • Explain why YAML needs quotes around no, 1.10 and 010, and why tabs are forbidden.
  • Load YAML safely with yaml.safe_load and explain why safe_load is the habit.
  • Read an XML tree: elements, attributes, root, and use find, findall, .text, .attrib.
  • Diagnose a file that will not load: parse error versus use error, validator, one fix at a time, re-run.

Side-by-side cheat-sheet

JSONYAMLXML
Container for a dictionary{"k": "v"}k: velement with child elements
Container for a list[1, 2]- 1 on each linerepeated child elements
Stringsalways double quotesquotes optionalelement text
Commentsnot allowed#allowed between comment markers
Booleanstrue / falsetrue / false (and yes / no in older loaders)text such as true
Python modulejsonyamlxml.etree.ElementTree
Network homeREST, RESTCONFAnsible, variablesNETCONF

Python quick reference

TaskCode
JSON text to datajson.loads(text)
JSON file to datajson.load(f)
Data to JSON textjson.dumps(data, indent=2)
Data to JSON filejson.dump(data, f, indent=2)
YAML to datayaml.safe_load(f)
Data to YAML textyaml.safe_dump(data)
XML text to treeET.fromstring(text)
Read XML valueroot.find("name").text

Mini glossary

structured data
Information with labels and a fixed shape, readable by programs.
object / mapping
Labelled values: a JSON object, a YAML mapping, a Python dictionary.
array / sequence
Ordered items: a JSON array, a YAML sequence, a Python list.
scalar
A single plain value such as text, a number, a boolean or null.
validator
A tool that checks a file against the format rules and names the first mistake.
jq
A command-line tool that extracts values from JSON using a path.
yamllint
A tool that checks YAML syntax and style.
namespace
A label in XML saying which vocabulary a tag belongs to.
source of truth
The one place where the intended state of the network is recorded.
round trip
Dump data to a format, load it back and check it is identical.

Most tested facts

  • JSON has no comments, no trailing commas, no single quotes; keys are double-quoted strings.
  • JSON true, false, null become Python True, False, None.
  • loads and dumps work with strings; load and dump work with files.
  • YAML structure comes from space indentation; tabs are not allowed; comments start with #.
  • Always load YAML with yaml.safe_load.
  • XML has opening and closing tags, one root, attributes in the opening tag, and is used by NETCONF.
  • KeyError points to a key mismatch; JSONDecodeError and YAMLError point to a broken file.
  • List indexes start at 0, so [1] is the second item.

The two labs in one paragraph each

Lab 1, JSON. inventory.json fails with a single-quote error on line 15 and then a trailing comma on line 24 before the closing brace on line 25; after both fixes show_inv.py fails on KeyError: 'mgmt', which you cure by using mgmt_ip. Then jq '.devices[0].vlans[1]' gives 20.

Lab 2, YAML. vlans.yml has a misaligned dash on line 11; after the fix build.py prints VLANs 10, 20 and 30 for both switches, you add VLAN 40 (GUEST) and switch WH-SW3 in the data only, convert to vlans.json with json.dump(plan, f, indent=2), and jq '.vlans[1].name' returns "SCANNERS".

What comes next

The next module, parsing, handles the opposite problem: devices that only give text. You will turn show output into the dictionaries and lists you now know how to save as JSON and YAML.

The pipeline that rejected a pull request

A team added a validation step to its repository: every JSON and YAML file must pass a validator before a merge. The first week it blocked three pull requests, one for a trailing comma and two for YAML indentation. None of the three would have reached production anyway, but each would have broken a nightly job and cost an hour. After a month nobody missed the manual checking.

Lesson: validate data files automatically, the same way you test code.

"Why does network automation use YAML for variables and JSON for APIs?"

YAML is easy for humans to write and review and allows comments, so it suits files that engineers maintain, such as Ansible variables and inventories. JSON is strict and simple for machines to produce and parse, so it suits API traffic. Both load into the same dictionaries and lists, so the code does not care which one it received.

Key takeaways

  • JSON, YAML and XML are three notations for dictionaries and lists.
  • JSON is strict; YAML is human-friendly and indentation-based; XML is tag-based and used by NETCONF.
  • Parse errors mean a broken file; KeyError means a mismatch between data and code.
  • Keep data in files and logic in code; load YAML with safe_load.
  • Validate, fix one error at a time, and re-run the whole script.
๐ŸŽ“ For educational purposes only โ€” all devices are simulationsTerms of UsePrivacy Policyยฉ 2026 Network Kings
CONFIG by Network Kings โ€” an educational IT simulation platform for learning purposes only. It is not Cisco IOS, Junos, FortiOS or PAN-OS and contains no Cisco, Juniper, Fortinet or Palo Alto Networks software. Cisco, IOS, CCNA, CCNP, Juniper, JNCIA, JNCIS, JNCIP, Fortinet, FortiGate, FortiOS, NSE, Palo Alto Networks, PAN-OS and PCNSE are trademarks of their respective owners. Network Kings is not affiliated with or endorsed by Cisco Systems, Inc., Juniper Networks, Inc., Fortinet, Inc. or Palo Alto Networks, Inc.