Jump to chapter (9)
Welcome to Ansible: what you will learn and the big picture
What you will learn in this module. This is your first step into network automation with Ansible, and you do not need any earlier experience. By the end you will have built a working control node, described three real devices in an inventory, talked to them with ad-hoc commands, written a playbook, dry-run it, applied it, and proved that running it again changes nothing. You will also learn the short list of errors every beginner meets and how to read them.
- What Ansible is, what "agentless" means, and why network engineers use it.
- The control node,
ansible-core, collections and fully qualified module names. - An inventory with groups and the variables that tell Ansible how to log in.
ansible.cfgand the SSH host key question.- Ad-hoc commands:
ping,ios_command,ios_facts,ios_config. - YAML and your first playbook, line by line, with
--check,--diffand idempotency.
Prerequisites. You should be able to type basic Linux commands (cd, ls, cat) and know what an SSH login and an ntp server line on a router are. You do not need Python, YAML or any Ansible knowledge. Everything is explained from zero.
Analogy: a tireless assistant with a checklist
Imagine you must add the same NTP server to fifty routers tonight. By hand you would open fifty SSH sessions, type the same line fifty times and hope you never mistype. Now imagine a tireless assistant who reads your checklist, logs in to each router exactly as you would, types the commands, and reports "done" or "already there" for every device. That assistant is Ansible. You write the checklist (a playbook) and the list of devices (an inventory); Ansible does the typing.
Ansible runs on one Linux box and logs in to every device over SSH, the same way you do.
The words you will meet
- Control node
- The Linux machine where Ansible is installed and where you run commands. In our lab it is AUTO01.
- Managed node
- Any device Ansible manages: here two routers and a switch.
- Agentless
- Nothing is installed on the managed node. Ansible only needs the SSH login that you already use.
- Inventory
- A file that lists the devices and how to reach them.
- Module
- A ready-made worker for one job, for example
cisco.ios.ios_configsends configuration lines. - Task, play, playbook
- A task calls one module. A play says which hosts run a list of tasks. A playbook is a file holding one or more plays.
- Idempotent
- Running the same playbook twice gives the same end state, and the second run changes nothing.
Worked example. The Network Kings lab has R1 in Pune (10.99.0.1), R2 in Bengaluru (10.99.0.2) and the access switch SW1 (10.99.0.10). Adding ntp server 10.99.0.60 by hand means three logins. With Ansible it is one command: ansible-playbook ntp.yml. With three devices the saving is small; with three hundred it is the difference between a night and a minute.
Why network engineers like Ansible
- No agent, no daemon: routers and switches cannot run agents anyway. Ansible uses SSH and the normal CLI.
- Readable: a playbook is plain YAML text that a colleague can review like a change ticket.
- Safe by design: you can dry-run, you can limit a run to one device, and good modules only send what is missing.
- One tool, many vendors: the same ideas work for Cisco IOS, Junos, NX-OS and Linux servers.
Common beginner mistake. Thinking Ansible must be installed on the routers, or that it "pushes a program" to them. It does not. If you can SSH to a device from the control node, Ansible can manage it. Another trap: running a playbook against all devices on day one. You will learn to limit and dry-run first.
Exam trap. Interviewers ask "is Ansible agent-based?" The answer is no: it is agentless and push-based over SSH (or an API). The control node must be Linux or macOS; Windows is not supported as a control node.
The 2 AM NTP change
A bank branch network of forty routers had a retired NTP server. The engineer on call logged in to each router by hand. By router 28 he mistyped an IP address, and logs on that router were stamped with the wrong time for a week. The team later wrote a playbook with one NTP variable. The next time the server changed, one edit and one run updated all forty devices and the recap showed which ones were already correct.
Lesson: automation is not about speed alone; it removes typing mistakes and gives you a record of what changed.
"What is Ansible and why is it agentless?"
Ansible is an open-source automation tool. You describe the desired state in YAML playbooks and an inventory; the control node connects over SSH (or an API), runs modules, and reports ok, changed or failed per device. It is agentless because network devices cannot host an agent, so it reuses the management access you already have. Mention idempotency and the inventory/playbook split to show you understand the model.
Key takeaways
- Ansible runs on a Linux control node and manages devices over SSH with nothing installed on them.
- The inventory says which devices; the playbook says what to do.
- A module does one job; a task calls a module; a play targets hosts; a playbook holds plays.
- Idempotent means a second run changes nothing.
- Our lab: AUTO01 manages R1, R2 and SW1 on 10.99.0.0/24.
Your control node: Linux, ansible-core, collections and module names
Every Ansible journey starts on one machine: the control node. Think of it as your workshop. The tools (Ansible), the instructions (playbooks) and the address book (inventory) all live here, and from here you reach out to the devices. In our lab the control node is a Linux server called AUTO01. You log in as netadmin, and the server already has Ansible installed.
What is installed, and what is not
The name "Ansible" covers two layers. ansible-core is the engine: the ansible and ansible-playbook commands, the YAML parser, the connection plugins and a small set of built-in modules such as debug and assert. The modules for specific vendors live in collections, which are downloadable bundles. For a network engineer the two that matter first are ansible.netcommon (generic network pieces such as the network_cli connection) and cisco.ios (modules for IOS devices).
The full module name (FQCN) tells Ansible exactly which collection a module comes from.
Check what you have
Two commands tell you everything. ansible --version shows the engine version, the Python it runs on and, importantly, which configuration file is active. ansible-galaxy collection list shows installed collections.
AUTO01$ ansible --version ansible [core 2.16.3] config file = /etc/ansible/ansible.cfg configured module search path = ['/home/netadmin/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules'] ansible python module location = /usr/lib/python3/dist-packages/ansible ansible collection location = /home/netadmin/.ansible/collections:/usr/share/ansible/collections executable location = /usr/bin/ansible python version = 3.12.3 (main, Sep 11 2026, 14:09:07) [GCC 13.3.0] (/usr/bin/python3) jinja version = 3.1.2 libyaml = True AUTO01$ ansible-galaxy collection list # /usr/lib/python3/dist-packages/ansible_collections Collection Version ----------------------------- ------- ansible.netcommon 5.3.0 ansible.utils 2.12.0 cisco.ios 5.3.0 community.general 8.3.0
Notice config file = /etc/ansible/ansible.cfg. This is the system-wide file. In the next chapters you will create one inside your own project folder, and the same command will show the new path.
Fully qualified collection names (FQCN)
Older tutorials write ios_config:. Current Ansible expects the fully qualified collection name: cisco.ios.ios_config. The three parts are the namespace (cisco), the collection (ios) and the module (ios_config). Using the full name avoids confusion when two collections ship a module with the same name, and it is what linters and the documentation use. Built-in modules use ansible.builtin, for example ansible.builtin.debug.
You can read the built-in help for any module without internet access:
AUTO01$ ansible-doc cisco.ios.ios_command | head -20 > CISCO.IOS.IOS_COMMAND (/usr/lib/python3/dist-packages/ansible_collections/cisco/ios/plugins/modules/ios_command.py) Run commands on remote devices running Cisco IOS OPTIONS (= is mandatory): - commands: list of show/exec commands; results in stdout (list) and stdout_lines (list of lists) All parameters: commands, wait_for, match, retries, interval
Installing on your own machine
You will not need this in the simulator, but on a real control node (a Linux VM or WSL on Windows) the usual steps are below.
# Reference configuration (not run in the simulator)
python3 -m venv ~/ansible-venv
. ~/ansible-venv/bin/activate
pip install ansible-core
ansible-galaxy collection install cisco.ios ansible.netcommon
A virtual environment keeps Ansible and its Python libraries separate from the operating system, which makes upgrades safe. Always pin versions in a real project so that your playbooks behave the same next month.
Your project folder habit
Keep each automation project in its own folder, for example ~/netauto. It will hold the inventory, ansible.cfg and playbooks together. Put the folder under version control (Git) as soon as it has a file worth keeping; the auto-git module shows how.
Common beginner mistake. Writing short module names such as ios_config copied from old blog posts. They may still work through redirects, but they hide which collection you depend on. Another trap is installing Ansible with the operating-system package and a separate pip copy; then two versions exist and ansible --version tells you which one wins.
Exam trap. "Where does Ansible look for its configuration file?" The first one found in this order: the ANSIBLE_CONFIG variable, ./ansible.cfg in the current folder, ~/.ansible.cfg, then /etc/ansible/ansible.cfg. It never merges them; only one is used.
Two Ansibles on one server
A new engineer ran a playbook that worked on a colleague's laptop but failed on the shared control node with "couldn't resolve module/action cisco.ios.ios_config". ansible --version showed an old distribution package, while the colleague had a newer one in a virtual environment. Running ansible-galaxy collection list showed that the cisco.ios collection was missing on the server.
Lesson: when behaviour differs between machines, compare ansible --version and the collection list first. Pin versions in a requirements file.
"What is the difference between ansible-core and a collection?"
ansible-core is the engine and the built-in modules. Collections are separately versioned bundles of modules, plugins and roles, for example cisco.ios for IOS devices. You call a module with its full name, namespace.collection.module, so Ansible knows where it comes from. Mentioning that collections can be updated without upgrading the engine shows real-world awareness.
Key takeaways
- The control node is a Linux machine with
ansible-coreand the collections you need. ansible --versionshows the version, Python and the active config file.ansible-galaxy collection listshows which collections are installed.- Use fully qualified names such as
cisco.ios.ios_configandansible.builtin.debug. ansible-doc MODULEgives offline help with the options.
The inventory: telling Ansible which devices exist and how to log in
Before Ansible can touch a router it must know three things: what the device is called, where it is (its IP address) and how to log in. All three go into the inventory. If you have ever kept a spreadsheet of devices with columns for name, management IP and username, you already understand an inventory. Ansible simply reads it as a text file.
Create the project folder and the file
Inside the lab terminal we create files with cat and a here-document: everything between <<'EOF' and the closing EOF line is written into the file. On a real server you would use nano or vim; the file content is the same. The quotes around 'EOF' stop the shell from changing characters such as $ or @.
AUTO01$ mkdir ~/netauto && cd ~/netauto AUTO01$ cat > inventory.ini <<'EOF' [core] pun-core-r1 ansible_host=10.99.0.1 [edge] blr-edge-r2 ansible_host=10.99.0.2 [switches] pun-acc-sw1 ansible_host=10.99.0.10 [routers:children] core edge [all:vars] ansible_connection=ansible.netcommon.network_cli ansible_network_os=cisco.ios.ios ansible_user=netops ansible_password=NK@2026 EOF
Reading it line by line
[core]starts a group named core. Lines below it are the members, until the next bracket line or a blank line.pun-core-r1 ansible_host=10.99.0.1is one host. The first word,pun-core-r1, is the inventory name: the label Ansible prints in its output and that you type in commands. It does not have to match the device hostname.ansible_host=10.99.0.1is the real address Ansible connects to. If you leave it out, Ansible tries to resolve the inventory name through DNS.[routers:children]means "this group is made of other groups". The lines below are group names, not hosts. Sorouterscontains everything incoreandedge.[all:vars]holds variables that apply to every host.allis a built-in group that always contains every host.
The four variables in [all:vars] are the heart of network automation:
| Variable | Meaning |
|---|---|
ansible_connection | How to talk to the device. ansible.netcommon.network_cli means "log in with SSH and type CLI commands". |
ansible_network_os | Which CLI dialect: prompts, paging and config mode. cisco.ios.ios is IOS. |
ansible_user | The SSH username on the device. |
ansible_password | The SSH password. In this lab the password is the lab-only value from the lab brief. |
Groups can contain hosts or other groups. Everything is inside the built-in group all.
Ask Ansible what it understood
Never assume the file says what you meant. ansible-inventory --graph draws the tree Ansible built:
AUTO01$ ansible-inventory -i inventory.ini --graph @all: |--@switches: | |--pun-acc-sw1 |--@routers: | |--@core: | | |--pun-core-r1 | |--@edge: | | |--blr-edge-r2
-i inventory.ini names the inventory file. In the next chapter a configuration file will remember it so that you can leave -i out. ansible-inventory --list prints every host with its variables as JSON:
AUTO01$ ansible-inventory --list
{
"_meta": {
"hostvars": {
"blr-edge-r2": {
"ansible_connection": "ansible.netcommon.network_cli",
"ansible_host": "10.99.0.2",
"ansible_network_os": "cisco.ios.ios",
"ansible_password": "NK@2026",
"ansible_user": "netops"
},
...
Every host received all four variables from [all:vars], plus its own ansible_host. That is variable inheritance at work, and a whole module of this course is devoted to it.
Common beginner mistake. Putting a real production password in a plain file. Here it is acceptable only because the lab password is a throw-away value. In real projects you keep secrets in an encrypted Ansible Vault file or fetch them from a secrets manager; the later modules show how. Also never commit such a file to a shared Git repository.
Exam trap. The inventory name and the device hostname are different things. pun-core-r1 is only a label. Another classic: [routers:children] lists groups, while [routers] lists hosts. Mixing them up makes Ansible report an unknown host.
The router that had two names
An engineer saw "pun-core-r1" in Ansible output but the device prompt said "PUN-CORE-R1". He feared Ansible was talking to a different box. A quick ansible-inventory --list showed ansible_host: 10.99.0.1 for that label, which is the address of the core router. The label in the inventory is chosen by humans; the device hostname comes from the router's own configuration.
Lesson: Ansible connects to ansible_host, not to the name. Pick short, consistent inventory names and document them.
"What is an Ansible inventory and which variables does a network device need?"
An inventory lists hosts and groups and carries their variables. A network device over SSH needs ansible_host, ansible_connection (network_cli), ansible_network_os (for example cisco.ios.ios), ansible_user and a password or key. Add that you check the result with ansible-inventory --graph and that secrets belong in Vault.
Key takeaways
- The inventory names the hosts, groups them, and holds the login variables.
ansible_hostis the address; the first word on the line is only a label.[group:children]builds a group from other groups;[all:vars]applies to everyone.- Network devices need
ansible_connection,ansible_network_os, user and password. - Verify with
ansible-inventory --graphand--list.
ansible.cfg and the host key question every beginner meets
You have an inventory. Typing -i inventory.ini in every command would be tiring, and you also want a few project settings kept in one place. That place is ansible.cfg, a small text file. In this chapter you create it, then you meet the first real error of almost every network automation learner: the host key failure. Understanding it once saves hours later.
Which configuration file does Ansible use?
Ansible looks for a configuration file in a fixed order and uses the first one it finds; it does not merge files.
- The path in the
ANSIBLE_CONFIGenvironment variable. ansible.cfgin the current folder (the folder you run the command from).~/.ansible.cfgin your home folder./etc/ansible/ansible.cfg, the system default (this is what we saw in chapter 2).
The habit that works best is one ansible.cfg per project folder. Then the project carries its own settings and you always run commands from inside that folder.
AUTO01$ cd ~/netauto AUTO01$ cat > ansible.cfg <<'EOF' [defaults] inventory = inventory.ini EOF AUTO01$ ansible --version
ansible [core 2.16.3]
config file = /home/netadmin/netauto/ansible.cfg
configured module search path = ['/home/netadmin/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules']
The path changed from /etc/ansible/ansible.cfg to your project file, which proves Ansible now uses it. The line inventory = inventory.ini under [defaults] means you can drop -i: a plain ansible-inventory --graph now reads the file for you. If you run a command from a different folder, Ansible will not find this file and you will see a warning such as the one below, which is a very common source of confusion.
AUTO01$ ansible all -m ping [WARNING]: provided hosts list is empty, only localhost is available. Note that the implicit localhost does not match 'all' [WARNING]: No hosts matched, nothing to do
"No hosts matched" almost always means one of three things: you are in the wrong folder, the inventory path is wrong, or the pattern names a group that does not exist.
The first ping fails: known_hosts
Now try the first real connection. ping here is not an ICMP ping; it logs in over SSH and checks that Ansible can talk to the device (more in the next chapter).
AUTO01$ ansible all -m ping
pun-core-r1 | FAILED! => {
"changed": false,
"msg": "Server '10.99.0.1' not found in known_hosts"
}
blr-edge-r2 | FAILED! => {
"changed": false,
"msg": "Server '10.99.0.2' not found in known_hosts"
}
pun-acc-sw1 | FAILED! => {
"changed": false,
"msg": "Server '10.99.0.10' not found in known_hosts"
}
Nothing is wrong with the network, the user or the password. The message is about trust.
The first connection to a device must be trusted by someone. Ansible cannot answer that question interactively.
Why it happens
The first time you SSH to any device, your SSH client asks "The authenticity of host cannot be established. Continue?" and, if you say yes, stores the device's host key in ~/.ssh/known_hosts. This protects you from someone impersonating the router. Ansible runs without a person to answer, so with host key checking turned on (the default) it simply refuses to continue.
Two fixes: lab and production
In a lab it is normal to turn checking off in the project file:
AUTO01$ echo 'host_key_checking = False' >> ansible.cfg AUTO01$ cat ansible.cfg [defaults] inventory = inventory.ini host_key_checking = False
AUTO01$ ansible all -m ping
pun-core-r1 | SUCCESS => {
"changed": false,
"ping": "pong"
}
blr-edge-r2 | SUCCESS => {
"changed": false,
"ping": "pong"
}
pun-acc-sw1 | SUCCESS => {
"changed": false,
"ping": "pong"
}
Three SUCCESS lines with pong: Ansible can log in to every device and talk to its CLI. Your control node is ready.
In production keep checking on and collect the keys once, after verifying the device fingerprint through a trusted channel:
# Reference configuration (not run in the simulator)
ssh-keyscan 10.99.0.1 >> ~/.ssh/known_hosts
ssh-keyscan 10.99.0.2 >> ~/.ssh/known_hosts
ssh-keyscan 10.99.0.10 >> ~/.ssh/known_hosts
Better still, a team can maintain one known_hosts file in its repository and review changes to it.
Common beginner mistake. Copying host_key_checking = False into a production project because "it fixed the error". It also removes protection against a man-in-the-middle device. Keep it for labs only. Another trap: forgetting that the setting must be under the [defaults] heading; a line placed above it is ignored.
Exam trap. Remember what ansible --version reports: the active configuration file. If your changes seem to have no effect, you are probably editing a file that is not the active one, for example the project file while running from your home folder.
It worked yesterday on the other server
A trainee copied a playbook project to a new control node. Every task failed with "not found in known_hosts" even though the same project worked on the old server. The old server had the device keys in its ~/.ssh/known_hosts from months of manual logins; the new one had never connected. The senior engineer verified the fingerprints, ran ssh-keyscan for each address and the next run was clean, with host key checking still on.
Lesson: the error is about trust, not connectivity. Fix it by establishing trust, not by switching security off.
"Ansible fails with 'not found in known_hosts'. What do you do?"
I explain that this is SSH host key verification, not a network fault. In a lab I may set host_key_checking = False. In production I verify the device fingerprint, add the key with ssh-keyscan or a managed known_hosts file, and keep checking enabled. That answer shows security awareness.
Key takeaways
- Ansible uses the first configuration file it finds;
ansible --versionshows which. - Keep one
ansible.cfgper project and run commands from that folder. inventory = inventory.iniunder[defaults]removes the need for-i.- "not found in known_hosts" is a trust problem; labs may disable checking, production should add the keys.
- Three SUCCESS lines with
pongprove login and CLI access to every device.
Ad-hoc commands: your first conversations with the devices
Before you write a playbook, it helps to say a single sentence to many devices and see their replies. That is what an ad-hoc command does: one module, one run, no file. Think of it as shouting one question across a room of routers: "Everyone, show me your interfaces", and each router answers. Ad-hoc commands are the quickest way to check connectivity, collect a fact, or make one tiny change.
The shape of an ad-hoc command
ansible PATTERN -m MODULE -a "ARGUMENTS"
# who which worker what to give it
- PATTERN chooses the hosts:
all, a group such asrouters, a single inventory name, or a combination such asrouters:!edge(routers except the edge group). The next module explores patterns in depth. - -m names the module, preferably by its full name.
- -a passes arguments to the module as
key=valuepairs inside quotes.
1. ping: can Ansible log in?
You met this in the last chapter. For a network device the ping module does not send ICMP; it opens the SSH session and checks the CLI. The flag -o prints each result on one line, handy for many devices.
AUTO01$ ansible all -m ping -o
pun-core-r1 | SUCCESS => { "changed": false, "ping": "pong" }
blr-edge-r2 | SUCCESS => { "changed": false, "ping": "pong" }
pun-acc-sw1 | SUCCESS => { "changed": false, "ping": "pong" }
2. ios_command: run show commands
cisco.ios.ios_command runs exec-level commands such as show and returns the text. It never changes configuration. We ask only the routers group, so the switch stays silent.
AUTO01$ ansible routers -m cisco.ios.ios_command -a "commands='show ip interface brief'"
pun-core-r1 | SUCCESS => {
"changed": false,
"stdout": [
"Interface IP-Address OK? Method Status Protocol\nGigabitEthernet0/0 10.99.0.1 YES manual up up\nGigabitEthernet0/1 unassigned YES unset administratively down down\n..."
],
"stdout_lines": [
[
"Interface IP-Address OK? Method Status Protocol",
"GigabitEthernet0/0 10.99.0.1 YES manual up up",
"GigabitEthernet0/1 unassigned YES unset administratively down down",
"GigabitEthernet0/2 unassigned YES unset administratively down down",
"GigabitEthernet0/3 unassigned YES unset administratively down down"
]
]
}
blr-edge-r2 | SUCCESS => { ... same shape, Gi0/0 is 10.99.0.2 ... }
(Output shortened: the stdout text is cut and the second router is summarised.) Two fields hold the text. stdout is a list with one string per command; stdout_lines is the same text split into lines, a list of lists. You will use both when you write playbooks that react to output.
3. ios_facts: collect facts
Facts are details about a device that Ansible gathers and names for you.
AUTO01$ ansible pun-core-r1 -m cisco.ios.ios_facts
pun-core-r1 | SUCCESS => {
"ansible_facts": {
"ansible_net_all_ipv4_addresses": [
"10.99.0.1"
],
"ansible_net_api": "cliconf",
"ansible_net_hostname": "PUN-CORE-R1",
...
"ansible_net_model": "ISR4331",
...
"ansible_net_system": "ios",
"ansible_net_version": "2.0"
},
"changed": false
}
(Output shortened with dots.) Every fact starts with ansible_net_. Later chapters use them in conditions, for example "only run this on newer software".
4. ios_config: make one small change
cisco.ios.ios_config sends configuration lines. Use parents for the interface or section the lines belong to. Here we label a spare port on the Pune router.
AUTO01$ ansible pun-core-r1 -m cisco.ios.ios_config -a "parents='interface GigabitEthernet0/3' lines='description SPARE-PORT'" pun-core-r1 | CHANGED => { "banners": {}, "changed": true, "commands": [ "interface GigabitEthernet0/3", "description SPARE-PORT" ], "updates": [ "interface GigabitEthernet0/3", "description SPARE-PORT" ] } AUTO01$ # run the exact same command again pun-core-r1 | SUCCESS => { "banners": {}, "changed": false, "commands": [], "updates": [] }
Look at the second run: the state is SUCCESS, changed is false and commands is empty. The module compared the running-config with your line, found it already present and sent nothing. This is idempotency in action, and it is why Ansible is safe to repeat.
Reading the result words
| Result | Meaning |
|---|---|
| SUCCESS | Ran, nothing changed (a read, or the device already matched). |
| CHANGED | Ran and modified the device. |
| FAILED | Logged in, but the task itself failed. Read msg. |
| UNREACHABLE | Could not connect at all (address, timeout, SSH). |
pun-core-r1 | FAILED! => {
"changed": false,
"msg": "command 'show ip bogus' failed: % Invalid input detected at '^' marker."
}
blr-edge-r2 | UNREACHABLE! => {
"changed": false,
"msg": "ssh connect failed: Connection timed out",
"unreachable": true
}
The first is a device-side error: the CLI rejected the command. The second came from a wrong ansible_host address, so nothing answered. Telling these two apart is the first step in every troubleshooting session.
Common beginner mistake. Using ad-hoc ios_config for big changes. It leaves no file to review, no history and no easy repeat. Use ad-hoc commands for checks and tiny fixes; put real changes in playbooks. Also remember quotes: the whole -a value goes in double quotes and the inner command in single quotes.
Exam trap. ansible all -m ping on a network device is a login test, not an ICMP reachability test, and the module that runs show commands is ios_command, while ios_config changes configuration.
Is the whole estate reachable before the change window?
Two hours before a maintenance window, a lead ran ansible all -m ping -o across eighty devices. Seventy-seven answered pong. Three printed UNREACHABLE with "Connection timed out"; all three belonged to one site whose firewall had a stale management rule. The network team fixed the rule, repeated the ping, and the change started on time with a clean baseline.
Lesson: a one-line ad-hoc ping is a cheap pre-flight check. Run it before every change.
"When would you use an ad-hoc command instead of a playbook?"
For quick, one-off, low-risk jobs: testing login with ping, collecting show output with ios_command, or reading facts. Anything that changes configuration repeatedly, needs review or must be repeatable belongs in a playbook under version control. Mention that ios_config is still idempotent even ad hoc.
Key takeaways
- Format:
ansible PATTERN -m MODULE -a "ARGS". pingtests SSH login and CLI, not ICMP;ios_commandreads;ios_configchanges;ios_factsgathers facts.stdoutis a list of strings;stdout_linesis a list of line lists.- CHANGED means it modified the device; a repeat gives SUCCESS with
changed: false. - FAILED means a task error after login; UNREACHABLE means no connection.
YAML and your first playbook, line by line
Ad-hoc commands are conversations. A playbook is a written plan. It is a text file that says "on these devices, do these things", and because it is a file you can read it, review it, store it in Git and run it again next month. Playbooks are written in YAML, a format designed to be easy for people to read. If you can read a shopping list with sub-items, you can read YAML.
YAML in five rules
- Key and value.
hosts: allis a key (hosts), a colon, a space, and a value (all). The space after the colon is required. - Indentation means structure. Lines indented under a key belong to that key. Use spaces only, never tabs, and keep the same number of spaces for items at the same level (two is the convention).
- Lists use a dash. Each item starts with
-(dash and space). A list underlines:holds the configuration lines. - Comments start with #. Ansible ignores everything after it on that line.
- Three dashes
---at the top mark the start of the document. Optional, but a good habit.
Your first playbook
Our goal: every device uses 10.99.0.60 (the control node) as NTP and syslog server. Create ntp.yml in ~/netauto:
AUTO01$ cat > ntp.yml <<'EOF'
---
- name: NTP and syslog baseline
hosts: all
gather_facts: false
tasks:
- name: NTP server
cisco.ios.ios_config:
lines:
- ntp server 10.99.0.60
- name: Syslog server
cisco.ios.ios_config:
lines:
- logging host 10.99.0.60
- logging trap informational
EOF
Reading it line by line
| Line | What it means |
|---|---|
--- | Start of the YAML document. |
- name: NTP and syslog baseline | The dash starts one play. name is a human label printed when the play runs. |
hosts: all | Which inventory hosts or groups this play targets. Try routers to touch only routers. |
gather_facts: false | Ansible normally collects facts about a Linux host first using Python on that host. Network devices cannot do that, so we turn it off. (To collect device facts, use the ios_facts module as a task.) |
tasks: | The list of jobs, run top to bottom, on every host, before moving to the next task. |
- name: NTP server | One task. A descriptive name is shown in the output, so write names you will understand at 2 AM. |
cisco.ios.ios_config: | The module the task calls. Everything indented below it is the module's arguments. |
lines: | An argument that takes a list of configuration lines. |
- ntp server 10.99.0.60 | One configuration line, typed in global configuration mode exactly as you would on the router. |
A playbook holds plays; a play targets hosts and lists tasks; each task calls one module with arguments.
Check the file before you trust it
Three read-only commands tell you what Ansible understood, without touching a device.
AUTO01$ ansible-playbook ntp.yml --syntax-check
playbook: ntp.yml
AUTO01$ ansible-playbook ntp.yml --list-tasks
playbook: ntp.yml
play #1 (all): NTP and syslog baseline TAGS: []
tasks:
NTP server TAGS: []
Syslog server TAGS: []
AUTO01$ ansible-playbook ntp.yml --list-hosts
playbook: ntp.yml
play #1 (all): NTP and syslog baseline TAGS: []
pattern: ['all']
hosts (3):
pun-core-r1
blr-edge-r2
pun-acc-sw1
--syntax-check only proves that the YAML and the structure are valid; it prints the playbook name when all is well. --list-tasks shows the plan, and --list-hosts shows exactly which devices would be touched. Get into the habit of reading --list-hosts before every real run; it would have stopped many outages.
Common beginner mistake. Wrong indentation. If lines: is aligned with the module name instead of indented below it, Ansible reads it as a separate key of the task and the playbook fails. Another classic: forgetting that - needs a space, or writing hosts all without the colon. The troubleshooting chapter shows the exact error text for these.
Exam trap. A playbook is a list of plays (the top-level dash), each play has hosts and tasks, and tasks run in order on all hosts before the next task starts (the default "linear" strategy). Also note YAML forbids tab characters for indentation.
The playbook that touched everything
A junior wrote a playbook meant for the Pune site and left hosts: all in it. Before the first run, a reviewer asked him to attach the output of --list-hosts to the change ticket. It listed 41 hosts including the Bengaluru edge. He changed the play to target the pune group, the list shrank to 12 and the change went ahead without surprises.
Lesson: hosts: is the blast radius. Check it with --list-hosts.
"Explain the structure of an Ansible playbook."
A playbook is a YAML list of plays. A play has a name, a hosts target and a list of tasks. Each task calls one module with arguments, for example cisco.ios.ios_config with lines. Tasks run in order on all targeted hosts. I also say that I validate with --syntax-check and confirm the target with --list-hosts.
Key takeaways
- YAML:
key: value, indentation with spaces,-for lists,#for comments. - Play =
hosts+tasks; task = one module + its arguments. - Set
gather_facts: falsefor network devices; useios_factswhen you need facts. ios_configwithlinessends configuration lines exactly as typed on the CLI.--syntax-check,--list-tasks,--list-hostsinspect a playbook safely.
Dry run, apply, run again: reading the recap, idempotency and saving
A good engineer never pushes a change blind. The safe rhythm for every playbook is look, rehearse, apply, repeat. First you look at the plan (last chapter). Now you rehearse with a dry run, apply for real, and run it again to prove that the devices match the playbook. At the end you will also learn why a router can lose your change on reload and how to prevent it.
Step 1: rehearse with --check and --diff
--check asks every module "what would you change?" without changing anything. --diff adds the configuration lines that would appear. Reading it is like proofreading before pressing send.
AUTO01$ ansible-playbook ntp.yml --check --diff PLAY [NTP and syslog baseline] ************************************************ TASK [NTP server] ************************************************************* --- before: running-config +++ after: running-config @@ -47,4 +47,5 @@ login local transport input ssh ! +ntp server 10.99.0.60 end changed: [pun-core-r1] ... (the same diff for blr-edge-r2 and pun-acc-sw1) TASK [Syslog server] ********************************************************** ... +logging host 10.99.0.60 +logging trap informational end changed: [pun-core-r1] ... PLAY RECAP ******************************************************************** pun-core-r1 : ok=2 changed=2 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0 blr-edge-r2 : ok=2 changed=2 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0 pun-acc-sw1 : ok=2 changed=2 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
(Output shortened with dots.) In check mode changed means "would change". The + lines are exactly what will be added. Nothing was sent: the next section proves it by running for real.
Step 2: apply for real
AUTO01$ ansible-playbook ntp.yml
PLAY [NTP and syslog baseline] ************************************************
TASK [NTP server] *************************************************************
changed: [pun-core-r1]
changed: [blr-edge-r2]
changed: [pun-acc-sw1]
TASK [Syslog server] **********************************************************
changed: [pun-core-r1]
changed: [blr-edge-r2]
changed: [pun-acc-sw1]
PLAY RECAP ********************************************************************
pun-core-r1 : ok=2 changed=2 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
blr-edge-r2 : ok=2 changed=2 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
pun-acc-sw1 : ok=2 changed=2 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
Output is grouped by task, one line per host, in the order tasks run. Every task changed every device, which is what we expected on a first run.
Reading the PLAY RECAP
- ok
- Tasks that completed without error. It includes the changed ones, so ok=2 changed=2 means two tasks ran and both changed something.
- changed
- Tasks that modified the device.
- unreachable
- Hosts Ansible could not connect to.
- failed
- Tasks that failed.
- skipped
- Tasks not run because a condition was false.
- rescued / ignored
- Failures handled by a
rescueblock or byignore_errors. You will use them in the Ansible logic module.
Step 3: run again, the idempotency test
AUTO01$ ansible-playbook ntp.yml PLAY [NTP and syslog baseline] ************************************************ TASK [NTP server] ************************************************************* ok: [pun-core-r1] ok: [blr-edge-r2] ok: [pun-acc-sw1] TASK [Syslog server] ********************************************************** ok: [pun-core-r1] ok: [blr-edge-r2] ok: [pun-acc-sw1] PLAY RECAP ******************************************************************** pun-core-r1 : ok=2 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0 blr-edge-r2 : ok=2 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0 pun-acc-sw1 : ok=2 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
Every line says ok and the recap shows changed=0. This is idempotency: the playbook describes a desired state, the devices already match, so nothing is sent. A clean second run is also your proof that the devices match the playbook. A playbook that keeps reporting changed on every run is a warning sign.
The safe rhythm for every change.
Running config is not startup config
Changes made through the CLI go to the running-config in RAM. If the router reloads before someone saves, they are lost. Ask the startup-config what it knows:
AUTO01$ ansible all -m cisco.ios.ios_command -a "commands='show startup-config | include ntp server|logging host'" -o
pun-core-r1 | SUCCESS => { "changed": false, "stdout": [ "" ], "stdout_lines": [ [ "" ] ] }
... (blr-edge-r2 and pun-acc-sw1 are also empty)
Empty: the changes exist only in RAM. The fix is save_when: modified on a task: Ansible compares running and startup configuration and saves (the equivalent of copy running-config startup-config) when they differ. Edit the Syslog task so it ends like this, then run again:
- name: Syslog server
cisco.ios.ios_config:
lines:
- logging host 10.99.0.60
- logging trap informational
save_when: modified # same indent as lines:
AUTO01$ ansible-playbook ntp.yml TASK [NTP server] ************************************************************* ok: [pun-core-r1] ... TASK [Syslog server] ********************************************************** changed: [pun-core-r1] changed: [blr-edge-r2] changed: [pun-acc-sw1] ... PLAY RECAP ******************************************************************** pun-core-r1 : ok=2 changed=1 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0 AUTO01$ ansible-playbook ntp.yml # third run: everything is saved now pun-core-r1 : ok=2 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0 AUTO01$ ansible all -m cisco.ios.ios_command -a "commands='show startup-config | include ntp server|logging host'" -o pun-core-r1 | SUCCESS => { "changed": false, "stdout": [ "logging host 10.99.0.60\nntp server 10.99.0.60" ], ... }
(Output shortened.) The task reported changed because the save itself is a change. After that, running and startup match and the run is clean again.
Common beginner mistake. Treating --check as a guarantee. It is a best-effort rehearsal: modules that cannot predict an outcome may report differently in a real run. Also, a first --check after a partial run can show changed for lines that depend on earlier tasks. Use it to catch obvious mistakes, then roll out to one device first with -l pun-core-r1.
Exam trap. In the recap, ok counts changed tasks too. Also, changed=0 on a second run is the definition of idempotent, and save_when belongs to ios_config with values such as modified or always.
The change that vanished after a power cut
A branch lost power an hour after an automated NTP rollout. When the router came back, it had no NTP server, and clocks drifted again. The playbook had worked perfectly, but nobody had asked it to save. The team added save_when: modified, re-ran, and used a startup-config check as part of the standard verification.
Lesson: "changed" in the recap describes the running-config. Decide deliberately whether and when to save.
"What does idempotent mean and how do you prove a playbook is idempotent?"
Running it repeatedly leaves the same state; the second run makes no changes. I prove it by running the playbook twice and checking that the second recap shows changed=0 for every host. For network modules this works because they compare the existing configuration and only send missing lines.
Key takeaways
- Safe rhythm: look (
--list-hosts), rehearse (--check --diff), apply, repeat. - The PLAY RECAP counts ok, changed, unreachable, failed, skipped, rescued, ignored per host.
- A clean second run with
changed=0proves idempotency and that devices match the playbook. - Running-config changes are lost on reload; use
save_when: modifiedto save. - Check mode is a best-effort rehearsal, not a guarantee.
Troubleshooting your first runs: a workflow for red text
Every beginner sees red error text within the first hour. That is normal, and it is also useful: Ansible errors are usually precise if you know where to look. The trick is to ask the right question in the right order, from the cheapest check to the most expensive. A doctor takes your temperature before ordering an MRI; do the same.
Work down the ladder; most problems are solved on rung 2, 3 or 4.
Rung 2: syntax and structure errors (before anything runs)
These appear instantly, before any device is contacted, and always name a file, a line and a column. Three samples from the lab:
AUTO01$ ansible-playbook bad2.yml --syntax-check
ERROR! We were unable to read either as JSON nor YAML, these are the errors we got from each:
JSON: Expecting value: line 1 column 1 (char 0)
Syntax Error while loading YAML.
could not find expected ':' (every line in a mapping needs 'key: value')
The error appears to be in '/home/netadmin/netauto/bad2.yml': line 3, column 12, but may
be elsewhere in the file depending on the exact syntax problem.
The offending line appears to be:
- name: Missing colon
hosts all
^ here
AUTO01$ ansible-playbook bad1.yml --syntax-check
ERROR! conflicting action statements: cisco.ios.ios_config, lines
AUTO01$ ansible-playbook bad3.yml --syntax-check
ERROR! couldn't resolve module/action 'cisco.ios.ios_conf'. This often indicates a misspelling, missing collection, or incorrect module path.
(Each message continues with the file name, line, column and an excerpt.) The first is a missing colon after hosts. The second means lines: was indented at the same level as the module name, so Ansible saw two actions in one task: indent it under the module. The third is a typo in the module name. The caret ^ here marks where the parser gave up; the real mistake is often on that line or the one above.
Rung 3: inventory problems
AUTO01$ ansible all -m ping [WARNING]: provided hosts list is empty, only localhost is available. Note that the implicit localhost does not match 'all' [WARNING]: No hosts matched, nothing to do
This means Ansible found no inventory: wrong folder (so ansible.cfg was not read), wrong path in inventory =, or a pattern that names a group that does not exist. Run ansible --version to see the active config file and ansible-inventory --graph to see the hosts.
Rung 4: connection problems
Here the exact word matters.
| What you see | Likely cause | Fix |
|---|---|---|
not found in known_hosts | First SSH connection not trusted | Add keys with ssh-keyscan; lab: host_key_checking = False |
Authentication failed for user netops | Wrong ansible_user or ansible_password, or the account does not exist | Fix the variable; test the login by hand with SSH |
UNREACHABLE ... Connection timed out | Wrong ansible_host, routing or firewall | Check the address, ping and SSH from the control node |
Unable to automatically determine the network OS | ansible_network_os missing | Add ansible_network_os=cisco.ios.ios |
AUTO01$ ansible all -m ping
pun-core-r1 | FAILED! => {
"changed": false,
"msg": "ssh connection failed: Authentication failed for user netops"
}
...
AUTO01$ ansible all -m ping
pun-core-r1 | SUCCESS => {
"changed": false,
"ping": "pong"
}
blr-edge-r2 | UNREACHABLE! => {
"changed": false,
"msg": "ssh connect failed: Connection timed out",
"unreachable": true
}
pun-acc-sw1 | SUCCESS => {
"changed": false,
"ping": "pong"
}
(First block: every device failed because the password in the inventory was wrong. Second block: only one device failed, because its ansible_host was a wrong address.) The wording of these messages differs between Ansible versions and connection types, so always read the msg text rather than relying on whether the word UNREACHABLE or FAILED appears.
Rung 5: module and device errors
When login works but a task fails, the message comes from the module or from the device CLI.
fatal: [pun-core-r1]: FAILED! => {"changed": false, "msg": "Unsupported parameters for (cisco.ios.ios_config) module: line. Supported parameters include: after, backup, backup_options, before, defaults, diff_against, diff_ignore_lines, intended_config, lines, match, multiline_delimiter, parents, replace, running_config, save_when, src."}
...
NO MORE HOSTS LEFT ************************************************************
PLAY RECAP ********************************************************************
pun-core-r1 : ok=0 changed=0 unreachable=0 failed=1 skipped=0 rescued=0 ignored=0
(Output shortened.) The option was spelled line instead of lines, and the module politely lists every valid option. When the message comes from the router, such as % Invalid input detected at '^' marker, type that command on the device by hand: the CLI is the final judge. ansible-doc -s cisco.ios.ios_config prints a snippet with the valid options, and adding -v to a run shows the full module result, including the commands that were sent.
Worked ticket. "Ansible cannot reach the Bengaluru router." You run ansible all -m ping: Pune answers, blr-edge-r2 is UNREACHABLE with a timeout. Rung 3: ansible-inventory --host blr-edge-r2 shows ansible_host: 10.99.0.99, but the router is 10.99.0.2. Someone mistyped the address. You correct the inventory, ping again and get pong. Total time: two minutes, no guessing.
Common mistake. Reading only the bottom of the output. The recap tells you which host failed, but the real reason is in the first fatal: line above it. Another: changing three things at once, then not knowing which one helped. Change one thing, re-run.
Exam trap. UNREACHABLE means Ansible could not establish the connection at all; FAILED means it connected and the task failed. Do not mix them when you describe a fault. And remember the order: syntax, inventory, connection, then module.
Works for Pune, silent for Bengaluru
On a Monday, a scheduled playbook skipped one router. The recap listed unreachable=1 for the edge router. The engineer followed the ladder: syntax was fine, the inventory listed the host, but a ping timed out. A manual SSH from the control node also hung. The cause was a firewall change on the WAN link, not Ansible. Once the network team restored the rule, the same playbook ran clean.
Lesson: if SSH by hand fails, Ansible will fail too. Test the lowest layer first.
"A playbook fails with UNREACHABLE. How do you troubleshoot?"
I confirm the inventory address with ansible-inventory --host, ping the device from the control node, and try SSH by hand to see whether it is routing, a firewall, or credentials. Only then do I look at Ansible variables such as ansible_host, ansible_user, ansible_network_os, and add -vvv for connection detail. Mention that unreachable and failed are different states.
Key takeaways
- Work from cheap to expensive: first error line, syntax, inventory, connection, module and device.
- Syntax errors name file, line and column; the real mistake is often one line above the caret.
- "No hosts matched" means wrong folder, wrong inventory path or wrong pattern.
- known_hosts, authentication and timeout errors are three different connection problems.
- If SSH by hand fails, fix that first; change one thing at a time.
Summary and exam checklist
You started this module never having used Ansible. You now own the whole beginner loop: set up a control node, describe devices, talk to them, write a playbook, rehearse it, apply it and prove it is repeatable. This chapter gathers everything you need before the labs, an interview or a certification-style question.
Can-do checklist
Tick each item only if you can do it without looking at the chapters.
- I can explain agentless, push-based automation and name the control node and managed nodes.
- I can read
ansible --versionand tell which configuration file is active. - I can write an INI inventory with groups, a group of groups and
[all:vars]for network_cli login. - I can verify an inventory with
ansible-inventory --graphand--list. - I can create
ansible.cfgand explain the known_hosts error and its two fixes. - I can run
ping,ios_command,ios_factsandios_configad hoc and read SUCCESS, CHANGED, FAILED and UNREACHABLE. - I can write a two-task playbook and explain every line.
- I can use
--syntax-check,--list-hosts,--check --diffand-l. - I can read the PLAY RECAP and prove idempotency with a second run.
- I can save the configuration with
save_when: modified. - I can follow the troubleshooting ladder from syntax to device.
Command cheat-sheet
| Goal | Command |
|---|---|
| Version and active config file | ansible --version |
| Installed collections | ansible-galaxy collection list |
| Module help | ansible-doc cisco.ios.ios_config or ansible-doc -s MODULE |
| Show the inventory tree | ansible-inventory --graph |
| Show every host variable | ansible-inventory --list |
| Login test | ansible all -m ping |
| Show command on a group | ansible routers -m cisco.ios.ios_command -a "commands='show ip interface brief'" |
| Facts from one device | ansible pun-core-r1 -m cisco.ios.ios_facts |
| Validate a playbook | ansible-playbook ntp.yml --syntax-check |
| Which hosts / tasks | --list-hosts, --list-tasks |
| Dry run with lines | ansible-playbook ntp.yml --check --diff |
| One device only | ansible-playbook ntp.yml -l pun-core-r1 |
| Show module results | ansible-playbook ntp.yml -v |
Mini glossary
- Control node
- Linux host where Ansible runs (AUTO01).
- Managed node
- A device Ansible configures; no agent on it.
- Inventory
- File of hosts, groups and variables.
- Group / children
- A named set of hosts;
:childrenbuilds a group from groups. - FQCN
- Fully qualified collection name, such as
cisco.ios.ios_config. - Module
- A unit of work called by a task.
- Play / playbook
- Hosts plus tasks / a file of plays.
- network_cli
- Connection type: SSH and the device CLI.
- Idempotent
- Repeating a run does not change the result.
- Check mode
- Dry run reporting what would change.
- Facts
- Device data gathered by
ios_facts, namedansible_net_*.
Most tested facts
- Ansible is agentless and push-based; the control node must be Linux or macOS.
- Configuration file order:
ANSIBLE_CONFIG,./ansible.cfg,~/.ansible.cfg,/etc/ansible/ansible.cfg; only the first found is used. - Network devices need
ansible_connection=ansible.netcommon.network_cli,ansible_network_os, user and password. gather_facts: falsefor network devices; collect facts withios_facts.pingon a network device is a login test, not ICMP.- In the PLAY RECAP,
okincludeschanged. - Second run
changed=0equals idempotent. - YAML: spaces not tabs,
-for lists, colon followed by a space. - UNREACHABLE: no connection; FAILED: connected, task failed.
What comes next on the Ansible for Network Engineers track.
Last reminders. Never trust a playbook you have not listed (--list-hosts) and rehearsed (--check --diff). Never keep real passwords in plain files. Always run the playbook a second time and look for changed=0.
Exam trap round-up. Inventory name is not the device hostname; [x:children] lists groups; ansible --version shows the active config; check mode is best effort; saving is a separate decision from changing.
Your first production-style change, end to end
A team needs a new syslog server on 60 devices tonight. The engineer lists hosts, runs the playbook with --check --diff against one canary router, applies to that router, runs it a second time to see changed=0, then runs the full group with save_when: modified. The ticket carries the dry-run output and the clean recap as proof.
Lesson: the habits you built in this module (look, rehearse, apply, repeat) scale from three lab devices to sixty production ones.
"Walk me through running your first Ansible change safely."
I confirm the inventory with ansible-inventory --graph, test login with ping, list the target hosts, rehearse with --check --diff, apply to one canary device with -l, then to the group, and repeat the run to confirm changed=0. I decide explicitly about saving the configuration and keep secrets out of plain text.
Key takeaways
- Control node plus inventory plus playbook is the whole model; devices need only SSH.
- Use FQCN module names and the network_cli connection variables.
- Look, rehearse, apply, repeat; a clean second run is proof.
- Troubleshoot from syntax to inventory to connection to module.
- Next: scale with YAML inventories, group_vars and host_vars in Ansible 2.