
There is a special kind of embarrassment in being the person who explains Ansible to other people during the week and then spends Sunday evening typing freebsd-update fetch install into a dozen SSH sessions.
My RHEL machines were fine. They have been patched from Satellite through Ansible for a long time. Everything else was not: FreeBSD routers, jail hosts with Bastille and with classic jail.conf jails, a mail server, a Proxmox cluster, a Proxmox Backup Server. Each one had its own little ritual, and the rituals lived in my head and in my shell history.
This article is about replacing those rituals with playbooks. Not with one clever generic “update everything” task, but with patching logic that knows the difference between freebsd-update and pkgbase, between a thin and a thick jail, between a router that may reboot and a hypervisor that must not.
All host names and addresses in this article are placeholders (example.com, RFC 5737 and RFC 3849 documentation ranges). The structure and the code are what I actually run.
Table of Contents
- Table of Contents
- The Fleet
- Do You Need Ansible Automation Platform?
- Inventory as Code
- One Service User, Bootstrapped Once
- RHEL: The Boring Part
- Debian and Proxmox: Patch, Never Reboot
- FreeBSD: Two Base Systems, One Playbook
- Jails: Two Kinds, Two Strategies
- Defensive Practices
- What Changed
- Conclusion
- References
The Fleet
| Host (placeholder) | OS | Notes |
|---|---|---|
edge1 … edge5.example.com |
FreeBSD 15.1, FRR | BGP routers, doas |
mail.example.com |
FreeBSD 15.1 | four thick jails in jail.conf, encrypted dataset |
web.example.com |
FreeBSD 15.1, pkgbase | nine thin Bastille jails |
social.example.com |
FreeBSD 15.1 | four thin Bastille jails (Mastodon) |
pve1, pve2.example.com |
Proxmox VE 9 (Debian 13) | two-node cluster with QDevice |
pbs.example.com |
Proxmox Backup Server 4 | Debian 13 |
rhel*.example.com |
RHEL 9/10 | Satellite-managed, already automated |
Three operating system families, four ways to update a base system (freebsd-update, pkgbase, apt and dnf), two jail managers, and one hypervisor cluster that happens to run the automation controller itself. That last detail will matter later.
So far, real runs cover the mail server, its jails, and both Bastille hosts. The router and Proxmox paths are written and reviewed, but have not had their first real run yet. I’ll say so where it matters.
Do You Need Ansible Automation Platform?
No.
I run this through Ansible Automation Platform (AAP), because I already have it and because a web UI with job history, credential storage, RBAC, and a “launch” button is genuinely nice for something you do every week.
But every playbook in this article is a plain playbook. ansible-playbook -i inventory/hosts patch_freebsd.yml from a laptop works the same, as long as it has the same collections, a compatible ansible-core, the same credentials, and the same inventory checkout. A systemd timer, a cron job, Semaphore, or a CI pipeline would all be fine schedulers. What AAP added for me was operational, not functional:
- Credentials stay in one place. The private key of the service user lives in AAP, encrypted, and never on my laptop.
- Job templates with prompts. One template, launched with a
limitand a job type ofcheckorrun, covers both the dry run and the real thing. - Collections only from my own Private Automation Hub. More on that in the defensive practices section.
If you don’t have AAP, skip the parts that mention it. You lose convenience, not capability.
Inventory as Code
The inventory lives in its own Git repository, separate from the playbooks. AAP syncs it as an inventory source, and the same checkout works locally.
external_inventory/
├── README.md
└── inventory/
├── hosts # INI: hosts and groups only, no inline variables
├── group_vars/
│ ├── all.yml
│ ├── freebsd.yml
│ ├── mail_jails.yml
│ └── proxmox.yml
└── host_vars/
├── edge1.example.com.yml
├── mail.example.com.yml
└── ...
I like INI for the host list because it reads like a table of contents, and YAML for everything else because variables deserve structure. The one rule: no variables in hosts. Group membership lives in one file, and host-specific overrides have exactly one predictable place: that host’s file in host_vars/.
[freebsd_routers]
edge1.example.com
edge2.example.com
# ... edge3 to edge5
[freebsd_mail]
mail.example.com
[bastille_hosts]
web.example.com
social.example.com
# base system updated via pkg, not freebsd-update
[freebsd_pkgbase]
web.example.com
[freebsd:children]
freebsd_routers
freebsd_mail
bastille_hosts
# thick jails on mail.example.com, reached through the jail host
[mail_jails]
mail-mta
mail-webmail
mail-client
# ... (excerpt)
# Proxmox VE nodes and Proxmox Backup Server: patched, never rebooted
[proxmox]
pve1.example.com
pve2.example.com
pbs.example.com
The groups describe behaviour, not just location. freebsd_pkgbase switches the base system update method. bastille_hosts adds release and jail handling. mail_jails is deliberately not a child of freebsd, so the host play never accidentally runs against a jail.
# group_vars/all.yml
ansible_user: ansible
ansible_port: 2222
ansible_python_interpreter: auto_silent
# group_vars/freebsd.yml
ansible_become_method: community.general.doas
# group_vars/proxmox.yml
ansible_become_method: sudo
Address Hosts the Way You Can Reach Them During an Outage
One host_vars entry per host turned out to be non-negotiable:
# host_vars/edge1.example.com.yml
# Provider address: SSH must not depend on the network this host is part of
ansible_host: 192.0.2.10
My routers’ DNS names point to addresses inside the network they route. That is perfectly fine until a playbook reboots the router that carries the active path: the SSH session to that address would disappear together with the route, the reboot module would time out, and the play would stop halfway through. Pinning every host to its provider-assigned address removes the dependency on the routes being patched. It doesn’t remove every dependency: if your automation controller’s own egress runs through that same network, the controller can still lose its path, so check where its traffic actually leaves.
One Service User, Bootstrapped Once
Every host gets the same service user, the same public key, and passwordless privilege escalation: doas on FreeBSD, sudo on Debian. Proxmox VE and Proxmox Backup Server don’t install sudo by default, so it has to be there before the bootstrap runs. A one-time bootstrap playbook, run as root from my workstation, sets it up:
- name: Ensure python is available
hosts: freebsd
gather_facts: false
tasks:
- name: Install python312 if no python3 is present
ansible.builtin.raw: ls /usr/local/bin/python3.* >/dev/null 2>&1 || pkg install -y python312
register: python_install
changed_when: "'Installing' in python_install.stdout"
raw first, because one of my routers had no Python at all. The same check, wrapped in jexec, puts Python into the thick mail jails, which jailexec needs because modules run inside the jail. After that, ordinary modules create the user, install the key, and add exactly one line to doas.conf:
- name: Check for existing doas.conf
ansible.builtin.stat:
path: /usr/local/etc/doas.conf
register: doas_conf
# Existing files keep their mode; only a newly created file gets one
- name: Allow the service user passwordless doas
ansible.builtin.lineinfile:
path: /usr/local/etc/doas.conf
line: permit nopass ansible as root
create: true
mode: "{{ omit if doas_conf.stat.exists else '0640' }}"
validate: doas -C %s
The stat dance exists because the first dry run showed me that a plain mode: "0640" would silently change the permissions of every existing doas.conf. That is exactly the kind of thing --check --diff is for.
Firewalls With Comments That Must Survive
All of my hosts filter SSH to a small list of trusted addresses, and the controller’s egress address had to be added. My firewall configs document those lists with a comment line underneath:
trusted_ipv4 = "{ 198.51.100.7, 203.0.113.42 }"
# ^VPN ^OFFICE
A replace with a regex would have appended the new address and left the comment line out of sync. So the bootstrap uses a small custom module from the playbook repository’s library/ directory: it appends the address to the first trusted list it finds, extends the annotation line with ^AAP, leaves every other byte alone, validates with pfctl -nf or sh -n, and supports check and diff mode:
-trusted_ipv4 = "{ 198.51.100.7, 203.0.113.42 }"
-# ^VPN ^OFFICE
+trusted_ipv4 = "{ 198.51.100.7, 203.0.113.42, 203.0.113.10 }"
+# ^VPN ^OFFICE ^AAP
Sixty lines of Python, and a configuration file that is still documentation afterwards.
RHEL: The Boring Part
RHEL was already automated, and it is the least interesting case on purpose. Satellite promotes content views, and the playbook just applies whatever the host is allowed to see:
- name: Patch all systems and reboot if required
hosts: "{{ host }}"
become: true
tasks:
- name: Ensure all updates are applied
ansible.builtin.package:
name: "*"
state: latest
update_cache: true
update_only: true
- name: Check whether a reboot is required
ansible.builtin.command: dnf needs-restarting -r
register: result
changed_when: false
check_mode: false
failed_when: result.rc not in [0, 1]
- name: Reboot if needed
ansible.builtin.reboot:
when: result.rc == 1
check_mode: false on the probe matters: without it, a check run skips the command, result has no rc, and the reboot condition fails. Even with it, a check run reports whether the host needs a reboot now, not whether it would after the updates it didn’t install.
The control is in Satellite, not in the playbook: what reaches production is decided when a content view is promoted. If you don’t run Satellite, the same pattern works with plain repositories, you just lose the gate.
Debian and Proxmox: Patch, Never Reboot
Proxmox VE is Debian with opinions. The most important one: always dist-upgrade (or apt full-upgrade), never a plain upgrade. Proxmox updates regularly change dependencies: new packages get pulled in, old ones are replaced or removed. apt-get upgrade does neither, and apt upgrade installs new packages but never removes any. Either can leave Proxmox packages held back or stuck halfway through a transition.
This playbook is one of the two that haven’t had a real run yet.
- name: Patch Proxmox VE and Backup Server hosts
hosts: proxmox
serial: 1
max_fail_percentage: 0
become: true
gather_facts: false
tasks:
- name: Detect Proxmox VE cluster node
ansible.builtin.stat:
path: /usr/bin/pvecm
register: pvecm_bin
- name: Check cluster quorum before upgrade
ansible.builtin.command:
argv: [/usr/bin/pvecm, status]
register: quorum_before
changed_when: false
check_mode: false
failed_when: quorum_before.stdout is not search('Quorate:\s+Yes')
when: pvecm_bin.stat.exists
- name: Apply all updates (Proxmox requires dist-upgrade)
ansible.builtin.apt:
update_cache: true
upgrade: dist
dpkg_options: force-confdef,force-confold
environment:
DEBIAN_FRONTEND: noninteractive
async: "{{ 0 if ansible_check_mode else 3600 }}"
poll: 15
- name: Check cluster quorum after upgrade
ansible.builtin.command:
argv: [/usr/bin/pvecm, status]
register: quorum_after
changed_when: false
check_mode: false
retries: 12
delay: 10
until: quorum_after.stdout is search('Quorate:\s+Yes')
when: pvecm_bin.stat.exists
Three details are worth explaining:
- Quorum before and after. A node is only touched if the cluster is quorate, and the next node only if it still is afterwards. On the Backup Server the check is skipped, because
pvecmdoesn’t exist there. asyncfor apt, except in check mode. The automation controller runs as a VM on this cluster. If a package upgrade restarts networking on the node, the SSH session can drop. Withasync,dpkgkeeps running on the node as a detached job instead of dying with the session. That covers exactly one failure mode: a package script can still fail, a node can still crash, and a job that outlives the async timeout gets killed. Ansible’s async mode doesn’t support check mode, so during a--checkrun the task runs synchronously (async: 0), whereaptonly simulates anyway.- No reboot. Hypervisors and the backup server are never rebooted by automation. The playbook compares the running kernel with the newest installed one, checks
/var/run/reboot-required, prints the result with adebugtask, and publishes it throughset_stats. In AAP, that lands in the job’s artifacts and is passed on to later workflow nodes. On the command line, custom stats only appear in the recap withshow_custom_stats = True:
- name: Record reboot status
ansible.builtin.set_fact:
reboot_required: "{{ reboot_flag.stat.exists or newest_kernel.stdout != running_kernel.stdout }}"
- name: Publish reboot status
ansible.builtin.set_stats:
data:
proxmox_reboot_required: "{{ {inventory_hostname: reboot_required | bool} }}"
aggregate: true
I’ll do the reboot myself, migrating VMs first, on an evening I choose.
FreeBSD: Two Base Systems, One Playbook
The FreeBSD playbook started as a quick sketch and ended up based on a much more paranoid version I had written earlier and never tested. The core idea: never trust a single return code.
The freebsd-update Path
- name: Host | freebsd-update fetch
ansible.builtin.command:
argv: [/usr/sbin/freebsd-update, --not-running-from-cron, fetch]
environment:
PAGER: /bin/cat
register: _fu_host_fetch
changed_when: false
check_mode: false
retries: 3
delay: 30
until: _fu_host_fetch is succeeded
# updatesready: rc 0 = updates staged, rc 2 = nothing to do
- name: Host | Check for staged updates
ansible.builtin.command:
argv: [/usr/sbin/freebsd-update, updatesready]
register: _fu_host_ready
changed_when: false
failed_when: _fu_host_ready.rc not in [0, 2]
check_mode: false
- name: Host | freebsd-update install
ansible.builtin.command:
argv: [/usr/sbin/freebsd-update, --not-running-from-cron, install]
environment:
PAGER: /bin/cat
register: _fu_host_install
changed_when: "'Installing updates' in _fu_host_install.stdout"
when: _fu_host_ready.rc == 0
# Patch-level updates: the only accepted leftover is a two-stage install
- name: Host | Verify the install is complete
ansible.builtin.command:
argv: [/usr/sbin/freebsd-update, updatesready]
register: _fu_host_ready_after
changed_when: false
failed_when: >-
_fu_host_ready_after.rc not in [0, 2]
or (_fu_host_ready_after.rc == 0
and 'Please reboot and run' not in _fu_host_install.stdout | default(''))
when: _fu_host_install is not skipped
fetch only stages files into /var/db/freebsd-update, exactly like the daily freebsd-update cron would. That is why it runs even in check mode: a --check run tells me which hosts actually have pending updates, without installing anything. updatesready turns “did anything happen?” into an explicit answer instead of output parsing. The verification step catches an install that silently left something behind.
For patch-level updates, install normally finishes in one pass. The two-stage case (kernel first, reboot, install again) is mostly what release upgrades do, and those add another pass to remove old shared libraries, which this playbook doesn’t support. If a patch install does ask for a reboot, the reboot block runs the second install pass once the host is back. A final updatesready afterwards must report nothing pending, otherwise the host is marked failed. If the host is not allowed to reboot, the summary reports the reboot as pending, which in that case also means the install isn’t finished.
--not-running-from-cron and PAGER=/bin/cat are the two flags every automated freebsd-update needs. Without a TTY and without them, it either refuses to run or waits for a pager that nobody will ever quit.
The pkgbase Path
One of my hosts runs pkgbase. There, freebsd-update refuses to work, and the base system is just another set of packages. The playbook detects it the same way freebsd-update itself does, by asking pkg who owns /usr/bin/uname:
- name: Host | Detect pkgbase
ansible.builtin.command:
argv: [/usr/local/sbin/pkg, which, /usr/bin/uname]
register: _fu_pkgbase_probe
changed_when: false
failed_when: false
check_mode: false
- name: Host | pkgbase detection must match the inventory
ansible.builtin.assert:
that:
- (_fu_pkgbase_probe.rc == 0) == (inventory_hostname in groups['freebsd_pkgbase'])
fail_msg: "pkgbase detection and inventory group disagree. Fix the inventory."
Detection alone would work. Detection plus an inventory group that has to agree with it means a host that changes its nature without anyone updating the inventory fails loudly instead of being patched the wrong way. On pkgbase hosts, community.general.pkgng with name: "*" and state: latest updates base and ports in one go, and a failure there is fatal, because it is the base system.
Rebooting, Carefully
After all updates, the playbook compares freebsd-version -k (installed kernel) with the running kernel. fbsd_update_reboot decides what happens next: auto (the default) reboots only if the two differ, always reboots regardless of the kernel, and never only reports a pending reboot. A reboot additionally requires that:
/var/run/noshutdowndoes not exist,- the run is not in check mode.
It reboots with shutdown -r rather than reboot, so rc.shutdown stops jails and services cleanly. Afterwards it verifies the running kernel, waits until every jail that was running before is running again, and on routers waits for service frr status. Those are startup checks: they prove that the jails and FRR came up, not that mail is delivered or that BGP sessions and routes are back. Everything runs with serial: 1 and max_fail_percentage: 0: the first broken host stops the rollout.
Errors that are not about the base system (a single jail failing pkg upgrade, a release that cannot be updated) are collected instead of aborting. The host is marked failed at the end with a summary, but a broken PHP package in one jail doesn’t leave the other eight unpatched.
What a Patch Doesn’t Restart
Installed is not the same as running. A kernel update ends in a reboot, and a changed jail base ends in a jail restart (see below). Two cases are not covered by the playbook today: a userland-only freebsd-update on the host, where the kernel version stays the same, and a pkg upgrade that replaces binaries or libraries on the host or in a jail. FreeBSD packages generally don’t restart running services on upgrade, so those daemons keep running the old code until someone restarts them. For now, the job summary lists which hosts and jails had packages upgraded, and fbsd_update_reboot: always is the blunt instrument for a host where I want everything restarted. Detecting the affected services and restarting exactly those is the next gap to close.
Jails: Two Kinds, Two Strategies
This is where most “patch everything” playbooks give up. I have two different kinds of jails, and each needs different handling.
Thin Bastille Jails
Thin Bastille jails nullfs-mount a shared, read-only base release. Updating the release updates every jail that uses it:
- name: Bastille | bastille update for releases
ansible.builtin.command:
argv: [/usr/local/bin/bastille, update, "{{ item.name }}"]
environment:
PAGER: /bin/cat
loop: "{{ _fu_releases }}"
loop_control:
index_var: _idx
register: _fu_rel_update
changed_when: >-
'Installing updates' in _fu_rel_update.stdout
or _fu_rel_update.stdout is search('Number of packages to be ')
failed_when: >-
_fu_rel_update.rc != 0
or '[ERROR]' in (_fu_rel_update.stdout ~ _fu_rel_update.stderr)
ignore_errors: true
when:
- item.usable
- item.pkgbase or (_fu_rel_ready.results[_idx].rc | default(2)) == 0
Before that, the playbook runs its own freebsd-update fetch and updatesready against each release directory, with the arguments bastille update uses (-b <release> -d <release>/var/db/freebsd-update --currently-running <release>), because the Bastille version I run doesn’t propagate freebsd-update‘s exit code. That gives three states per release: nothing staged (the update is skipped), updated (changed, recognised from the output of freebsd-update or, for pkgbase releases, pkg), and failed (collected as an error and excluded from the restart below). For freebsd-update releases, a final updatesready afterwards must report nothing pending, so “updated” means installed, not just attempted. Then bastille pkg <jail> upgrade -y runs per running jail, skipping jails without pkg.
The step I missed in the first version: an updated release only takes effect for running services after the jail restarts. On my first real run, all nine jails reported the new patch level and kept running the old libraries. Now the playbook finds the jails whose fstab mounts an updated release and restarts exactly those, unless the host reboots anyway:
- name: Bastille | Find jails using an updated release
ansible.builtin.find:
paths: "{{ _fu_jailsdir }}"
patterns: fstab
recurse: true
depth: 2
contains: '^{{ (_fu_releasesdir ~ "/" ~ item.item.name) | regex_escape }}\s'
loop: "{{ _fu_rel_update.results | select('changed') | reject('failed') | list }}"
register: _fu_rel_jails
Thick Jails With jailexec
My mail server runs classic thick jails from jail.conf. Each has its own base system and needs freebsd-update inside the jail. For that I use jailexec, my connection plugin that SSHes to the jail host and wraps every command in jexec. The jails don’t need SSH or an IP address, only Python.
# group_vars/mail_jails.yml
ansible_connection: chofstede.jailexec.jailexec
ansible_jail_host: 192.0.2.25
ansible_jail_privilege_escalation: doas
ansible_become: false
# host_vars/mail-webmail.yml
ansible_jail_name: webmail
Each jail is a normal inventory host. The jail play runs first, before the host play, with serial: 1:
- name: Jail | Build freebsd-update arguments
ansible.builtin.set_fact:
_fj_fu_args:
- /usr/sbin/freebsd-update
- --not-running-from-cron
- --currently-running
- "{{ _fj_userland_before.stdout | trim | regex_replace('-p[0-9]+$', '') }}"
--currently-running is the important part. Inside a jail, freebsd-update would otherwise derive the version from the host’s kernel, which is not what the jail’s userland is running. The playbook passes the jail’s own userland release instead, then runs the same fetch, updatesready, install, and verify sequence as on the host.
If the jail’s base changed and the host does not reboot, the host play restarts that jail afterwards with service jail restart. If the host reboots, the jails restart anyway.
That coupling has a consequence for --limit. The jails are separate inventory hosts, and Ansible knows nothing about their relationship to mail.example.com beyond a connection variable. Limiting a run to the mail server skips the jail play; limiting it to one jail skips the host play that would restart it. A run for the mail server therefore always selects both:
ansible-playbook -i inventory/hosts patch_freebsd.yml \
--limit 'mail.example.com:mail_jails' --check --diff
In AAP, the same pattern goes into the job template’s limit prompt.
The CVE I wrote about two weeks ago is why this is safe to do at all: since 2.0.0, every file transfer is resolved inside the jail, and jailexec itself only needs one doas rule on the host, for jexec. In this setup the service user has full doas anyway, because the host play patches the host, but a jail-only setup can stay at that single rule.
How a Connection Plugin Ends Up in a Container
When I first ran this in AAP, I wondered how jailexec could work at all: the job runs inside an execution environment container that has never heard of my plugin.
The answer is that Ansible loads plugins from a connection_plugins/ directory next to the playbook. AAP mounts the project checkout into the container, so a copy of jailexec.py in the playbook repository is enough. That worked on the first try, but a vendored copy is exactly what my CVE article warned about: a file that never updates itself and that none of my dependency scanning would ever flag.
So jailexec is now also an Ansible collection, chofstede.jailexec, built from the same single source file that the PyPI package uses. On my side it lives in my Private Automation Hub, and the playbook repository just lists it:
# collections/requirements.yml
collections:
- name: community.general
version: ">=13.0.0,<14.0.0"
- name: chofstede.jailexec
version: ">=2.0.3,<3.0.0"
The connection becomes chofstede.jailexec.jailexec, and there is exactly one place to release it. jailexec is already on GitHub, Codeberg and PyPI. The collection is coming to Ansible Galaxy soon.
Defensive Practices
Most of the work in these playbooks isn’t updating packages. It’s reducing the ways updating packages can hurt, and noticing quickly when it does anyway.
Dry run first, one host next, the rest last. Every change started with --check --diff against everything, then a real run against one unimportant host (a jail I only use to read mail), then the rest. The dry run caught the doas.conf permission change before it happened. Read-only tasks carry check_mode: false, so the check run actually knows something.
Know which hosts must never reboot unattended. My mail server’s jails live on an encrypted dataset that I unlock by hand after every boot. I knew that. I still launched a patch run with a kernel update against it, and the playbook did exactly what it was told: it rebooted the host, waited for the jails, and reported success, because I happened to unlock the dataset within its five-minute window. The fix is one line of inventory:
# host_vars/mail.example.com.yml
# Never reboot automatically: the jail dataset needs a manual "zfs load-key" after boot
fbsd_update_reboot: never
The deeper fix is to ask “which host reboots?” before every run, and to say the answer out loud. I’m also thinking about a pre-reboot check that refuses on its own when a ZFS dataset has keylocation=prompt.
Don’t let the automation patch its own floor. The controller runs on the Proxmox cluster. Proxmox nodes are therefore never rebooted by automation, apt runs asynchronously, and the workflow patches them last.
Constrain collections to a curated source. AAP in my setup pulls collections only from my Private Automation Hub. Public Galaxy is not configured at all. A new collection version reaches production when I approve it in the hub, not when someone publishes it. The requirements file only constrains the major version (>=13.0.0,<14.0.0). That is a compatibility guard, not a lock: the approval step in the hub decides which versions are eligible at all. Exact reproducibility would take pinned versions or a fixed execution environment. A missing collection fails the project sync instead of quietly coming from the internet. On my workstation, ansible-galaxy uses the same hub through environment variables (a project’s ansible.cfg would shadow ~/.ansible.cfg completely, because Ansible never merges config files).
Check requires_ansible before you blame the plugin. community.general 13 requires ansible-core 2.18 or later. The default execution environment in my AAP still ships 2.16. Ansible only warns about that and runs anyway, which is a fine way to get mysterious failures in the doas become plugin months later. The patch job templates use an EE with ansible-core 2.20, and the requirements file stays on the 13.x major.
Ad-hoc commands are not a test. An AAP ad-hoc ping failed against every FreeBSD host and every jail, because ad-hoc commands run without the project and therefore without its collections and plugins. A job template launched in check mode against two hosts is the real smoke test.
Inventory and playbooks are code. Both repositories get reviewed changes and signed commits only. I pushed unsigned commits once during this project and was reminded, firmly, by myself, that this is not how my repositories work. git commit --amend -S and --force-with-lease fixed it before anyone else noticed.
What Changed
| Before | After |
|---|---|
| A weekly SSH ritual per host | One workflow, launched when I choose |
| Patch knowledge in my head | Two Git repositories, reviewed and signed |
freebsd-update copy-pasted |
fetch, updatesready, install, verify, with check mode |
| Jails patched when I remembered | Thin jails via release + restart, thick jails via jailexec |
| Reboots whenever it felt right | Explicit per host: auto, always or never; Proxmox only reports |
| Collections from wherever | A curated hub, constrained versions |
The first real runs updated the mail server and its jails, the pkgbase host with all nine Bastille jails, and confirmed that the Mastodon host and its jails were already current. The routers and the Proxmox cluster are next, one at a time. Their playbooks are written and reviewed, but haven’t touched a production router or hypervisor yet.
Conclusion
None of the individual pieces in this article is clever. freebsd-update has had updatesready for years, Bastille has had bastille update since forever, and dist-upgrade is in every Proxmox upgrade guide. The value is in writing all of it down once, in the right order, with the right exceptions, and in a form that a dry run can check.
If you also have a fleet that is “mostly automated, except for the weird ones”, start with the inventory. Once every host’s groups describe how it must be treated, not just where it lives, the playbooks get much easier to write. What remains are the exceptions, and those are exactly what is worth writing down. And before the first real run, find out which of your hosts cannot boot on its own. Ask me how I know.
References
- Managing FreeBSD Jails with Ansible: The jailexec Connection Plugin
- jailexec on GitHub · Codeberg · PyPI
- Upgrading FreeBSD 15.0-RELEASE to 15.1-RELEASE: The Official Paths
- Two Sites, One Cluster: My Hetzner Proxmox VE Setup
- Reproducible Ansible with Execution Environments
- freebsd-update(8)
- Bastille documentation
- Proxmox VE: System Software Updates
Comments
You can use your Mastodon or other ActivityPub account to comment on this article by replying to the associated post.
Search for the copied link on your Mastodon instance to reply.
Loading comments...