
The previous article followed FreeBSD from power-on through the loader and the kernel, and stopped at a box in the diagram that just said rc(8) scripts. That box is where your machine actually becomes useful: networking comes up, filesystems get mounted, jails start, and every daemon you care about gets launched. This article opens that box.
This is the fourth article in the FreeBSD Foundationals series. The first covered Jails, the second covered ZFS, and the third covered the boot process. The boot article promised that packages and ports would come next. I changed the order, because rc.d is the direct continuation of the boot story, and because nearly every FreeBSD article on this blog ends with “and here’s the rc.d script” without ever explaining how rc.d works. Packages are next, I promise.
rc.d is also one of the best arguments for FreeBSD’s design approach. The whole service framework is a few thousand lines of POSIX shell that you can read in an afternoon. When something goes wrong, the answer is in a file you can open with less, not hidden in a binary.
By the end of this article you’ll understand how rc decides what starts and in what order, where configuration really comes from, what rc.subr does for you, how to write a service script that behaves correctly on start, stop, status and shutdown, and which handful of mistakes account for most broken rc.d scripts I’ve seen.
All examples were checked against the FreeBSD 15.1 sources.
Table of Contents
- Table of Contents
- From init to rc: What Actually Runs
- rc.conf Is a Shell Script
- rcorder: Dependency Ordering in Four Comment Lines
- Anatomy of an rc.d Script
- What rc.subr Gives You for Free
- service and sysrc: The Day-to-Day Interface
- Writing Your Own: A Supervised Service Done Right
- Case Study: Auditing My Own Scripts
- Service Jails: Isolation as an rc.conf One-Liner
- Debugging rc.d
- Coming From systemd
- Common Pitfalls: The Checklist
- Conclusion
- References
From init to rc: What Actually Runs
When the kernel finishes booting, it starts init(8) as PID 1. In multi-user mode, init does almost nothing itself. It runs /bin/sh /etc/rc autoboot and waits for it to finish. Once /etc/rc exits, init spawns getty on the terminals listed in /etc/ttys, and you get a login prompt.
So “the boot sequence” from here on is simply /etc/rc, which is a short shell script. Stripped of comments, it does this:
- Sources
/etc/rc.subr, the function library every rc.d script uses. - Loads the system configuration (
/etc/defaults/rc.conf, then/etc/rc.conf, and so on; more on that shortly). - Works out which scripts to skip: anything tagged
nostart, plusnojailscripts if running inside a jail, plusnojailvnetif it’s a non-VNET jail, plusfirstbootscripts unless the/firstbootsentinel file exists. - Asks
rcorder(8)to sort all the executable scripts in/etc/rc.d/into dependency order. - Runs them in that order, but stops at
FILESYSTEMS. - Now that all local filesystems are mounted (and
/usr/localis therefore guaranteed to be readable), adds the scripts fromlocal_startup, which is/usr/local/etc/rc.dby default, and runsrcorderagain over the combined set. - Runs everything that hasn’t run yet.
The two-pass design is the detail people don’t know about. Scripts from packages live under /usr/local, which on some systems is a separate filesystem that isn’t mounted yet during the first pass. So rc runs the base system up to the point where filesystems are mounted (the early_late_divider, which defaults to FILESYSTEMS), and only then looks at /usr/local/etc/rc.d. In practice this means a package’s script can never run before the local filesystems are mounted, whatever its REQUIRE line says.
init(8)
└── /etc/rc autoboot
├── . /etc/rc.subr
├── load_rc_config (defaults → rc.conf → rc.conf.local)
│
├── pass 1: rcorder /etc/rc.d/*
│ ... hostid, zfs, zvol, fsck, root, mountcritlocal ...
│ └── FILESYSTEMS ← early_late_divider: stop here
│
├── pass 2: rcorder /etc/rc.d/* /usr/local/etc/rc.d/*
│ (skips what already ran)
│ ... NETWORKING, SERVERS, DAEMON, LOGIN, jail, your apps ...
│
└── exit 0 → init starts getty → login:
Shutdown is the same story in reverse. shutdown -r now makes init run /etc/rc.shutdown, which calls rcorder -k shutdown to select only the scripts carrying the shutdown keyword, reverses the list, and invokes each script with faststop: the ordinary stop method with the fast prefix (more on prefixes below). A watchdog (rcshutdown_timeout, 90 seconds by default) kills the whole thing if a script hangs. After that, init sends SIGTERM and then SIGKILL to whatever is left. Keep that in mind: it becomes important when we write our own script.
rc.conf Is a Shell Script
This is the most important mental shift for people coming from other systems: /etc/rc.conf is not a config file in the INI sense. It is shell code, and it is sourced by /etc/rc and by every single rc.d script when it runs.
That has a few consequences:
sshd_enable="YES"is a shell variable assignment. No spaces around=.- A syntax error in
rc.conf(an unbalanced quote, say) doesn’t break one service. It breaks all of them, because every script sources the same file. - You can put logic in there. You shouldn’t, but you can, and some people do.
- The last assignment wins. If
ifconfig_vtnet0is set twice, the second one is what rc sees.
The Configuration Layers
load_rc_config in rc.subr reads configuration in this order, each layer overriding the one before:
| Order | File | Purpose |
|---|---|---|
| 1 | /etc/defaults/rc.conf |
FreeBSD defaults. Never edit; upgrades overwrite it. |
| 2 | /etc/defaults/vendor.conf |
Optional vendor overrides of the FreeBSD defaults, for appliances and derived products. |
| 3 | /etc/rc.conf |
Your system configuration. |
| 4 | /etc/rc.conf.local |
Historical local override file, read because it’s listed in rc_conf_files. |
| 5 | /etc/rc.conf.d/<name> |
Per-service configuration, read only by the service <name>. |
| 6 | <prefix>/rc.conf.d/<name> |
Same, for every local_startup entry with its trailing /rc.d removed. With the default, that’s /usr/local/etc/rc.conf.d/<name>. |
Layers 5 and 6 are underused. /etc/rc.conf.d/<name> can be a file, or a directory whose files are all sourced. Two things make this useful:
- Configuration management. A tool can own
/etc/rc.conf.d/nginxcompletely without touching (or templating) the sharedrc.conf. One file per service means no merge conflicts and no line-editing. - Scoping. A per-service file is only sourced when that service runs. The rest of the system never sees those variables.
The defaults file is worth reading at least once. less /etc/defaults/rc.conf is the best documentation of what knobs exist for base-system services. The boot article already covered sysrc for editing all of this safely, and everything there applies here: use sysrc instead of an editor, and the “breaks every service” risk mostly disappears.
# Write into a per-service file instead of rc.conf
install -d /etc/rc.conf.d
sysrc -f /etc/rc.conf.d/nginx nginx_enable=YES nginx_flags="-q"
# Quick syntax check of the shared file after manual edits
sh -n /etc/rc.conf && echo OK
YES, yes, true, 1
checkyesno, the function every script uses to test an _enable variable, accepts YES, TRUE, ON and 1 (case-insensitive) as true and NO, FALSE, OFF and 0 as false. Anything else (an empty string, a typo like YSE) produces a warning at boot: $foo_enable is not set properly - see rc.conf(5). If you ever see that line, you have a typo.
rcorder: Dependency Ordering in Four Comment Lines
There’s no manually maintained central list of what starts when. Instead, every rc.d script declares its relationship to other scripts in specially formatted comments, and rcorder(8) topologically sorts them. Here’s the header of /etc/rc.d/sshd:
#!/bin/sh
# PROVIDE: sshd
# REQUIRE: LOGIN FILESYSTEMS
# KEYWORD: shutdown
The four directives (strictly speaking, KEYWORD is a selection tag, not a dependency):
| Directive | Meaning |
|---|---|
PROVIDE: |
Names this script provides. Usually just its own name. |
REQUIRE: |
Names that must run before this script. |
BEFORE: |
Names this script must run before. The inverse of REQUIRE, useful for inserting yourself early without editing other scripts. |
KEYWORD: |
Tags used to include or exclude scripts: shutdown, nojail, nojailvnet, firstboot, nostart. |
The Placeholder Scripts
You’ll notice sshd requires LOGIN and FILESYSTEMS, which aren’t services. They’re placeholders: empty scripts in /etc/rc.d that exist only to provide a name and mark a milestone in the boot sequence. Here’s /etc/rc.d/DAEMON in full:
#!/bin/sh
# PROVIDE: DAEMON
# REQUIRE: NETWORKING SERVERS
# This is a dummy dependency, to ensure that general purpose daemons
# are run _after_ the above are.
The milestones, in order:
| Placeholder | Reached when |
|---|---|
FILESYSTEMS |
All local filesystems are mounted. Also the early/late divider. |
NETWORKING |
Interfaces are configured, routing is up. |
SERVERS |
Basic system servers (syslogd, etc.) are running. |
DAEMON |
General-purpose daemons can start. |
LOGIN |
Everything is up; users may log in. |
For a typical third-party daemon, REQUIRE: DAEMON is a good starting point. Add NETWORKING, LOGIN or a concrete provider (say, REQUIRE: postgresql) when the service actually depends on that milestone. LOGIN in particular is the checkpoint for login services and things that run commands as users, such as cron and inetd. It isn’t a catch-all for “anything from a package”.
Seeing the Order
You never need to guess. service -r prints the exact order rc would use at boot:
service -r | grep -n -E 'FILESYSTEMS|NETWORKING|DAEMON|LOGIN|nginx|jail'
38:/etc/rc.d/FILESYSTEMS
71:/etc/rc.d/NETWORKING
92:/etc/rc.d/DAEMON
133:/etc/rc.d/LOGIN
141:/usr/local/etc/rc.d/nginx
157:/etc/rc.d/jail
(Line numbers will differ on your system.) If you’ve just written a script and want to know where it lands, this is the command.
What rcorder Is Not
Here’s a point that confuses anyone coming from systemd: REQUIRE is about ordering, not activation. If your script says REQUIRE: postgresql and PostgreSQL isn’t enabled, nothing starts PostgreSQL for you, and your script still runs, just in the position where PostgreSQL would have run. Likewise, if PostgreSQL fails to start, rc carries on and starts your script anyway.
rcorder decides the order. Whether something runs is decided by its _enable variable, and that’s all. If your service really can’t work without another one, check for it in a start_precmd (shown later) and fail with a clear message.
Anatomy of an rc.d Script
Let’s read the real sshd script from 15.1, trimmed to the parts that matter:
#!/bin/sh
# PROVIDE: sshd
# REQUIRE: LOGIN FILESYSTEMS
# KEYWORD: shutdown
. /etc/rc.subr
name="sshd"
desc="Secure Shell Daemon"
rcvar="sshd_enable"
command="/usr/sbin/${name}"
pidfile="/var/run/${name}.pid"
start_precmd="sshd_precmd"
reload_precmd="sshd_configtest"
restart_precmd="sshd_configtest"
configtest_cmd="sshd_configtest"
extra_commands="configtest keygen reload"
sshd_configtest()
{
echo "Performing sanity check on ${name} configuration."
eval ${command} ${sshd_flags} -t
}
sshd_precmd()
{
run_rc_command keygen
run_rc_command configtest
}
load_rc_config $name
run_rc_command "$1"
The useful mental model, and the thesis of this whole article: an rc.d script normally declares a service rather than implementing its lifecycle. rc.subr owns the lifecycle; the script supplies metadata, configuration, and the exceptions.
Which is why there’s surprisingly little here. The script doesn’t contain any code to start a process, find its PID, stop it, or report status. It declares a handful of variables, defines two small helper functions, and hands control to run_rc_command. Everything else comes from rc.subr.
The structure is always the same:
- Header with
PROVIDE/REQUIRE/KEYWORD. . /etc/rc.subrto load the library.- Declarations:
name,rcvar,command,pidfile, and hooks. load_rc_config $nameto read rc.conf and the per-service files. This must come after settingname, and before you use any config variable.run_rc_command "$1"to dispatchstart,stop,statusor whatever the user asked for.
What rc.subr Gives You for Free
run_rc_command implements start, stop, restart, status, poll, rcvar, enabled, describe, enable, disable and delete, plus anything listed in extra_commands. What each of them does is driven by variables. These are the ones worth knowing:
| Variable | What it does |
|---|---|
name |
Service name. Required. Prefix for all ${name}_* variables. |
rcvar |
The enable variable, conventionally ${name}_enable. |
desc |
One-line description, shown by service foo describe. |
command |
Full path of the program to run. |
command_args |
Arguments set by the script. |
pidfile |
Where to find the PID. Strongly recommended. |
procname |
Process name to verify against the PID. Defaults to command. |
required_files, required_dirs |
Refuse to start if these don’t exist. |
required_modules |
Kernel modules to kldload before starting. |
sig_stop, sig_reload |
Signals for stop (default TERM) and reload (default HUP). |
extra_commands |
Additional verbs, each implemented by a <verb>_cmd function. |
<verb>_precmd, <verb>_postcmd, <verb>_cmd |
Hooks before, after, or instead of the default for any verb. |
And these per-service variables are set by the admin in rc.conf, not in the script, and are honoured automatically by any script built on rc.subr:
| rc.conf variable | Effect |
|---|---|
${name}_enable |
Start at boot. |
${name}_flags |
Extra arguments, added before command_args. |
${name}_user, ${name}_group |
Run as this user (via su -m). |
${name}_chdir, ${name}_chroot |
Working directory, or chroot. |
${name}_env, ${name}_env_file |
Environment variables, inline or from a file. |
${name}_nice |
Scheduling priority. |
${name}_fib |
Routing table, via setfib(1). Ties in nicely with dual-FIB policy routing. |
${name}_limits, ${name}_login_class |
Resource limits via limits(1). |
${name}_oomprotect |
Protect from the OOM killer (YES or ALL to include children). |
${name}_svcj, ${name}_svcj_options |
Run inside a service jail. Covered below. |
Read that list again with a specific service in mind. You can put any service on the second FIB, renice it, protect it from the OOM killer, give it an environment file, or jail it, without touching its script, just by adding a line to rc.conf. A port author doesn’t have to anticipate any of this; rc.subr handles it uniformly for every service.
# Run unbound on FIB 1, with OOM protection, without editing its script
sysrc unbound_fib=1 unbound_oomprotect=YES
service unbound restart
The Verb Prefixes
Every verb accepts a prefix that changes its checks:
| Prefix | Effect | Typical use |
|---|---|---|
one |
Ignore _enable; run anyway, with all other checks. |
service nginx onestart to test a service you haven’t enabled yet. |
force |
Ignore _enable and failing precmd/required_* checks. |
Last resort. |
fast |
Skip the “already running?” check. | Used by rc at boot for speed. |
quiet |
Suppress “Starting foo.” and not-enabled messages. | Used by rc for non-autoboot. |
onestart/onestop/onestatus are the ones you’ll use daily. If you get Cannot 'start' nginx. Set nginx_enable to YES in /etc/rc.conf or use 'onestart' instead of 'start'., that’s rc.d telling you exactly this.
service and sysrc: The Day-to-Day Interface
service(8) is a small wrapper that finds the script in /etc/rc.d or local_startup and runs it. Two details matter.
First, it deliberately runs the script with the same restricted environment it gets at boot: HOME=/ and a minimal PATH=/sbin:/bin:/usr/sbin:/usr/bin. Don’t rely on your interactive shell’s PATH or exported variables. (On 15.1, service.sh implements this by starting the script through env -i, so nothing else from your shell gets through either.) This matches the boot environment, so always test with service, not by running the script directly from your shell. A script that only works when run directly is a script that will fail at 3 AM after a reboot.
Second, it can do most things you’d otherwise do by hand:
service nginx start | stop | restart | status | reload
service nginx configtest # any extra_command works
service nginx describe # prints $desc
service nginx enabled && echo on # exit status only, for scripts
service nginx enable # sets nginx_enable=YES via sysrc
service nginx disable # sets nginx_enable=NO
service nginx delete # removes nginx_enable entirely
service -e # list enabled services, in boot order
service -l # list all available scripts
service -r # full boot order (rcorder output)
service -R # restart all enabled local services
service -e is underrated. It’s the one-line answer to “what does this machine actually run?”, and it’s the first thing I type on a box I didn’t build.
Inside Jails
Both tools speak jail natively, which ties straight back to the first article in this series:
# Manage services in a jail from the host - no ssh, no jexec shell
service -j www nginx restart
service -j www -e
# Edit the jail's rc.conf from the host
sysrc -j www nginx_enable=YES
A full jail runs its own /etc/rc when it starts (via exec.start = "/bin/sh /etc/rc" in the jail config), with all the same machinery. The only difference is the skip list: scripts tagged nojail don’t run, which is why you don’t see devd or ntpd or disk-related scripts trying to start inside a jail. Non-VNET jails additionally skip nojailvnet scripts, so there’s no attempt to configure interfaces they don’t own.
Writing Your Own: A Supervised Service Done Right
Time to write one. The scenario is the most common one on my servers: a modern application that runs in the foreground, logs to stdout, writes no pidfile of its own, and should be restarted if it crashes. Think of a Go binary, a Node app, or a Python service. I’ll call it webhookd.
FreeBSD’s answer to “supervise a foreground process” is daemon(8), which is in base. It detaches from the terminal, drops privileges, redirects output to syslog or a file, writes pidfiles, and optionally restarts the child when it exits. Combined with rc.subr, you get roughly what a systemd unit with Restart=on-failure gives you, in about 40 lines of shell:
#!/bin/sh
# PROVIDE: webhookd
# REQUIRE: DAEMON NETWORKING
# KEYWORD: shutdown
. /etc/rc.subr
name="webhookd"
desc="Webhook receiver"
rcvar="webhookd_enable"
load_rc_config $name
# Defaults; override any of these in rc.conf or /etc/rc.conf.d/webhookd
: ${webhookd_enable:="NO"}
: ${webhookd_user:="webhookd"}
: ${webhookd_config:="/usr/local/etc/webhookd.yml"}
: ${webhookd_args:="--listen [::1]:9000"}
# Clear webhookd_user before rc.subr sees it: daemon(8) handles the
# privilege drop, and rc.subr would otherwise su the supervisor itself.
_webhookd_runas="${webhookd_user}"
webhookd_user=""
pidfile="/var/run/${name}.pid"
procname="/usr/sbin/daemon"
command="/usr/sbin/daemon"
command_args="-r -R 5 -P ${pidfile} -t ${name} -S -T ${name} \
-u ${_webhookd_runas} \
/usr/local/bin/webhookd --config ${webhookd_config} ${webhookd_args}"
required_files="${webhookd_config}"
start_precmd="webhookd_precmd"
webhookd_precmd()
{
if ! id "${_webhookd_runas}" >/dev/null 2>&1; then
err 1 "user ${_webhookd_runas} does not exist"
fi
}
run_rc_command "$1"
Install it as /usr/local/etc/rc.d/webhookd, make it executable, and enable it:
install -m 0555 webhookd.rc /usr/local/etc/rc.d/webhookd
service webhookd enable
service webhookd start
service webhookd status
# webhookd is running as pid 4182.
Almost every line in that script is there to avoid a specific bug. Let’s go through them.
The daemon(8) Flags
| Flag | Why |
|---|---|
-r -R 5 |
Supervise: restart the child if it exits, after 5 seconds (plain -r uses 1 second). The delay stops a crash loop from pinning a CPU. |
-P ${pidfile} |
Write the supervisor’s PID. This is the important one, see below. |
-t ${name} |
Process title, so ps shows daemon: webhookd[4183] instead of a long command line. |
-S -T ${name} |
Send the child’s stdout/stderr to syslog with the tag webhookd (facility daemon, priority notice by default). |
-u ${_webhookd_runas} |
Run the child as an unprivileged user; the supervisor stays root. |
If you’d rather have a log file than syslog, replace -S -T with -o /var/log/webhookd.log -H. The -H makes daemon reopen the file on SIGHUP, so newsyslog can rotate it.
The Pidfile Trap: -p vs -P
This is the bug I see most often, and I’ve seen it in production scripts written by experienced people. daemon(8) has two pidfile options:
-p child_pidfilerecords the PID of the application.-P supervisor_pidfilerecords the PID of thedaemonprocess itself.
Now combine -p with -r. You run service webhookd stop. rc.subr reads the pidfile, finds the application’s PID, and sends it SIGTERM. The application exits cleanly. The supervisor, which is still running, does exactly what you told it to: it notices the child died and restarts it. service stop reported success, and the service is running again five seconds later.
With -P, stop signals the supervisor instead. daemon forwards SIGTERM to the child, waits for it to exit, removes the pidfiles, and exits itself. That’s what you actually wanted. The daemon(8) man page says as much in its own words: --supervisor-pidfile “is especially important if you use --restart in an rc script”.
Because the pidfile now points at /usr/sbin/daemon, procname has to match. rc.subr doesn’t blindly trust a pidfile: it checks that the PID belongs to a process whose name matches procname, so a stale pidfile pointing at a recycled PID doesn’t make it kill some unrelated process. daemon’s retitled process shows up as daemon: webhookd[4183], and rc.subr explicitly accepts the daemon: form, so procname="/usr/sbin/daemon" is correct. (It’s also the default when command is daemon, but I set it explicitly so the next person reading the script doesn’t have to know that.)
You can see both processes:
ps -o pid,ppid,user,command -p $(cat /var/run/webhookd.pid) -p $(pgrep -P $(cat /var/run/webhookd.pid))
PID PPID USER COMMAND
4182 1 root daemon: webhookd[4183] (daemon)
4183 4182 webhookd /usr/local/bin/webhookd --config /usr/local/etc/webhookd.yml --listen [::1]:9000
Why webhookd_args and Not webhookd_flags
rc.subr builds the start command as $command ${name}_flags $command_args. Normally command is the application, so _flags lands where you’d expect. But here command is /usr/sbin/daemon. Anything an admin puts in webhookd_flags ends up between daemon and its own options, and daemon tries to parse --listen as one of its own flags and fails.
So for a daemon-wrapped service, don’t use _flags for application arguments. Use your own variable (webhookd_args above) and put it at the end of command_args, after the application path, where it belongs. Many ports follow the same convention for the same reason.
Why webhookd_user Is Cleared
${name}_user is one of the variables rc.subr honours automatically: it would wrap the whole start command in su -m webhookd. That would run daemon itself unprivileged, which then can’t write its pidfile to /var/run, and can’t call -u because only root can change user. The script keeps the familiar webhookd_user knob for the admin, copies it, and clears the original so rc.subr doesn’t act on it. Privilege dropping happens exactly once, in daemon(8).
An alternative is to keep ${name}_user, drop -u, and put the pidfile in a directory the service user owns (e.g. /var/run/webhookd/). Both work. The important thing is to choose one mechanism and not both.
The Rest
KEYWORD: shutdownputs the script inrc.shutdown‘s list, sostopis called in reverse order on shutdown, before init sends SIGTERM to everything left. Without it your service gets no ordered stop, only init’s blanket SIGTERM and SIGKILL at the very end. For anything that needs an orderly shutdown (databases, queues, stateful applications), treat this keyword as mandatory.: ${var:="default"}afterload_rc_configsets a default only if the admin hasn’t. This is the standard idiom; the:is the shell’s no-op command, used only for its side effect of assigning the default.required_filesmakesstartfail fast withwebhookd: required file /usr/local/etc/webhookd.yml does not existinstead of starting a supervisor whose child crash-loops forever.start_precmdis where you check preconditions or prepare the environment (create a runtime directory, verify a dependency). If it returns non-zero, the start is aborted.
And don’t forget the script has to be executable, and it must contain a # PROVIDE: line. At boot, rc only picks up files in local_startup that are both. A non-executable script or one without PROVIDE is silently skipped. No error, no log, nothing.
Case Study: Auditing My Own Scripts
Writing a clean example is easy. It’s more honest to look at what’s actually running on my servers. I went through the hand-written rc.d scripts in my jails while writing this article, and they make a good tour of everything above: one is nearly perfect, one is a single edit away from the pidfile trap, and one has grown into a small shell program that works around rc.subr instead of using it.
mastodon_web: Working Around rc.subr
The multi-jail Mastodon article from last December published the first version of this script. Its stop function looked like this:
mastodon_web_stop() {
kill -9 `cat /var/run/mastodon/${name}_supervisor.pid` 2>/dev/null
kill -15 `cat /var/run/mastodon/${name}.pid` 2>/dev/null
}
Look at the order: SIGKILL the supervisor first, so it can’t restart anything, then SIGTERM the child. That’s the workaround you end up with when you hit the pidfile trap without knowing what it is. daemon -r kept bringing puma back, so the supervisor got shot before it could. It does work. But the supervisor dies without cleaning up its pidfiles, and the function returns immediately without waiting for puma to exit.
The version running in production today has no -r at all, and a gentler stop function. Why -r disappeared, I honestly can’t tell you any more. These scripts are old enough that the reasoning is lost, which is a good argument for writing it down in a comment. Here’s the core of it, with the addresses replaced by documentation ones:
command="/usr/local/bin/ruby33"
procname="/usr/local/bin/ruby33"
start_cmd="mastodon_start"
stop_cmd="mastodon_stop"
mastodon_start() {
/bin/mkdir -p /var/run/mastodon /var/log/mastodon
/usr/sbin/chown mastodon:mastodon /var/run/mastodon /var/log/mastodon
export PATH=/sbin:/bin:/usr/sbin:/usr/bin:/usr/local/sbin:/usr/local/bin:~/bin
export RAILS_ENV=production
export BIND=198.51.100.9
export PORT=3000
cd /home/mastodon/live && \
/usr/sbin/daemon -u mastodon -H \
-T ${name} \
-P ${supervisor_pidfile} \
-p ${pidfile} \
-f \
-o /var/log/mastodon/puma.log \
/usr/local/bin/bundle exec puma -C config/puma.rb
}
mastodon_stop() {
if [ -f "${supervisor_pidfile}" ]; then
kill -TERM $(cat "${supervisor_pidfile}")
fi
if [ -f "${pidfile}" ]; then
if kill -0 $(cat "${pidfile}") >/dev/null 2>&1; then
kill -TERM $(cat "${pidfile}")
fi
fi
}
It runs, and it has done for months. But measured against what rc.subr offers, it has five problems:
stopdoesn’t wait. rc.subr’s defaultstopsends the signal and then waits for the PIDs to actually exit. A customstop_cmdreplaces all of that. Soservice mastodon_web restartsends SIGTERM and immediately starts a new puma while the old one is still finishing requests and holding port 3000. If the new one loses the race for the port, it exits, and without-rnothing brings it back. On shutdown,rc.shutdownsimilarly moves on before puma is gone.- No supervision. Without
-r, a crashed puma or Sidekiq stays dead until someone notices. -Tquietly enables syslog.-Tsets the syslog tag, and daemon(8) turns on syslog output whenever any syslog option is given. So every line puma writes goes to/var/log/mastodon/puma.logand to syslog. I wanted-t, the process title.- Everything is hard-coded. Addresses, ports and environment sit in the script, three times over (web, streaming, sidekiq), instead of in rc.conf where rc.subr would apply them.
-
statuslies. While writing this, I ran it on the production jail, with Mastodon happily serving requests:text root@mastodonweb:~ # service mastodon_web status mastodon_web is not running.The pidfile is the child’s (
-p) andprocnameis/usr/local/bin/ruby33. rc.subr only believes a pidfile if the process behind that PID matchesprocname, and here it doesn’t, so it concludes nothing is running. That’s not just cosmetic. The same lookup drives the “already running?” check instart, soservice mastodon_web starton a running system will start a second puma next to the first one.
Here’s the rewrite. The script now just declares the service:
#!/bin/sh
# PROVIDE: mastodon_web
# REQUIRE: DAEMON
# KEYWORD: shutdown
. /etc/rc.subr
name="mastodon_web"
desc="Mastodon Web Service"
rcvar="mastodon_web_enable"
load_rc_config $name
: ${mastodon_web_enable:="NO"}
: ${mastodon_web_chdir:="/home/mastodon/live"}
pidfile="/var/run/mastodon/${name}.pid"
command="/usr/sbin/daemon"
command_args="-r -R 5 -P ${pidfile} -t ${name} -u mastodon \
-f -o /var/log/mastodon/puma.log -H \
/usr/local/bin/bundle exec puma -C config/puma.rb"
start_precmd="mastodon_web_precmd"
mastodon_web_precmd()
{
install -d -o mastodon -g mastodon /var/run/mastodon /var/log/mastodon
}
run_rc_command "$1"
And the instance-specific settings move to /etc/rc.conf.d/mastodon_web, where rc.subr’s ${name}_env picks them up:
mastodon_web_enable="YES"
mastodon_web_env="PATH=/usr/local/bin:/usr/bin:/bin \
RAILS_ENV=production BIND=198.51.100.9 PORT=3000 \
PROMETHEUS_EXPORTER_LOCAL=false PROMETHEUS_EXPORTER_HOST=2001:db8:9000::100"
There’s no custom start, stop or status code left. One pidfile, pointing at the supervisor, with procname defaulting to /usr/sbin/daemon, which matches. So status tells the truth and start refuses to start a second copy. stop is rc.subr’s own, so it waits, and restart no longer races. -r -R 5 brings puma back after a crash. ${name}_chdir replaces the cd, install -d replaces the mkdir/chown pair, and the streaming and sidekiq scripts become the same twenty lines with a different name and command.
One more thing from the Sidekiq script: it says REQUIRE: DAEMON postgresql. But in this setup PostgreSQL runs in the separate database jail. Inside the Sidekiq jail there is no script that provides postgresql, so that dependency matches nothing, and nobody complains. rcorder only sees the scripts of the system it’s running on. Ordering between jails is the jail manager’s job (depend in jail.conf(5)), not rc.d’s. If Sidekiq must not start before the database is reachable, that belongs in a start_precmd that waits for the port.
zigbee2mqtt: Nearly Right
From the Home Assistant setup:
: ${zigbee2mqtt_enable:="NO"}
: ${zigbee2mqtt_user:="z2m"}
pidfile="/var/run/zigbee2mqtt/zigbee2mqtt.pid"
command=/usr/sbin/daemon
command_args="-f -P ${pidfile} /usr/local/bin/zigbee2mqtt-start"
load_rc_config $name
run_rc_command "$1"
This one is close to ideal. It uses rc.subr’s default methods, the pidfile is the supervisor’s (-P), and procname defaults to /usr/sbin/daemon, which matches. It also uses the other privilege-dropping approach from earlier: zigbee2mqtt_user makes rc.subr su the whole daemon invocation to z2m, which works because the pidfile lives in a directory z2m owns.
The only thing missing is the reason to use daemon(8) in the first place: -r. Zigbee2MQTT is the service in my house most likely to exit on its own, typically when the USB coordinator hiccups. Because the pidfile is already -P, adding -r -R 10 is safe, and service zigbee2mqtt stop will still stop it.
(The defaults are also set before load_rc_config. That happens to work, because rc.conf assigns its values unconditionally and overwrites them. Putting the defaults after load_rc_config is the convention, and it’s correct whatever the config file does.)
netbox: One Flag Away from the Trap
pidfile="/var/run/${name}.pid"
procname="/opt/netbox/venv/bin/python3.11"
command="/usr/sbin/daemon"
command_args="-u netbox -p ${pidfile} /opt/netbox/venv/bin/gunicorn \
--pythonpath /opt/netbox/netbox --config /opt/netbox/gunicorn.py netbox.wsgi"
This is correct as it stands. The pidfile is the child’s (-p), and procname is set to the Python interpreter, because the gunicorn master shows up in ps as /opt/netbox/venv/bin/python3.11 /opt/netbox/venv/bin/gunicorn .... stop sends SIGTERM to gunicorn, gunicorn shuts down gracefully, and daemon, which has nothing to restart, exits too.
But it’s exactly one change away from the trap. The obvious improvement is “add -r so it restarts on crash”. Do only that, and service netbox stop will stop gunicorn and daemon will immediately start it again. The correct change is three edits together: add -r, switch -p to -P, and remove procname so it defaults to the daemon binary.
It has a second issue too: no -f, -o or -S. When none of those are given, daemon(8) writes the child’s output to its own stdout, which at boot is the console. Inside a jail, that means gunicorn’s error log ends up in the jail’s console log, or nowhere. Adding -S -T netbox sends it to syslog instead.
command_args="-r -R 5 -P ${pidfile} -t ${name} -S -T ${name} -u netbox \
/opt/netbox/venv/bin/gunicorn \
--pythonpath /opt/netbox/netbox --config /opt/netbox/gunicorn.py netbox.wsgi"
The Pattern
Put the three side by side and one rule stands out: the more a script relies on rc.subr, the fewer bugs it has. The script with custom start_cmd and stop_cmd had the most problems, because each custom function has to reimplement something rc.subr already does correctly: waiting on stop, checking PIDs, applying the environment. If you find yourself writing a stop_cmd, check first whether the real problem is a pidfile pointing at the wrong process.
One last thing about jails: when a jail stops, it runs its own /etc/rc.shutdown, so KEYWORD: shutdown matters inside jails just as much as on the host. All three scripts got that right.
Service Jails: Isolation as an rc.conf One-Liner
FreeBSD 15.0 added service jails to the rc framework, and they’re the most interesting change to rc.d in years. Set ${name}_svcj=YES and rc.subr starts the service inside a dedicated, throwaway jail named svcj-<name>, without you writing a jail.conf entry:
sysrc webhookd_svcj=YES
sysrc webhookd_svcj_options="net_basic"
service webhookd restart
jls name path
# svcj-webhookd /
_svcj_options grants specific capabilities: net_basic (IPv4/IPv6 with reserved ports), netv4, netv6, net_raw, net_all, sysvipc, sysvipcnew, mlock, routing, settime, vmm and nfsd. By default the service gets none of them. ${name}_svcj_ipaddrs restricts it to specific addresses.
Be clear about what this does and doesn’t buy you. A service jail uses path=/: it shares the host’s filesystem. It does not hide /etc or /home from the service. What it does isolate is the process table (the service can’t see or signal other processes) and the network privileges it’s allowed. SysV IPC is off unless you ask for it: sysvipc shares the host’s IPC namespace, while sysvipcnew gives the service a separate one of its own. Think of it as a cheap reduction in what a compromised daemon can do, not as a replacement for the full jails from the first article. That’s exactly why base sshd opts out of svcj_all_enable by default (: ${sshd_svcj:="NO"}): an sshd that can’t see other processes isn’t much use as a login service. You can still put it in a service jail explicitly with sshd_svcj=YES.
The framework is also still under active development. At the end of August 2026 a batch of fixes landed in main, and they were merged to stable/15 on 6 September. They include making svcj_all_enable actually work, running a script’s own restart_cmd and status_cmd inside its service jail, and sending stop and reload signals from inside the jail. That last one matters if you combine ${name}_svcj with ${name}_user: before the fix, stop and reload could fail and leave both the service and its jail running. None of this is in 15.1-RELEASE. If you rely on service jails, test stop, reload, and any custom status or restart methods on the exact FreeBSD build you deploy.
For services that just listen on a port and talk to a database, adding one rc.conf line costs nothing, and it’s worth doing.
Debugging rc.d
When a service won’t start and the error message isn’t helpful:
# 1. Does rc.d even think it's enabled, and under which variable?
service webhookd rcvar
# 2. Trace the script under the same environment service(8) uses
env -i -L -/daemon HOME=/ PATH=/sbin:/bin:/usr/sbin:/usr/bin \
/bin/sh -x /usr/local/etc/rc.d/webhookd start 2>&1 | less
# (sh -x shows every expansion; you'll see the exact command line rc.subr builds)
# 3. Turn on rc.subr's own debug output, for this run only
service -E rc_debug=YES webhookd start
# 4. Same, but for the whole boot: set it, reboot, read the console/messages
sysrc rc_debug=YES # remember to remove it afterwards
# 5. Where does it land in the boot order?
service -r | grep -n webhookd
Step 2 is exactly what service does internally, plus sh -x. Running plain sh -x from your shell is fine for debugging shell expansion, but it traces the script with your environment, which is the thing we just said not to trust. Step 2 resolves most problems on its own, because it shows you the exact command line that rc.subr builds, including every _flags, _env and _user expansion. When the generated command is wrong, sh -x shows you how it got that way.
For boot-time-only failures (works with service start, fails at boot), the cause is almost always ordering. The service starts before something it needs: a network address that isn’t configured yet, a ZFS dataset that isn’t mounted, a database that’s still starting. Check service -r, fix REQUIRE, and if the dependency takes a while to become ready (not just started), wait for it in start_precmd. For network readiness specifically, base has netwait (netwait_enable, netwait_ip, netwait_if in rc.conf), which blocks boot until an address answers ping or an interface has link.
Coming From systemd
Many readers of this blog run Linux, mostly RHEL and Fedora, and I do too, as the Podman deep dive shows. So here’s the mapping, without the flame war:
| systemd | rc.d |
|---|---|
Unit file in /etc/systemd/system/ |
Script in /usr/local/etc/rc.d/ |
systemctl enable --now foo |
service foo enable && service foo start |
systemctl status foo |
service foo status + syslog |
Drop-in foo.service.d/override.conf |
/etc/rc.conf.d/foo |
After= / Before= |
REQUIRE: / BEFORE: |
Requires= / Wants= (activation) |
No equivalent; ordering only |
Restart=on-failure |
daemon -r |
User= |
${name}_user or daemon -u |
Environment= / EnvironmentFile= |
${name}_env / ${name}_env_file |
Nice=, LimitNOFILE= |
${name}_nice, ${name}_limits |
StandardOutput=journal |
daemon -S → syslog (a rough equivalent; journald does far more) |
Sandboxing (PrivateTmp=, ProtectSystem=…) |
Service jails (narrower), or a full jail |
| cgroup resource control | rctl per jail/user/process |
| Socket activation, timers, parallel start | inetd, cron, sequential start |
They’re built on different philosophies, and each has real strengths. systemd is a supervisor with a dependency graph, parallel startup, and a lot of sandboxing and resource control built in. rc.d is a sequential, ordered set of shell scripts with a shared library, plus daemon(8) when you need supervision. rc.d boots more slowly on a machine with hundreds of services, and it has nothing like systemd’s built-in sandboxing knobs. In exchange, the whole thing is a few files of shell you can read and trace with sh -x, and a server with fifteen services, which is most of mine, starts in seconds anyway.
Common Pitfalls: The Checklist
Everything below is covered above. This section is the short version, for when you come back to this article with a broken script.
daemon -r with -p instead of -P. service stop kills the child, the supervisor restarts it, and stop “succeeds” while the service keeps running. With -r, the rc.d pidfile must be the supervisor’s (-P).
Application arguments in ${name}_flags for daemon-wrapped services. They end up as arguments to daemon, not the application. Use a custom ${name}_args variable at the end of command_args.
Running both ${name}_user and daemon -u. The supervisor gets su‘d to an unprivileged user, can’t write its pidfile, and can’t drop privileges again. Pick one mechanism.
A custom stop_cmd that doesn’t wait. rc.subr’s default stop waits for the process to exit; your replacement probably doesn’t, and then restart races the old process for its port. Usually the fix is to delete the custom function and correct the pidfile.
daemon -T when you meant -t. -T is the syslog tag and implicitly turns on syslog output, so logs get duplicated into syslog. -t is the process title.
REQUIRE across jails. rcorder only sees scripts on its own system. A REQUIRE: postgresql in an application jail does nothing if PostgreSQL runs in another jail.
Missing KEYWORD: shutdown. The service gets no ordered stop. On shutdown it just receives init’s SIGTERM along with everything else, and a SIGKILL shortly afterwards. Databases do not appreciate this.
Assuming REQUIRE starts dependencies. It only orders. An un-enabled or failed dependency is silently skipped, and your service starts anyway.
Testing by running the script from your shell. Your shell has a full PATH and environment; the service at boot doesn’t. Test with service, which reproduces the boot environment. Use absolute paths for everything in a script.
Non-executable script, or no # PROVIDE: line. Silently ignored at boot. service -r | grep yourservice tells you immediately whether rc can see it.
Editing scripts in /etc/rc.d. These belong to the base system and are replaced on upgrade (and with pkgbase, they’re files owned by a package). Change behaviour through rc.conf variables, /etc/rc.conf.d/<name>, or ${name}_prepend, not by editing the script. Same rule as never editing /etc/defaults/rc.conf.
A slow-stopping service on shutdown. rc.shutdown has a global 90-second watchdog (rcshutdown_timeout). If a database needs longer to flush, raise the timeout, or the stop gets killed halfway through.
A typo in rc.conf. It’s a shell script sourced by everything. Use sysrc, and if you edited by hand, sh -n /etc/rc.conf before rebooting.
Conclusion
The rc.d system is small enough to understand completely, and that’s its biggest strength. init runs /etc/rc. /etc/rc asks rcorder to sort scripts by their PROVIDE/REQUIRE comments and runs them in two passes, base first and /usr/local after the filesystems are mounted. Every script sources the same layered shell configuration, declares a few variables, and lets rc.subr do the work. Shutdown is the same list, filtered by KEYWORD: shutdown and run in reverse.
The mental model to carry away: an rc.d script declares a service; it doesn’t implement one. Declare name, rcvar, command and pidfile, add hooks only where you need behaviour, and every admin knob (_user, _env, _fib, _nice, _svcj) works for free. When you need supervision, daemon(8) provides it, and as long as the pidfile points at the supervisor (-P), start and stop behave as expected.
If you take four habits from this article: test with service, never by running the script directly; use -P whenever you use daemon -r; put KEYWORD: shutdown on anything that holds state; and when in doubt, sh -x the script and read what rc.subr actually built. It’s all shell. You can always just read it.
The next article in this series is the one I promised last time: the base system versus ports and packages, including where the software comes from, why base and third-party software are deliberately separate, and how freebsd-update, pkgbase, pkg and poudriere fit together.
References
- FreeBSD Handbook - Managing Services in FreeBSD
- Practical rc.d scripting in BSD - the canonical tutorial
- rc(8) man page
- rc.subr(8) man page - every variable
run_rc_commandunderstands - rc.conf(5) man page
- rcorder(8) man page
- service(8) man page
- sysrc(8) man page
- daemon(8) man page - see the note on
--supervisor-pidfile - FreeBSD Handbook - Service Jails
Comments
You can use your Mastodon or other ActivityPub account to comment on this article by replying to the associated post.
Search for the copied link on your Mastodon instance to reply.
Loading comments...