Many CLIs grow an “agent” mode: mytool agent syncs files every minute, mytool watch reacts to changes, mytool serve exposes a small API. Running it in a terminal tab works for a day; for real use it should start at login or boot, restart after crashes, log somewhere searchable and stop cleanly at shutdown. On Linux, systemd provides all of that, and a CLI can integrate with it with very little code. The integration has two halves: a unit file that tells systemd how to run the command, and a few lines in the program that tell systemd when the service is ready, what it is doing, and that it is still alive. This guide writes both — a Type=notify unit with a watchdog, a dependency-free sd_notify implementation, a main loop that uses it, and a service install command for per-user services — and tests the protocol without systemd running the code. It belongs to the long-running and watch-mode CLIs topic; scheduled one-shot jobs are covered in running a CLI on a schedule with cron and systemd.
Prerequisites
- A Linux system with systemd (any current distribution); user services need a logged-in session or lingering enabled.
- Graceful shutdown handling as in handling SIGTERM and graceful shutdown.
How systemd and the service talk
With Type=notify, systemd starts the process and passes the path of a Unix datagram socket in the NOTIFY_SOCKET environment variable. The service sends short text messages to it: READY=1 once start-up is complete (systemd considers the unit started only then, so units ordered after it wait correctly), STATUS=… with a human-readable line shown in systemctl status, WATCHDOG=1 periodically to prove it is not hung, and STOPPING=1 when shutting down. Everything else — logging, restarts, stopping — uses mechanisms the program already has: stdout and stderr go to the journal, and stopping sends SIGTERM.
The recipe
sd_notify in fifteen lines
# src/mytool/sdnotify.py
"""Minimal sd_notify: tell systemd we are ready, alive, reloading or stopping."""
from __future__ import annotations
import os
import socket
def notify(*fields: str) -> bool:
"""Send fields such as "READY=1" to $NOTIFY_SOCKET; a no-op outside systemd."""
address = os.environ.get("NOTIFY_SOCKET")
if not address:
return False
if address.startswith("@"): # abstract namespace socket
address = "\0" + address[1:]
with socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM | socket.SOCK_CLOEXEC) as sock:
sock.connect(address)
sock.sendall("\n".join(fields).encode())
return True
def watchdog_interval() -> float | None:
"""Half of WatchdogSec, in seconds, if the watchdog is enabled for this process."""
usec = os.environ.get("WATCHDOG_USEC")
pid = os.environ.get("WATCHDOG_PID")
if not usec or (pid and int(pid) != os.getpid()):
return None
return int(usec) / 1_000_000 / 2
The protocol is simple enough that a dependency is not needed: connect a datagram socket to the given address and send newline-separated KEY=value fields. An address starting with @ is in Linux’s abstract socket namespace, written with a leading NUL byte in Python. When NOTIFY_SOCKET is not set — running in a terminal, in tests, under another supervisor — notify does nothing, so the same code runs everywhere. watchdog_interval reads WATCHDOG_USEC, which systemd sets from WatchdogSec=, and returns half of it: pinging at half the timeout is the recommended margin.
The service loop
# src/mytool/service.py
from __future__ import annotations
import shutil
import signal
import sys
import threading
import time
from pathlib import Path
from mytool.sdnotify import notify, watchdog_interval
UNIT = """\
[Unit]
Description=mytool sync agent
Documentation=https://example.com/mytool/docs
After=network-online.target
Wants=network-online.target
[Service]
Type=notify
ExecStart={exe} agent
Restart=on-failure
RestartSec=5
WatchdogSec=30
TimeoutStopSec=20
Environment=PYTHONUNBUFFERED=1
NoNewPrivileges=yes
PrivateTmp=yes
[Install]
WantedBy=default.target
"""
def render_unit(exe: str | None = None) -> str:
exe = exe or shutil.which("mytool") or f"{sys.executable} -m mytool"
return UNIT.format(exe=exe)
def user_unit_path() -> Path:
return Path.home() / ".config/systemd/user/mytool.service"
def run_agent(work, *, interval: float = 5.0, stop: threading.Event | None = None) -> None:
"""Main loop: signal readiness, do work, pet the watchdog, stop cleanly on SIGTERM."""
stop = stop or threading.Event()
if threading.current_thread() is threading.main_thread():
signal.signal(signal.SIGTERM, lambda *_: stop.set())
ping = watchdog_interval()
notify("READY=1", "STATUS=waiting for first run")
last_ping = time.monotonic()
while not stop.is_set():
count = work()
notify(f"STATUS=synced {count} files at {time.strftime('%H:%M:%S')}")
# Wait in short slices so the watchdog keeps being fed during long intervals.
deadline = time.monotonic() + interval
while not stop.is_set() and time.monotonic() < deadline:
if ping and time.monotonic() - last_ping >= ping:
notify("WATCHDOG=1")
last_ping = time.monotonic()
stop.wait(min(1.0, max(0.0, deadline - time.monotonic())))
notify("STOPPING=1")
The loop sends READY=1 only after setup (configuration loaded, connections established in a real agent), reports a short status after each iteration, and feeds the watchdog while waiting. The inner wait uses short slices so a long interval — an hourly sync — does not starve a 30-second watchdog. SIGTERM sets an event, the loop finishes its current iteration and reports STOPPING=1, and TimeoutStopSec=20 tells systemd how long to wait before escalating to SIGKILL.
The watchdog is the feature that justifies Type=notify for most CLIs. A process that has deadlocked on a network call or a stuck lock is still running, so Restart=on-failure alone would never notice. With WatchdogSec=30, systemd kills and restarts the service if no WATCHDOG=1 arrives for 30 seconds. Ping from the place that proves real progress — the main loop, not a separate thread that keeps ticking while the work is stuck.
The unit file and an install command
# src/mytool/cli.py
import subprocess
from typing import Annotated
import typer
from mytool.service import render_unit, run_agent, user_unit_path
app = typer.Typer()
service = typer.Typer(help="Run mytool as a background service (systemd).")
app.add_typer(service, name="service")
@app.command()
def agent(interval: float = 60.0) -> None:
"""Run the sync loop in the foreground (systemd starts this)."""
run_agent(lambda: 3, interval=interval)
@service.command("unit")
def print_unit() -> None:
"""Print the systemd unit file."""
typer.echo(render_unit(), nl=False)
@service.command()
def install(start: Annotated[bool, typer.Option(help="Enable and start it now.")] = True) -> None:
"""Install a systemd user service for the current user."""
path = user_unit_path()
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(render_unit(), encoding="utf-8")
typer.echo(f"wrote {path}", err=True)
subprocess.run(["systemctl", "--user", "daemon-reload"], check=True)
if start:
subprocess.run(["systemctl", "--user", "enable", "--now", "mytool.service"], check=True)
typer.echo("started; follow logs with: journalctl --user -u mytool -f", err=True)
service unit prints the unit for inspection or for packagers; service install writes it to ~/.config/systemd/user/ and enables it with systemctl --user. The unit’s settings each have a job: After=/Wants=network-online.target delays start until the network is up; Restart=on-failure with RestartSec=5 restarts after crashes without spinning; PYTHONUNBUFFERED=1 makes log lines appear in the journal immediately rather than in 8 KB bursts; NoNewPrivileges and PrivateTmp are cheap hardening that rarely affect a CLI. ExecStart must be an absolute path, so the generator resolves the installed mytool executable — for a pipx or uv tool install that is the shim in ~/.local/bin.
User services or system services
A user service runs as the user, starts when they log in, and is managed with systemctl --user — no root needed, which suits personal agents such as sync tools. To keep it running when the user is logged out (on a server, for instance), an administrator enables lingering once: loginctl enable-linger username. A system service in /etc/systemd/system/ starts at boot and should run as a dedicated unprivileged user (User=mytool, or DynamicUser=yes for stateless services); packaging the unit with a .deb or .rpm is the usual way to ship one. The program code is identical for both.
UX considerations
- Log to stdout/stderr, not files. The journal captures both with timestamps and metadata;
journalctl --user -u mytool -ffollows them. Structured fields are covered in sending CLI logs to journald and syslog. - Keep
STATUS=short and current. It is the line users see first insystemctl status; “synced 3 files at 14:02:11” beats “running”. - Print the next step. After installing, tell users how to see logs and how to stop the service.
- Offer an uninstall.
service uninstallshoulddisable --now, remove the file anddaemon-reload. - Support reload. If the agent reads configuration, implement
SIGHUPreloading as in reloading config on SIGHUP.
Testing the behaviour
The notify protocol is easy to test without systemd: bind a Unix datagram socket in a temporary directory, point NOTIFY_SOCKET at it, run the loop and read what arrived. A real systemd-analyze verify checks the unit file where systemd is available:
# tests/test_service.py
import os
import shutil
import socket
import subprocess
import threading
import pytest
from mytool.sdnotify import notify, watchdog_interval
from mytool.service import render_unit, run_agent
@pytest.fixture
def notify_socket(tmp_path, monkeypatch):
path = str(tmp_path / "notify.sock")
sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
sock.bind(path)
sock.settimeout(2)
monkeypatch.setenv("NOTIFY_SOCKET", path)
yield sock
sock.close()
def received(sock) -> list[str]:
messages = []
sock.setblocking(False)
while True:
try:
messages.append(sock.recv(4096).decode())
except BlockingIOError:
return messages
def test_notify_is_a_noop_outside_systemd(monkeypatch):
monkeypatch.delenv("NOTIFY_SOCKET", raising=False)
assert notify("READY=1") is False
def test_agent_reports_ready_status_and_stopping(notify_socket):
stop = threading.Event()
calls = []
def work():
calls.append(1)
stop.set() # one iteration, then shut down
return 7
run_agent(work, interval=0.1, stop=stop)
messages = received(notify_socket)
assert messages[0].startswith("READY=1\n")
assert any(m.startswith("STATUS=synced 7 files") for m in messages)
assert messages[-1] == "STOPPING=1"
def test_watchdog_is_fed_during_long_waits(notify_socket, monkeypatch):
monkeypatch.setenv("WATCHDOG_USEC", "400000") # WatchdogSec=0.4 → ping every 0.2 s
monkeypatch.setenv("WATCHDOG_PID", str(os.getpid()))
assert watchdog_interval() == 0.2
stop = threading.Event()
threading.Timer(1.5, stop.set).start()
run_agent(lambda: 0, interval=10, stop=stop)
assert received(notify_socket).count("WATCHDOG=1") >= 1
def test_watchdog_ignores_other_pids(monkeypatch):
monkeypatch.setenv("WATCHDOG_USEC", "1000000")
monkeypatch.setenv("WATCHDOG_PID", "1")
assert watchdog_interval() is None
@pytest.mark.skipif(not shutil.which("systemd-analyze"), reason="needs systemd")
def test_unit_file_is_valid(tmp_path):
unit = tmp_path / "mytool.service"
unit.write_text(render_unit(exe="/bin/true"))
result = subprocess.run(["systemd-analyze", "--user", "verify", str(unit)],
capture_output=True, text=True)
assert result.returncode == 0, result.stderr
The watchdog test sets a 0.4-second watchdog and a 10-second work interval, then stops the loop after 1.5 seconds; at least one WATCHDOG=1 must have arrived, which proves long waits do not starve the watchdog. systemd-analyze verify catches typos in directive names and invalid values — the kind of error that otherwise appears only as a cryptic failure when a user installs the service.
Conclusion
A long-running CLI becomes a reliable service with a small unit file and a few notifications. Use Type=notify; send READY=1 after setup, a short STATUS= after each unit of work, WATCHDOG=1 from the main loop at half of WatchdogSec, and STOPPING=1 on SIGTERM; implement sd_notify with a datagram socket and make it a no-op elsewhere; generate the unit with an absolute ExecStart, restart and hardening options; install user services with systemctl --user; log to stdout; and test the protocol against a temporary socket.
Frequently asked questions
Do I need the systemd or sdnotify Python packages?
No. The protocol is a single datagram per message, and the fifteen-line function above covers it. The packages are convenient, but every dependency is something to install on servers.
What if my CLI also runs on macOS or Windows?
notify is a no-op without NOTIFY_SOCKET, so the agent runs anywhere. Service installation differs: launchd property lists on macOS, a scheduled task or a service wrapper on Windows. Keep service install Linux-only and say so in its help.
Why Type=notify rather than Type=simple?
With Type=simple, systemd considers the service started the moment the process is forked, even if it then fails while loading configuration. notify reports readiness precisely and enables the watchdog and status line. Type=exec is a middle ground that at least waits for the program to start executing.
How do I pass secrets to the service?
Not in the unit file, which is world-readable for system units. Use EnvironmentFile= with a 0600 file, or systemd credentials (LoadCredential=), which appear as files under $CREDENTIALS_DIRECTORY; see reading secrets from env and files.
Can the service restart itself after an upgrade?
The upgrade process should run systemctl --user restart mytool after replacing the files; a running Python process keeps using the modules it already imported, so a restart is needed for new code to take effect.