Runtime

Reloading Configuration on SIGHUP in a Python CLI

Let a long-running Python CLI pick up config changes without restarting: SIGHUP handlers that set flags, validate-then-swap reloads, systemd reloads, tests.

Updated

A long-running agent reads its configuration at start-up. Then someone adds a target to the config file, and the choice is between restarting the service — dropping in-flight work, reconnecting, losing caches — or living with stale settings. Unix daemons have solved this for decades with a convention: send the process SIGHUP, and it re-reads its configuration in place. nginx, sshd, rsyslog and PostgreSQL all do it, systemctl reload triggers it, and administrators expect it. Implementing it in Python is short but has three traps: doing real work inside the signal handler, replacing working settings with broken ones, and telling the supervisor when the reload has finished. This guide builds an agent with a safe reload path — the handler only sets a flag, the main loop loads and validates the new file, a bad file keeps the old settings running — integrates it with systemd’s reload protocol, and tests it by sending real signals. It belongs to the long-running and watch-mode CLIs topic.

Prerequisites

The reload sequence

A safe reload Sequence of an administrator reloading a command line tool agent: systemd sends SIGHUP, the handler sets a flag, the main loop validates the new file and swaps settings. A safe reload admin systemd handler main loop systemctl reload mytool SIGHUP set reload + wake events RELOADING=1, load + validate READY=1 (new or previous settings) The handler notices; the loop does the work at a safe point.

The signal arrives at an arbitrary moment — in the middle of a network call, halfway through writing a file. Python runs signal handlers in the main thread between bytecode instructions, so a handler that reloads configuration directly could swap settings while code is using them. The safe structure separates noticing from doing: the handler sets an event, and the main loop, at a well-defined point between units of work, performs the reload. The new configuration is loaded and validated completely before it replaces the old one; if anything is wrong, the old settings stay in force and the error is logged.

The recipe

# src/mytool/reloadable.py
from __future__ import annotations

import logging
import signal
import threading
import time
import tomllib
from dataclasses import dataclass
from pathlib import Path

from mytool.sdnotify import notify

log = logging.getLogger("mytool.agent")


@dataclass(frozen=True)
class Settings:
    interval: float = 60.0
    targets: tuple[str, ...] = ()


def load_settings(path: Path) -> Settings:
    data = tomllib.loads(path.read_text(encoding="utf-8"))
    interval = float(data.get("interval", 60))
    if interval <= 0:
        raise ValueError("interval must be positive")
    targets = data.get("targets", [])
    if not isinstance(targets, list) or not all(isinstance(t, str) for t in targets):
        raise ValueError("targets must be a list of strings")
    return Settings(interval, tuple(targets))


class Agent:
    def __init__(self, config_path: Path) -> None:
        self.config_path = config_path
        self.settings = load_settings(config_path)       # fail fast on a bad config at start-up
        self.reload_requested = threading.Event()
        self.stop_requested = threading.Event()
        self.wake = threading.Event()

    def install_signal_handlers(self) -> None:
        # Handlers only set flags; the real work happens in the main loop, between iterations.
        signal.signal(signal.SIGHUP, lambda *_: (self.reload_requested.set(), self.wake.set()))
        signal.signal(signal.SIGTERM, lambda *_: (self.stop_requested.set(), self.wake.set()))

    def reload(self) -> bool:
        notify("RELOADING=1", f"MONOTONIC_USEC={time.monotonic_ns() // 1000}")
        try:
            new = load_settings(self.config_path)
        except (OSError, ValueError, tomllib.TOMLDecodeError) as exc:
            log.error("reload failed, keeping previous settings: %s", exc)
            notify("READY=1", "STATUS=reload failed; running with previous settings")
            return False
        old, self.settings = self.settings, new
        log.info("reloaded settings: interval %s -> %s, %d targets",
                 old.interval, new.interval, len(new.targets))
        notify("READY=1", f"STATUS=reloaded; {len(new.targets)} targets")
        return True

    def run(self, work) -> None:
        notify("READY=1")
        while not self.stop_requested.is_set():
            if self.reload_requested.is_set():
                self.reload_requested.clear()
                self.reload()
            work(self.settings)
            self.wake.wait(self.settings.interval)
            self.wake.clear()
        notify("STOPPING=1")

Handlers that only set flags

Both handlers do two things: set their own event and set wake. The main loop sleeps on wake.wait(interval) rather than time.sleep, so a signal interrupts the wait immediately — a reload takes effect within milliseconds instead of at the end of an hourly interval. Everything else, including logging, happens in the loop. A handler that logs, writes files or touches shared state runs in the middle of whatever the main thread was doing — possibly halfway through writing a log record or updating that same state — which is the classic reason to keep handlers trivial.

Validate, then swap

reload() builds a completely new Settings object before touching the running one. Parsing errors, missing files and failed validation all leave self.settings unchanged. This is the behaviour administrators rely on: a typo in a config file followed by systemctl reload must not take the service down. The settings object is frozen and replaced as a whole, so code that grabbed self.settings at the start of an iteration sees one consistent version even if a reload happens later.

At start-up the rule is the opposite: Agent.__init__ lets a bad configuration raise, so the service fails immediately and visibly instead of running with nothing. Validation rules are best shared with your other configuration code, as in validating config files with JSON Schema.

A typo does not take the service down Terminal session reloading an agent with a broken configuration file, seeing the failure in the status line, then fixing and reloading successfully. A typo does not take the service down bash $ systemctl --user reload mytool $ systemctl --user status mytool | grep Status Status: "reload failed; running with previous settings" $ vim ~/.config/mytool/agent.toml && systemctl --user reload mytool $ journalctl --user -u mytool -n1 mytool[812]: reloaded settings: interval 60.0 -> 30.0, 2 targets Validate first, swap second, report either way.

What can be reloaded

Not every setting can change in place. Intervals, target lists, log levels, feature flags and rate limits are easy: the next iteration reads the new value. A listening port, a database path or a worker count requires tearing down and rebuilding resources; either implement that explicitly (open the new resource, switch, close the old) or log that the setting “requires a restart” and keep the old value. Log what actually changed — the example logs the interval and target count — so the effect of a reload is visible in the journal.

systemd integration

With the unit from the service guide, add a reload command. For any systemd version:

[Service]
Type=notify
ExecStart=/home/ann/.local/bin/mytool agent
ExecReload=/bin/kill -HUP $MAINPID

systemctl reload mytool then sends SIGHUP. On its own, systemd cannot tell when the reload has finished. The notify protocol fixes that: the agent sends RELOADING=1 together with MONOTONIC_USEC (the current CLOCK_MONOTONIC time in microseconds) when it starts reloading, and READY=1 when done, as reload() does. On systemd 253 and later, Type=notify-reload goes one step further — systemd sends SIGHUP itself and waits for that handshake, so no ExecReload= line is needed and systemctl reload returns only once the new settings are live.

What a reload can change Kinds of settings in a long-running command line tool and whether they can be applied on reload or need a restart. What a reload can change Setting On reload How Interval, targets, flags yes next iteration reads it Log level yes logger.setLevel Connection pool, workers with work open new, switch, close old Listening port, env vars no log "requires restart" Log what actually changed so the effect is visible.

UX considerations

  • Document the signal. A line in the help or man page — “send SIGHUP to reload the configuration” — is how administrators find out it exists.
  • Offer a command too. mytool agent reload that finds the PID (from a PID file or systemctl show -p MainPID) and sends the signal is friendlier than kill.
  • Report failures where people look. Log the error and put it in the STATUS= line, so systemctl status shows “reload failed; running with previous settings”.
  • Check before reloading. A mytool config validate command lets administrators test a file before sending the signal.
  • Consider watching the file. For desktop agents, reloading automatically when the file changes is friendlier; building a watch mode with watchfiles shows how. Servers usually prefer the explicit signal.

Testing the behaviour

The tests run the agent in the main thread — Python only delivers signals there — while a helper thread edits the config file and sends real signals to the test process with os.kill. A semaphore released after each iteration keeps the steps in order without sleeps:

# tests/test_reload.py
import os
import signal
import threading

import pytest

from mytool.reloadable import Agent

pytestmark = pytest.mark.skipif(not hasattr(signal, "SIGHUP"), reason="POSIX signals")


@pytest.fixture
def config(tmp_path):
    path = tmp_path / "agent.toml"
    path.write_text('interval = 30\ntargets = ["a"]\n')
    return path


def run_with_signals(agent, actions):
    """Run the agent in the main thread; a helper thread performs each action after an iteration."""
    seen = []
    iterated = threading.Semaphore(0)
    agent.install_signal_handlers()

    def driver():
        for action in actions:
            iterated.acquire()                         # wait for an iteration to finish
            action()
        iterated.acquire()                             # let the loop pick the action up
        os.kill(os.getpid(), signal.SIGTERM)

    def work(settings):
        seen.append(settings)
        iterated.release()

    threading.Thread(target=driver, daemon=True).start()
    try:
        agent.run(work)
    finally:
        signal.signal(signal.SIGHUP, signal.SIG_DFL)
        signal.signal(signal.SIGTERM, signal.SIG_DFL)
    return seen


def test_sighup_applies_new_settings(config):
    agent = Agent(config)

    def edit_and_hup():
        config.write_text('interval = 30\ntargets = ["a", "b"]\n')
        os.kill(os.getpid(), signal.SIGHUP)

    seen = run_with_signals(agent, [edit_and_hup])
    assert seen[0].targets == ("a",)
    assert seen[-1].targets == ("a", "b")


def test_invalid_config_keeps_previous_settings(config, caplog):
    agent = Agent(config)

    def break_and_hup():
        config.write_text("interval = -5\n")
        os.kill(os.getpid(), signal.SIGHUP)

    seen = run_with_signals(agent, [break_and_hup])
    assert all(s.targets == ("a",) for s in seen)
    assert "keeping previous settings" in caplog.text


def test_bad_config_at_startup_fails_fast(tmp_path):
    path = tmp_path / "agent.toml"
    path.write_text("interval = 0\n")
    with pytest.raises(ValueError):
        Agent(path)

The ordering matters: a first version of the driver sent SIGHUP and SIGTERM back to back, and the loop, seeing the stop flag, exited before applying the reload — a real race the tests caught. Waiting for an iteration between the two signals models how a reload is used in practice. Restoring default handlers in finally keeps one test’s handlers from leaking into the rest of the suite.

Conclusion

Reloading on SIGHUP lets a long-running CLI pick up configuration changes without dropping work. Keep handlers to setting events, wake the main loop with an event instead of sleeping, load and validate the complete new configuration before swapping a frozen settings object, keep the old settings and report the error when the new file is bad, fail fast on bad configuration at start-up, tell systemd about reloads with RELOADING=1/READY=1 (or use Type=notify-reload), and test with real signals sent from a helper thread.

Frequently asked questions

Why SIGHUP?

SIGHUP originally meant “the terminal hung up”. Daemons have no terminal, so the signal was free, and reloading on it became a convention. A process attached to a terminal still receives it when the terminal closes — another reason a foreground CLI should only install the handler in agent mode.

Can I use SIGUSR1 instead?

You can, and some programs use SIGUSR1 for other actions such as reopening log files or dumping state. Use SIGHUP for reloading configuration; it is what administrators and systemctl reload expect.

How do asyncio programs handle this?

Use loop.add_signal_handler(signal.SIGHUP, callback), which runs the callback inside the event loop rather than in an arbitrary place, then schedule the reload as a task. The validate-then-swap rule stays the same.

What happens to a reload during shutdown?

In the recipe, the stop flag wins: if both arrive, the loop exits without reloading. Reloading a service that is about to stop has no benefit.

Should a reload re-read environment variables too?

A process cannot see changes to its environment made elsewhere, so only the file changes. Settings that come from the unit’s Environment= need a restart.