Architecture

Avoiding Import-Time Side Effects in a Python CLI

Keep importing a CLI free of work: move config loading, logging setup, directory creation and network calls into functions, and catch regressions with an audit-hook test.

Updated

Every line at the top level of a Python module runs the moment the module is imported. In a CLI that is almost every module, on every invocation — for --help, for --version, for each Tab press in shell completion, and once per test that imports the package. Code that does things at import time — loads configuration, configures logging, creates directories, opens a database, calls an API, reads environment variables into constants — therefore runs far more often than anyone intended, slows every one of those paths down, and fails in surprising places. A CLI whose import creates ~/.mytool crashes in a container with no writable home before it can even print help. This guide catalogues the common side effects, shows how to move each into explicit code paths, and adds a test that uses Python's audit hooks to catch new ones. It belongs to the startup performance topic.

Prerequisites

What counts as a side effect

Defining things is fine at import time: functions, classes, constants computed from literals, the Typer app object and its command registrations. Doing things is not:

What may happen at import time Things that are fine to do at module level in a command line tool and side effects that belong in functions instead. What may happen at import time Fine at import ✓ Define functions, classes, constants ✓ Create the Typer app, register commands ✓ logging.getLogger(__name__) ✓ Cheap, standard-library imports Move into functions ✗ logging.basicConfig, signal handlers ✗ mkdir, file writes, opening databases ✗ Reading config or env into constants ✗ Network calls and heavy imports Importing should define things, never do things.
# src/badtool/cli.py — everything below the imports runs on every import
import logging
import os
from pathlib import Path

import typer

logging.basicConfig(level=logging.INFO)                 # reconfigures the host's logging
CONFIG_DIR = Path(os.environ.get("HOME", "/tmp")) / ".badtool"
CONFIG_DIR.mkdir(exist_ok=True)                         # writes to disk
(CONFIG_DIR / "last-start").write_text("x")             # ...on every --help and Tab press
app = typer.Typer()

Run with a HOME that does not exist — common in containers and CI sandboxes — and importing the module fails:

File ".../badtool/cli.py", line 9, in <module>
    CONFIG_DIR.mkdir(exist_ok=True)
FileNotFoundError: [Errno 2] No such file or directory: '/nonexistent/.badtool'

The user never got as far as running a command. The same pattern with a config file that has a syntax error means mytool --help cannot print help explaining how to fix the config.

The recipe

Move work into functions that run when needed

# src/goodtool/config.py
from __future__ import annotations

import functools
import os
from dataclasses import dataclass
from pathlib import Path


@dataclass(frozen=True)
class Settings:
    config_dir: Path
    verbose: bool


@functools.cache
def settings() -> Settings:
    """Read the environment once, on first use."""
    base = Path(os.environ.get("GOODTOOL_HOME") or Path.home() / ".goodtool")
    return Settings(config_dir=base, verbose=os.environ.get("GOODTOOL_VERBOSE") == "1")


def ensure_config_dir() -> Path:
    path = settings().config_dir
    path.mkdir(parents=True, exist_ok=True)
    return path
# src/goodtool/cli.py
import logging

import typer

from goodtool.config import ensure_config_dir, settings

app = typer.Typer()
log = logging.getLogger(__name__)          # getting a logger is fine; configuring it is not


@app.callback()
def main(verbose: bool = typer.Option(False, "--verbose", "-v")) -> None:
    """Good tool."""
    logging.basicConfig(level=logging.DEBUG if verbose or settings().verbose else logging.WARNING)


@app.command()
def init() -> None:
    """Create the configuration directory."""
    typer.echo(f"configuration in {ensure_config_dir()}")

Three patterns cover most cases. Cached accessors (functools.cache) replace module-level constants computed from the environment or files: the work happens once, on first use, and tests can call settings.cache_clear() after changing the environment. Explicit setup in the callback replaces module-level logging and signal configuration: it runs only when the CLI actually runs, after options such as --verbose are known — see adding verbose and quiet logging flags. Commands own their resources: the command that needs a directory, a database connection or an HTTP client creates it, as in dependency injection patterns for CLI commands.

Moving side effects to the right place Common import-time side effects in command line tools and where each should move. Moving side effects to the right place At import Move to CONFIG = load_config() @functools.cache accessor logging.basicConfig(...) root callback, after --verbose is known CONFIG_DIR.mkdir() the command that needs it client = httpx.Client() created per command, injected import pandas inside the function that uses it --help and --version then work even with a broken config or no home directory.

Imports of heavy modules are a side effect too

import pandas at the top of a module is not wrong in itself, but it costs hundreds of milliseconds on every import of that module. In a CLI, import heavy libraries inside the function that uses them, and keep command modules light enough that building the command tree is cheap. Lazy-loading subcommands for faster startup takes this to the command level.

UX considerations

  • --help and --version must always work. They are what a confused user runs first. If they depend on a readable config file or a writable home directory, the user is stuck.
  • Fail with context, not at import. A config error discovered inside a command can name the file and the line; one discovered at import is a traceback from deep inside your package.
  • Respect the host's logging. A CLI that is also used as a library must never call logging.basicConfig at import; it would reconfigure the importing application's logging.
  • Keep completion fast. Shell completion imports your CLI on every Tab; anything done at import time is paid for on every keystroke. See enabling tab completion in Click and Typer.

Testing the behaviour

Python's audit hooks (sys.addaudithook) report low-level events — files opened for writing, directories created, sockets connected, processes started — as they happen. Importing the CLI in a subprocess with a hook installed shows exactly what the import does:

# tests/import_probe.py
"""Import a module and print the side effects observed by an audit hook, as JSON."""
import importlib
import json
import sys

WATCH = {"open", "os.mkdir", "os.remove", "os.rename", "socket.connect", "subprocess.Popen"}
events: list[list[str]] = []


def hook(event: str, args: tuple) -> None:
    if event not in WATCH:
        return
    if event == "open":
        mode = args[1]
        if not isinstance(mode, str) or not any(c in mode for c in "wax+"):
            return                                  # reads are fine
    events.append([event, str(args[0]) if args else ""])


sys.dont_write_bytecode = True                      # Python's own .pyc writes are not ours
sys.addaudithook(hook)
importlib.import_module(sys.argv[1])
print(json.dumps(events))
# tests/test_import_side_effects.py
import json
import os
import subprocess
import sys
from pathlib import Path

PROBE = Path(__file__).with_name("import_probe.py")


def import_effects(module: str, tmp_path: Path) -> list[list[str]]:
    env = {**os.environ, "HOME": str(tmp_path / "missing-home"),
           "PYTHONPATH": os.pathsep.join(sys.path)}
    result = subprocess.run([sys.executable, str(PROBE), module], env=env,
                            capture_output=True, text=True)
    assert result.returncode == 0, f"importing {module} failed:\n{result.stderr}"
    return json.loads(result.stdout)


def test_importing_the_cli_does_nothing(tmp_path):
    assert import_effects("goodtool.cli", tmp_path) == []


def test_help_works_without_a_home_directory(tmp_path):
    env = {**os.environ, "HOME": str(tmp_path / "missing-home"),
           "PYTHONPATH": os.pathsep.join(sys.path)}
    result = subprocess.run([sys.executable, "-m", "goodtool", "--help"], env=env,
                            capture_output=True, text=True)
    assert result.returncode == 0 and "Usage" in result.stdout
What the audit-hook probe sees Terminal output of the import probe reporting a directory creation and a file write for a module with side effects, and nothing for the refactored module. What the audit-hook probe sees bash $ HOME=/tmp/h2 python tests/import_probe.py badtool.cli [["os.mkdir", "/tmp/h2/.badtool"], ["open", "/tmp/h2/.badtool/last-start"]] $ HOME=/tmp/h2 python tests/import_probe.py goodtool.cli [] Each event names the exact path the import touched.

The sys.dont_write_bytecode line is not optional. On a cold import, Python itself writes .pyc files into __pycache__ directories, and the hook dutifully reports them — the first version of this test failed on a clean checkout and passed on every run after, which is the worst kind of flaky. Disabling bytecode writing for the probe process removes that noise without hiding anything your code does. Running the probe in a subprocess matters too: by the time the test runner itself is going, your package has probably been imported already, and an audit hook cannot be removed once added. Pointing HOME at a directory that does not exist reproduces the container case. Run against the "bad" module above, the probe reports the os.mkdir and the write to last-start; against the refactored one it reports nothing. Add every command module to the test — a parametrised list of module names — so a new side effect in any of them fails CI with the exact file it touched.

Conclusion

Importing a CLI should define things, never do things. Move configuration loading into cached accessors, logging and signal setup into the root callback, resource creation into the commands that need them, and heavy imports into the functions that use them. Then keep it that way with an audit-hook probe that imports each module in a subprocess and fails on any write, network connection or process launch — and a test that --help works without a home directory.

Frequently asked questions

Is reading an environment variable at import time a side effect?

It does not change anything, but it freezes the value at import time, so tests that set the variable later and code that changes it at runtime see a stale value. A cached accessor with cache_clear() keeps the convenience without the trap.

What about registering plugins at import time?

Discovering entry points and importing plugins on every import is expensive and can run third-party code before your CLI has parsed its options. Discover plugins lazily, when the command tree is built or a plugin command runs, as in discovering plugins with entry points.

Does if __name__ == "__main__": solve this?

Only for code run as a script. Installed CLIs run through a console-script entry point that imports your module and calls main(), so the guard never fires; the fix is to keep top-level code free of work and put everything in main() or the commands.

Are audit hooks safe to use in tests?

Yes — they are a standard, documented interface designed for exactly this kind of observation. They cannot be removed once installed, which is why the probe runs in its own process.

How do I find existing side effects in a large codebase?

Run the probe against every module in the package — pkgutil.walk_packages lists them — and print the events per module. Then search for the usual suspects at module level: basicConfig, mkdir, open(, connect(, load_dotenv, signal.signal and environment reads assigned to constants. Fix the modules imported by the CLI's entry point first; they are the ones paid for on every invocation.