Architecture

Guarding CLI Startup with Import Tests

Stop heavy modules creeping into a Python CLI’s startup path: a deterministic test of what importing the CLI loads, a forbidden-modules list, and a tool that shows the import chain.

Updated

Startup time regresses one innocent line at a time. Someone adds from rich.table import Table to the top of a command module to format a report, and every invocation of the CLI — including --help and shell completion — now loads Rich's table machinery whether or not it prints a table. A wall-clock budget test, as in profiling Python CLI startup time, catches the big regressions, but timing tests are noisy on shared CI runners and say nothing about what went wrong. A complementary, deterministic guard asks a simpler question: after importing the CLI's entry point, which modules are loaded? If anything on a forbidden list is among them, the test fails — every time, on every machine — and a small helper prints the exact chain of imports responsible. This guide builds both. It belongs to the startup performance topic.

Prerequisites

What to guard

The idea is to write down the modules that must not be imported just to start the CLI, with the reason for each:

A forbidden-imports list Example heavy modules a command line tool must not import at startup, with the reason each is deferred. A forbidden-imports list Module Why it is deferred pandas only the export command needs it boto3 only the s3 backend needs it httpx most commands work offline rich.table (plain import only) Typer loads Rich for help anyway The reason string becomes the failure message — it tells contributors where the import belongs.

Good candidates are packages that are large, slow to import and needed by only some commands: data libraries (pandas, numpy, polars), cloud SDKs (boto3, google-cloud-*), HTTP clients when most commands work offline, template engines, YAML parsers, Rich's heavier modules (rich.table, rich.syntax, rich.markdown) and Textual. Measuring a fresh interpreter is instructive: importing Typer itself takes around 25 ms and does not import Rich — Typer loads it only when it renders help or an error — so a command module that imports rich.table at the top undoes that work for every run.

The recipe

A deterministic import test

Run the import in a fresh interpreter — the test process has long since imported everything — and compare sys.modules against the forbidden list:

# tests/test_startup_imports.py
import json
import subprocess
import sys

import pytest

HEAVY = {
    "pandas": "only the export command needs it",
    "boto3": "only the s3 backend needs it",
    "httpx": "most commands work offline",
    "yaml": "only YAML output needs it",
}
RICH = {"rich.table": "import inside the commands that print tables"}

# What each startup path may not load. Rendering Typer's help legitimately loads Rich.
ENTRY_POINTS = {
    "import mytool.cli": {**HEAVY, **RICH},
    "from mytool.cli import app; app(['--help'], standalone_mode=False)": HEAVY,
}


def loaded_modules(code: str) -> set[str]:
    probe = f"import json, sys\n{code}\nprint(json.dumps(sorted(sys.modules)))"
    result = subprocess.run([sys.executable, "-c", probe], capture_output=True, text=True,
                            check=True)
    return set(json.loads(result.stdout.splitlines()[-1]))


@pytest.mark.parametrize("code", ENTRY_POINTS)
def test_startup_does_not_import_heavy_modules(code):
    loaded = loaded_modules(code)
    offenders = {m: why for m, why in ENTRY_POINTS[code].items() if m in loaded}
    assert not offenders, (
        "heavy modules imported at startup: "
        + "; ".join(f"{m} ({why})" for m, why in offenders.items())
        + " — run scripts/import_chain.py to see who imported them"
    )

Testing two entry points matters: a plain import catches module-level imports, while invoking --help also exercises whatever the framework does to render help, including command registration code that runs at that point. The two paths get different lists for a reason found the hard way: the first version of this test forbade rich.table everywhere and failed on --help even after the offending command was fixed, because Typer renders its help panels with Rich tables. That is fine — Rich is only loaded when help is shown — so Rich modules are forbidden for the plain import and allowed for help, while the genuinely heavy packages are forbidden on both paths. standalone_mode=False keeps the help call from exiting the probe before it prints the module list; the last line of output is the JSON, so anything the help printed is ignored.

Find out who imported it

When the test fails, the next question is always "which import pulled this in?" Python's -X importtime option prints every import as a tree, indented by depth, with children listed before their parents. Walking that output backwards from the forbidden module up to the shallower entries gives the chain:

# scripts/import_chain.py
"""Show the chain of imports that loads TARGET when importing MODULE."""
from __future__ import annotations

import re
import subprocess
import sys

LINE = re.compile(r"import time:\s+\d+ \|\s+\d+ \|( *)(\S+)")


def import_chain(module: str, target: str, python: str = sys.executable) -> list[str]:
    stderr = subprocess.run([python, "-X", "importtime", "-c", f"import {module}"],
                            capture_output=True, text=True, check=True).stderr
    entries = [(len(m.group(1)), m.group(2)) for m in map(LINE.match, stderr.splitlines()) if m]
    hits = [i for i, (_, name) in enumerate(entries) if name == target] or \
           [i for i, (_, name) in enumerate(entries) if name.startswith(target + ".")]
    if not hits:
        return []
    depth, chain = entries[hits[0]][0], [entries[hits[0]][1]]
    for d, name in entries[hits[0] + 1:]:
        if d < depth:                       # a shallower entry after us is our importer
            chain.append(name)
            depth = d
    return chain[::-1]


if __name__ == "__main__":
    chain = import_chain(sys.argv[1], sys.argv[2])
    print(" -> ".join(chain) if chain else f"{sys.argv[2]} is not imported")

Against a CLI whose report command imports rich.table at the top:

$ python scripts/import_chain.py mytool.cli rich.table
mytool.cli -> mytool.commands.report -> rich.table

The fix is then obvious and local: move from rich.table import Table inside the function in mytool/commands/report.py that builds the table.

From failure to fix Terminal session where the import test fails, the import chain script shows which module imported Rich tables, and the test passes after moving the import. From failure to fix bash $ pytest -q tests/test_startup_imports.py E heavy modules imported at startup: rich.table (import inside the commands…) $ python scripts/import_chain.py mytool.cli rich.table mytool.cli -> mytool.commands.report -> rich.table # move the import inside build() in mytool/commands/report.py $ pytest -q tests/test_startup_imports.py 2 passed Deterministic, explained, and fixed in one file.

Choosing between budgets and import tests

The two kinds of guard complement each other:

Time budgets and import tests A comparison of wall-clock startup budget tests with deterministic forbidden-import tests. Time budgets and import tests Property Time budget Import test Catches any slowness listed modules only Noise on CI yes — needs margin none Explains the cause no yes, with the chain Measures what users feel yes indirectly Use the import test as the gate and a loose budget as the backstop.

A time budget measures what users feel and catches slowness from any cause — a slow plugin discovery, an expensive computation at import — but it is noisy, needs generous margins on CI, and says nothing about the cause. An import test is exact and explains itself, but only guards against the modules you thought to list. Use the import test as the everyday gate and keep a loose time budget as a backstop.

UX considerations

The users of this test are contributors adding features:

  • Explain every entry. The reason string appears in the failure message; "only the export command needs it" tells a contributor exactly where the import belongs.
  • Point at the tool. The assertion message names scripts/import_chain.py, so nobody has to know about -X importtime.
  • Keep the list short and meaningful. Five or ten genuinely heavy modules. A list of fifty standard-library modules turns the test into noise.
  • Review additions to the list. Adding a module to FORBIDDEN is a performance decision worth a sentence in the pull request.

Testing the behaviour

The import-chain parser is pure string processing on real -X importtime output, so test it with a captured sample:

# tests/test_import_chain.py
from scripts import import_chain as ic

SAMPLE = """\
import time: self [us] | cumulative | imported package
import time:        65 |         65 |   mytool
import time:        68 |         68 |   mytool.commands
import time:       164 |        259 |       rich
import time:      1022 |       8598 |     rich.table
import time:        66 |       9356 |   mytool.commands.report
import time:        94 |      31732 | mytool.cli
"""


class FakeResult:
    stderr = SAMPLE


def test_chain_walks_up_to_the_entry_module(monkeypatch):
    monkeypatch.setattr(ic.subprocess, "run", lambda *a, **k: FakeResult())
    assert ic.import_chain("mytool.cli", "rich") == [
        "mytool.cli", "mytool.commands.report", "rich.table", "rich"]


def test_missing_target_returns_empty(monkeypatch):
    monkeypatch.setattr(ic.subprocess, "run", lambda *a, **k: FakeResult())
    assert ic.import_chain("mytool.cli", "pandas") == []

Then see the real failure once: add a top-level import yaml to a command module, run the suite, and follow the message to the chain script.

Conclusion

A forbidden-imports test turns "startup got slower" into a precise, deterministic failure: import the CLI in a fresh interpreter (and render --help), compare sys.modules with a short list of heavy modules and the reasons they are deferred, and point contributors at a script that prints the import chain from -X importtime. Keep a loose time budget alongside it, and lazy imports stay lazy.

Frequently asked questions

Why not just check sys.modules inside the test process?

Because pytest, plugins and other tests have already imported half the world by then. A fresh interpreter in a subprocess is the only reliable view of what the CLI itself loads.

Should the standard library be on the list?

Rarely. Most standard-library modules are cheap. A few are worth deferring in very startup-sensitive tools — asyncio, email (pulled in by importlib.metadata in some versions), decimal — but measure before adding them.

Does this work for plugin-based CLIs?

Yes, and it is especially useful there: installed plugins can import heavy modules at registration time. Run the test with your bundled plugins installed, and document the expectation for third-party plugin authors, as in writing a plugin for an existing CLI.

How does this relate to import-time side effects?

They are siblings: this test checks what importing loads, the audit-hook probe in avoiding import-time side effects checks what importing does. Run both in the same test module.

Should --version and completion get their own entries?

If they have their own code paths, yes. --version should be the cheapest path of all — ideally nothing beyond the framework and importlib.metadata — and shell completion runs on every Tab press, so its path deserves the strictest list. Add each as another key in ENTRY_POINTS, with the completion environment variables your framework uses set in the probe.