Architecture

Testing Python CLI Commands That Call Subprocesses

Test CLI commands that run git, docker or other programs: fake processes with pytest-subprocess, fake executables on PATH, missing tools, failures, timeouts and argument checks.

Updated

Many CLIs are orchestrators: they run git, docker, kubectl, terraform or a compiler, and their real job is deciding which commands to run and what to do with the results. Testing that against the real programs is slow, depends on what is installed on the machine, and cannot easily produce the interesting cases — the tool missing from PATH, a non-zero exit with a particular error message, a hang that should hit a timeout. Mocking subprocess.run with unittest.mock works but couples tests to the exact call style. Two better techniques cover almost everything: fake processes registered with pytest-subprocess, which intercepts subprocess calls and returns scripted results, and fake executables — tiny scripts placed first on PATH — for end-to-end tests that cross a real process boundary. This guide uses both on a command that wraps git. It belongs to the CLI testing topic.

Prerequisites

The code under test

# src/mytool/gitinfo.py
from __future__ import annotations

import shutil
import subprocess
from dataclasses import dataclass


class GitError(Exception):
    pass


@dataclass(frozen=True)
class RepoState:
    branch: str
    dirty: bool


def repo_state(timeout: float = 5.0) -> RepoState:
    if shutil.which("git") is None:
        raise GitError("git is not installed or not on PATH")
    try:
        branch = subprocess.run(["git", "rev-parse", "--abbrev-ref", "HEAD"],
                                capture_output=True, text=True, timeout=timeout)
        status = subprocess.run(["git", "status", "--porcelain"],
                                capture_output=True, text=True, timeout=timeout)
    except subprocess.TimeoutExpired:
        raise GitError(f"git did not answer within {timeout:.0f}s") from None
    for result in (branch, status):
        if result.returncode != 0:
            raise GitError(result.stderr.strip() or f"git exited with {result.returncode}")
    return RepoState(branch=branch.stdout.strip(), dirty=bool(status.stdout.strip()))
# src/mytool/cli.py
import typer

from mytool.gitinfo import GitError, repo_state

app = typer.Typer()


@app.callback()
def main() -> None:
    """Repository helper."""


@app.command()
def status() -> None:
    """Show the current branch and whether there are uncommitted changes."""
    try:
        state = repo_state()
    except GitError as exc:
        typer.echo(f"error: {exc}", err=True)
        raise typer.Exit(1)
    typer.echo(f"{state.branch}{' (dirty)' if state.dirty else ''}")

The interesting behaviours are all about the subprocess: the happy path, git missing, git failing outside a repository, and git hanging.

Scenarios for a git wrapper Test scenarios for a command line tool that wraps git, the fake set up for each and the expected result. Scenarios for a git wrapper Scenario Fake Expect Clean repository branch "main", empty status main Dirty repository status lists a file feature/x (dirty) Not a repository exit 128 + real stderr exit 1, message git missing which() returns None install hint git hangs TimeoutExpired callback "did not answer" Each row is one short test that reads like the situation it covers.

The recipe: fake processes

pytest-subprocess replaces the machinery behind subprocess.run, Popen and friends for the duration of a test. You register the commands you expect and what they should produce; anything unregistered raises an error, so unexpected commands cannot slip through.

# tests/test_gitinfo.py
import subprocess

import pytest
from typer.testing import CliRunner

from mytool import gitinfo
from mytool.cli import app

runner = CliRunner()
BRANCH = ["git", "rev-parse", "--abbrev-ref", "HEAD"]
STATUS = ["git", "status", "--porcelain"]


@pytest.fixture(autouse=True)
def git_on_path(monkeypatch):
    monkeypatch.setattr(gitinfo.shutil, "which", lambda name: f"/usr/bin/{name}")


def test_clean_repository(fp):
    fp.register(BRANCH, stdout="main\n")
    fp.register(STATUS, stdout="")
    result = runner.invoke(app, ["status"])
    assert result.exit_code == 0 and result.stdout == "main\n"
    assert fp.call_count(BRANCH) == 1


def test_dirty_repository(fp):
    fp.register(BRANCH, stdout="feature/x\n")
    fp.register(STATUS, stdout=" M src/mytool/cli.py\n")
    assert runner.invoke(app, ["status"]).stdout == "feature/x (dirty)\n"


def test_not_a_repository(fp):
    fp.register(BRANCH, returncode=128,
                stderr="fatal: not a git repository (or any of the parent directories): .git\n")
    fp.register(STATUS, returncode=128, stderr="fatal: not a git repository\n")
    result = runner.invoke(app, ["status"])
    assert result.exit_code == 1
    assert "not a git repository" in result.stderr


def test_timeout_is_reported(fp):
    def hang(process):
        raise subprocess.TimeoutExpired(BRANCH, 5)

    fp.register(BRANCH, callback=hang)
    with pytest.raises(gitinfo.GitError, match="did not answer"):
        gitinfo.repo_state()


def test_git_missing(monkeypatch):
    monkeypatch.setattr(gitinfo.shutil, "which", lambda name: None)
    result = runner.invoke(app, ["status"])
    assert "git is not installed" in result.stderr

Each test reads as a scenario: these commands, these outputs, this result. fp.call_count() verifies a command actually ran, which catches code paths that silently skip the subprocess. The timeout case uses a callback that raises TimeoutExpired, so the test does not actually wait five seconds; register(..., wait=…) can simulate a slow process when you need real elapsed time.

Fakes catching an unplanned command Terminal session where tests fail because the code under test started running a git command that was not registered with the fake process fixture. Fakes catching an unplanned command bash $ pytest -q tests/test_gitinfo.py E pytest_subprocess.exceptions.ProcessNotRegisteredError: E The process 'git branch --show-current' was not registered. # the code changed which git command it runs — update the fakes on purpose Unregistered commands fail loudly, so no subprocess call goes untested.

Matching arguments flexibly

Exact argument lists keep tests honest — if the code starts passing --no-optional-locks, the test fails and someone decides whether that was intended. Where a part of the command legitimately varies, fp.any() matches any arguments in that position: fp.register(["git", "log", fp.any()], stdout=...). Use it sparingly; the arguments are usually the behaviour under test.

Scripted sequences for retries

Registrations are consumed in order, so registering the same command twice scripts a sequence — the natural way to test retry logic, such as a git fetch that fails on a flaky network and succeeds on the second attempt:

def test_fetch_retries_after_network_error(fp):
    cmd = ["git", "fetch", "origin"]
    fp.register(cmd, returncode=128, stderr="fatal: unable to access: Could not resolve host\n")
    fp.register(cmd, returncode=0)
    first = subprocess.run(cmd, capture_output=True, text=True)
    second = subprocess.run(cmd, capture_output=True, text=True)
    assert (first.returncode, second.returncode) == (128, 0)
    assert fp.call_count(cmd) == 2

In a real test, the two subprocess.run calls are inside the code under test — a retry helper like the one in retries and backoff for CLI HTTP calls, adapted for commands — and the test asserts that the command succeeded overall and ran exactly twice. A third, unregistered attempt would fail the test, which is precisely how you catch a retry loop that does not stop.

The recipe: fake executables

Fake processes never leave the Python process, so they cannot catch problems in how the command line is actually executed — quoting on Windows, environment variables passed to the child, a cwd that does not exist. For a handful of end-to-end tests, put a fake git first on PATH instead:

# tests/test_fake_git_on_path.py
import os
import stat
import subprocess
import sys
import textwrap

import pytest


@pytest.fixture
def fake_git(tmp_path, monkeypatch):
    bin_dir = tmp_path / "bin"
    bin_dir.mkdir()
    script = bin_dir / "git"
    script.write_text(textwrap.dedent(f"""\
        #!{sys.executable}
        import sys
        args = sys.argv[1:]
        with open({str(tmp_path / 'calls.log')!r}, "a") as log:
            log.write(" ".join(args) + "\\n")
        if args[:1] == ["rev-parse"]:
            print("release/2.0")
        """))
    script.chmod(script.stat().st_mode | stat.S_IEXEC)
    monkeypatch.setenv("PATH", f"{bin_dir}{os.pathsep}{os.environ['PATH']}")
    return tmp_path / "calls.log"


@pytest.mark.skipif(sys.platform == "win32", reason="shebang scripts need POSIX")
def test_status_end_to_end(fake_git):
    out = subprocess.run([sys.executable, "-c", "from mytool.cli import app; app(['status'])"],
                         capture_output=True, text=True, env=os.environ.copy())
    assert out.stdout == "release/2.0\n"
    assert fake_git.read_text().splitlines() == ["rev-parse --abbrev-ref HEAD", "status --porcelain"]

The fake is a real executable, found through PATH exactly as the real git would be, and it logs the arguments it received, so the test can assert on what actually crossed the process boundary. On Windows, the equivalent is a git.cmd or git.bat wrapper; mark these tests per platform as in running CLI tests on Windows and macOS runners.

Fake processes or fake executables? A comparison of pytest-subprocess fake processes and fake executables placed on PATH for testing commands that call subprocesses. Fake processes or fake executables? Property pytest-subprocess Fake on PATH Speed in-process, instant real process spawn Timeouts, errors scripted easily script it yourself Real argv, env, quoting not exercised exercised Windows works needs .cmd wrappers Most tests use fake processes; a few end-to-end tests use fake executables.

UX considerations

The users of these tests are maintainers reading a failure:

  • Name scenarios, not mocks. test_not_a_repository tells the reader what situation is covered; test_rev_parse_returns_128 does not.
  • Use real error text. Copy the actual stderr of the real tool into the fake, so the test proves your error handling works with what users will really see.
  • Keep unregistered commands failing. pytest-subprocess's default — raising on unexpected commands — is a feature; resist fp.allow_unregistered(True) except in narrowly scoped tests.
  • Have a few real integration tests too. A CI job with the real tool installed, running against a temporary repository, catches version-specific behaviour no fake will.

Testing the behaviour

The guide's tests are the behaviour; run them twice to see the safety net work. Change repo_state to call ["git", "branch", "--show-current"] instead, and every fake-process test fails with "process not registered", naming the new command — exactly the prompt to update the fakes deliberately. And remove the shutil.which check: test_git_missing fails, reminding you that the missing-tool path needs its own message rather than a raw FileNotFoundError from subprocess.

Conclusion

Commands that orchestrate other programs are best tested by scripting those programs' behaviour. Use pytest-subprocess to register expected commands with their output, exit codes, callbacks and timeouts, keep unexpected commands failing, and assert on call counts; add a few fake-executable tests on PATH for what only a real process boundary reveals; and keep a small set of integration tests with the real tool. Each scenario then reads like the situation it covers — clean, dirty, not a repository, missing, hung.

Frequently asked questions

Why not just patch subprocess.run with unittest.mock?

It works for one call shape, but the test then breaks when the code switches to Popen, adds check=True, or calls run twice. pytest-subprocess fakes the process layer, so tests describe what runs rather than how it is called.

Does pytest-subprocess work with asyncio subprocesses?

Yes, it supports asyncio.create_subprocess_exec as well; register commands the same way. See running async code in Typer and Click for the command side.

How do I test streaming output from a subprocess?

Register the fake with a list of lines as stdout and read it the way your code does — line by line from Popen.stdout. The patterns in streaming subprocess output in real time work unchanged against fakes.

Should I fake shutil.which or put a fake on PATH?

For unit tests, fake which (or inject the executable path) — it is quick and explicit. For end-to-end tests, a fake executable on PATH exercises the real lookup.