Architecture

Transcript Tests for Python CLI Commands

Write CLI tests as terminal transcripts — commands, expected output and exit codes in plain text files — collected by a small pytest plugin with readable diffs on failure.

Updated

Most CLI tests are written as Python: invoke the command with CliRunner, assert on result.output, assert on result.exit_code. That works, but for commands whose whole point is their output, the Python gets in the way. The expected text is buried in string literals, a ten-line table becomes ten in assertions, and nobody reading the test can see at a glance what the user would see. Transcript tests flip it around: the test is a terminal session — a $ mytool … line, followed by exactly what should appear, followed by the exit code if it is not zero. Tools like cram and prysk popularised the format; with pytest's collection hooks, a CLI can support it in about eighty lines of conftest.py, running commands in-process and showing a diff when output changes. This guide builds that collector and covers when transcripts beat ordinary tests and when they do not. It belongs to the CLI testing topic.

Prerequisites

The format

# tests/transcripts/greet.t — greeting basics
$ mytool greet Ann
hello Ann

$ mytool greet Bo --times 2
hello Bo
hello Bo

$ mytool greet
Usage: mytool greet [OPTIONS] {name}
Try 'mytool greet --help' for help.
╭─ Error ──────────────────────────────────────────────────────────────────────╮
│ Missing argument 'name'.                                                     │
╰──────────────────────────────────────────────────────────────────────────────╯
[exit 2]
The transcript format The line types of a command line tool transcript test file and what each means. The transcript format Line Meaning $ mytool greet Ann a command, split like a shell would hello Ann expected output, line by line [exit 2] expected non-zero exit code # comment ignored — explain why output matters A transcript reads like the terminal session it tests.

Lines starting with $ are commands; everything up to the next command is the expected combined output; an [exit N] line states a non-zero exit code; # lines are comments; trailing blank lines are ignored. The transcript reads like documentation — which it partly is — and a reviewer sees immediately what changed in a pull request that alters output.

The recipe

The collector parses .t files under tests/transcripts/, turns each command into a pytest item, runs it in-process with CliRunner, and reports a unified diff when output or exit code differs:

# tests/conftest.py
from __future__ import annotations

import difflib
import shlex
from dataclasses import dataclass, field
from pathlib import Path

import pytest
from typer.testing import CliRunner

from mytool.cli import app


@dataclass
class Step:
    line: int
    args: list[str]
    expected: list[str] = field(default_factory=list)
    exit_code: int = 0


def parse_transcript(text: str) -> list[Step]:
    steps: list[Step] = []
    for number, raw in enumerate(text.splitlines(), start=1):
        if raw.startswith("$ "):
            args = shlex.split(raw[2:])
            if args[:1] != ["mytool"]:
                raise ValueError(f"line {number}: commands must start with 'mytool'")
            steps.append(Step(number, args[1:]))
        elif raw.startswith("[exit ") and steps:
            steps[-1].exit_code = int(raw[6:-1])
        elif raw.startswith("#") or not steps:
            continue
        else:
            steps[-1].expected.append(raw)
    for step in steps:
        while step.expected and step.expected[-1] == "":
            step.expected.pop()
    return steps


def pytest_collect_file(parent, file_path: Path):
    if file_path.suffix == ".t" and "transcripts" in file_path.parts:
        return TranscriptFile.from_parent(parent, path=file_path)


class TranscriptFile(pytest.File):
    def collect(self):
        for step in parse_transcript(self.path.read_text(encoding="utf-8")):
            yield TranscriptItem.from_parent(
                self, name=f"line{step.line}: {shlex.join(step.args)}", step=step)


class TranscriptMismatch(AssertionError):
    def __init__(self, step: Step, actual: list[str], code: int) -> None:
        diff = "\n".join(difflib.unified_diff(step.expected, actual, "expected", "actual",
                                              lineterm=""))
        super().__init__(f"$ mytool {shlex.join(step.args)}\n"
                         f"exit code: expected {step.exit_code}, got {code}\n{diff}")


class TranscriptItem(pytest.Item):
    def __init__(self, *, step: Step, **kwargs) -> None:
        super().__init__(**kwargs)
        self.step = step

    def runtest(self) -> None:
        result = CliRunner().invoke(app, self.step.args, prog_name="mytool",
                                    env={"COLUMNS": "80", "NO_COLOR": "1"})
        actual = result.output.rstrip("\n").splitlines()
        if actual != self.step.expected or result.exit_code != self.step.exit_code:
            raise TranscriptMismatch(self.step, actual, result.exit_code)

    def repr_failure(self, excinfo):
        if isinstance(excinfo.value, TranscriptMismatch):
            return str(excinfo.value)
        return super().repr_failure(excinfo)

    def reportinfo(self):
        return self.path, self.step.line - 1, self.name

Three details make it pleasant to use. Each command is its own pytest item, named after its line number and arguments, so failures point straight at the right place and -k greet selects transcripts like any other test. repr_failure replaces pytest's default traceback with the command and a diff — the only things you need to see. And the environment is pinned: COLUMNS=80 makes wrapping deterministic and NO_COLOR=1 keeps escape codes out of the comparison, the same precautions as in snapshot testing CLI output. prog_name="mytool" makes usage lines show the real command name instead of the runner's default.

When output drifts, the failure reads like a code review:

$ mytool greet
exit code: expected 2, got 2
--- expected
+++ actual
@@ -1,5 +1,5 @@
-Usage: mytool greet [OPTIONS] NAME
+Usage: mytool greet [OPTIONS] {name}

That diff is what you get when a transcript written for an older Typer runs against a release that renders required arguments as {name} in usage lines instead of NAME. Nothing in your code changed, yet users will see different output — and whether that matters is a human decision, which is exactly what a transcript test is for.

A transcript failure Terminal output of a failing transcript test showing the command, exit codes and a unified diff of expected versus actual output. A transcript failure bash $ pytest -q tests/transcripts FAILED tests/transcripts/greet.t::line9: greet $ mytool greet exit code: expected 2, got 2 -Usage: mytool greet [OPTIONS] NAME +Usage: mytool greet [OPTIONS] {name} The failure is the review: a human decides whether the new output is acceptable.

When transcripts fit, and when they do not

Where transcripts fit Kinds of command line tool output and whether transcript tests suit them. Where transcripts fit Output Transcript? Instead Help and error messages ideal — README walkthroughs ideal — Timestamps, temp paths no unit tests, placeholders Business logic no tests of the core Transcripts pin user-facing text; unit tests pin behaviour.

Transcripts are excellent for user-visible behaviour that should change rarely and deliberately: help output, error messages, table layouts, the walkthrough in your README. They double as executable documentation — some projects generate a "Usage examples" page straight from the transcript files.

They are a poor fit for output that varies: timestamps, durations, temporary paths, random IDs. Either keep those out of transcript tests, or extend the format with a placeholder such as <ANY> that matches any text on a line. They also exercise only the command line, so they complement rather than replace unit tests of the logic underneath. And because they pin exact text, they need updating when the framework changes its formatting — a small, honest cost that is usually worth paying for output users depend on.

UX considerations

The users of transcripts are developers and reviewers:

  • Keep files short and themed — greet.t, errors.t, json-output.t — so a failing file name already says what broke.
  • Comment the intent. A # line saying why an error message matters ("scripts grep for 'not found'") stops someone from "fixing" it casually.
  • Treat a transcript diff as a product change. If the diff is intended, update the transcript in the same pull request, and mention user-visible output changes in the changelog.
  • Run transcripts in CI on every supported platform if output includes paths or line endings, as in running CLI tests on Windows and macOS runners.

Testing the behaviour

The parser is the part worth unit-testing, since a parsing bug would make transcripts silently test the wrong thing:

# tests/test_transcript_parser.py
import pytest

from conftest import parse_transcript

SAMPLE = """\
# comment
$ mytool greet Ann
hello Ann

$ mytool greet "Ann Lee" --times 2
hello Ann Lee
hello Ann Lee
[exit 0]
$ mytool greet
error
[exit 2]
"""


def test_steps_outputs_and_exit_codes():
    steps = parse_transcript(SAMPLE)
    assert [s.args for s in steps] == [["greet", "Ann"], ["greet", "Ann Lee", "--times", "2"],
                                       ["greet"]]
    assert steps[0].expected == ["hello Ann"]
    assert steps[2].expected == ["error"] and steps[2].exit_code == 2


def test_commands_must_name_the_tool():
    with pytest.raises(ValueError, match="must start with 'mytool'"):
        parse_transcript("$ othertool run\n")

Importing from conftest works because pytest puts the test directory on sys.path; if you prefer, move the parser into a small tests/transcripts.py module and import it from both places.

Conclusion

Transcript tests make CLI output reviewable: a plain-text file of commands, expected output and exit codes, collected by a short pytest plugin that runs each command in-process and shows a diff when anything changes. Use them for help, errors and documented examples — the output users depend on — pin the terminal width and colour, keep varying values out, and keep unit tests for the logic underneath. A changed transcript then becomes what it should be: a visible, deliberate product decision.

Frequently asked questions

Why not use cram or prysk directly?

They are mature and language-agnostic, and run commands through a real shell. The in-process collector here is faster, needs no installed console script, works on Windows without a POSIX shell, and integrates with pytest's selection and reporting. Use the external tools when you specifically want to test through a shell.

Can transcripts include stdin?

Extend the format with a < text line that feeds input to the command, and pass it as input= to CliRunner.invoke. Keep it simple: a line or two of input per command.

How do I update many transcripts after an intended change?

Add an opt-in update mode — for example an environment variable that makes runtest write the actual output back into the file instead of failing — and review the resulting diff with git before committing. Snapshot tools such as syrupy offer the same workflow for Python-based tests.

Should stderr and stdout be separated in transcripts?

The runner here compares combined output, which matches what a user sees in a terminal. If stream separation matters for a command — machine output on stdout, diagnostics on stderr — test that with an ordinary pytest test, as in testing Click commands with CliRunner.

Do transcripts slow the test suite down?

Barely. Each command runs in-process through CliRunner, so a transcript step costs about as much as an ordinary CLI test — milliseconds. Only commands that genuinely do slow work (network calls, large files) are slow, and those should use the same fakes and fixtures as your other tests, injected through environment variables the transcript runner sets.