Most CLI tests are written as Python: invoke the command with CliRunner, assert on result.output, assert on result.exit_code. That works, but for commands whose whole point is their output, the Python gets in the way. The expected text is buried in string literals, a ten-line table becomes ten in assertions, and nobody reading the test can see at a glance what the user would see. Transcript tests flip it around: the test is a terminal session — a $ mytool … line, followed by exactly what should appear, followed by the exit code if it is not zero. Tools like cram and prysk popularised the format; with pytest's collection hooks, a CLI can support it in about eighty lines of conftest.py, running commands in-process and showing a diff when output changes. This guide builds that collector and covers when transcripts beat ordinary tests and when they do not. It belongs to the CLI testing topic.
Prerequisites
- pytest and a Typer or Click CLI; testing Click commands with CliRunner covers the runner used underneath.
- A rough idea of pytest plugins:
conftest.pyhooks are enough here.
The format
# tests/transcripts/greet.t — greeting basics
$ mytool greet Ann
hello Ann
$ mytool greet Bo --times 2
hello Bo
hello Bo
$ mytool greet
Usage: mytool greet [OPTIONS] {name}
Try 'mytool greet --help' for help.
╭─ Error ──────────────────────────────────────────────────────────────────────╮
│ Missing argument 'name'. │
╰──────────────────────────────────────────────────────────────────────────────╯
[exit 2]
Lines starting with $ are commands; everything up to the next command is the expected combined output; an [exit N] line states a non-zero exit code; # lines are comments; trailing blank lines are ignored. The transcript reads like documentation — which it partly is — and a reviewer sees immediately what changed in a pull request that alters output.
The recipe
The collector parses .t files under tests/transcripts/, turns each command into a pytest item, runs it in-process with CliRunner, and reports a unified diff when output or exit code differs:
# tests/conftest.py
from __future__ import annotations
import difflib
import shlex
from dataclasses import dataclass, field
from pathlib import Path
import pytest
from typer.testing import CliRunner
from mytool.cli import app
@dataclass
class Step:
line: int
args: list[str]
expected: list[str] = field(default_factory=list)
exit_code: int = 0
def parse_transcript(text: str) -> list[Step]:
steps: list[Step] = []
for number, raw in enumerate(text.splitlines(), start=1):
if raw.startswith("$ "):
args = shlex.split(raw[2:])
if args[:1] != ["mytool"]:
raise ValueError(f"line {number}: commands must start with 'mytool'")
steps.append(Step(number, args[1:]))
elif raw.startswith("[exit ") and steps:
steps[-1].exit_code = int(raw[6:-1])
elif raw.startswith("#") or not steps:
continue
else:
steps[-1].expected.append(raw)
for step in steps:
while step.expected and step.expected[-1] == "":
step.expected.pop()
return steps
def pytest_collect_file(parent, file_path: Path):
if file_path.suffix == ".t" and "transcripts" in file_path.parts:
return TranscriptFile.from_parent(parent, path=file_path)
class TranscriptFile(pytest.File):
def collect(self):
for step in parse_transcript(self.path.read_text(encoding="utf-8")):
yield TranscriptItem.from_parent(
self, name=f"line{step.line}: {shlex.join(step.args)}", step=step)
class TranscriptMismatch(AssertionError):
def __init__(self, step: Step, actual: list[str], code: int) -> None:
diff = "\n".join(difflib.unified_diff(step.expected, actual, "expected", "actual",
lineterm=""))
super().__init__(f"$ mytool {shlex.join(step.args)}\n"
f"exit code: expected {step.exit_code}, got {code}\n{diff}")
class TranscriptItem(pytest.Item):
def __init__(self, *, step: Step, **kwargs) -> None:
super().__init__(**kwargs)
self.step = step
def runtest(self) -> None:
result = CliRunner().invoke(app, self.step.args, prog_name="mytool",
env={"COLUMNS": "80", "NO_COLOR": "1"})
actual = result.output.rstrip("\n").splitlines()
if actual != self.step.expected or result.exit_code != self.step.exit_code:
raise TranscriptMismatch(self.step, actual, result.exit_code)
def repr_failure(self, excinfo):
if isinstance(excinfo.value, TranscriptMismatch):
return str(excinfo.value)
return super().repr_failure(excinfo)
def reportinfo(self):
return self.path, self.step.line - 1, self.name
Three details make it pleasant to use. Each command is its own pytest item, named after its line number and arguments, so failures point straight at the right place and -k greet selects transcripts like any other test. repr_failure replaces pytest's default traceback with the command and a diff — the only things you need to see. And the environment is pinned: COLUMNS=80 makes wrapping deterministic and NO_COLOR=1 keeps escape codes out of the comparison, the same precautions as in snapshot testing CLI output. prog_name="mytool" makes usage lines show the real command name instead of the runner's default.
When output drifts, the failure reads like a code review:
$ mytool greet
exit code: expected 2, got 2
--- expected
+++ actual
@@ -1,5 +1,5 @@
-Usage: mytool greet [OPTIONS] NAME
+Usage: mytool greet [OPTIONS] {name}
That diff is what you get when a transcript written for an older Typer runs against a release that renders required arguments as {name} in usage lines instead of NAME. Nothing in your code changed, yet users will see different output — and whether that matters is a human decision, which is exactly what a transcript test is for.
When transcripts fit, and when they do not
Transcripts are excellent for user-visible behaviour that should change rarely and deliberately: help output, error messages, table layouts, the walkthrough in your README. They double as executable documentation — some projects generate a "Usage examples" page straight from the transcript files.
They are a poor fit for output that varies: timestamps, durations, temporary paths, random IDs. Either keep those out of transcript tests, or extend the format with a placeholder such as <ANY> that matches any text on a line. They also exercise only the command line, so they complement rather than replace unit tests of the logic underneath. And because they pin exact text, they need updating when the framework changes its formatting — a small, honest cost that is usually worth paying for output users depend on.
UX considerations
The users of transcripts are developers and reviewers:
- Keep files short and themed —
greet.t,errors.t,json-output.t— so a failing file name already says what broke. - Comment the intent. A
#line saying why an error message matters ("scripts grep for 'not found'") stops someone from "fixing" it casually. - Treat a transcript diff as a product change. If the diff is intended, update the transcript in the same pull request, and mention user-visible output changes in the changelog.
- Run transcripts in CI on every supported platform if output includes paths or line endings, as in running CLI tests on Windows and macOS runners.
Testing the behaviour
The parser is the part worth unit-testing, since a parsing bug would make transcripts silently test the wrong thing:
# tests/test_transcript_parser.py
import pytest
from conftest import parse_transcript
SAMPLE = """\
# comment
$ mytool greet Ann
hello Ann
$ mytool greet "Ann Lee" --times 2
hello Ann Lee
hello Ann Lee
[exit 0]
$ mytool greet
error
[exit 2]
"""
def test_steps_outputs_and_exit_codes():
steps = parse_transcript(SAMPLE)
assert [s.args for s in steps] == [["greet", "Ann"], ["greet", "Ann Lee", "--times", "2"],
["greet"]]
assert steps[0].expected == ["hello Ann"]
assert steps[2].expected == ["error"] and steps[2].exit_code == 2
def test_commands_must_name_the_tool():
with pytest.raises(ValueError, match="must start with 'mytool'"):
parse_transcript("$ othertool run\n")
Importing from conftest works because pytest puts the test directory on sys.path; if you prefer, move the parser into a small tests/transcripts.py module and import it from both places.
Conclusion
Transcript tests make CLI output reviewable: a plain-text file of commands, expected output and exit codes, collected by a short pytest plugin that runs each command in-process and shows a diff when anything changes. Use them for help, errors and documented examples — the output users depend on — pin the terminal width and colour, keep varying values out, and keep unit tests for the logic underneath. A changed transcript then becomes what it should be: a visible, deliberate product decision.
Frequently asked questions
Why not use cram or prysk directly?
They are mature and language-agnostic, and run commands through a real shell. The in-process collector here is faster, needs no installed console script, works on Windows without a POSIX shell, and integrates with pytest's selection and reporting. Use the external tools when you specifically want to test through a shell.
Can transcripts include stdin?
Extend the format with a < text line that feeds input to the command, and pass it as input= to CliRunner.invoke. Keep it simple: a line or two of input per command.
How do I update many transcripts after an intended change?
Add an opt-in update mode — for example an environment variable that makes runtest write the actual output back into the file instead of failing — and review the resulting diff with git before committing. Snapshot tools such as syrupy offer the same workflow for Python-based tests.
Should stderr and stdout be separated in transcripts?
The runner here compares combined output, which matches what a user sees in a terminal. If stream separation matters for a command — machine output on stdout, diagnostics on stderr — test that with an ordinary pytest test, as in testing Click commands with CliRunner.
Do transcripts slow the test suite down?
Barely. Each command runs in-process through CliRunner, so a transcript step costs about as much as an ordinary CLI test — milliseconds. Only commands that genuinely do slow work (network calls, large files) are slow, and those should use the same fakes and fixtures as your other tests, injected through environment variables the transcript runner sets.