Project Setup

Testing a CLI Project Template with pytest

Test a Cookiecutter CLI template like code: bake it in tmp_path, check variables and hooks, run each generated project’s tests and linters in CI.

Updated

A project template is code that writes code, and it breaks the same way code does — except that nobody notices until a colleague generates a new CLI and finds that pyproject.toml has a stray comma, the Click variant does not import, or the CI workflow references a job that was renamed. The template's own repository has no failing build, because nothing in it ever runs. The fix is to give the template a test suite: generate ("bake") projects with representative answers into temporary directories, check the files, and then run each generated project's own tests and linters, exactly as a new user would on day one. This guide builds that suite with pytest and Cookiecutter's Python API, and wires it into CI. It belongs to the project scaffolding topic.

Prerequisites

  • A Cookiecutter template in its own repository, such as the one in building a Cookiecutter template for Typer CLIs.
  • pytest and cookiecutter in the template repository's development environment (uv add --dev pytest cookiecutter), and uv available on PATH for the slow tests.

Three layers of template tests

Three layers of template tests Layers of tests for a project template, from fast rendering checks to slow generated-project checks. Three layers of template tests Rendering tests fast parse TOML, check files, no leftover {{ }} — milliseconds Hook tests fast validation rejects bad names, optional files removed Generated-project tests slow install, run its tests, lint, build a wheel The slow layer is the one that catches what users actually hit.
  1. Rendering tests bake the template with specific answers and inspect the output: file names, file contents, parsed TOML. They run in milliseconds and catch Jinja mistakes and wrong conditionals.
  2. Hook tests check that pre-generation validation rejects bad input and that post-generation hooks add or remove what they should.
  3. Generated-project tests treat each baked project as a real project: install it, run its test suite, run its linters. They are slower — seconds to tens of seconds per combination — and they catch the bugs users actually hit.

The recipe

A baking helper

# tests/conftest.py
from pathlib import Path

import pytest
from cookiecutter.main import cookiecutter

TEMPLATE = Path(__file__).resolve().parents[1]


@pytest.fixture
def bake(tmp_path):
    def _bake(**context) -> Path:
        return Path(cookiecutter(str(TEMPLATE), no_input=True, output_dir=str(tmp_path),
                                 extra_context=context))
    return _bake


def pytest_configure(config):
    config.addinivalue_line("markers", "slow: generates and tests a whole project")

cookiecutter() returns the path of the generated project, no_input=True uses defaults for every unspecified variable, and extra_context overrides specific answers — the same as passing key=value on the command line. Writing into tmp_path keeps every test isolated and cleans up automatically. (The pytest-cookies plugin provides a similar cookies fixture; the hand-written version above has no extra dependency and makes the mechanics visible.)

Rendering tests

# tests/test_rendering.py
import tomllib

import pytest


def test_defaults_produce_valid_metadata(bake):
    project = bake()
    meta = tomllib.loads((project / "pyproject.toml").read_text())
    assert meta["project"]["name"] == "acme-deploy"
    assert meta["project"]["scripts"]["acme-deploy"] == "acme_deploy.cli:main"


@pytest.mark.parametrize("framework,dependency", [("typer", "typer"), ("click", "click")])
def test_framework_choice(bake, framework, dependency):
    project = bake(cli_framework=framework)
    deps = tomllib.loads((project / "pyproject.toml").read_text())["project"]["dependencies"]
    assert [d.split(">")[0] for d in deps] == [dependency]
    compile((project / "src/acme_deploy/cli.py").read_text(), "cli.py", "exec")


def test_no_unrendered_jinja_left(bake):
    project = bake()
    for path in project.rglob("*"):
        if path.is_file() and ".github" not in path.parts:
            assert "{{" not in path.read_text(errors="ignore"), path

Parsing pyproject.toml with tomllib catches syntax errors that a string comparison would miss. compile() checks that every generated Python file is at least syntactically valid, which is a cheap guard for templates with if/else blocks spanning code. The last test catches the most common template bug of all — a variable typo that leaves {{ cookiecutter.pakage_name }} in the output — while skipping files deliberately copied without rendering.

Hook tests

# tests/test_hooks.py
import pytest
from cookiecutter.exceptions import FailedHookException


@pytest.mark.parametrize("use_docker", [True, False])
def test_dockerfile_is_conditional(bake, use_docker):
    assert (bake(use_docker=use_docker) / "Dockerfile").exists() is use_docker


@pytest.mark.parametrize("name", ["9 Lives", "Acme!", "class"])
def test_invalid_names_are_rejected(bake, tmp_path, name):
    with pytest.raises(FailedHookException):
        bake(project_name=name)
    assert not any(tmp_path.iterdir())     # nothing half-generated is left behind

Generated-project tests

# tests/test_generated_projects.py
import subprocess

import pytest


@pytest.mark.slow
@pytest.mark.parametrize("framework", ["typer", "click"])
def test_generated_project_passes_its_own_checks(bake, framework):
    project = bake(cli_framework=framework)
    for command in (["uv", "run", "--quiet", "--with", "pytest", "pytest", "-q"],
                    ["uvx", "ruff", "check", "--quiet", "."],
                    ["uv", "build", "--quiet"]):
        result = subprocess.run(command, cwd=project, capture_output=True, text=True,
                                timeout=300)
        assert result.returncode == 0, f"{' '.join(command)} failed:\n{result.stdout}{result.stderr}"

This is the test that matters most. It installs the generated project with its declared dependencies, runs the tests the template ships, lints the code the template produced, and builds a wheel — a miniature version of the new user's first hour. Including the command output in the assertion message makes CI failures readable without re-running anything. The very first run of this test on a real template often finds something: a lint rule the template's own code violates, a test that imports the wrong module name in one variant, or a dependency missing from one branch of a conditional.

Running the template suite Terminal session running fast template tests separately from slow generated-project tests. Running the template suite bash $ uv run pytest -q -m "not slow" 13 passed, 2 deselected in 0.96s $ uv run pytest -q -m slow .. [100%] 2 passed in 1.9s Separate steps make it obvious which layer failed.

Choosing combinations

With three choices and two booleans, a template has twelve possible outputs; real templates have hundreds. Do not test them all. Test:

  • the defaults, because most projects use them;
  • each value of each choice once, varying one at a time from the defaults;
  • combinations that interact — two options that touch the same file, such as a framework choice and an optional "rich output" flag that both change cli.py.

Pairwise testing tools can generate a minimal set covering every pair of values, but for most CLI templates the list above, written as explicit parametrize cases, is enough and easier to read.

Which combinations to test A minimal set of template answer combinations: the defaults, each choice varied once, and options that touch the same file together. Which combinations to test Case framework use_docker Why defaults typer false most projects vary framework click false other branch of cli.py vary docker typer true hook keeps the file interaction click true both change the same files Four cases instead of every combination — and each one has a reason.

Snapshotting a reference project

Rendering tests check details; a reference snapshot shows the whole picture in code review. Commit one generated project — defaults only — under tests/reference/, and add a test that regenerates it and fails with a diff when anything changes:

# tests/test_reference.py
import filecmp
from pathlib import Path

REFERENCE = Path(__file__).parent / "reference" / "acme_deploy"


def differences(a: Path, b: Path) -> list[str]:
    cmp = filecmp.dircmp(a, b, ignore=["__pycache__"])
    out = [f"only in {a.name}: {n}" for n in cmp.left_only]
    out += [f"only in generated: {n}" for n in cmp.right_only]
    out += [f"changed: {n}" for n in cmp.diff_files]
    for sub in cmp.common_dirs:
        out += differences(a / sub, b / sub)
    return out


def test_defaults_match_reference(bake):
    assert differences(REFERENCE, bake()) == []

When a template change is intentional, regenerate the reference (cookiecutter --no-input . -o tests/reference -f) and commit it in the same pull request. Reviewers then see the effect of a template change on a real project as an ordinary diff — far easier to judge than Jinja edits — while the structural tests above keep checking the variants the reference does not cover.

UX considerations

The users of these tests are the template's maintainers:

  • Keep the fast tests fast. Rendering and hook tests should run in a second or two so maintainers run them on every change; mark the generated-project tests slow and run them in CI and before tagging a release.
  • Make failures point at the template, not the output. Name parameters in test IDs ([click]) and include file paths in assertion messages, so the fix location is obvious.
  • Test from a clean environment. The generated-project tests use uv run and uvx so they never accidentally rely on packages installed in the template's own development environment.
  • Pin tool versions the template uses. If the template's pre-commit config pins Ruff 0.6, test with that version — a newer Ruff in CI that flags new rules is a template update, not a test failure to ignore.

Testing the behaviour in CI

# .github/workflows/template.yml
name: template
on: [push, pull_request]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: astral-sh/setup-uv@v6
      - run: uv sync --locked
      - run: uv run pytest -q -m "not slow"
      - run: uv run pytest -q -m slow

Running the slow tests as a separate step makes it obvious from the job summary whether a failure is a rendering problem or a generated-project problem. For templates used on Windows, add a Windows runner as described in running CLI tests on Windows and macOS runners — path and line-ending bugs in templates are common.

Conclusion

Treat a template as software with users: bake it into temporary directories with the Cookiecutter API, assert on rendered files and parsed metadata, check that hooks validate and prune correctly, and run each generated project's own tests, linters and build. Cover the defaults, each choice once and the combinations that interact, keep the slow tests in their own marker, and run everything in CI so the template's next user gets a project that works on day one.

Frequently asked questions

Should I use pytest-cookies?

It is a convenient plugin that provides a cookies fixture with a bake() method and a result object. The hand-written fixture above does the same in a few lines without a dependency; choose whichever your team prefers.

How do I test a Copier template?

The same layers apply. Use copier.run_copy(template, tmp_path, data={...}, defaults=True, unsafe=True) instead of cookiecutter(), and add update tests from previous tags, as shown in updating generated projects with copier update.

The generated-project tests are slow — what can I do?

Run them in parallel with pytest -n auto (pytest-xdist), share uv's cache across runs in CI, and test fewer combinations. A handful of well-chosen combinations catches nearly everything.

Should the template's tests check exact file contents?

Rarely. Exact snapshots of generated files make every intentional template change a test update. Assert on structure and meaning — parsed TOML, presence of files, successful builds — and use snapshots only for small, stable files.