A CLI tested only on Linux is a CLI whose Windows users are its testers. The bugs are predictable: a path joined with "/", a snapshot test that fails because of \r\n, an é that turns into UnicodeEncodeError on a legacy console, a temporary file that cannot be deleted because it is still open, a signal.SIGTERM handler that does not exist. None of these show up on Linux, and all of them show up within minutes on a Windows runner. Adding Windows and macOS to CI is cheap with hosted runners; the real work is making the test suite honest across platforms. This guide adds the operating-system matrix, walks through the failures it typically exposes and how to fix each properly, and shows how to mark the few tests that genuinely are platform-specific. It belongs to the CI/CD topic, and extends testing a CLI across Python versions with GitHub Actions.
Prerequisites
- A CLI with a pytest suite running in GitHub Actions (other CI systems have equivalent runners).
- uv for environment setup; the patterns do not depend on it.
The matrix
# .github/workflows/test.yml
name: test
on: [push, pull_request]
jobs:
test:
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, windows-latest, macos-latest]
python: ["3.10", "3.13"]
runs-on: ${{ matrix.os }}
defaults:
run:
shell: bash # same shell syntax on every OS
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v6
with:
enable-cache: true
- run: uv sync --locked --python ${{ matrix.python }}
- run: uv run --locked pytest -q
Three choices keep this maintainable. fail-fast: false lets every combination finish, so one Windows failure does not hide a macOS failure behind it. shell: bash makes run: steps use Git Bash on Windows, so you write one set of shell commands — keep PowerShell for the steps that genuinely need it. And testing the oldest and newest supported Python on every OS, rather than every version everywhere, keeps the matrix at six jobs instead of fifteen while still catching nearly everything.
What breaks, and the real fix
The first Windows run of a Linux-only suite usually fails for a handful of recurring reasons. Each has a proper fix in the code — not a skip.
Paths built from strings. f"{root}/config.toml" and path.split("/") work by accident on POSIX. Use pathlib everywhere, compare paths as Path objects rather than strings, and print paths with str(path) only at the edge. Cross-platform paths with pathlib has the details. In tests, build expected values with Path too: assert result == tmp_path / "out" / "a.txt".
Line endings. Git may check out text fixtures with \r\n on Windows (depending on core.autocrlf), and text written in text mode on Windows uses \r\n. Snapshot tests then fail on every line. Add a .gitattributes with * text=auto eol=lf so fixtures are identical everywhere, and compare output with splitlines() or normalised newlines rather than raw strings.
Encoding. Windows' default encoding for files and pipes has historically been the ANSI code page, so open(path) without encoding= reads UTF-8 files wrongly and printing non-ASCII to a redirected stdout can raise UnicodeEncodeError. Always pass encoding="utf-8" when reading and writing text, and set PYTHONUTF8=1 in CI to match how most modern setups behave. Ruff's PLW1514 rule (currently a preview rule) flags open() calls without an explicit encoding. The user-facing side is covered in fixing Unicode and encoding errors on Windows.
Open files cannot be deleted or replaced. On Windows, deleting or renaming over a file that is still open fails with PermissionError. Close files before cleaning up (context managers), and on Windows retry os.replace briefly if an antivirus scanner holds the target — the atomic-write helper in writing files atomically in Python CLIs shows the pattern.
Executables and shells. The console script is mytool.exe on Windows; subprocess.run(["mytool", ...]) finds it through PATH, but tests that build paths to it by hand must add the suffix. Shell built-ins such as echo and true do not exist as programs on Windows — use sys.executable -c "…" as the portable "external command" in tests.
Signals. signal.SIGTERM can be registered on Windows but is never delivered the POSIX way, and SIGHUP and SIGKILL do not exist. Graceful-shutdown code needs a Windows branch, as discussed in handling SIGTERM and graceful shutdown, and its tests need platform markers.
macOS specifics are fewer: the default filesystem is case-insensitive (so Config.toml and config.toml are the same file), /tmp is a symlink to /private/tmp (so resolved paths differ from the ones you created), and runners are slower and more expensive than Linux ones.
Marking genuinely platform-specific tests
Some behaviour really is platform-specific. Register markers once and use them instead of scattering skipif expressions:
# tests/conftest.py
import sys
import pytest
PLATFORMS = {"windows": "win32", "macos": "darwin", "linux": "linux"}
def pytest_configure(config):
for name in PLATFORMS:
config.addinivalue_line("markers", f"{name}: only runs on {name}")
config.addinivalue_line("markers", "posix: only runs on Linux and macOS")
def pytest_runtest_setup(item):
wanted = [name for name in PLATFORMS if item.get_closest_marker(name)]
if wanted and not any(sys.platform.startswith(PLATFORMS[n]) for n in wanted):
pytest.skip(f"only on {', '.join(wanted)}")
if item.get_closest_marker("posix") and sys.platform == "win32":
pytest.skip("POSIX only")
# tests/test_platform_bits.py
import os
import signal
import subprocess
import sys
import pytest
@pytest.mark.posix
def test_sigterm_handler_registered():
assert signal.getsignal(signal.SIGTERM) is not None
@pytest.mark.windows
def test_console_script_has_exe_suffix(tmp_path):
assert os.path.splitext(sys.executable)[1].lower() == ".exe"
def test_portable_external_command():
out = subprocess.run([sys.executable, "-c", "print('ok')"],
capture_output=True, text=True, check=True).stdout
assert out.splitlines() == ["ok"]
def test_paths_compare_as_paths(tmp_path):
produced = tmp_path / "out" / "a.txt"
assert produced.relative_to(tmp_path).parts == ("out", "a.txt")
The last two tests show the portable style that most tests should follow: use the current interpreter as the external program, compare lines rather than raw text, and compare path parts rather than strings with separators in them.
UX considerations
The "users" of the matrix are contributors who mostly develop on one platform:
- Make failures reproducible locally. Document how to run the Windows job's equivalent — a Windows VM, or at minimum
PYTHONUTF8=0and CRLF fixtures on Linux to catch encoding and line-ending assumptions early. - Keep the matrix small but meaningful. Oldest and newest Python on all three systems, plus every Python on Linux, catches nearly every platform bug at a fraction of the cost.
- Never "fix" a platform failure with a skip unless the feature is truly unavailable there. A skipped test is a bug report you have decided not to read.
- Show the platform in test names or output when it matters, so a failure email says "windows-latest, 3.10" without anyone opening the logs.
- Mind the clock. macOS runners are slower and more expensive; consider running them only on pushes to the main branch and on release tags if minutes are scarce.
Testing the behaviour
The best check that a suite is portable is the matrix itself, but two cheap habits catch most problems before CI does. Run the suite on Linux with PYTHONUTF8=0 LANG=C occasionally — it surfaces many missing encoding= arguments. And add a lint rule for encodings in the project configuration, as in configuring Ruff for a CLI project:
[tool.ruff.lint]
preview = true # PLW1514 is a preview rule
explicit-preview-rules = true # ...so enable only the preview rules named here
extend-select = ["PLW1514", "PTH"] # open() without encoding; os.path instead of pathlib
PTH rules nudge code from os.path string handling to pathlib, which removes most path-separator bugs at the source. Together, the lint rules and the matrix keep a Linux-developed CLI working for the people on the other two platforms.
Conclusion
Adding Windows and macOS to CI is a few lines of YAML; the value is in what it exposes. Fix the recurring causes at the source — pathlib instead of strings, explicit UTF-8, .gitattributes for line endings, closing files before deleting them, the interpreter as a portable external command — and mark the few genuinely platform-specific tests with registered markers instead of ad hoc skips. Then your Windows users stop being your testers.
Frequently asked questions
Do I need Windows if my users are all on Linux?
Check before assuming: developers on Windows often run internal tools from WSL and from PowerShell. If you truly have no Windows users, say so in the README and skip it — but keep macOS if any developer uses a Mac.
Should I use shell: pwsh or shell: bash on Windows runners?
bash for portability of your workflow steps; pwsh when testing PowerShell-specific behaviour such as completion scripts or quoting. Both are available on GitHub's Windows runners.
Why do my tests pass locally on Windows but fail in CI?
Usually environment differences: CI runs non-interactively (no TTY, so colours and prompts behave differently), uses a different code page, or checks out files with different line endings. Print sys.stdout.encoding, sys.stdout.isatty() and locale.getpreferredencoding() in a debug step to compare.
How do I test the installed executable on each OS?
Build the wheel once, then install and run it on every runner in a separate job, as in smoke-testing the built wheel in CI. That catches entry-point and packaging problems that only appear on one platform.