Input & UX

Exporting CLI Results to Files from a Python CLI

Add an --output-file option that infers the format from the extension, treats - as stdout, refuses to clobber, writes atomically and compresses .gz files.

Updated

"Just redirect stdout" is the right answer for most CLI output, and it stops being right in a handful of common situations. On Windows PowerShell 5, mytool report > out.csv re-encodes the text as UTF-16, and the CSV opens as garbage. A command that shows a progress bar while exporting cannot also put the data on stdout without mixing the two. A nightly job that writes a report should never leave a half-written file when it is killed. And users simply expect mytool export --to report.xlsx-style ergonomics from a data tool. This guide adds an --output-file option that handles all of these, built on the record-and-renderer layer from the output formats topic.

Prerequisites

The rules an export option should follow

Rules for an export option The behaviours a file export option should have, and the shortcuts that cause trouble. Rules for an export option Do ✓ Treat - as stdout ✓ Infer the format from the extension ✓ Refuse to overwrite without --force ✓ Write to a temp file, then rename Avoid ✗ Guessing the format of report.txt ✗ Truncating the target before writing ✗ Status messages on stdout ✗ One flag meaning path and format A trailing .gz is the one extension that changes how the bytes are written, not which format.
  1. - means stdout. Scripts can then pass a variable that is sometimes a path and sometimes -, with one code path.
  2. The extension picks the format when --output/-o is not given: .csv, .tsv, .json, .jsonl/.ndjson. An explicit format always wins, and an unknown extension without one is an error rather than a guess.
  3. A trailing .gz compresses, so report.jsonl.gz is JSON Lines inside gzip.
  4. Existing files are not overwritten silently. Refuse unless --force is given — the same principle as in adding dry run and confirmation to destructive commands.
  5. Writes are atomic. The file appears complete or not at all.
  6. Feedback goes to stderr. "Wrote 1,204 rows to report.csv" helps the human and never pollutes a pipe.

The recipe

# src/fleet/export.py
from __future__ import annotations

import gzip
import os
import sys
import tempfile
from pathlib import Path

FORMAT_BY_SUFFIX = {".csv": "csv", ".tsv": "tsv", ".json": "json",
                    ".jsonl": "jsonl", ".ndjson": "jsonl"}


class ExportError(Exception):
    pass


def infer_format(path: Path, explicit: str | None) -> str:
    if explicit:
        return explicit
    suffixes = [s.lower() for s in path.suffixes]
    if suffixes and suffixes[-1] == ".gz":
        suffixes.pop()
    fmt = FORMAT_BY_SUFFIX.get(suffixes[-1] if suffixes else "")
    if fmt is None:
        known = ", ".join(sorted(FORMAT_BY_SUFFIX))
        raise ExportError(f"cannot tell the format of {path.name!r}; use -o or one of {known}")
    return fmt


def write_output(text: str, target: str, *, force: bool = False,
                 encoding: str = "utf-8") -> Path | None:
    """Write rendered text to stdout ('-') or atomically to a file. Returns the path."""
    if target == "-":
        sys.stdout.write(text)
        sys.stdout.flush()
        return None
    path = Path(target).expanduser()
    if path.is_dir():
        raise ExportError(f"{path} is a directory")
    if path.exists() and not force:
        raise ExportError(f"{path} already exists; pass --force to overwrite")
    path.parent.mkdir(parents=True, exist_ok=True)
    data = text.encode(encoding)
    if path.suffix.lower() == ".gz":
        data = gzip.compress(data, mtime=0)      # mtime=0 keeps output reproducible
    fd, tmp = tempfile.mkstemp(dir=path.parent, prefix=f".{path.name}.", suffix=".tmp")
    try:
        with os.fdopen(fd, "wb") as fh:
            fh.write(data)
            fh.flush()
            os.fsync(fh.fileno())
        os.replace(tmp, path)
    except BaseException:
        Path(tmp).unlink(missing_ok=True)
        raise
    return path

Writing bytes rather than text sidesteps two traps at once: there is no newline translation on Windows, and the encoding is explicit. Passing encoding="utf-8-sig" for an Excel-bound CSV is then a one-word change, as discussed in writing CSV and TSV output correctly.

The command layer resolves the format, renders, writes, and reports:

# src/fleet/cli.py
import json
from pathlib import Path
from typing import Annotated, Optional

import typer

from fleet.export import ExportError, infer_format, write_output

app = typer.Typer()
ROWS = [{"name": "web-1", "cpu": 0.42}, {"name": "web-2", "cpu": 0.91}]


def render(rows, fmt: str) -> str:
    if fmt == "json":
        return json.dumps(rows, indent=2) + "\n"
    if fmt == "jsonl":
        return "".join(json.dumps(r) + "\n" for r in rows)
    sep = "," if fmt == "csv" else "\t"
    cols = list(rows[0]) if rows else []
    return "\n".join([sep.join(cols)] + [sep.join(str(r[c]) for c in cols) for r in rows]) + "\n"


@app.callback()
def main() -> None:
    """Manage the fleet."""


@app.command()
def export(
    output_file: Annotated[str, typer.Option("--output-file", "-O",
        help="Where to write: a path, or '-' for stdout.")] = "-",
    output: Annotated[Optional[str], typer.Option("--output", "-o",
        help="Format; inferred from the file extension when omitted.")] = None,
    force: Annotated[bool, typer.Option("--force", help="Overwrite an existing file.")] = False,
) -> None:
    """Export servers."""
    try:
        fmt = output or ("jsonl" if output_file == "-" else None)
        fmt = infer_format(Path(output_file), fmt)
        path = write_output(render(ROWS, fmt), output_file, force=force)
    except ExportError as exc:
        typer.echo(f"error: {exc}", err=True)
        raise typer.Exit(1)
    if path is not None:
        typer.echo(f"wrote {len(ROWS)} rows to {path} ({fmt})", err=True)


if __name__ == "__main__":
    app()

(The tiny render stands in for the shared renderers; use the real ones in your tool.)

Exporting to files Terminal session exporting to a CSV file, being refused when the file exists, and overwriting with force. Exporting to files bash $ fleet export -O report.csv wrote 2 rows to report.csv (csv) $ fleet export -O report.csv error: report.csv already exists; pass --force to overwrite $ fleet export -O nightly.jsonl.gz --force wrote 2 rows to nightly.jsonl.gz (jsonl) Every line here is on stderr; stdout stays empty when the data goes to a file.

Default file names and timestamps

Scheduled exports usually want a fresh, sortable file name on every run. Rather than making every cron job assemble one in shell, accept a small set of placeholders in the path and expand them before writing:

from datetime import datetime, timezone


def expand_target(target: str, *, now: datetime | None = None) -> str:
    """Expand {date} and {time} placeholders in an output path."""
    now = now or datetime.now(timezone.utc)
    return target.format(date=now.strftime("%Y-%m-%d"), time=now.strftime("%H%M%S"))


# fleet export -O 'exports/servers-{date}.csv'  ->  exports/servers-2026-10-02.csv

ISO dates sort correctly as plain strings, so ls exports/ lists runs in order, and UTC avoids two runs an hour apart producing the same name on the night clocks change. Keep the placeholder set tiny and documented; it is a convenience, not a template language. If two runs could share a name, combine the date with the time, or let --force decide whether the second run replaces the first.

For exports that run unattended, also consider what happens to old files. A --keep N option that deletes all but the newest N matching files is simple to add and saves someone writing a cleanup script — but make it opt-in, and only ever delete files that match the pattern your tool created.

UX considerations

  • Pick an unambiguous name. Many tools use --output for the format and -O/--output-file for the path; curl uses -o for the path. Whatever you choose, never let one flag mean both things in different commands.
  • Confirm writes on stderr with a count and a path. It is the only feedback a user gets, and the absolute path removes doubt about which directory it landed in.
  • Show progress only for real files. When the target is -, stdout is the data stream; keep progress bars on stderr and switch them off when stderr is not a terminal, as in adding progress bars and spinners.
  • Do not create surprise directories deep in the tree. Creating a missing parent is convenient; creating ~/reprots/2026/ because of a typo is not. Some tools only create one level, or ask first in interactive mode.
  • Respect permissions. A file that may contain sensitive data should be created with 0o600 — tempfile.mkstemp already does that, and the rename keeps it.
Naming the two output options Common conventions for naming the format option and the output file option in command line tools. Naming the two output options Convention Format File path kubectl-style -o / --output redirect only This guide -o / --output -O / --output-file curl-style (none) -o / --output Alternative --format --output Pick one convention and use it in every command; mixing them is worse than any single choice.

Testing the behaviour

tmp_path makes file exports easy to test. Cover inference, refusal, force, compression, and the stdout path:

# tests/test_export.py
import gzip
import json

import pytest
from typer.testing import CliRunner

from fleet.cli import app
from fleet.export import ExportError, infer_format

runner = CliRunner()


@pytest.mark.parametrize("name,fmt", [("a.csv", "csv"), ("a.JSONL", "jsonl"),
                                      ("a.ndjson.gz", "jsonl"), ("a.tar.json", "json")])
def test_infer_format(tmp_path, name, fmt):
    assert infer_format(tmp_path / name, None) == fmt


def test_unknown_extension_needs_explicit_format(tmp_path):
    with pytest.raises(ExportError, match="cannot tell the format"):
        infer_format(tmp_path / "report.txt", None)


def test_refuses_to_overwrite_without_force(tmp_path):
    target = tmp_path / "out.json"
    target.write_text("keep me")
    result = runner.invoke(app, ["export", "-O", str(target)])
    assert result.exit_code == 1
    assert target.read_text() == "keep me"
    assert runner.invoke(app, ["export", "-O", str(target), "--force"]).exit_code == 0
    assert json.loads(target.read_text())[0]["name"] == "web-1"


def test_gzip_and_message_on_stderr(tmp_path):
    target = tmp_path / "out.jsonl.gz"
    result = runner.invoke(app, ["export", "-O", str(target)])
    assert result.stdout == ""
    assert "wrote 2 rows" in result.stderr
    assert gzip.decompress(target.read_bytes()).decode().count("\n") == 2


def test_dash_means_stdout():
    result = runner.invoke(app, ["export", "-O", "-"])
    assert result.stdout.splitlines()[0] == '{"name": "web-1", "cpu": 0.42}'

Note the assertion that stdout is empty when writing to a file: it is the regression test for the most common export bug, a status message that ends up inside the data stream.

Conclusion

An export option is a few dozen lines once records and renderers are separated: infer the format from the extension, treat - as stdout, refuse to clobber, write atomically, compress on .gz, and report on stderr. Those rules make file output as safe to script as stdout and much kinder to Windows and spreadsheet users.

Frequently asked questions

Should I support writing Excel .xlsx files?

If your users live in Excel, yes — it avoids the type-guessing problems of CSV because cells can be typed. Use openpyxl as an optional extra (pip install mytool[excel]) and import it only when the extension is .xlsx, so the base install stays light.

How do I stream a huge export without building the whole string?

Change the writer to accept an iterator of text chunks and write them into the temporary file as they arrive, then rename at the end. For .gz, wrap the temporary file in gzip.GzipFile(fileobj=fh, mode="wb", mtime=0). Atomicity is unchanged because the rename still happens last.

What if the user passes a directory?

Either refuse, as above, or treat it as "write a default file name inside it" (exports/servers-2026-10-02.csv). The second is friendly but surprising when someone meant a file; if you do it, print the resulting path.

Should --force also create missing parent directories?

Keep them separate. --force answers "may I overwrite?"; creating parents is a convenience you can either always do or never do. Tying them together makes the flag mean two things.