Sooner or later every CLI that lists things gets the same three feature requests. Someone wants to paste the output into a spreadsheet, so they ask for CSV. Someone wants to feed it into another tool, so they ask for JSON. Someone wants only two of the eleven columns in a shell loop, so they ask for "a way to pick fields". Each request is reasonable on its own, and each one bolted on separately — a --csv flag here, a --json flag there, a --fields option that only works with one of them — produces a command surface nobody can predict. The fix is to treat output format as one concern with one design: a single format option, a rendering layer that every command shares, and a small set of rules about what each format guarantees.
This topic builds that design. It sits in the Advanced Input Parsing & User Experience section next to working with stdin, stdout and pipes, which covers the stream mechanics, and interactive terminal UI with Rich, which covers the human-facing table. Here the question is narrower and more practical: when the same records must reach a person, a spreadsheet, a jq pipeline and a shell loop, how do you serve all four without four code paths?
TL;DR
- One option, many formats. Expose
--output/-owith a fixed set of choices (table,json,jsonl,csv,tsv,template) on every listing command, never a separate boolean flag per format. - Commands return records; renderers print them. A command produces a list of plain dictionaries and a column list. A shared renderer table turns that into text. No command ever formats its own output.
- Pick the default by audience. A table on a terminal; a stable machine format when stdout is a pipe. Let an environment variable override the default for people who always want JSON.
- Treat machine formats as an API. Field names, types and ordering in JSON and CSV are a contract that scripts depend on. Version them like any other interface.
- Selection and templates sit on top.
--fields name,statustrims columns for every format;--format '{name}\t{status}'covers the one-off shell loop without inventing a mini language. - Test every format with snapshots of the exact bytes, because the bugs live in quoting, escaping and line endings.
Who reads your output
Before deciding which formats to offer, it helps to name the readers. A command's output has at least four distinct consumers, and they want contradictory things.
A person at a terminal wants alignment, truncation to fit the window, colour for status, and human units ("3 minutes ago", "1.2 GB"). A program — jq, a Python script, a CI step — wants exact values: ISO timestamps, byte counts, stable keys, nothing truncated, no colour codes. A spreadsheet user wants CSV with a header row and quoting that Excel and LibreOffice both understand. A shell loop wants one record per line with fields split by a character that cannot appear inside a field — usually a tab.
The mistake is to optimise one format for two readers. A table "that is also easy to parse" ends up with fixed-width columns that break on long names. JSON "that is also readable" grows pretty-printed whitespace and human units that scripts then have to un-format. Give each reader its own format and let each format be uncompromising about its job.
That split also decides where data gets converted. Human formatting — relative times, rounding, truncation, colour — happens only inside the table renderer. Every other renderer receives the raw values. If a command pre-formats a timestamp as "3 minutes ago" before handing records to the renderer, the JSON output is ruined for every script that consumes it.
The shape of the solution
The design that keeps this manageable has three layers, and the boundaries matter more than any individual piece of code.
- The command does the work and returns records: a list of dictionaries with raw, JSON-serialisable values, plus the list of default columns in display order. It knows nothing about formats.
- The selection step applies the cross-format options:
--fieldsto choose and order columns, and optionally--sortand--limit. Because it runs before rendering, every format gets the same rows and columns. - The renderer for the chosen format turns rows and columns into text, and the command layer writes that text to stdout — or to a file when
--output-fileis given.
In code the renderer layer is a dictionary from an Enum to a function, which keeps the set of formats closed and lets the framework validate the option:
# src/mytool/output.py
from __future__ import annotations
import csv
import io
import json
import sys
from collections.abc import Callable, Sequence
from enum import Enum
from typing import Any
Row = dict[str, Any]
class Format(str, Enum):
table = "table"
json = "json"
jsonl = "jsonl"
csv = "csv"
def render_table(rows: Sequence[Row], columns: Sequence[str]) -> str:
from rich.console import Console # imported lazily: only tables need Rich
from rich.table import Table
table = Table(*columns, box=None, pad_edge=False)
for row in rows:
table.add_row(*(str(row.get(c, "")) for c in columns))
buf = io.StringIO()
Console(file=buf, width=100).print(table)
return buf.getvalue()
def render_json(rows: Sequence[Row], columns: Sequence[str]) -> str:
return json.dumps([{c: r.get(c) for c in columns} for r in rows], indent=2) + "\n"
def render_jsonl(rows: Sequence[Row], columns: Sequence[str]) -> str:
return "".join(json.dumps({c: r.get(c) for c in columns}) + "\n" for r in rows)
def render_csv(rows: Sequence[Row], columns: Sequence[str]) -> str:
buf = io.StringIO(newline="")
writer = csv.DictWriter(buf, fieldnames=columns, extrasaction="ignore", lineterminator="\n")
writer.writeheader()
writer.writerows(rows)
return buf.getvalue()
RENDERERS: dict[Format, Callable[[Sequence[Row], Sequence[str]], str]] = {
Format.table: render_table,
Format.json: render_json,
Format.jsonl: render_jsonl,
Format.csv: render_csv,
}
def default_format() -> Format:
return Format.table if sys.stdout.isatty() else Format.jsonl
A command then looks like this:
# src/mytool/cli.py
import sys
import typer
from mytool.output import RENDERERS, Format, default_format
app = typer.Typer()
@app.callback()
def main() -> None:
"""Fleet tool."""
@app.command("list")
def list_servers(
output: Format | None = typer.Option(
None, "--output", "-o",
help="Output format (default: table on a terminal, jsonl in a pipe).",
),
) -> None:
"""List servers."""
rows = fetch_servers() # list[dict] with raw values
columns = ["name", "region", "cpu", "status"]
sys.stdout.write(RENDERERS[output or default_format()](rows, columns))
Because Format is an Enum, Typer lists the choices in --help and rejects -o xml with a clear usage error and exit code 2. Adding a format later means one new function and one dictionary entry — no command changes. Adding a format flag for table, JSON and CSV walks through this recipe end to end, including the Click version and the tests.
Choosing the default format
The default is the format a user gets without asking, and it is the one most people will see. There are three defensible policies.
Always a table. Simple and predictable: the output never changes shape depending on where it goes. The cost is that every script must remember -o json, and a forgotten flag produces a table that a script then parses with awk — fragile code that breaks the first time a column widens.
Table on a terminal, machine format in a pipe. This is what tools such as ls (one column when piped) and many modern CLIs do. It gives humans the table and scripts something parseable without any flag. The cost is surprise: mytool list | less shows JSON Lines instead of the table the user just saw. Mitigate it by making the rule visible in --help, and by honouring an explicit -o table even in a pipe.
Configurable. Let MYTOOL_OUTPUT=json or a config file key set the default, with the flag still winning. This suits teams where most users are scripting. It slots into the same precedence order the rest of your settings use; see config precedence: flags, env, files, defaults.
The second policy plus the third's override is a good fit for most data CLIs. Whatever you choose, apply it in one function — default_format() above — rather than in each command, and use the same TTY detection as the rest of the tool, described in detecting a TTY and adapting output.
What each format guarantees
Offering a format is a promise, and the details of the promise are what scripts break on. Write the guarantees down — in the help text or the docs — so that both you and your users know what may change.
tableguarantees nothing to programs. Columns may be added, reordered, truncated or reformatted in any release. Say so in the docs, and point scripters at the machine formats.jsonis a single JSON document — an array of objects, or an object with aitemskey if you need room for metadata such as a pagination cursor. Keys are stable; new keys may appear; removing or renaming a key is a breaking change. Timestamps are ISO 8601 strings in UTC, sizes are integers in bytes, andnullmeans "no value", not "unknown".jsonl(JSON Lines, also called NDJSON) is one compact JSON object per line. It streams — a consumer can process line one before line ten thousand exists — which makes it the right default for pipes. The object shape is identical to the items injson.csvhas a header row, RFC 4180 quoting and a fixed column order matching the selected fields. Nested values must be flattened or serialised; writing CSV and TSV output correctly covers quoting, encodings and the spreadsheet quirks.tsvis the shell's format: one record per line, tab-separated, no quoting, with tabs and newlines inside values escaped. It is whatcut -f2andwhile IFS=$'\t' read -r name statusexpect.
Machine formats are part of your public interface in the same way flags are. Treat a renamed JSON key like a removed flag: deprecate, warn, and remove in a major version, as in versioning and deprecating CLI flags.
Choosing fields and shaping records
Once every command returns records, a second set of options comes almost for free. --fields (some tools call it --columns) chooses which keys to emit and in what order; it applies to every format, so -o csv --fields name,cpu and -o table --fields name,cpu show the same two columns.
def select_columns(available: Sequence[str], requested: str | None) -> list[str]:
if not requested:
return list(available)
wanted = [f.strip() for f in requested.split(",") if f.strip()]
unknown = [f for f in wanted if f not in available]
if unknown:
raise typer.BadParameter(
f"unknown field(s): {', '.join(unknown)}; choose from {', '.join(available)}",
param_hint="--fields",
)
return wanted
Validating against the known columns matters: a typo that silently yields an empty column is worse than an error, because the script that depends on it keeps running with blanks. Nested data needs a rule too — either flatten with dotted names (owner.email) or refuse nested fields in flat formats. Selecting fields and columns from CLI output covers dotted paths, wildcards and sorting.
For the shell loop that wants exactly one string per record, a template format beats any amount of field selection. Docker, kubectl and the GitHub CLI all offer a variant of --format '{{.Name}}'. In Python, str.format_map with a restricted mapping gives the same power with a syntax your users already know from f-strings — without opening the door to arbitrary attribute access. Custom output templates with a format string shows how to build it safely.
Writing to files as well as stdout
Most of the time stdout is the right destination: the shell already knows how to redirect it. But two cases justify an explicit --output-file (or --export) option. The first is Windows, where redirecting with > in PowerShell 5 re-encodes output to UTF-16 and mangles CSV for every consumer. The second is any command that also prints progress or prompts: once the data goes to a named file, the terminal is free for the human-facing feedback.
When a file path is given, three extra rules apply. Infer the format from the extension (report.csv → csv) unless -o says otherwise. Write atomically so a crash never leaves half a CSV, as in writing files atomically in Python CLIs. And treat - as stdout, so scripts can use one code path for both. Exporting CLI results to files implements all three.
Large result sets and streaming
The renderer signature above takes a fully materialised list. That is fine for a few thousand rows and simplifies tables, which need every row to compute column widths. For commands that can return millions of records, change the contract for the streaming formats: give jsonl, csv and tsv renderers an iterator and have them write each row as it arrives.
from collections.abc import Iterable
from typing import TextIO
def stream_jsonl(rows: Iterable[Row], columns: Sequence[str], out: TextIO) -> None:
for row in rows:
out.write(json.dumps({c: row.get(c) for c in columns}) + "\n")
With that split, memory stays flat for the machine formats and only table and json (which must close the array) buffer. If a command can produce unbounded output, consider refusing -o table above a row count, or truncating with a clear "showing 1,000 of 2.3 million rows; use -o jsonl for all" footer. The NDJSON patterns in processing large files and NDJSON streams apply directly, and so does the broken pipe handling that keeps mytool list -o jsonl | head from printing a traceback.
Testing every format
Output formats fail in the details: an unquoted comma, a \r\n that a Unix test never sees, a float printed as 1e-05, a key renamed by an innocent refactor. Unit-test the renderers directly with awkward data, and snapshot-test each command once per format.
import csv
import io
import json
from mytool.output import render_csv, render_jsonl
AWKWARD = [{"name": 'a "quoted", name', "note": "line1\nline2", "n": 3}]
def test_csv_round_trips_awkward_values():
text = render_csv(AWKWARD, ["name", "note", "n"])
assert list(csv.DictReader(io.StringIO(text))) == [
{"name": 'a "quoted", name', "note": "line1\nline2", "n": "3"}
]
def test_jsonl_is_one_object_per_line():
lines = render_jsonl(AWKWARD * 3, ["name", "n"]).splitlines()
assert len(lines) == 3
assert all(json.loads(line)["n"] == 3 for line in lines)
Round-trip tests — render, then parse with the standard library and compare — catch quoting bugs without hard-coding the escaping rules. For the command level, store one snapshot per format so that a changed key shows up as a diff in review; see snapshot testing CLI output for the tooling.
Common pitfalls
- Colour codes in machine formats. Rich and Click both strip colour when stdout is not a TTY, but only if you write through them. A JSON renderer must never use the coloured console; write plain strings to
sys.stdout. - Logs on stdout. Progress messages or warnings printed to stdout corrupt every machine format. Send all diagnostics to stderr, always, as described in structured logging for CLI apps.
- Locale-dependent numbers. Never format numbers with the locale in machine output;
1.234,5is valid in a German table and a parse error in a JSON consumer. - Different columns per format. If
-o tableshowsageand-o jsonhascreated_at, scripts cannot use the table to discover field names. Keep column names identical; let only the values be humanised in the table. - An empty result that is not valid output. An empty
jsonresult must be[], an emptycsvresult should still have the header row, and an empty table should print a short message to stderr rather than nothing at all.
Key takeaways
- Model output as records plus a column list, and put every formatting decision in one shared renderer layer.
- Offer one
--outputoption with a closed set of choices, validated by anEnum. - Default to a table for terminals and JSON Lines for pipes, and let users override the default.
- Treat JSON, JSON Lines, CSV and TSV as a stable interface with documented guarantees.
- Build
--fields, templates and file export on top of the same records so every format benefits. - Snapshot and round-trip test each format with deliberately awkward data.
Frequently asked questions
Should I use --json or --output json?
Use --output json (with -o as the short form) when you offer more than one machine format, because it scales to new formats without new flags and makes the choices mutually exclusive by construction. A single --json boolean is fine for a tool that will only ever have one machine format, and many tools offer it as an alias for -o json.
Is YAML worth offering?
Only if your users already live in YAML — Kubernetes and Ansible tooling, for example. YAML output is mostly for reading, and JSON is valid YAML anyway. If you add it, render it with yaml.safe_dump(..., sort_keys=False) and keep the same keys as JSON.
Should machine output include metadata such as totals or a next-page cursor?
In json, wrap the list in an object ({"items": [...], "next_cursor": "..."}) from the very first release, because moving from a bare array to an object later is a breaking change. In jsonl and csv, keep the stream to pure records and put metadata on stderr or behind a separate flag.
How do I keep the table readable on a narrow terminal?
Let the table renderer — and only the table renderer — truncate and wrap. Rich does this automatically when it knows the console width. Adapting output to terminal width covers choosing which columns to drop first.
Does every command need every format?
Every command that returns a list of records should support the same set, so users can rely on it. Commands that perform an action and print a confirmation need at most a json result for scripts; a table makes no sense there.
Related
- Up: Advanced Input Parsing & User Experience
- Down: Adding a format flag for table, JSON and CSV
- Down: Writing CSV and TSV output correctly
- Down: Selecting fields and columns from CLI output
- Down: Custom output templates with a format string
- Down: Exporting CLI results to files
- Sideways: Working with stdin, stdout and pipes
- Sideways: Interactive terminal UI with Rich
- Sideways: Emitting JSON output for scripting