A listing command that returns twelve fields is helpful in JSON and overwhelming in a table. Users quickly want control: just name and status for a quick look, owner.email for a script that sends reminders, the five busiest servers sorted by CPU. Without built-in options they reach for jq, cut and sort, which works until a column moves. This guide adds three cross-format options — --fields, --sort and --limit — that run before rendering, so the table, JSON, CSV and TSV all honour them identically. It is part of the output formats topic and builds on the record-and-renderer split from adding a format flag.
Prerequisites
- Python 3.10+ and Typer or Click.
- Commands that return lists of dictionaries plus a list of default columns.
- An idea of which fields are available per command — the selection step validates against it.
Where selection sits
Selection is a pure function between the command and the renderer: it takes all records and returns a narrower, reordered, possibly shorter list plus the final column list. Nothing upstream changes, and renderers do not know selection happened.
Keeping it in one place has a useful consequence: a user who learns --fields on one command can use it on every listing command, with every format.
The recipe
Start with a helper that reads a dotted path from a nested record. Dotted names (owner.email) are the convention users know from jq, kubectl's custom columns and most APIs.
# src/fleet/select.py
from __future__ import annotations
from collections.abc import Sequence
from typing import Any
_MISSING = object()
def get_path(record: dict[str, Any], path: str) -> Any:
"""Return record['a']['b'] for path 'a.b', or None if any step is missing."""
current: Any = record
for part in path.split("."):
if not isinstance(current, dict):
return None
current = current.get(part, _MISSING)
if current is _MISSING:
return None
return current
def parse_fields(raw: str | None, available: Sequence[str], default: Sequence[str]) -> list[str]:
"""Turn 'name,owner.email' into a validated column list."""
if not raw:
return list(default)
if raw.strip() == "*":
return list(available)
fields = [f.strip() for f in raw.split(",") if f.strip()]
unknown = [f for f in fields if f not in available]
if unknown:
raise ValueError(
f"unknown field(s): {', '.join(unknown)}. Available: {', '.join(available)}"
)
return fields
def parse_sort(raw: str | None, available: Sequence[str]) -> list[tuple[str, bool]]:
"""'-cpu,name' -> [('cpu', True), ('name', False)] where True means descending."""
keys = []
for item in (raw or "").split(","):
item = item.strip()
if not item:
continue
descending = item.startswith("-")
name = item.lstrip("+-")
if name not in available:
raise ValueError(f"cannot sort by unknown field {name!r}")
keys.append((name, descending))
return keys
def apply(records: Sequence[dict[str, Any]], fields: Sequence[str],
sort: Sequence[tuple[str, bool]] = (), limit: int | None = None) -> list[dict[str, Any]]:
rows = list(records)
# Stable sorts applied from the last key to the first give multi-key ordering.
for name, descending in reversed(sort):
present = [r for r in rows if get_path(r, name) is not None]
missing = [r for r in rows if get_path(r, name) is None]
present.sort(key=lambda r: get_path(r, name), reverse=descending)
rows = present + missing # None always sorts last
if limit is not None:
rows = rows[:limit]
return [{f: get_path(r, f) for f in fields} for r in rows]
Three design choices are worth calling out. Rows are projected to flat dictionaries keyed by the dotted name, so a CSV header reads owner.email and JSON gets {"owner.email": ...} — predictable, and the same key in every format. Sorting uses Python's stable sort applied once per key from last to first, the documented way to sort by several keys with mixed directions. And missing values sort last in both directions, which is what people expect from "sort by CPU, busiest first".
Now wire the options into a command. The available fields come from the data model, not from the first record, so validation does not depend on what the API happened to return:
# src/fleet/cli.py
from typing import Annotated, Optional
import typer
from fleet import select
app = typer.Typer()
SERVERS = [
{"name": "web-1", "cpu": 0.42, "owner": {"team": "web", "email": "web@example.com"}},
{"name": "web-2", "cpu": 0.91, "owner": {"team": "web", "email": "web@example.com"}},
{"name": "db-1", "cpu": None, "owner": {"team": "data", "email": "data@example.com"}},
]
AVAILABLE = ["name", "cpu", "owner.team", "owner.email"]
DEFAULT = ["name", "cpu"]
def complete_fields(incomplete: str) -> list[str]:
head, _, last = incomplete.rpartition(",")
prefix = f"{head}," if head else ""
return [prefix + f for f in AVAILABLE if f.startswith(last)]
@app.callback()
def main() -> None:
"""Manage the fleet."""
@app.command("list")
def list_servers(
fields: Annotated[Optional[str], typer.Option(
"--fields", "-f", autocompletion=complete_fields,
help="Comma-separated fields to show, or '*' for all.")] = None,
sort: Annotated[Optional[str], typer.Option(
"--sort", help="Sort keys, e.g. '-cpu,name' (prefix '-' for descending).")] = None,
limit: Annotated[Optional[int], typer.Option("--limit", min=1)] = None,
) -> None:
"""List servers."""
try:
cols = select.parse_fields(fields, AVAILABLE, DEFAULT)
keys = select.parse_sort(sort, AVAILABLE)
except ValueError as exc:
raise typer.BadParameter(str(exc)) from None
for row in select.apply(SERVERS, cols, keys, limit):
typer.echo("\t".join("" if row[c] is None else str(row[c]) for c in cols))
if __name__ == "__main__":
app()
In a real tool the last loop is the shared emit() call from the format-flag guide; it is TSV here to keep the example short.
UX considerations
- Show what is available. The error for an unknown field lists the valid names;
--fields '*'shows everything; and a--list-fieldsflag (or a line in--help) helps discovery. Completion helps most of all — thecomplete_fieldscallback above completes each comma-separated item; see dynamic completion values. - Respect the order the user typed.
--fields status,namemeans status first. That is the main reason people use the option in shell pipelines. - Sort server-side when the API can. If the backend supports ordering and limits, pass them through rather than downloading ten thousand rows to keep five — but keep client-side sorting as the fallback so the option behaves the same everywhere.
- Reject rather than ignore. A typo in
--fieldsor--sortmust be a usage error with exit code 2, never an empty column; scripts would otherwise run with blank data. - Mind the short flag.
-foften means--forceor--file. If either exists in your tool, give--fieldsno short form; naming commands and flags consistently has the wider argument.
Saved column sets
Users who always want the same columns should not have to type them on every run. Two conventions work well and compose with --fields. A named view — -o wide in kubectl, for example — is a preset column list the tool ships with: default, wide and all cover most needs. A config default lets users store their own: [output] fields = "name,status,owner.email" in the config file, overridden by an explicit --fields. Resolve both in the same function that parses --fields, so precedence stays obvious: explicit flag, then config, then the command's default list.
Testing the behaviour
Test the pure functions exhaustively and the command lightly. The pure layer is where ordering and missing-value rules live:
# tests/test_select.py
import pytest
from typer.testing import CliRunner
from fleet import select
from fleet.cli import app
ROWS = [
{"name": "b", "cpu": 0.5, "owner": {"team": "x"}},
{"name": "a", "cpu": None, "owner": {"team": "y"}},
{"name": "c", "cpu": 0.9, "owner": {}},
]
AVAILABLE = ["name", "cpu", "owner.team"]
def test_dotted_paths_and_missing_values():
out = select.apply(ROWS, ["name", "owner.team"])
assert out[2] == {"name": "c", "owner.team": None}
def test_descending_sort_puts_missing_last():
keys = select.parse_sort("-cpu", AVAILABLE)
assert [r["name"] for r in select.apply(ROWS, ["name"], keys)] == ["c", "b", "a"]
def test_multi_key_sort_and_limit():
keys = select.parse_sort("owner.team,name", AVAILABLE)
assert [r["name"] for r in select.apply(ROWS, ["name"], keys, limit=2)] == ["b", "a"]
def test_unknown_field_is_rejected():
with pytest.raises(ValueError, match="unknown field"):
select.parse_fields("name,nope", AVAILABLE, ["name"])
def test_command_level_usage_error():
result = CliRunner().invoke(app, ["list", "--fields", "nme"])
assert result.exit_code == 2
assert "Available: name" in result.output
The last test pins the user-facing behaviour: a bad field name is a usage error, and the message lists the alternatives.
Conclusion
--fields, --sort and --limit turn a fixed listing into something users can shape without an extra tool, and because they run on records before any renderer sees them, they work identically for tables, JSON, CSV and TSV. Validate names against a declared list, keep the user's order, push sorting to the server when you can, and test the pure functions thoroughly. For the remaining "one string per record" use case, add a format-string template.
Frequently asked questions
Should --fields change the JSON shape or just filter keys?
Just filter and reorder. Projecting to flat dotted keys keeps every format consistent; if users need the original nesting, they can omit --fields and use jq. Some tools offer both — a flat projection for --fields and the raw object otherwise — which is reasonable as long as it is documented.
How do I support wildcards like owner.*?
Expand them against the available list before validation: [f for f in available if fnmatch(f, pattern)] for each pattern containing *. Keep the expansion deterministic — in the order of the available list — so output columns do not shuffle between runs.
Is it worth adding a filter option such as --where status=hot?
Only for simple equality on a few fields, and only if the backend cannot filter. A filter mini-language grows quickly into a query parser you have to maintain. Equality filters plus JSON output and jq cover most needs.
Why not just tell users to use jq?
Many will, and JSON output should make that easy. But --fields works for tables and CSV too, needs no extra install, and survives on Windows machines where jq is not available.