File names may contain spaces, tabs, quotes, non-UTF-8 bytes and even newlines. The only bytes a Unix file name cannot contain are / and NUL. That is why the classic pipeline idiom for file lists is find . -print0 | xargs -0 …: separating names with NUL instead of newline is the only format that cannot be fooled by any name. A CLI that reads paths from stdin one per line works fine — until somebody has a file called notes\nold.txt, and the tool sees two paths that do not exist. If your CLI reads or prints lists of paths or other arbitrary strings, supporting NUL-delimited records is how it joins these pipelines safely. This guide adds the conventional -0 (input) and -z (output) flags, streams records without loading everything into memory, preserves undecodable file names byte for byte, and tests the whole thing against real find and xargs. It belongs to the stdin, stdout and pipes topic.
Prerequisites
- A CLI that reads from stdin, as in reading piped input in Python CLIs.
- A Unix-like shell with
findandxargsto try the pipelines (GNU and BSD versions both support-print0and-0).
Why newline is not enough
With newline-separated input, the separator is a character that can appear inside a record, so a single name with a newline is indistinguishable from two names. Quoting schemes (ls -b, shell escaping) solve this only if both sides agree on them. NUL works because it is guaranteed never to appear in a path, nor in most text arguments — C strings, and therefore every Unix command-line argument and environment variable, cannot contain it. The ecosystem converged on it: find -print0, xargs -0, git ls-files -z, grep -lZ, sort -z, fd -0, rg --null and du -0 all speak the format.
The recipe
# src/mytool/records.py
from __future__ import annotations
import sys
from collections.abc import Iterable, Iterator
from typing import BinaryIO
def read_records(stream: BinaryIO, *, null: bool, chunk_size: int = 65536) -> Iterator[str]:
"""Yield records separated by NUL (null=True) or newline, decoded like file names."""
sep = b"\0" if null else b"\n"
buffer = b""
while chunk := stream.read(chunk_size):
buffer += chunk
*complete, buffer = buffer.split(sep)
for raw in complete:
yield from _decode(raw, null)
if buffer:
yield from _decode(buffer, null)
def _decode(raw: bytes, null: bool) -> Iterator[str]:
if not null:
raw = raw.removesuffix(b"\r") # tolerate CRLF input on newline mode
if raw: # empty records (blank lines, "\0\0") are skipped
yield raw.decode(sys.getfilesystemencoding(), "surrogateescape")
def write_records(records: Iterable[str], stream: BinaryIO, *, null: bool) -> None:
sep = b"\0" if null else b"\n"
for record in records:
data = record.encode(sys.getfilesystemencoding(), "surrogateescape")
if not null and b"\n" in data:
raise ValueError(f"{record!r} contains a newline; use -0/--null")
stream.write(data + sep)
stream.flush()
# src/mytool/cli.py
import sys
from pathlib import Path
from typing import Annotated
import typer
from mytool.records import read_records, write_records
app = typer.Typer()
NullIn = Annotated[bool, typer.Option("-0", "--null", help="Input is NUL-separated (find -print0).")]
NullOut = Annotated[bool, typer.Option("-z", "--print0", help="Separate output with NUL, for xargs -0.")]
@app.callback()
def main() -> None:
"""File tool."""
@app.command()
def stale(null: NullIn = False, print0: NullOut = False,
days: Annotated[int, typer.Option(help="Older than this many days.")] = 30) -> None:
"""Read paths from stdin and print those not modified for DAYS days."""
import time
cutoff = time.time() - days * 86400
paths = read_records(sys.stdin.buffer, null=null)
old = (p for p in paths if Path(p).is_file() and Path(p).stat().st_mtime < cutoff)
try:
write_records(old, sys.stdout.buffer, null=print0)
except ValueError as exc:
typer.echo(f"error: {exc}", err=True)
raise typer.Exit(1)
The command reads paths, keeps the ones not modified recently, and writes them out — a typical filter in a pipeline:
find ~/Downloads -type f -print0 | mytool stale -0 -z --days 90 | xargs -0 rm -i --
git ls-files -z | mytool stale -0
Reading records in binary
sys.stdin is a text stream that decodes with the locale encoding and translates newlines; neither is what you want for file names. sys.stdin.buffer gives the raw bytes. The reader pulls fixed-size chunks, splits on the separator and keeps the incomplete tail for the next round, so memory stays flat whether the input has ten names or ten million, and the first results flow downstream immediately. A record that straddles a chunk boundary is reassembled correctly — the test with chunk_size=3 exists to prove it.
Empty records are skipped: blank lines in newline mode, and doubled NULs, which some generators produce. In newline mode, a trailing \r is removed, so input generated on Windows does not produce names ending in an invisible carriage return.
Decoding like the operating system does
File names on Linux are bytes. Most are UTF-8, but a disk copied from an old system may have Latin-1 names that are not valid UTF-8. Decoding with errors="strict" crashes on them; errors="replace" turns them into �, a different name that does not exist. Python’s own answer — the one os.listdir and sys.argv use — is the file system encoding with the surrogateescape error handler: undecodable bytes become lone surrogate code points in the str, and encoding back with the same handler restores the original bytes. A Path built from such a string opens the right file, and writing it back out reproduces it exactly. Using the same encoding and handler on both sides makes the tool transparent to any name.
Writing records
-z writes NUL after each record, which is what xargs -0, sort -z and the next tool’s -0 expect. In newline mode, a record containing a newline cannot be written unambiguously, so the writer refuses with an error that names the fix instead of producing output that another program will silently misread. Writing through sys.stdout.buffer avoids text-mode newline translation on Windows and encoding errors for surrogate-escaped names.
Flag conventions
There is no single standard, but there are strong habits worth following so users can guess your flags:
-0/--nullfor input, as inxargs -0.-z/--print0(or--zero) for output, as ingit ls-files -z,grep -zandfind -print0. Some tools use-zfor both directions, likesort -z; if your command only ever reads or only ever writes lists, one flag for both is fine.- Document them together in the help, with a pipeline example — users who need them search for “print0”.
UX considerations
- Keep newline mode the default. It is what people type interactively and read on screen; NUL output on a terminal looks like names glued together.
- Use
--before paths when passing names to other commands (xargs -0 rm --), so a file called-rfis not read as an option. Your own CLI should accept--too; Click and Typer do. - Stream, do not collect. Emit each result as soon as it is known; downstream
xargsstarts work in parallel andhead -zcan stop early. Handle the resulting broken pipe as in handling broken pipe and SIGPIPE. - Accept arguments as well as stdin.
mytool stale a.txt b.txtshould work too; read stdin only when no paths are given, so the tool is equally comfortable as anxargstarget. - Apply it to JSON as well. NDJSON has the same property NUL gives paths — one record per line, with newlines inside values escaped — which is why it suits structured records; see processing large files and NDJSON streams.
Testing the behaviour
Unit tests feed BytesIO objects with the nastiest names you can think of; one end-to-end test runs a real find -print0 | mytool | xargs -0 pipeline:
# tests/test_records.py
import io
import os
import shlex
import shutil
import subprocess
import sys
import pytest
from mytool.records import read_records, write_records
AWKWARD = ["plain.txt", "with space.txt", "new\nline.txt", "tab\there.txt", "café.txt"]
def test_null_round_trip_survives_any_name():
buf = io.BytesIO()
write_records(AWKWARD, buf, null=True)
assert list(read_records(io.BytesIO(buf.getvalue()), null=True)) == AWKWARD
def test_records_split_across_chunks():
data = b"\0".join(name.encode() for name in AWKWARD) + b"\0"
assert list(read_records(io.BytesIO(data), null=True, chunk_size=3)) == AWKWARD
def test_newline_mode_skips_blanks_and_crlf():
data = b"a.txt\r\n\nb.txt\n"
assert list(read_records(io.BytesIO(data), null=False)) == ["a.txt", "b.txt"]
def test_newline_output_refuses_ambiguous_names():
with pytest.raises(ValueError, match="use -0"):
write_records(["new\nline.txt"], io.BytesIO(), null=False)
def test_undecodable_bytes_round_trip():
raw = b"bad-\xff.txt\0"
[name] = read_records(io.BytesIO(raw), null=True)
out = io.BytesIO()
write_records([name], out, null=True)
assert out.getvalue() == raw
@pytest.mark.skipif(not shutil.which("find") or not shutil.which("xargs"), reason="needs find/xargs")
def test_find_print0_to_xargs0(tmp_path):
for name in AWKWARD:
(tmp_path / name).write_text("x")
os.utime(tmp_path / name, (0, 0)) # very old
tool = f"{shlex.quote(sys.executable)} -c 'from mytool.cli import app; app()'"
cmd = f"find {shlex.quote(str(tmp_path))} -type f -print0 | {tool} stale -0 -z | xargs -0 ls -1d"
out = subprocess.run(cmd, shell=True, capture_output=True, check=True)
for name in ["with space.txt", "tab\there.txt", "café.txt"]:
assert str(tmp_path / name) in out.stdout.decode()
The round-trip test with \xff is the one most implementations fail: it proves a name that is not valid UTF-8 comes out exactly as it went in. The pipeline test runs only where find and xargs exist, which covers Linux and macOS CI runners; on Windows, the unit tests still cover the logic.
Conclusion
Paths can contain anything except / and NUL, so NUL is the only safe separator for lists of them. Add -0 to read NUL-separated input and -z to write it, read and write the binary buffers, stream records in chunks while handling records split across chunk boundaries, decode with the file system encoding and surrogateescape so every name round-trips, refuse to write newline-containing names in newline mode, and test with awkward names and a real find/xargs pipeline.
Frequently asked questions
Does this matter on Windows?
Less: Windows file names cannot contain newlines, and xargs is not part of the system. Supporting -0 costs nothing there, and users of Git Bash, WSL and MSYS2 pipelines benefit from it.
How do I pass NUL-separated names to a subprocess from Python?
You do not need a separator at all — pass the list as arguments: subprocess.run(["rm", "--", *paths]). Arguments are separate strings, so no name can be split. Mind the operating system’s argument length limit for very long lists, as discussed in calling external commands safely with subprocess.
Can Click or Typer read NUL-separated stdin directly?
They provide click.File("rb") and - for stdin, but not record splitting. Read sys.stdin.buffer (or the opened binary file) with a reader like the one above.
What about NUL inside data, not file names?
If your records are arbitrary binary data that may contain NUL, no separator is safe; use a length-prefixed format or a structured encoding such as JSON lines with escaping.
Is surrogateescape safe to print to a terminal?
Printing a surrogate-escaped name through a text stream raises UnicodeEncodeError. Write such names through the binary buffer, as the recipe does, or show them with repr() in messages meant for people.