Architecture

Designing Idempotent Commands in a Python CLI

Make CLI commands safe to run twice: ensure-style commands, created/updated/unchanged results, exit codes for already-done work, idempotency keys for API retries, and tests.

Updated

Commands get run twice. A CI job is retried after a network blip. A deployment script is re-run after fixing the step that failed halfway. An operator is not sure whether the last attempt went through and presses Enter again. If mytool user add ann fails the second time with "user already exists", every script that calls it needs special-case error handling; if mytool invoice send 42 sends a second invoice, someone has a very bad afternoon. An idempotent command produces the same end state whether it runs once or five times. It is the single most useful property for a CLI that changes things, because it makes retries — by people, scripts and schedulers — safe by default. This guide shows how to design commands that way, how to report what happened, how to keep retried API calls idempotent, and how to test it. It belongs to the CLI interface design topic.

Prerequisites

Imperative versus declarative commands

Actions versus desired states A comparison of imperative commands that describe actions with declarative commands that describe a desired state, when run a second time. Actions versus desired states Command Second run Safe to retry? user add ann error: already exists no user ensure ann unchanged, exit 0 yes invoice send 42 second invoice sent no send 42 + idempotency key server returns first result yes Declarative commands make retries by people, scripts and schedulers safe by default.

An imperative command describes an action: add, create, send, append. Running it again repeats the action, so it is idempotent only by accident. A declarative command describes a desired state: ensure user ann exists with role admin, apply config.yaml, set retention 30d. Running it again finds the state already correct and does nothing. Tools like Terraform, Ansible and kubectl apply are built entirely on the second model, which is why they can be re-run safely.

A CLI does not need to become a configuration-management system to benefit. Most "create" commands can be made declarative with a small change in semantics: instead of "create, failing if it exists", the command means "make sure this exists and looks like this".

The recipe

Report what happened, not just success

An idempotent command still needs to tell the user and scripts what it did. Model the outcome explicitly:

# src/mytool/users.py
from __future__ import annotations

from dataclasses import dataclass
from enum import Enum


class Outcome(str, Enum):
    created = "created"
    updated = "updated"
    unchanged = "unchanged"


@dataclass
class User:
    name: str
    role: str


class Directory:
    """Stand-in for an API or database."""

    def __init__(self) -> None:
        self.users: dict[str, User] = {}

    def ensure_user(self, name: str, role: str) -> Outcome:
        existing = self.users.get(name)
        if existing is None:
            self.users[name] = User(name, role)
            return Outcome.created
        if existing.role != role:
            existing.role = role
            return Outcome.updated
        return Outcome.unchanged

    def ensure_absent(self, name: str) -> Outcome:
        if self.users.pop(name, None) is None:
            return Outcome.unchanged
        return Outcome.updated
# src/mytool/cli.py
import json

import typer

from mytool.users import Directory, Outcome

app = typer.Typer()
DIRECTORY = Directory()


@app.callback()
def main() -> None:
    """User management."""


@app.command()
def ensure(
    name: str,
    role: str = typer.Option("member", help="Role the user should have."),
    as_json: bool = typer.Option(False, "--json"),
) -> None:
    """Make sure NAME exists with ROLE. Safe to run repeatedly."""
    outcome = DIRECTORY.ensure_user(name, role)
    if as_json:
        typer.echo(json.dumps({"user": name, "role": role, "outcome": outcome.value}))
    else:
        typer.echo({Outcome.created: f"created {name} ({role})",
                    Outcome.updated: f"updated {name} to role {role}",
                    Outcome.unchanged: f"{name} already has role {role}"}[outcome])


@app.command()
def remove(name: str) -> None:
    """Make sure NAME does not exist. Succeeds if it is already gone."""
    outcome = DIRECTORY.ensure_absent(name)
    typer.echo(f"removed {name}" if outcome is Outcome.updated else f"{name} was not present")

All three outcomes exit with status 0, because all three end in the desired state. That is the key decision. "Already exists" is not an error when the user asked for something to exist, and "not found" is not an error when they asked for it to be gone. Scripts that need to know whether anything changed read the outcome field from --json output — or a --exit-code-style flag, as git diff offers, can map "changed" to a distinct exit status for scripts that want it.

Running the same command three times Terminal session running an ensure command repeatedly, reporting created, unchanged and updated outcomes, all with exit status zero. Running the same command three times bash $ mytool ensure ann --role admin created ann (admin) $ mytool ensure ann --role admin ann already has role admin $ mytool ensure ann --role member --json {"user": "ann", "role": "member", "outcome": "updated"} Every run ends in the desired state, so every run exits 0.

Keep strict variants for people who want them

Sometimes a user really does want "fail if it exists" — to catch a typo that would otherwise silently match an existing record. Offer it as an explicit opt-in, such as --create-only (fail if present) or --must-exist for updates, rather than as the default. The defaults should make retries safe; the flags make intent strict.

Idempotency across the network

Local checks are not enough when the change happens on a server and the response can be lost. If POST /invoices/42/send times out, the client cannot know whether the server sent the invoice. Retrying blindly may send it twice; not retrying may not send it at all. APIs that support idempotency keys solve this: the client generates a unique key per logical operation and sends it with every attempt, and the server remembers the result for that key and returns it again instead of repeating the work.

# src/mytool/api.py
import uuid

import httpx


def send_invoice(client: httpx.Client, invoice_id: int, *, attempts: int = 3) -> dict:
    key = str(uuid.uuid4())                       # one key for the whole logical operation
    last_error: Exception | None = None
    for _ in range(attempts):
        try:
            response = client.post(f"/invoices/{invoice_id}/send",
                                   headers={"Idempotency-Key": key}, timeout=10)
            response.raise_for_status()
            return response.json()
        except (httpx.TransportError, httpx.HTTPStatusError) as exc:
            if isinstance(exc, httpx.HTTPStatusError) and exc.response.status_code < 500:
                raise                             # 4xx: retrying will not help
            last_error = exc
    raise RuntimeError(f"giving up after {attempts} attempts") from last_error

The crucial detail is that the key is created once, outside the retry loop. A new key per attempt defeats the whole mechanism. Combine this with the backoff strategy from retries and backoff for CLI HTTP calls; for APIs without idempotency keys, check the resulting state (GET the invoice and look at its status) before retrying a non-idempotent call.

UX considerations

  • Name commands after outcomes. ensure, apply, set, sync and remove suggest idempotent behaviour; add, create and append suggest the opposite. If a command is idempotent, let its name say so.
  • Say "unchanged" clearly. A silent success leaves people unsure whether anything happened; "ann already has role admin" confirms both the state and the fact that nothing changed.
  • Make partial failures resumable. A bulk command that processes 500 items and fails at item 301 should, when re-run, skip the 300 already done — naturally true if each item is handled idempotently, and easy with state tracking as in storing CLI state in SQLite.
  • Pair with --dry-run. For declarative commands, a dry run can show exactly the planned outcome per item — created, updated or unchanged — before anything happens.
  • Document the contract. "Running this command twice is safe" in the help text is a promise scripts will rely on; make it explicit.
Three outcomes, one exit status The outcomes an idempotent command reports and how scripts learn which one happened. Three outcomes, one exit status created exit 0 the thing did not exist and now does updated exit 0 it existed but differed; now it matches unchanged exit 0 it already matched — nothing was done strict mode opt-in --create-only fails if present, for people who want that Scripts read the outcome from --json instead of guessing from exit codes.

Testing the behaviour

The defining test of an idempotent command is running it twice and checking the second run changes nothing and still succeeds:

# tests/test_idempotency.py
import json

import httpx
import pytest
import respx
from typer.testing import CliRunner

from mytool import api, cli

runner = CliRunner()


@pytest.fixture(autouse=True)
def fresh_directory(monkeypatch):
    monkeypatch.setattr(cli, "DIRECTORY", cli.Directory())


def outcome(*args):
    result = runner.invoke(cli.app, [*args, "--json"])
    assert result.exit_code == 0, result.output
    return json.loads(result.stdout)["outcome"]


def test_ensure_twice_is_unchanged():
    assert outcome("ensure", "ann", "--role", "admin") == "created"
    assert outcome("ensure", "ann", "--role", "admin") == "unchanged"
    assert outcome("ensure", "ann", "--role", "member") == "updated"


def test_remove_missing_user_succeeds():
    result = runner.invoke(cli.app, ["remove", "nobody"])
    assert result.exit_code == 0
    assert "was not present" in result.stdout


@respx.mock
def test_retries_reuse_the_idempotency_key():
    route = respx.post("https://api.example.com/invoices/42/send").mock(side_effect=[
        httpx.ConnectTimeout("lost"),
        httpx.Response(200, json={"status": "sent"}),
    ])
    with httpx.Client(base_url="https://api.example.com") as client:
        assert api.send_invoice(client, 42) == {"status": "sent"}
    keys = {call.request.headers["Idempotency-Key"] for call in route.calls}
    assert len(route.calls) == 2 and len(keys) == 1

The last test is the one that matters for network operations: it proves two attempts carried the same key, so a server that supports idempotency keys performs the action once. The general respx patterns are in mocking HTTP in CLI tests with respx.

Conclusion

Design commands that change things around desired states rather than actions: ensure instead of create, remove that succeeds when the thing is already gone. Return an explicit outcome — created, updated or unchanged — with exit status 0 for all of them, offer strict variants as opt-in flags, reuse one idempotency key across retries of network operations, and test every command by running it twice. Retries then stop being a risk and become the normal way to recover.

Frequently asked questions

Is "already exists" ever a legitimate error?

Yes, when creating a new thing is the whole point — reserving a unique name, for example — and silently reusing an existing one would be wrong. In that case, keep the error but give it a distinct exit code and a clear message, so scripts can tell it apart from real failures.

How do idempotency keys work if the CLI process crashes between attempts?

The key lives only in memory, so a re-run after a crash generates a new one. For operations where that matters, derive the key from the operation itself — for example a hash of the invoice ID and the date — or store it in local state before the first attempt, so a re-run can reuse it.

Are DELETE requests already idempotent?

By HTTP semantics, yes: deleting something twice leaves it deleted. Many APIs still return 404 on the second call; treat that as "unchanged" in a remove command rather than as an error.

How does this relate to dry runs?

They answer different questions. A dry run says what would change; idempotency guarantees that running for real a second time changes nothing more. Declarative commands make both easy, because the planned outcome is exactly what the real run would report.