Input & UX

Validating Config Files with JSON Schema in a Python CLI

Describe a CLI config file once with JSON Schema: report every problem by key path, flag unknown keys, ship the schema for editor completion, and test it.

Updated

Configuration errors are the most frustrating kind a CLI can produce, because the user did not type them a second ago — they are in a file written weeks earlier, maybe copied from a colleague. A loader that raises KeyError: 'url' or silently ignores timout = 5 leaves the user guessing. A good configuration check reports every problem in one pass, points at the exact key (api.timeout), says what was expected, and catches misspelt keys instead of ignoring them. JSON Schema is a standard way to express exactly those rules, and it brings a bonus no hand-written check can: the same schema file gives users autocompletion, hover documentation and inline errors in VS Code, JetBrains IDEs and any editor using the Taplo TOML language server. This guide writes a schema for a TOML config, validates it with the jsonschema library, turns raw validation errors into messages users understand, adds config validate and config schema commands, and tests the lot. It belongs to the configuration files and environment variables topic.

Prerequisites

  • uv add jsonschema typer (examples checked with jsonschema 4.26, Typer 0.27 and Python 3.13).
  • A CLI that loads TOML, as in reading TOML config with tomllib. YAML works the same way once parsed.

Why a schema, and not just code

One schema, many consumers A JSON Schema file shipped in the package is used by the command line tool’s loader, its validate command, users’ editors and CI checks. One schema, many consumers config.schema.json in the package Loader + validate all problems at once Editors completion, hover, errors CI check shared configs iter_errors #:schema config schema Rules written as data travel further than rules written as if statements.

JSON Schema validates data, not JSON text — any parsed TOML, YAML or JSON document is a tree of dictionaries, lists, strings and numbers, which is exactly what a schema describes. Writing the rules as data rather than if statements has three effects. The rules are complete and declarative, so “all problems at once” comes for free. They are portable: editors, CI linters such as check-jsonschema, and documentation generators can all read the same file. And they are reviewable: a pull request that adds a setting shows the new key, its type, its range and its description in one place.

If your CLI already uses a typed settings model, you may not need to write the schema by hand — see the section on pydantic below. The validation and reporting code is the same either way.

The recipe

The schema

The schema lives inside the package as src/mytool/config.schema.json, next to the code that uses it, so it ships in the wheel:

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://example.com/mytool/config.schema.json",
  "title": "mytool configuration",
  "type": "object",
  "additionalProperties": false,
  "properties": {
    "api": {
      "type": "object",
      "additionalProperties": false,
      "properties": {
        "url": {"type": "string", "format": "uri", "pattern": "^https?://"},
        "timeout": {"type": "number", "exclusiveMinimum": 0, "maximum": 600}
      },
      "required": ["url"]
    },
    "output": {
      "type": "object",
      "additionalProperties": false,
      "properties": {
        "format": {"enum": ["table", "json", "csv"]},
        "color": {"type": "boolean"}
      }
    },
    "profiles": {
      "type": "object",
      "additionalProperties": {
        "type": "object",
        "properties": {"region": {"type": "string", "minLength": 1}},
        "required": ["region"]
      }
    }
  }
}

Three choices make it strict in the ways that help users. additionalProperties: false on each table turns a misspelt key from a silently ignored setting into an error. Ranges and enums (exclusiveMinimum, maximum, enum) catch values that parse fine but make no sense. required documents which keys have no default. The profiles table shows the pattern for user-named sections: additionalProperties holds a schema rather than false, so any profile name is allowed but every profile must have a region.

Validating and reporting

# src/mytool/config_schema.py
from __future__ import annotations

import json
import tomllib
from dataclasses import dataclass
from functools import cache
from importlib.resources import files
from pathlib import Path

from jsonschema import Draft202012Validator
from jsonschema.exceptions import ValidationError


@dataclass(frozen=True)
class Problem:
    location: str
    message: str

    def __str__(self) -> str:
        return f"{self.location}: {self.message}"


@cache
def validator() -> Draft202012Validator:
    schema = json.loads(files("mytool").joinpath("config.schema.json").read_text(encoding="utf-8"))
    Draft202012Validator.check_schema(schema)        # a broken schema is our bug, not the user's
    return Draft202012Validator(schema, format_checker=Draft202012Validator.FORMAT_CHECKER)


def _location(error: ValidationError) -> str:
    parts = [str(p) for p in error.absolute_path]
    if error.validator == "additionalProperties":   # point at the unknown key itself
        extra = sorted(set(error.instance) - set(error.schema.get("properties", {})))
        parts += extra[:1]
    return ".".join(parts) or "(top level)"


def _message(error: ValidationError) -> str:
    match error.validator:
        case "additionalProperties":
            allowed = ", ".join(sorted(error.schema.get("properties", {})))
            return f"unknown key (allowed: {allowed})"
        case "enum":
            return "must be one of " + ", ".join(map(repr, error.validator_value))
        case "required":
            return error.message.replace("is a required property", "is required")
        case _:
            return error.message


def check(data: dict) -> list[Problem]:
    errors = sorted(validator().iter_errors(data), key=lambda e: list(map(str, e.absolute_path)))
    return [Problem(_location(e), _message(e)) for e in errors]


def check_file(path: Path) -> list[Problem]:
    try:
        data = tomllib.loads(path.read_text(encoding="utf-8"))
    except tomllib.TOMLDecodeError as exc:
        return [Problem(str(path), f"not valid TOML: {exc}")]
    return check(data)

iter_errors yields every violation rather than stopping at the first, and sorting by path groups problems by table. absolute_path is the list of keys leading to the failing value, which joins naturally into the dotted notation users see in their file (api.timeout). The two helpers then translate the library’s messages where they read poorly. An additionalProperties error normally says “Additional properties are not allowed ('timout' was unexpected)” and points at the parent table; the helper points at the unknown key itself and lists the keys that are allowed, which is usually enough for the user to spot the typo. Enum errors become a plain list of choices. Everything else uses the library’s message, which is already clear for types and ranges.

$ mytool config validate ~/.config/mytool/config.toml
~/.config/mytool/config.toml: api.retries: unknown key (allowed: timeout, url)
~/.config/mytool/config.toml: api: 'url' is required
~/.config/mytool/config.toml: api.timeout: -1 is less than or equal to the minimum of 0
~/.config/mytool/config.toml: output.format: must be one of 'table', 'json', 'csv'

The validator is built once and cached. check_schema validates the schema against the JSON Schema meta-schema first, so a mistake in your schema fails loudly in your tests rather than producing confusing results for users.

Commands for users and editors

# src/mytool/validate_cmd.py
import sys
from importlib.resources import files
from pathlib import Path
from typing import Annotated

import typer

from mytool.config_schema import check_file

app = typer.Typer(help="Configuration commands.")


@app.command()
def validate(path: Annotated[Path, typer.Argument(exists=True, dir_okay=False)]) -> None:
    """Check a configuration file against the schema and report every problem."""
    problems = check_file(path)
    for problem in problems:
        typer.echo(f"{path}: {problem}", err=True)
    if problems:
        raise typer.Exit(1)
    typer.echo(f"{path}: ok", err=True)


@app.command()
def schema() -> None:
    """Print the JSON Schema for editors and other tools."""
    sys.stdout.write(files("mytool").joinpath("config.schema.json").read_text(encoding="utf-8"))

config validate is useful on its own — in CI for teams that keep a shared config in a repository, or after hand-editing — and the runtime loader should call the same check_file and refuse to run with an invalid file, printing the same messages. config schema prints the schema on stdout so other tools can consume it.

Every problem, by key Terminal session validating a configuration file and receiving four problems at once, each with a dotted key path, then a clean run. Every problem, by key bash $ mytool config validate config.toml config.toml: api.retries: unknown key (allowed: timeout, url) config.toml: api.timeout: -1 is less than or equal to the minimum of 0 config.toml: output.format: must be one of 'table', 'json', 'csv' $ mytool config validate config.toml # after fixing config.toml: ok One run, one list — not one error per attempt.

Editor support for free

The schema becomes much more valuable once users’ editors know about it. For TOML, the Taplo language server (used by the “Even Better TOML” VS Code extension and others) reads a directive on the first line of the file:

#:schema https://example.com/mytool/config.schema.json
[api]
url = "https://api.example.com"
timeout = 30

With that line, the editor completes key names, shows the description fields on hover and underlines invalid values as the user types. Put the directive in the template written by config init, as in writing a config init and edit command, and publish the schema at a stable URL with each release. For YAML configuration, the equivalent is # yaml-language-server: $schema=<url>; for a project-level config, registering the schema in the community SchemaStore catalogue makes editors pick it up by filename without any directive at all. Adding "description" to every property is the single most useful improvement for that experience.

Generating the schema from a settings model

If configuration is loaded into a pydantic model, as in typed settings with pydantic-settings, the model already contains the rules, and Settings.model_json_schema() produces a schema from it. Generate the file in a small script, commit it, and add a test that regenerates it and compares — the same “generated file must match” pattern used for documentation. You keep a single source of truth while still shipping a static schema for editors. Pydantic’s own validation errors are good, so many tools use pydantic for runtime checks and the generated schema only for editors.

Schema keywords that help users JSON Schema keywords used in a command line tool configuration schema and the mistake each one catches. Schema keywords that help users Keyword Catches Example additionalProperties: false misspelt keys timout = 5 enum unsupported choices format = "xml" exclusiveMinimum / maximum nonsense numbers timeout = 0 required missing settings no api.url pattern wrong shape url = "ftp://…" Add a description to every property too — editors show it on hover.

UX considerations

  • Report everything at once. Fixing one error per run, five runs in a row, is the experience to avoid.
  • Use the user’s vocabulary. Dotted TOML paths and the file name; never Python reprs of internal objects or JSON Pointer syntax.
  • Be strict about keys, lenient about extras you expect. Unknown keys should fail; if you support plugin sections, give them their own additionalProperties schema rather than turning strictness off.
  • Mind format checks. "format": "uri" is only enforced when the optional rfc3987 package is installed (and other formats have their own extras); the pattern above enforces the important part regardless.
  • Keep old keys working. When renaming a setting, keep the old name in the schema for a release with "deprecated": true, and warn — see versioning and deprecating CLI flags.

Testing the behaviour

Test the schema itself, one representative problem per rule, and the command-line behaviour:

# tests/test_config_schema.py
import json
from importlib.resources import files

import pytest
from jsonschema import Draft202012Validator

from mytool.config_schema import check, check_file

VALID = {"api": {"url": "https://api.example.com", "timeout": 30},
         "output": {"format": "json"}, "profiles": {"eu": {"region": "eu-west-1"}}}


def test_schema_itself_is_valid():
    schema = json.loads(files("mytool").joinpath("config.schema.json").read_text())
    Draft202012Validator.check_schema(schema)


def test_valid_config_has_no_problems():
    assert check(VALID) == []


@pytest.mark.parametrize("data,location,fragment", [
    ({"api": {"url": "https://x", "timout": 5}}, "api.timout", "unknown key"),
    ({"api": {"url": "https://x", "timeout": 0}}, "api.timeout", "less than or equal to the minimum"),
    ({"api": {"url": "ftp://x"}}, "api.url", "does not match"),
    ({"api": {}}, "api", "'url' is required"),
    ({"output": {"format": "xml"}}, "output.format", "must be one of 'table'"),
    ({"profiles": {"eu": {}}}, "profiles.eu", "'region' is required"),
])
def test_problems_point_at_the_key(data, location, fragment):
    problems = check(data)
    assert len(problems) == 1
    assert problems[0].location == location
    assert fragment in problems[0].message


def test_all_problems_are_reported_at_once():
    data = {"api": {"url": "https://x", "timeout": -1}, "output": {"format": "xml", "colour": True}}
    assert [p.location for p in check(data)] == ["api.timeout", "output.colour", "output.format"]


def test_broken_toml_is_reported(tmp_path):
    path = tmp_path / "config.toml"
    path.write_text("[api\n")
    [problem] = check_file(path)
    assert "not valid TOML" in problem.message
# tests/test_validate_cmd.py
import json

from typer.testing import CliRunner

from mytool.validate_cmd import app

runner = CliRunner()


def test_validate_reports_problems_and_exit_code(tmp_path):
    path = tmp_path / "config.toml"
    path.write_text('[api]\nurl = "https://x"\ntimeout = 0\n[output]\nformat = "xml"\n')
    result = runner.invoke(app, ["validate", str(path)])
    assert result.exit_code == 1
    assert result.stderr.count(str(path)) == 2


def test_validate_ok(tmp_path):
    path = tmp_path / "config.toml"
    path.write_text('[api]\nurl = "https://api.example.com"\n')
    assert runner.invoke(app, ["validate", str(path)]).exit_code == 0


def test_schema_is_printed_as_json():
    result = runner.invoke(app, ["schema"])
    assert json.loads(result.stdout)["title"] == "mytool configuration"

The parametrised test doubles as a specification of the messages users see; when someone loosens a rule by accident, a case fails. Keep a valid example config in the tests (or in your documentation, tested from there) so a schema change that rejects real configurations is caught before release.

Conclusion

Describe the configuration file with a JSON Schema shipped inside the package: additionalProperties: false to catch typos, ranges and enums for values, required for keys without defaults. Validate with Draft202012Validator.iter_errors to collect every problem, convert error paths to dotted keys and rewrite the few unfriendly messages, expose config validate and config schema commands, put a #:schema directive in generated files so editors autocomplete and check as users type, and test both the schema and the messages.

Frequently asked questions

Which JSON Schema draft should I use?

Draft 2020-12 is current and supported by the jsonschema library and the major editor integrations. Declare it with $schema and use the matching validator class; mixing drafts is a common source of keywords being silently ignored.

Is jsonschema fast enough to run on every invocation?

For a config file of a few dozen keys, validation takes well under a millisecond once the validator exists. The import costs a few tens of milliseconds, so import it lazily inside the validation function if startup time matters, as discussed in CLI startup performance and lazy loading.

Can I report line numbers?

tomllib does not keep positions. tomlkit preserves them internally but does not expose a stable API for it, so most CLIs report dotted key paths, which users can search for. Editors using the schema show the exact position anyway.

Should environment variables and flags be validated with the same schema?

Validate the merged configuration with the same rules, so a bad MYTOOL_API_TIMEOUT=-1 is caught too — but report the source (“from MYTOOL_API_TIMEOUT”) rather than a file path.

What about defaults?

JSON Schema’s default keyword is documentation; validators do not fill values in. Apply defaults in the loader (or the settings model) and keep the default annotations in the schema so editors can show them.