Input & UX

Validating URLs, Hosts and Ports in Python CLI Arguments

Validate --url, --host and HOST:PORT arguments strictly: urlsplit pitfalls, allowed schemes, IPv6 brackets, port ranges, internationalised hostnames, refusing credentials in URLs, and tests.

Updated

Network options look trivial: --api-url, --host, --proxy, --listen 0.0.0.0:8080. Validate them with "does it start with http?" and the bugs arrive later and far away: a missing scheme that turns api.example.com/v1 into a relative path, a space in a hostname that fails deep inside an HTTP library with an unhelpful message, a port of 80800, an IPv6 address split on the wrong colon, or a URL with an embedded password that ends up in logs and shell history. Python's urllib.parse does the parsing but deliberately validates very little. This guide builds three small validators — for URLs, hostnames, and HOST:PORT pairs — that check what a CLI actually needs, normalise values once, explain mistakes clearly, and stay out of the business of checking reachability. It belongs to the argument validation topic.

Prerequisites

What urlsplit does and does not check

What urlsplit lets through Inputs that urllib.parse.urlsplit accepts without complaint, and what a command line tool should do with each. What urlsplit lets through Input urlsplit says Validator says api.example.com path only, no scheme did you mean https://…? https://exa mple.com hostname with a space contains whitespace https://user:pw@host credentials kept refuse — use a token https://host:99999 ValueError on .port port out of range urlsplit parses; it is deliberately not a validator.

urllib.parse.urlsplit splits a string into scheme, network location, path, query and fragment. It is lenient on purpose: urlsplit("api.example.com") succeeds with no scheme and the whole string as a path; urlsplit("https://exa mple.com") happily returns a hostname containing a space; user:password@ is parsed and kept. It does raise for one thing — accessing .port on https://host:99999 raises ValueError: Port out of range — which is a useful check to trigger deliberately. Everything else is your job.

The recipe

# src/mytool/netparse.py
from __future__ import annotations

import ipaddress
import re
from dataclasses import dataclass
from urllib.parse import urlsplit, urlunsplit

_LABEL = re.compile(r"^(?!-)[a-z0-9-]{1,63}(?<!-)$")


def validate_hostname(text: str) -> str:
    """A DNS name or IP address, normalised (lower case, IDNA, no brackets)."""
    host = text.strip().rstrip(".")
    if host.startswith("[") and host.endswith("]"):
        host = host[1:-1]
    try:
        return str(ipaddress.ip_address(host))
    except ValueError:
        pass
    try:
        ascii_host = host.encode("idna").decode("ascii").lower()
    except UnicodeError:
        raise ValueError(f"{text!r} is not a valid hostname") from None
    labels = ascii_host.split(".")
    if not ascii_host or len(ascii_host) > 253 or not all(_LABEL.match(l) for l in labels):
        raise ValueError(f"{text!r} is not a valid hostname")
    return ascii_host


def validate_port(value: str | int) -> int:
    try:
        port = int(value)
    except (TypeError, ValueError):
        raise ValueError(f"{value!r} is not a port number") from None
    if not 1 <= port <= 65535:
        raise ValueError(f"port {port} is out of range 1-65535")
    return port


def validate_url(text: str, *, schemes: tuple[str, ...] = ("https", "http"),
                 allow_credentials: bool = False) -> str:
    """Return a normalised absolute URL or raise ValueError with a helpful message."""
    raw = text.strip()
    if any(c.isspace() for c in raw):
        raise ValueError(f"{text!r} contains whitespace")
    parts = urlsplit(raw)
    if not parts.scheme or not parts.netloc:
        raise ValueError(f"{text!r} is not an absolute URL; did you mean https://{raw}?")
    if parts.scheme.lower() not in schemes:
        raise ValueError(f"scheme {parts.scheme!r} is not allowed; use {' or '.join(schemes)}")
    if (parts.username or parts.password) and not allow_credentials:
        raise ValueError("do not put credentials in the URL; use --token or the keyring")
    host = validate_hostname(parts.hostname or "")
    port = parts.port                                   # raises ValueError if out of range
    netloc = f"[{host}]" if ":" in host else host
    if port is not None:
        netloc += f":{port}"
    return urlunsplit((parts.scheme.lower(), netloc, parts.path.rstrip("/"), parts.query, ""))


@dataclass(frozen=True)
class HostPort:
    host: str
    port: int

    def __str__(self) -> str:
        return f"[{self.host}]:{self.port}" if ":" in self.host else f"{self.host}:{self.port}"


def parse_host_port(text: str, *, default_port: int | None = None) -> HostPort:
    """'host:port', '[::1]:8080', or 'host' with a default port."""
    s = text.strip()
    if s.startswith("["):
        host, sep, rest = s[1:].partition("]")
        port = rest[1:] if rest.startswith(":") else None
        if not sep or (rest and port is None):
            raise ValueError(f"{text!r} is not HOST:PORT; bracket IPv6 like [::1]:8080")
    elif s.count(":") == 1:
        host, port = s.split(":")
    elif ":" in s:
        raise ValueError(f"{text!r} looks like an IPv6 address; write it as [{s}]:PORT")
    else:
        host, port = s, None
    if port is None:
        if default_port is None:
            raise ValueError(f"{text!r} needs a port, like {s}:8080")
        port = default_port
    return HostPort(validate_hostname(host), validate_port(port))

The validators normalise as they check: scheme and hostname are lower-cased, internationalised names are converted to their ASCII (IDNA) form, IPv6 addresses lose their brackets for storage and regain them for display, trailing slashes are dropped. The rest of the program then compares and logs canonical values. Each error message says what was wrong and what to type instead — did you mean https://api.example.com? resolves the most common mistake of all.

Normalised and rejected values Terminal session showing a URL normalised to lower case without a trailing slash, an IPv6 listen address accepted with brackets, and a missing scheme rejected with a suggestion. Normalised and rejected values bash $ mytool serve --listen '[::1]:9000' --api-url HTTPS://Api.Example.com/ proxying [::1]:9000 -> https://api.example.com $ mytool serve --api-url api.example.com Invalid value for '--api-url': 'api.example.com' is not an absolute URL; did you mean https://api.example.com? Canonical values once, at the boundary — the rest of the program compares strings safely.

Wiring into a command

# src/mytool/cli.py
from typing import Annotated

import typer

from mytool.netparse import HostPort, parse_host_port, validate_url

app = typer.Typer()


def _url(text: str) -> str:
    try:
        return validate_url(text)
    except ValueError as exc:
        raise typer.BadParameter(str(exc)) from None


def _listen(text: str) -> HostPort:
    try:
        return parse_host_port(text, default_port=8080)
    except ValueError as exc:
        raise typer.BadParameter(str(exc)) from None


@app.callback()
def main() -> None:
    """Service tool."""


@app.command()
def serve(
    api_url: Annotated[str, typer.Option(parser=_url, metavar="URL", envvar="MYTOOL_API_URL")]
        = "https://api.example.com",
    listen: Annotated[HostPort, typer.Option(parser=_listen, metavar="HOST[:PORT]")]
        = "127.0.0.1:8080",
) -> None:
    """Run the local proxy."""
    typer.echo(f"proxying {listen} -> {api_url}")

Because the option uses envvar=, a malformed MYTOOL_API_URL in CI is rejected with the same message as a malformed flag, at startup rather than at the first request.

Validate syntax, not reachability

It is tempting to check that the host resolves or the port is open while validating. Do not, at least not by default: DNS may be slow, the network may be down, a VPN may not be connected yet, and the command may never need to connect at all (--dry-run, --help, writing a config file). Validation should be instant and offline. Connectivity checks belong in the code that connects — with a clear error there — or in an explicit mytool doctor or --check command. The debugging guide's doctor command is a natural home for "can we reach the API?".

Network option validation Checks to perform when validating URL, host and port options, and checks to avoid during validation. Network option validation Check ✓ Absolute URL with an allowed scheme ✓ Hostname labels or an IP address ✓ Port in 1–65535; IPv6 in brackets ✓ No credentials embedded in the URL Do not ✗ Resolve DNS while parsing options ✗ Open connections to "test" the URL ✗ Echo a rejected URL with its password ✗ Default listeners to 0.0.0.0 Validation is instant and offline; connectivity is checked where you connect.

UX considerations

  • Suggest the fix. A missing scheme is the most frequent error; offering the https:// version saves a round trip.
  • Refuse credentials in URLs. They leak into logs, process listings and shell history. Point users at a token option or storing tokens with keyring, and never echo the rejected URL with its password.
  • Default to safe binds. 127.0.0.1 rather than 0.0.0.0 for anything that listens; make exposure on all interfaces an explicit choice.
  • Accept what people paste. Trailing slashes, upper-case schemes and IPv6 with brackets should all just work.
  • Show the normalised value in verbose output, so users can see what the tool will actually use.

Testing the behaviour

# tests/test_netparse.py
import pytest

from mytool.netparse import HostPort, parse_host_port, validate_hostname, validate_url


@pytest.mark.parametrize("text,expected", [
    ("https://API.Example.com/v1/", "https://api.example.com/v1"),
    ("HTTP://example.com:8080", "http://example.com:8080"),
    ("https://[::1]:8443/x", "https://[::1]:8443/x"),
    ("https://bücher.example/", "https://xn--bcher-kva.example"),
])
def test_urls_are_normalised(text, expected):
    assert validate_url(text) == expected


@pytest.mark.parametrize("text,message", [
    ("api.example.com", "did you mean https://api.example.com"),
    ("ftp://example.com", "scheme 'ftp' is not allowed"),
    ("https://exa mple.com", "contains whitespace"),
    ("https://user:pw@example.com", "do not put credentials"),
    ("https://example.com:99999", "Port out of range"),
    ("https://-bad-.example.com", "not a valid hostname"),
])
def test_bad_urls_are_explained(text, message):
    with pytest.raises(ValueError, match=message):
        validate_url(text)


@pytest.mark.parametrize("text,expected", [
    ("example.com:443", HostPort("example.com", 443)),
    ("[::1]:8080", HostPort("::1", 8080)),
    ("10.0.0.5", HostPort("10.0.0.5", 8080)),
])
def test_host_port(text, expected):
    assert parse_host_port(text, default_port=8080) == expected


def test_bare_ipv6_needs_brackets():
    with pytest.raises(ValueError, match=r"write it as \[::1\]:PORT"):
        parse_host_port("::1")


def test_ip_addresses_are_canonical():
    assert validate_hostname("[0:0:0:0:0:0:0:1]") == "::1"

Parametrised tables keep the accepted and rejected forms visible in one place; adding a newly reported user mistake is a one-line change.

Conclusion

urlsplit parses but barely validates, so CLIs need their own checks: require an absolute URL with an allowed scheme, reject whitespace and embedded credentials, validate hostnames as IP addresses or DNS labels (with IDNA for international names), check port ranges, and handle IPv6 brackets in HOST:PORT values. Normalise while validating, explain every error with the fix, keep validation offline, and test with tables of real user input.

Frequently asked questions

Should I use pydantic's HttpUrl instead?

If your CLI already uses Pydantic for settings, HttpUrl and friends are a good choice and cover many of these checks. The standalone validators here avoid the dependency and let you tailor messages and rules — such as refusing credentials — exactly.

How do I accept URLs with paths that must keep a trailing slash?

Some APIs treat /v1 and /v1/ differently. If yours does, remove the rstrip("/") and document the expected form; normalisation should never change meaning.

What about localhost and single-label names?

They pass the hostname check, which is right for development. If a production option must not point at localhost, enforce that as a separate rule with its own message rather than weakening the general validator.

Can I validate CIDR ranges the same way?

Yes — ipaddress.ip_network(text, strict=False) parses and normalises them, and a parser= wrapper turns its ValueError into a usage error, just like the validators above.

Should validation reject private or internal addresses?

Only when the option's purpose demands it — a webhook URL that your server will call, for example, should not point at 169.254.169.254 or 10.x addresses, to prevent server-side request forgery. ipaddress.ip_address(host).is_private and .is_link_local make the check easy, but remember that a hostname can resolve to a private address too; for that case, check after resolution, at connection time.