Network options look trivial: --api-url, --host, --proxy, --listen 0.0.0.0:8080. Validate them with "does it start with http?" and the bugs arrive later and far away: a missing scheme that turns api.example.com/v1 into a relative path, a space in a hostname that fails deep inside an HTTP library with an unhelpful message, a port of 80800, an IPv6 address split on the wrong colon, or a URL with an embedded password that ends up in logs and shell history. Python's urllib.parse does the parsing but deliberately validates very little. This guide builds three small validators — for URLs, hostnames, and HOST:PORT pairs — that check what a CLI actually needs, normalise values once, explain mistakes clearly, and stay out of the business of checking reachability. It belongs to the argument validation topic.
Prerequisites
- Python 3.10+; everything here is standard library (
urllib.parse,ipaddress). - Typer or Click; the
parser=pattern from validating dates and durations in CLI arguments is reused.
What urlsplit does and does not check
urllib.parse.urlsplit splits a string into scheme, network location, path, query and fragment. It is lenient on purpose: urlsplit("api.example.com") succeeds with no scheme and the whole string as a path; urlsplit("https://exa mple.com") happily returns a hostname containing a space; user:password@ is parsed and kept. It does raise for one thing — accessing .port on https://host:99999 raises ValueError: Port out of range — which is a useful check to trigger deliberately. Everything else is your job.
The recipe
# src/mytool/netparse.py
from __future__ import annotations
import ipaddress
import re
from dataclasses import dataclass
from urllib.parse import urlsplit, urlunsplit
_LABEL = re.compile(r"^(?!-)[a-z0-9-]{1,63}(?<!-)$")
def validate_hostname(text: str) -> str:
"""A DNS name or IP address, normalised (lower case, IDNA, no brackets)."""
host = text.strip().rstrip(".")
if host.startswith("[") and host.endswith("]"):
host = host[1:-1]
try:
return str(ipaddress.ip_address(host))
except ValueError:
pass
try:
ascii_host = host.encode("idna").decode("ascii").lower()
except UnicodeError:
raise ValueError(f"{text!r} is not a valid hostname") from None
labels = ascii_host.split(".")
if not ascii_host or len(ascii_host) > 253 or not all(_LABEL.match(l) for l in labels):
raise ValueError(f"{text!r} is not a valid hostname")
return ascii_host
def validate_port(value: str | int) -> int:
try:
port = int(value)
except (TypeError, ValueError):
raise ValueError(f"{value!r} is not a port number") from None
if not 1 <= port <= 65535:
raise ValueError(f"port {port} is out of range 1-65535")
return port
def validate_url(text: str, *, schemes: tuple[str, ...] = ("https", "http"),
allow_credentials: bool = False) -> str:
"""Return a normalised absolute URL or raise ValueError with a helpful message."""
raw = text.strip()
if any(c.isspace() for c in raw):
raise ValueError(f"{text!r} contains whitespace")
parts = urlsplit(raw)
if not parts.scheme or not parts.netloc:
raise ValueError(f"{text!r} is not an absolute URL; did you mean https://{raw}?")
if parts.scheme.lower() not in schemes:
raise ValueError(f"scheme {parts.scheme!r} is not allowed; use {' or '.join(schemes)}")
if (parts.username or parts.password) and not allow_credentials:
raise ValueError("do not put credentials in the URL; use --token or the keyring")
host = validate_hostname(parts.hostname or "")
port = parts.port # raises ValueError if out of range
netloc = f"[{host}]" if ":" in host else host
if port is not None:
netloc += f":{port}"
return urlunsplit((parts.scheme.lower(), netloc, parts.path.rstrip("/"), parts.query, ""))
@dataclass(frozen=True)
class HostPort:
host: str
port: int
def __str__(self) -> str:
return f"[{self.host}]:{self.port}" if ":" in self.host else f"{self.host}:{self.port}"
def parse_host_port(text: str, *, default_port: int | None = None) -> HostPort:
"""'host:port', '[::1]:8080', or 'host' with a default port."""
s = text.strip()
if s.startswith("["):
host, sep, rest = s[1:].partition("]")
port = rest[1:] if rest.startswith(":") else None
if not sep or (rest and port is None):
raise ValueError(f"{text!r} is not HOST:PORT; bracket IPv6 like [::1]:8080")
elif s.count(":") == 1:
host, port = s.split(":")
elif ":" in s:
raise ValueError(f"{text!r} looks like an IPv6 address; write it as [{s}]:PORT")
else:
host, port = s, None
if port is None:
if default_port is None:
raise ValueError(f"{text!r} needs a port, like {s}:8080")
port = default_port
return HostPort(validate_hostname(host), validate_port(port))
The validators normalise as they check: scheme and hostname are lower-cased, internationalised names are converted to their ASCII (IDNA) form, IPv6 addresses lose their brackets for storage and regain them for display, trailing slashes are dropped. The rest of the program then compares and logs canonical values. Each error message says what was wrong and what to type instead — did you mean https://api.example.com? resolves the most common mistake of all.
Wiring into a command
# src/mytool/cli.py
from typing import Annotated
import typer
from mytool.netparse import HostPort, parse_host_port, validate_url
app = typer.Typer()
def _url(text: str) -> str:
try:
return validate_url(text)
except ValueError as exc:
raise typer.BadParameter(str(exc)) from None
def _listen(text: str) -> HostPort:
try:
return parse_host_port(text, default_port=8080)
except ValueError as exc:
raise typer.BadParameter(str(exc)) from None
@app.callback()
def main() -> None:
"""Service tool."""
@app.command()
def serve(
api_url: Annotated[str, typer.Option(parser=_url, metavar="URL", envvar="MYTOOL_API_URL")]
= "https://api.example.com",
listen: Annotated[HostPort, typer.Option(parser=_listen, metavar="HOST[:PORT]")]
= "127.0.0.1:8080",
) -> None:
"""Run the local proxy."""
typer.echo(f"proxying {listen} -> {api_url}")
Because the option uses envvar=, a malformed MYTOOL_API_URL in CI is rejected with the same message as a malformed flag, at startup rather than at the first request.
Validate syntax, not reachability
It is tempting to check that the host resolves or the port is open while validating. Do not, at least not by default: DNS may be slow, the network may be down, a VPN may not be connected yet, and the command may never need to connect at all (--dry-run, --help, writing a config file). Validation should be instant and offline. Connectivity checks belong in the code that connects — with a clear error there — or in an explicit mytool doctor or --check command. The debugging guide's doctor command is a natural home for "can we reach the API?".
UX considerations
- Suggest the fix. A missing scheme is the most frequent error; offering the
https://version saves a round trip. - Refuse credentials in URLs. They leak into logs, process listings and shell history. Point users at a token option or storing tokens with keyring, and never echo the rejected URL with its password.
- Default to safe binds.
127.0.0.1rather than0.0.0.0for anything that listens; make exposure on all interfaces an explicit choice. - Accept what people paste. Trailing slashes, upper-case schemes and IPv6 with brackets should all just work.
- Show the normalised value in verbose output, so users can see what the tool will actually use.
Testing the behaviour
# tests/test_netparse.py
import pytest
from mytool.netparse import HostPort, parse_host_port, validate_hostname, validate_url
@pytest.mark.parametrize("text,expected", [
("https://API.Example.com/v1/", "https://api.example.com/v1"),
("HTTP://example.com:8080", "http://example.com:8080"),
("https://[::1]:8443/x", "https://[::1]:8443/x"),
("https://bücher.example/", "https://xn--bcher-kva.example"),
])
def test_urls_are_normalised(text, expected):
assert validate_url(text) == expected
@pytest.mark.parametrize("text,message", [
("api.example.com", "did you mean https://api.example.com"),
("ftp://example.com", "scheme 'ftp' is not allowed"),
("https://exa mple.com", "contains whitespace"),
("https://user:pw@example.com", "do not put credentials"),
("https://example.com:99999", "Port out of range"),
("https://-bad-.example.com", "not a valid hostname"),
])
def test_bad_urls_are_explained(text, message):
with pytest.raises(ValueError, match=message):
validate_url(text)
@pytest.mark.parametrize("text,expected", [
("example.com:443", HostPort("example.com", 443)),
("[::1]:8080", HostPort("::1", 8080)),
("10.0.0.5", HostPort("10.0.0.5", 8080)),
])
def test_host_port(text, expected):
assert parse_host_port(text, default_port=8080) == expected
def test_bare_ipv6_needs_brackets():
with pytest.raises(ValueError, match=r"write it as \[::1\]:PORT"):
parse_host_port("::1")
def test_ip_addresses_are_canonical():
assert validate_hostname("[0:0:0:0:0:0:0:1]") == "::1"
Parametrised tables keep the accepted and rejected forms visible in one place; adding a newly reported user mistake is a one-line change.
Conclusion
urlsplit parses but barely validates, so CLIs need their own checks: require an absolute URL with an allowed scheme, reject whitespace and embedded credentials, validate hostnames as IP addresses or DNS labels (with IDNA for international names), check port ranges, and handle IPv6 brackets in HOST:PORT values. Normalise while validating, explain every error with the fix, keep validation offline, and test with tables of real user input.
Frequently asked questions
Should I use pydantic's HttpUrl instead?
If your CLI already uses Pydantic for settings, HttpUrl and friends are a good choice and cover many of these checks. The standalone validators here avoid the dependency and let you tailor messages and rules — such as refusing credentials — exactly.
How do I accept URLs with paths that must keep a trailing slash?
Some APIs treat /v1 and /v1/ differently. If yours does, remove the rstrip("/") and document the expected form; normalisation should never change meaning.
What about localhost and single-label names?
They pass the hostname check, which is right for development. If a production option must not point at localhost, enforce that as a separate rule with its own message rather than weakening the general validator.
Can I validate CIDR ranges the same way?
Yes — ipaddress.ip_network(text, strict=False) parses and normalises them, and a parser= wrapper turns its ValueError into a usage error, just like the validators above.
Should validation reject private or internal addresses?
Only when the option's purpose demands it — a webhook URL that your server will call, for example, should not point at 169.254.169.254 or 10.x addresses, to prevent server-side request forgery. ipaddress.ip_address(host).is_private and .is_link_local make the check easy, but remember that a hostname can resolve to a private address too; for that case, check after resolution, at connection time.