Runtime

Caching HTTP Responses on Disk in a Python CLI

Cache API responses in SQLite from a Python CLI: TTLs, ETag revalidation, stale fallback when offline, size limits, refresh and no-cache flags, and tests.

Updated

A CLI that talks to an API repeats itself constantly. mytool show 42 fetches ticket 42; a minute later mytool show 42 --json | jq .status fetches it again; the shell completion for ticket numbers fetches the whole list on every Tab. Each request costs latency the user feels, counts against a rate limit, and fails outright on a train with no signal. A small disk cache fixes all three: answers for recently fetched URLs come from local storage in a millisecond, revalidation with ETag makes refreshes nearly free, and stale data can stand in when the network is gone. This guide builds that cache on SQLite with httpx, with the expiry, size limits and user controls that keep it from becoming a source of confusing stale results. It belongs to the SQLite state topic.

Prerequisites

  • Python 3.10+, httpx, platformdirs; respx and pytest for tests.
  • An API client along the lines of building an API client CLI with httpx.
  • A clear idea of which endpoints are safe to cache. Only cache GET requests whose answers are the same for every call with the same URL and credentials.

How the cache decides

For each GET, the cache has four possible answers, and the order of checks matters:

Four answers to a GET How the disk cache answers a request: a fresh hit, a revalidation with an ETag, a normal fetch, or stale data when offline. Four answers to a GET 1. Fresh hit ~1 ms younger than the TTL — no network at all 2. Revalidate tiny request stale but has an ETag — If-None-Match, 304 refreshes it 3. Miss full request fetch, store only successful and storable responses 4. Offline fallback warning network error with stale data — return it, labelled stale Errors are never cached, so an outage does not outlive itself.
  1. Fresh hit — a stored response younger than its time-to-live. Return it without touching the network.
  2. Stale with a validator — the stored response is old, but it has an ETag (or Last-Modified). Send a conditional request; a 304 Not Modified answer refreshes the timestamp and returns the stored body, at the cost of a tiny response.
  3. Miss — nothing stored, or stale without a validator. Fetch normally and store the result if it is cacheable.
  4. Network failure with stale data — return the stale response and tell the user it is stale. Without stale data, the error propagates as usual.

The recipe

# src/tickets/httpcache.py
from __future__ import annotations

import json
import sqlite3
import time
from dataclasses import dataclass
from pathlib import Path

import httpx

SCHEMA = """
CREATE TABLE IF NOT EXISTS responses (
    key TEXT PRIMARY KEY,
    status INTEGER NOT NULL,
    headers TEXT NOT NULL,
    body BLOB NOT NULL,
    etag TEXT,
    fetched_at REAL NOT NULL,
    size INTEGER NOT NULL
)"""


@dataclass
class CachedResponse:
    status: int
    headers: dict[str, str]
    body: bytes
    age: float
    stale: bool = False

    def json(self):
        return json.loads(self.body)


class HttpCache:
    def __init__(self, path: Path, *, max_bytes: int = 50_000_000) -> None:
        path.parent.mkdir(parents=True, exist_ok=True)
        try:
            self.conn = self._open(path)
        except sqlite3.DatabaseError:            # corrupt cache: throw it away
            path.unlink(missing_ok=True)
            self.conn = self._open(path)
        self.max_bytes = max_bytes

    @staticmethod
    def _open(path: Path) -> sqlite3.Connection:
        conn = sqlite3.connect(path, timeout=5.0)
        conn.execute("PRAGMA journal_mode = WAL")
        conn.execute(SCHEMA)
        return conn

    def get(self, key: str) -> CachedResponse | None:
        row = self.conn.execute(
            "SELECT status, headers, body, fetched_at FROM responses WHERE key = ?", (key,)
        ).fetchone()
        if row is None:
            return None
        status, headers, body, fetched_at = row
        return CachedResponse(status, json.loads(headers), body, time.time() - fetched_at)

    def etag(self, key: str) -> str | None:
        row = self.conn.execute("SELECT etag FROM responses WHERE key = ?", (key,)).fetchone()
        return row[0] if row else None

    def put(self, key: str, response: httpx.Response) -> None:
        body = response.content
        headers = {k: v for k, v in response.headers.items() if k.lower() in
                   ("content-type", "etag", "last-modified")}
        with self.conn:
            self.conn.execute(
                "INSERT OR REPLACE INTO responses VALUES (?, ?, ?, ?, ?, ?, ?)",
                (key, response.status_code, json.dumps(headers), body,
                 response.headers.get("etag"), time.time(), len(body)),
            )
        self.prune()

    def touch(self, key: str) -> None:
        with self.conn:
            self.conn.execute("UPDATE responses SET fetched_at = ? WHERE key = ?",
                              (time.time(), key))

    def prune(self) -> None:
        """Evict the oldest entries until the cache fits in max_bytes."""
        total = self.conn.execute("SELECT COALESCE(SUM(size), 0) FROM responses").fetchone()[0]
        if total <= self.max_bytes:
            return
        with self.conn:
            for key, size in self.conn.execute(
                    "SELECT key, size FROM responses ORDER BY fetched_at").fetchall():
                self.conn.execute("DELETE FROM responses WHERE key = ?", (key,))
                total -= size
                if total <= self.max_bytes:
                    break

    def clear(self) -> int:
        with self.conn:
            return self.conn.execute("DELETE FROM responses").rowcount


def cached_get(client: httpx.Client, cache: HttpCache, url: str, *, ttl: float = 300,
               refresh: bool = False, use_cache: bool = True) -> CachedResponse:
    key = str(client.build_request("GET", url).url)          # normalised absolute URL
    if not use_cache:
        r = client.get(url)
        r.raise_for_status()
        return CachedResponse(r.status_code, dict(r.headers), r.content, 0.0)
    stored = cache.get(key)
    if stored and stored.age < ttl and not refresh:
        return stored                                         # 1. fresh hit
    headers = {}
    if stored and (tag := cache.etag(key)):
        headers["If-None-Match"] = tag                        # 2. revalidate
    try:
        r = client.get(url, headers=headers)
    except httpx.TransportError:
        if stored:
            stored.stale = True                               # 4. offline fallback
            return stored
        raise
    if r.status_code == 304 and stored:
        cache.touch(key)
        stored.age = 0.0
        return stored
    r.raise_for_status()
    if "no-store" not in r.headers.get("cache-control", ""):
        cache.put(key, r)                                     # 3. store the miss
    return CachedResponse(r.status_code, dict(r.headers), r.content, 0.0)

The cache key is the fully built URL, including query parameters, so ?page=2 and ?page=3 are separate entries. Only successful responses are stored, because raise_for_status() runs first — caching a 500 would make an outage last as long as the TTL. Responses marked Cache-Control: no-store are never written, which respects servers that send sensitive data. And a corrupt cache file is deleted and recreated rather than reported, because nothing in a cache is worth an error.

The cache from the user’s side Terminal session showing a cached ticket lookup, a forced refresh and an offline fallback with a staleness warning. The cache from the user’s side bash $ time tickets show 42 Fix login redirect real 0.31s $ time tickets show 42 Fix login redirect real 0.06s $ tickets --refresh show 42 # on a train warning: offline, showing data from 14 minutes ago Fix login redirect The second call never left the machine; the third admits its data is old.

Credentials and the cache key

If different users or profiles can see different data at the same URL, the URL alone is not a safe key — one profile would be served another's cached answer. Include the profile or account name in the key (f"{profile}|{url}"), or keep one cache file per profile. Never include the token itself; the key is stored in plain text.

Wiring it into commands

# src/tickets/cli.py
import httpx
import typer
from platformdirs import user_cache_path

from tickets.httpcache import HttpCache, cached_get

app = typer.Typer()
API = "https://api.example.com"


@app.callback()
def main(ctx: typer.Context,
         refresh: bool = typer.Option(False, "--refresh", help="Revalidate cached data."),
         no_cache: bool = typer.Option(False, "--no-cache", help="Bypass the cache.")) -> None:
    """Ticket tool."""
    ctx.obj = {"refresh": refresh, "use_cache": not no_cache,
               "cache": HttpCache(user_cache_path("tickets") / "http.db"),
               "client": httpx.Client(base_url=API, timeout=10)}


@app.command()
def show(ctx: typer.Context, ticket: int) -> None:
    """Show one ticket."""
    o = ctx.obj
    r = cached_get(o["client"], o["cache"], f"/tickets/{ticket}", ttl=120,
                   refresh=o["refresh"], use_cache=o["use_cache"])
    if r.stale:
        typer.echo(f"warning: offline, showing data from {int(r.age // 60)} minutes ago",
                   err=True)
    typer.echo(r.json()["title"])


cache_app = typer.Typer(help="Manage the local HTTP cache.")
app.add_typer(cache_app, name="cache")


@cache_app.command("clear")
def clear(ctx: typer.Context) -> None:
    typer.echo(f"removed {ctx.obj['cache'].clear()} cached responses")

UX considerations

  • Choose TTLs per endpoint. A list of projects can be cached for an hour; a build status for ten seconds; anything the user just changed should not be cached at all. A single global TTL is always wrong for some endpoint.
  • Invalidate after writes. After mytool close 42, delete the cached entry for ticket 42, or the next show will contradict the action the user just performed.
  • Say when data is stale. The offline fallback must announce itself on stderr. Silent stale data is worse than an error.
  • Offer --refresh and --no-cache. Refresh revalidates (cheap with ETags); no-cache bypasses the store entirely, which is what users reach for when debugging.
  • Put it in the cache directory. The file may be deleted by the user or a cleaner at any time — that is the definition of a cache. Show its location and size in a cache info command.
  • Cache completion lookups aggressively. Shell completion runs on every Tab and must answer in tens of milliseconds; a long TTL for completion data is fine, as discussed in dynamic completion values from APIs and files.
Different data, different lifetimes Example time-to-live values in seconds for different kinds of API data cached by a command line tool. Different data, different lifetimes Project list 3600 s Completion candidates 900 s Ticket details 120 s Build status 10 s just-changed data: invalidate instead of waiting for expiry One global TTL is always wrong for some endpoint.

Testing the behaviour

respx counts calls, which makes "did this hit the network?" a direct assertion. Patch time.time to age entries instead of sleeping:

# tests/test_httpcache.py
import httpx
import pytest
import respx

from tickets import httpcache
from tickets.httpcache import HttpCache, cached_get

URL = "https://api.example.com/tickets/42"


@pytest.fixture
def setup(tmp_path):
    client = httpx.Client(base_url="https://api.example.com")
    return client, HttpCache(tmp_path / "http.db", max_bytes=1000)


@respx.mock
def test_fresh_hit_skips_network(setup):
    client, cache = setup
    route = respx.get(URL).respond(json={"title": "Fix login"})
    cached_get(client, cache, "/tickets/42")
    assert cached_get(client, cache, "/tickets/42").json()["title"] == "Fix login"
    assert route.call_count == 1


@respx.mock
def test_stale_entry_revalidates_with_etag(setup, monkeypatch):
    client, cache = setup
    route = respx.get(URL).mock(side_effect=[
        httpx.Response(200, json={"title": "A"}, headers={"ETag": '"v1"'}),
        httpx.Response(304),
    ])
    cached_get(client, cache, "/tickets/42", ttl=60)
    real = httpcache.time.time
    monkeypatch.setattr(httpcache.time, "time", lambda: real() + 3600)
    r = cached_get(client, cache, "/tickets/42", ttl=60)
    assert r.json()["title"] == "A"
    assert route.calls[1].request.headers["If-None-Match"] == '"v1"'


@respx.mock
def test_offline_falls_back_to_stale(setup):
    client, cache = setup
    respx.get(URL).mock(side_effect=[httpx.Response(200, json={"title": "A"}),
                                     httpx.ConnectError("offline")])
    cached_get(client, cache, "/tickets/42")
    r = cached_get(client, cache, "/tickets/42", refresh=True)
    assert r.stale and r.json()["title"] == "A"


@respx.mock
def test_errors_are_not_cached(setup):
    client, cache = setup
    respx.get(URL).respond(500)
    with pytest.raises(httpx.HTTPStatusError):
        cached_get(client, cache, "/tickets/42")
    assert cache.get(URL) is None


@respx.mock
def test_size_limit_evicts_oldest(setup):
    client, cache = setup
    for n in range(5):
        respx.get(f"https://api.example.com/t/{n}").respond(content=b"x" * 300)
        cached_get(client, cache, f"/t/{n}")
    assert cache.get("https://api.example.com/t/0") is None
    assert cache.get("https://api.example.com/t/4") is not None


def test_corrupt_file_is_replaced(tmp_path):
    path = tmp_path / "http.db"
    path.write_bytes(b"this is not sqlite")
    assert HttpCache(path).get("anything") is None

Each test pins one row of the decision diagram, so a refactor that breaks revalidation or starts caching errors fails loudly. The general respx techniques are covered in mocking HTTP in CLI tests with respx.

Conclusion

An HTTP cache makes a CLI faster, kinder to rate limits and usable offline — provided it is honest about freshness. Store successful GET responses in SQLite keyed by URL (and profile), serve them within a per-endpoint TTL, revalidate with ETag, fall back to clearly labelled stale data when the network is gone, cap the size, and give users --refresh, --no-cache and cache clear. Test each path with respx call counts rather than timing.

Frequently asked questions

Should I use hishel or requests-cache instead?

Both are solid: hishel adds RFC 9111 caching to httpx, requests-cache does the same for requests, and both can store in SQLite. Use one if you want standards-complete HTTP caching semantics. The hand-written version here is about 100 lines, has no extra dependency, and makes the CLI-specific policies — per-endpoint TTLs, stale fallback, profile-aware keys — explicit.

Is it safe to cache responses that contain personal data?

Only with care. The cache file sits unencrypted in the user's cache directory, readable by the user's account. Respect Cache-Control: no-store, keep TTLs short for sensitive endpoints, create the file with owner-only permissions, and offer cache clear. Never cache responses from authentication endpoints.

How do I cache POST requests that are really queries?

Some APIs (search, GraphQL) use POST for reads. Cache them only if you are certain they have no side effects, and include a hash of the request body in the key. Make it opt-in per call rather than automatic.

Why not let the server's Cache-Control headers set the TTL?

You can: parse max-age and use it when present, falling back to your per-endpoint default. Many internal APIs send no caching headers at all, which is why the explicit TTL is the primary control here.