Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Introduction

acme-proxy is an ACME server: the thing certbot, acme.sh, lego, Traefik and Caddy talk to when they ask for a certificate. It serves the whole RFC 8555 flow — account, order, authorization, challenge, finalize, certificate — plus revocation, account key rollover (§7.3.5) and Renewal Information (RFC 9773). What it deliberately does not implement is listed on Protocol Support.

What it does behind that interface is the point. Clients prove control of their names to acme-proxy, under whatever policy you configure, and it decides how the certificate is actually produced — signing locally, relaying to a public CA, or handing the request to a script.

The problem

Internal servers and IoT devices need TLS certificates, and cannot easily get them:

  1. Public CAs (like Let’s Encrypt) require DNS or HTTP validation, which internal servers hidden behind firewalls cannot easily satisfy.
  2. Handing out the company’s global DNS API credentials to every internal server to perform DNS-01 challenges is a massive security risk.
  3. Legacy internal PKI systems often do not speak the ACME protocol, forcing operators to write custom bash scripts to rotate certificates.
  4. When using external or commercial CAs, organizations are sometimes restricted to a single account or a limited number of External Account Binding (EAB) credentials validated for specific domains. Distributing these scarce upstream credentials directly to hundreds of internal servers is both impractical and risky.

The solution

acme-proxy solves this by standing between your internal clients and the actual Certificate Authority. By registering a single account with the upstream CA, it acts as an account multiplexer—allowing you to issue unlimited local EAB credentials to your internal teams without exhausting your upstream limits.

graph LR
    subgraph clients["Your network"]
        A["Internal server<br/>certbot"]
        B["IoT device<br/>acme.sh"]
        C["Traefik / Caddy"]
    end

    P{{"acme-proxy"}}

    subgraph backends["One of these actually signs"]
        LOCAL["Embedded CA<br/>key on disk or in an HSM"]
        UP["Upstream ACME CA<br/>Let's Encrypt, ZeroSSL, commercial"]
        SCRIPT["Your script<br/>legacy PKI, internal API"]
    end

    A --> P
    B --> P
    C --> P
    P --> LOCAL
    P --> UP
    P --> SCRIPT

Clients speak ordinary ACME to acme-proxy and prove control of their names to it. Which of the three backends produces the certificate is a configuration choice they never see — and it can differ per endpoint, so one process can serve a local CA at /profile/dev and a Let’s Encrypt relay at /profile/prod.

  • As a Proxy: It intercepts ACME requests from internal clients, opens a corresponding order with Let’s Encrypt, and safely solves the external DNS-01 challenges on the client’s behalf using a single, centrally secured DNS TSIG key.
  • As a Filter: It inspects the client’s IP and requested DNS names against strict policies — including asking your IPAM (NetBox or phpIPAM) whether that address owns those names — before allowing the request to proceed.
  • As a Local CA: It can operate entirely offline, issuing from an embedded ECDSA CA whose key is a file on disk or a PKCS#11 token.
  • As a Multi-Tenant Server: Through its Profile system, a single binary can host a Local CA on /profile/dev and a strictly-filtered Let’s Encrypt relay on /profile/prod.
  • As a Record: Every issuance and every refusal is written to an append-only audit trail, naming the actor, the address it came from and the names it asked for — the question “who got this certificate, and from where” has an answer.

It is built on axum, tokio and sqlx, and stores everything in one SQLite file in WAL mode. There is no second datastore, no message queue and no scheduler to operate.

It is published on crates.io, so cargo install acme-proxy gets you the whole thing — server and admin CLI in one binary. See Installation.

Core Concepts & Glossary

Eight words carry most of the meaning in the rest of this book. This page defines them once, in the order you meet them, so that every other page can use them without re-explaining.

Profile

A profile is an independent ACME endpoint, and the isolation boundary everything else sits inside. Rather than running one process per environment, you define several profiles in one; each is served at /profile/<name>/directory.

Accounts and orders are isolated per profile — the same client key at two profiles is two unrelated accounts — and each profile carries its own signer, filters, challenge validation and EAB policy. A dev profile backed by a local CA can sit beside a prod profile relaying to Let’s Encrypt under strict NetBox filtering, in one process, over one socket and one database.

See Profiles & Routing.

Signer

A signer is what actually produces the certificate once a client has been authorized. Which one runs is a per-profile configuration choice, and the client never sees the difference.

  • Local CA — an embedded certificate authority signing directly. The issuing key is a file, or a PKCS#11 token.
  • Relay — opens its own order with an upstream ACME CA and returns what that CA signs.
  • Custom script — anything else: a legacy PKI, an internal API, a CA that does not speak ACME.

See Signers.

Filter

A filter is a policy applied to a request before anything is signed. Filters answer “may this client ask for this?”, which challenge validation does not: a client can genuinely control a name and still have no business holding a certificate for it from you.

[filter] is a small policy engine rather than a list of switches. A check is one named question about a request — “is this address in the management network?” — and a rule is a boolean expression over check names plus what a match means. filter.rules says which rules run and in what order, and the first match wins; a stage where a rule was applicable and none matched falls to filter.default.

Rules act at two points — on the connection, and on the identifiers, the latter running again at finalize against the names in the CSR.

See Filters.

EAB (External Account Binding)

External Account Binding (RFC 8555 §7.3.4) makes newAccount require a credential you minted out of band — a key identifier and an HMAC secret. Reaching the directory is then no longer enough to register: an operator has to have issued that client a credential first.

It runs in the other direction too. A commercial CA that granted you one scarce EAB credential is exactly the case the relay backend exists for: one upstream credential, any number of local ones.

See External Account Binding.

ARI (ACME Renewal Information)

ACME Renewal Information (RFC 9773) lets the CA tell a client when to renew, rather than leaving it to guess from the expiry date. Two things follow: a fleet spreads its renewals across a window instead of stampeding at the same moment, and a CA that needs certificates replaced early can say so and be listened to.

See Renewal Information.

Order

An order is a client’s request for a certificate. It names the identifiers wanted and progresses through the states RFC 8555 defines:

  • pending — created; one or more authorizations still need to be satisfied.
  • ready — every authorization is valid; the client may now finalize.
  • processing — issuance is under way but not finished.
  • valid — the certificate is available.
  • invalid — terminal failure.
stateDiagram-v2
    [*] --> pending: newOrder
    pending --> ready: every authorization valid
    ready --> pending: an authorization is deactivated (§7.5.2)
    ready --> processing: finalize
    processing --> valid: the worker signed
    processing --> invalid: signer refused or gave up
    pending --> invalid: an authorization failed, or expires passed
    valid --> [*]
    invalid --> [*]

acme-proxy enforces these states and transitions in the database itself. The set of states is a CHECK constraint on the status column — see Database Schema — and every transition is an UPDATE guarded on the state it leaves. A validation or a signing that finishes after the order moved on, because the client deactivated an authorization or a sibling challenge already decided it, therefore changes nothing above its own challenge.

ready → pending is the one backwards edge, and it exists only so §7.5.2 can hold: deactivating an authorization on an order that already reached ready has to demote it, or the order would be finalizable for a name no longer authorized.

Two details are easy to trip on:

  • Every finalize answers processing. Signing needs the CA key, which only the worker role holds, so finalize checks the CSR, claims the order and queues the signing, and the client polls until the order is valid. With local_ca that is a moment; with relay it is as long as the upstream CA takes. A CSR the backend itself rejects makes the order invalid with a badCSR error, since the client is already polling by then; a CSR finalize can refuse on its own leaves the order ready for a corrected one.
  • Revocation is orthogonal to this machine. RFC 8555 defines no “revoked” order status, so a revoked order’s status stays valid. The revocation timestamp and reason are recorded separately, and both admin front ends show them and can revoke — acme-proxy order show/order revoke, and the order detail page in the panel. See Revocation & CRL.

Job

A job is one unit of work the server owes itself: a relayed issuance to finish, a notification to deliver, a table to sweep. Jobs are rows in the same SQLite file as everything else, drained by one runner per process, so they survive a restart and need no scheduler beside the server.

What is worth carrying away is how a handler reports failure. Retry says the attempt decided nothing — a refused connection, a proxy, a 503 — and the job goes back in the queue under a growing backoff; Failed says the other side stated a reason and is believed at once. That split is what keeps a client’s order processing through a five-second upstream blip rather than terminally invalid, and it is why an order that is not progressing is a question for acme-proxy jobs list before it is a question for anything else.

See Admin CLI → Job queue.

Challenge

A challenge is the concrete proof that a client controls an identifier: serving a token over HTTP, publishing a DNS TXT record, or presenting a special certificate in a TLS handshake.

Each authorization carries one challenge per enabled type, and satisfying any one of them makes the authorization valid — the others stay pending for ever, which is correct rather than a stuck state.

Triggering one queues the check rather than performing it, so a triggered challenge answers processing and the client polls it.

See Challenge Validation.

Quick Start

Get acme-proxy running in a couple of minutes with an auto-generated Local CA.

Step 1 — Write a configuration file

acme-proxy serves ACME only through profiles, so at least one enabled profile is required — the server refuses to start without one. Everything else has a working default.

By default the server also performs real domain-control validation (HTTP-01). To test a client against the proxy locally without routing port 80 or configuring DNS, this quick start turns validation off with challenge.bypass.

Create config.toml in your current directory:

[challenge]
# Testing only. See "Moving to production" below.
bypass = true

[profiles.default]

[profiles.default] is not empty by accident — a profile’s enabled key defaults to true, so naming the profile is all that is required. The profile name becomes part of the URL, and of every kid a client stores.

Step 2 — Run the server

acme-proxy serve

serve is also the default, so a bare acme-proxy does the same thing.

There is no --config flag. The server reads config.toml from the current working directory; point it elsewhere with ACME_PROXY_CONFIG:

ACME_PROXY_CONFIG=/etc/acme-proxy/config.toml acme-proxy serve

(The extension may be omitted — the format is then inferred.) A missing configuration file is not an error: the server falls back to defaults, which is why an environment-only deployment works.

Individual keys can also be overridden with ACME_PROXY_* environment variables; see the Configuration Reference.

The server binds [::]:3000 and serves the ACME directory at http://localhost:3000/profile/default/directory. On first run it creates sqlite.db, ca.pem and ca.key in the working directory.

Step 3 — Request a certificate

Point any standard ACME client at the profile’s directory URL. Examples for internal.example.com:

certbot

certbot certonly \
  --server http://localhost:3000/profile/default/directory \
  --standalone \
  --domain internal.example.com \
  --email admin@example.com \
  --agree-tos \
  --no-eff-email

acme.sh

acme.sh --issue \
  --server http://localhost:3000/profile/default/directory \
  -d internal.example.com \
  --standalone

lego

lego --server http://localhost:3000/profile/default/directory \
  --email admin@example.com \
  --domains internal.example.com \
  --http \
  run

Traefik

In Traefik’s static configuration (traefik.yml):

certificatesResolvers:
  myresolver:
    acme:
      caServer: http://localhost:3000/profile/default/directory
      email: admin@example.com
      httpChallenge:
        entryPoint: web

Caddy

In your Caddyfile:

internal.example.com {
    tls admin@example.com {
        ca http://localhost:3000/profile/default/directory
    }
    respond "Hello, world!"
}

What just happened?

  1. The client fetched the directory from acme-proxy.
  2. It registered an account and created a new order.
  3. acme-proxy offered an HTTP-01 challenge.
  4. Because challenge.bypass = true, the proxy marked the challenge valid the moment the client triggered it, without making any network request back to the client’s responder.
  5. The proxy signed the CSR with its auto-generated local ECDSA CA and returned the certificate.

The certificate is signed by a CA nothing trusts yet. See Trusting the CA for how to install ca.pem where your clients will accept it.

Step 4 — Moving to production

The defaults are safe; this quick start deliberately relaxed them. Before exposing the server:

  1. Remove the bypass. Delete challenge.bypass = true so acme-proxy actually validates domain control. With bypass on, [filter] is the only access control there is — which is exactly why bypass is not the default. See Challenge Validation.
  2. Enable filters. Configure allowed_ip, identifiers or ipam to restrict which clients may request which names. See Filters & Policies.
  3. Choose where state lives. The defaults write sqlite.db and the CA material to the current working directory. Set database.url and signer.local_ca.cert_path / signer.local_ca.key_path to permanent paths — note these are the CA’s files, distinct from server.tls.cert_path / server.tls.key_path, which belong to the HTTPS listener.
  4. Serve over HTTPS. RFC 8555 §6.1 expects ACME over HTTPS: either put the server behind a reverse proxy, or turn on TLS termination.

Installation

acme-proxy is a Rust application, published on crates.io. There are four ways to get it: cargo install, a source build, the published container image, or a container image you build yourself. Only the published image comes prebuilt; no standalone prebuilt binaries are published.

The result is a single binary carrying both the server and the admin CLI, so a deployment never needs a second tool.

Prerequisites

  • Rust toolchain: the crate is edition 2024 and declares a rust-version in Cargo.toml (currently 1.97). That file is the source of truth; cargo refuses to build with anything older.
  • Cargo: the Rust package manager.

SQLite is not a prerequisite: the driver is bundled with sqlx, the database file is created automatically, and DATABASE_URL is not needed to compile. The PostgreSQL driver is bundled the same way, so the same binary speaks both — what it does need is a server and a database that already exist. A sqlite3 binary is only useful if you want to inspect the database by hand.

You can install Rust via rustup:

curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh

From crates.io

cargo install acme-proxy

This fetches the published crate, compiles it, and puts the binary in ~/.cargo/bin — which needs to be on your PATH. Confirm it with:

acme-proxy --version

Tab completion and a man page come out of the binary itself, so there is nothing extra to download:

acme-proxy completions zsh > ~/.zfunc/_acme-proxy
acme-proxy man | sudo tee /usr/share/man/man1/acme-proxy.1 > /dev/null

Both are generated from the command tree, so regenerate them when you upgrade. The per-shell paths are in the admin CLI chapter.

To pin a version, or to move to a specific one later, name it:

cargo install acme-proxy --version 0.6.0

Upgrading is cargo install acme-proxy --force. Before doing so across a minor version, read the ### Breaking section of the changelog. Before 1.0.0 the database schema is the only compatibility guarantee, so the data survives an upgrade but a configuration key may have been renamed.

Building from source

Prefer this if you intend to change anything, or want the test suite and the book sources alongside the binary.

  1. Clone the repository:

    git clone https://github.com/acme-proxy/acme-proxy.git
    cd acme-proxy
    
  2. Build the project in release mode for production use:

    cargo build --release
    

The binary will be located at target/release/acme-proxy.

Optional features

The default build has no optional features. One is available:

FeatureWhat it adds
hsmPKCS#11 support for the Local CA’s issuing key, so it can live in a YubiKey or an HSM instead of a file — see Hardware Keys.
cargo install acme-proxy --features hsm    # or, from a clone:
cargo build --release --features hsm

It is off by default because it pulls in cryptoki and its bindings, which a deployment signing with an on-disk key does not need. The PKCS#11 module itself is loaded at runtime, so enabling this adds no build-time C toolchain requirement. Configuring signer.local_ca.key_source = "pkcs11" on a binary built without it is a startup error naming the feature, never a silent fallback to the file key.

The published container image is the default build, without hsm. A deployment that needs it builds its own binary.

Container (Docker / Podman)

Every release is published to the GitHub Container Registry as a multi-architecture image, for linux/amd64 and linux/arm64:

podman pull ghcr.io/acme-proxy/acme-proxy:0.6.0
TagPoints at
0.6.0That release. It never moves.
0.6The newest 0.6.x release: its fixes, never a new minor.
latestThe highest release, whatever its minor.
edgeThe head of main, rebuilt on every merge. Not a release.
sha-…One commit of main, as edge was when it was built.

Run a version tag, or 0.6 to take patch releases without a change on your side. Avoid latest for an unattended deployment. Before 1.0.0 a minor release may rename a configuration key, so a pull of latest can stop a server from starting; the ### Breaking sections of the changelog list every such change. edge is for trying what the next release will hold, never for a certificate authority anyone depends on.

Every image holds the release build with the default features: the binary cargo install acme-proxy produces.

To build the image yourself instead, use the Containerfile in a clone of the repository:

podman build -t acme-proxy .

That is the same release build as the published image, fat LTO included, so it takes tens of minutes.

The image’s working directory is /data and its entrypoint is the acme-proxy binary, so mount a volume there for the SQLite database, the configuration and the CA key material — all of which default to paths relative to the working directory. The image runs as a non-root user, so the mounted directory must be writable by it — the :U flag below is the rootless-Podman way; see Deployment for Docker.

podman run -d \
  -p 3000:3000 \
  -v ./data:/data:U \
  ghcr.io/acme-proxy/acme-proxy:0.6.0

Drop a config.toml into ./data (it must define at least one profile — see the Quick Start), or configure the container entirely through ACME_PROXY_* environment variables:

podman run -d \
  -p 3000:3000 \
  -v ./data:/data:U \
  -e ACME_PROXY_PROFILES__DEFAULT__ENABLED=true \
  -e ACME_PROXY_SERVER__BASE_URL=https://acme.example.com \
  ghcr.io/acme-proxy/acme-proxy:0.6.0

Verifying the image

Each published image carries a signed build provenance attestation. It records the repository, the commit and the workflow run that built the image. Check it with the GitHub CLI before you run the image:

gh attestation verify oci://ghcr.io/acme-proxy/acme-proxy:0.6.0 \
  --repo acme-proxy/acme-proxy

The attestation is on the multi-architecture index, the object a tag resolves to. The digest of one architecture’s image, on its own, has no attestation to find.

Trusting the CA

With the local_ca signer, acme-proxy mints certificates from a CA it generated itself. Those certificates are perfectly valid, but nothing on your network trusts the CA that signed them yet — so browsers, curl, and every TLS library will reject them until you install the root.

This page covers distributing that root. It does not apply when you use the relay backend to relay to a public CA, whose roots are already trusted everywhere.

Getting the root certificate

Each profile serves its own CA material unauthenticated at {base_url}/profile/<name>/ca.pem, as application/x-pem-file:

curl -o internal-root.pem https://acme.internal/profile/default/ca.pem

These are exactly the bytes appended to every certificate that profile issues, so a client that fetches them here and one that reads the tail of its own chain end up trusting the same anchor.

Two things worth knowing about the route. It is not advertised in the ACME directory — it is CA infrastructure rather than an ACME resource, so a client will not find it on its own and you distribute the URL yourself. And it is served inside the profile router, which means it sits behind that profile’s filter policy: if you restrict the endpoint by address, the hosts that most need the root — the ones that do not have it yet — may be exactly the ones refused. Add a path check allowing /ca.pem if so.

The same file is on the server’s disk at signer.local_ca.cert_path, ca.pem by default in the working directory, which is the way to get it when the server is not reachable or is not running:

scp acme-host:/var/lib/acme-proxy/ca.pem ./internal-root.pem

Inspect it before distributing it:

openssl x509 -in internal-root.pem -noout -subject -issuer -dates -ext basicConstraints

A freshly generated root is self-signed (subject equals issuer) and carries CA:TRUE, pathlen:0.

If cert_path holds a bundle — an intermediate followed by a root, as in the multi-tier setup — then the last certificate in the file is the root, and it is the only one your clients need to trust. The intermediate is shipped with every issued certificate and does not need installing.

# Split a bundle into its constituent certificates.
csplit -z -f cert- -b '%02d.pem' ca_bundle.pem '/-----BEGIN CERTIFICATE-----/' '{*}'

Installing it

Debian / Ubuntu

The file must have a .crt extension, and must be PEM despite the name.

sudo cp internal-root.pem /usr/local/share/ca-certificates/acme-proxy-root.crt
sudo update-ca-certificates

RHEL / Fedora / CentOS

sudo cp internal-root.pem /etc/pki/ca-trust/source/anchors/acme-proxy-root.pem
sudo update-ca-trust extract

Alpine

sudo cp internal-root.pem /usr/local/share/ca-certificates/acme-proxy-root.crt
sudo update-ca-certificates

Verify

curl -v https://internal.example.com 2>&1 | grep -i 'SSL certificate verify'
# or, without a server:
openssl verify -CAfile internal-root.pem issued-cert.pem

Applications with their own trust store

Updating the system store is not enough for everything. These maintain their own:

RuntimeHow to add the root
FirefoxIts own store, always. Settings → Privacy & Security → Certificates → View Certificates → Authorities → Import. Enterprise deployments can use the Certificates policy in policies.json.
Chrome / EdgeUses the system store on Windows and macOS; on Linux it reads the NSS database — certutil -d sql:$HOME/.pki/nssdb -A -t "C,," -n acme-proxy-root -i internal-root.pem.
Java / JVMkeytool -importcert -trustcacerts -alias acme-proxy-root -file internal-root.pem -keystore "$JAVA_HOME/lib/security/cacerts".
Node.jsIgnores the system store by default. Set NODE_EXTRA_CA_CERTS=/path/to/internal-root.pem.
Python requestsUses certifi, not the system store. Set REQUESTS_CA_BUNDLE (or SSL_CERT_FILE for ssl/urllib).
GoUses the system store on Linux; no action needed after update-ca-certificates.
ContainersEach image has its own store. Mount the root in and run the distribution’s update command in your Dockerfile, or bake it into a base image.

Distributing at scale

Installing a root by hand does not survive a fleet. In practice:

  • Ansible / Puppet / Chef — ship the file and run the update command as a handler. This is the common approach for Linux estates.
  • Active Directory Group Policy — Computer Configuration → Windows Settings → Security Settings → Public Key Policies → Trusted Root Certification Authorities.
  • MDM (Jamf, Intune, …) — deploy as a certificate payload.
  • Golden images — bake the root into your base image so new hosts trust it from first boot.

Whichever you use, deploy the root before you start issuing certificates from it, or the first clients to renew will break.

Revocation

If you revoke certificates, clients need to be able to see the CRL. It is served unauthenticated at {base_url}/profile/<name>/crl as application/pkix-crl:

curl -o internal.crl https://acme.internal/profile/default/crl
openssl crl -in internal.crl -inform DER -noout -text

Like /ca.pem, the CRL is not advertised in the ACME directory and sits behind the profile’s filter policy. Issued certificates carry a CRL distribution point only when you set signer.local_ca.crl_distribution_points; leave it unset and a client will not find the CRL automatically, so distribute the URL alongside the root if your validation policy needs it. See Revocation & CRL.

Planning ahead

The root’s validity is finite, and replacing it later means touching every host that trusts it. Two things make that easier:

  • Use an intermediate. Keep an offline root and hand acme-proxy only an intermediate. The root you distribute then long outlives any single signing key, and a compromised proxy costs you an intermediate rather than your whole trust anchor. See Multi-Tier PKI.
  • Distribute early, rotate overlapping. Trust stores accept multiple roots, so push a replacement root well before it is needed and remove the old one only after nothing is signed by it.

Deployment

While acme-proxy can be run in a container, many organizations prefer running infrastructure components directly on standard Linux VMs using systemd.

systemd service setup

Below is an example systemd service file that runs acme-proxy securely.

  1. Create a dedicated user:

    sudo useradd -r -s /bin/false acme-proxy
    
  2. Prepare directories:

    sudo mkdir -p /etc/acme-proxy
    sudo mkdir -p /var/lib/acme-proxy
    sudo chown acme-proxy:acme-proxy /var/lib/acme-proxy
    
  3. Create the service file: Create /etc/systemd/system/acme-proxy.service:

    [Unit]
    Description=ACME Proxy Server
    After=network.target
    
    [Service]
    Type=simple
    User=acme-proxy
    Group=acme-proxy
    ExecStart=/usr/local/bin/acme-proxy serve
    ExecReload=/bin/kill -HUP $MAINPID
    WorkingDirectory=/var/lib/acme-proxy
    
    # Configuration. The extension may be omitted, in which case the format
    # is inferred.
    Environment="ACME_PROXY_CONFIG=/etc/acme-proxy/config.toml"
    Environment="ACME_PROXY_DATABASE__URL=sqlite:///var/lib/acme-proxy/acme.db"
    
    # Security / Sandboxing
    ProtectSystem=strict
    ReadWritePaths=/var/lib/acme-proxy
    ProtectHome=true
    PrivateTmp=true
    NoNewPrivileges=true
    
    Restart=on-failure
    RestartSec=5
    
    [Install]
    WantedBy=multi-user.target
    

ProtectSystem=strict makes the whole filesystem read-only except ReadWritePaths, so everything the server writes must land in /var/lib/acme-proxy. With WorkingDirectory set there, the defaults already do: signer.local_ca.cert_path (ca.pem), key_path (ca.key), crl_path (ca.crl) and the lock beside it are all resolved relative to the working directory, as are server.tls.cert_path / key_path if you enable TLS. If you set any of them to an absolute path, add that path to ReadWritePaths too.

acme-proxy shuts down gracefully on SIGTERM (and on Ctrl+C when run in a terminal), so systemctl restart and systemctl stop let in-flight requests finish rather than cutting them off. Both listeners stop together. It also reloads its configuration on SIGHUP without restarting, which is what the ExecReload line above wires up — see Reloading the Configuration for what a reload may change and what it refuses. One case still deserves a quiet period: a request that waits — a custom script’s crl or renewal_info hook, or a relay revocation waiting on its job — can take up to server.request_timeout_ms. If systemd’s TimeoutStopSec (90 s by default) is shorter than that, systemd sends SIGKILL first and the graceful path is skipped — raise it, or lower the request timeout. Challenge validation and issuance are not among these: they run in the job queue, and a job left unfinished by a restart is reclaimed by lease expiry.

  1. Enable and start the service:
    sudo systemctl daemon-reload
    sudo systemctl enable --now acme-proxy
    

Where each socket belongs

acme-proxy opens two listeners when the web admin is enabled, and they belong on different sides of your boundary. The ACME listener answers unauthenticated clients by design; the admin listener has no filter chain and no admission control, and its only access controls are the bind address, TLS and the session.

graph TD
    subgraph internal["Internal network"]
        CLIENTS["ACME clients<br/>certbot, acme.sh, Traefik"]
        OPS["Operator workstation"]
    end

    subgraph host["The acme-proxy host"]
        RP["Reverse proxy (optional)<br/>sets X-Forwarded-For"]
        ACME[":3000 — ACME listener<br/>filters, admission, nonces"]
        ADMIN[":3001 — admin listener<br/>loopback by default"]
        DB[("sqlite.db + WAL")]
        CAKEY[["ca.key — 0600, or a PKCS#11 token"]]
    end

    CLIENTS --> RP --> ACME
    OPS -.->|"SSH tunnel or VPN,<br/>NOT an open port"| ADMIN
    ACME --> DB
    ADMIN --> DB
    ACME --> CAKEY

    ACME -->|"challenge validation:<br/>back to the client, :80 / :443 / DNS"| CLIENTS
    ACME -->|"upstream ACME, DNS updates, SMTP"| OUT(["Egress"])

Two edges are the ones people get wrong:

  • The dotted one. If the admin listener is reachable from anywhere but loopback, startup refuses unless admin.tls.enabled is on — and even then, a tunnel is the better answer. See below.
  • The validation edge points back at the client. With challenge.bypass = false, the server opens connections to the machines asking for certificates. A firewall that only permits inbound traffic leaves orders sitting at pending.

Reverse proxy (optional)

acme-proxy acts as an HTTP server, typically binding to port 3000. You can bind it directly to 80 (requires CAP_NET_BIND_SERVICE) or place it behind a reverse proxy like Nginx or Traefik, which can provide TLS termination for the ACME API itself.

Two things to get right when proxying:

  • Set server.base_url to the public URL. It is what the directory advertises and what every signed request is checked against (RFC 8555 §6.4), so a mismatch rejects every client. It is never derived from the request.
  • If you want IP-based filters to see the real client rather than the proxy, set filter.trusted_proxies to the proxy’s addresses and, if it is not x-forwarded-for, filter.forwarded_header. Note these are [filter] keys, not [server] keys.

Alternatively, skip the reverse proxy and let acme-proxy terminate TLS itself — see TLS Termination.


Exposing the web admin (or rather, not)

The Web Admin is a second listener and is off by default. When you turn it on, it binds 127.0.0.1:3001 and stays there unless you say otherwise.

The recommended way to reach it is an SSH tunnel, which needs no configuration change and no second certificate:

$ ssh -N -L 3001:127.0.0.1:3001 ca.example.com

Then open http://localhost:3001. admin.base_url stays at its default, because from the browser’s point of view the panel really is on localhost.

If you must bind it to a real interface, TLS is mandatory — startup refuses a non-loopback bind while admin.tls.enabled is false, because the session cookie is sent Secure and a browser silently declines to store one over plain HTTP anywhere but localhost:

[admin]
enabled      = true
bind_address = "0.0.0.0:3001"
base_url     = "https://admin.example.com:3001"

[admin.tls]
enabled = true

Two things this listener does not have, deliberately: admission control and a filter chain. Access control here is the bind address, TLS, and the session. Note also that it does not honour X-Forwarded-For — behind a reverse proxy the sign-in rate limiter counts the proxy, which is one more reason to prefer the tunnel.

Under systemd, nothing extra is needed: the panel shares the process, the unit and the database. Bootstrap the first operator once, before or after enabling it:

$ printf '%s' "$PASSWORD" | acme-proxy admin user create alice

Container deployments (Docker / Podman)

For containerized environments, you can run acme-proxy using Docker Compose or Podman.

Each release is published as ghcr.io/acme-proxy/acme-proxy:<version>, for linux/amd64 and linux/arm64, and the examples below pin one. Installation covers why to pin, building the image yourself from the Containerfile, and verifying its provenance:

podman pull ghcr.io/acme-proxy/acme-proxy:0.6.0

The image’s working directory is /data and its entrypoint is the binary itself, so /data is where the database, the CA key material and the CRL land unless you override their paths.

The image runs as a non-root user (acme-proxy, uid/gid 1000), so the directory you mount at /data must be writable by that uid. How you arrange that depends on the runtime:

  • Rootless Podman — add U to the mount flags (-v ./data:/data:U, or :Z,U on SELinux systems). Podman then chowns the volume’s contents to the user the container runs as. Alternatives: podman unshare chown -R 1000:1000 ./data beforehand, or use a named volume (-v acme-proxy-data:/data), which Podman initializes with the right owner.
  • Docker / Docker Compose — create the directory owned by uid 1000 before the first run: mkdir -p ./data && sudo chown 1000:1000 ./data. Or override the uid to your own (--user "$(id -u):$(id -g)", or a Compose user: line) and own ./data yourself — the server only needs to read and write that one directory.

Docker Compose

Create a docker-compose.yml file:

services:
  acme-proxy:
    image: ghcr.io/acme-proxy/acme-proxy:0.6.0
    container_name: acme-proxy
    restart: unless-stopped
    ports:
      - "3000:3000"
    volumes:
      - ./data:/data
    environment:
      - ACME_PROXY_PROFILES__DEFAULT__ENABLED=true
      - ACME_PROXY_DATABASE__URL=sqlite:///data/acme.db
      - RUST_LOG=acme_proxy=info

Create the data directory with the right owner before the first run:

mkdir -p ./data && sudo chown 1000:1000 ./data

Run the stack using:

docker compose up -d

ACME_PROXY_PROFILES__DEFAULT__ENABLED=true is what defines the profile when there is no configuration file — the server serves nothing without at least one. For anything beyond a single default profile, mount a config.toml into /data instead.

Podman (rootless)

Under rootless Podman you can run the container directly. The U flag chowns the mounted directory to the non-root user the container runs as; add Z as well on SELinux-enabled systems (RHEL/Fedora) for the mount label.

podman run -d --name acme-proxy \
  -p 3000:3000 \
  -v ./data:/data:U \
  -e ACME_PROXY_PROFILES__DEFAULT__ENABLED=true \
  -e ACME_PROXY_DATABASE__URL=sqlite:///data/acme.db \
  ghcr.io/acme-proxy/acme-proxy:0.6.0

Running the roles as separate processes

acme-proxy serve runs three jobs in one process: serving ACME, serving the web admin, and draining the job queue. --role splits them across processes of the same binary, reading the same configuration.

RoleDoesHolds
acmeServes ACME to certificate clientsThe ACME listener
adminServes /ui and /apiThe admin listener
workerDrains the job queue; signs and revokes; owns the schema and the first-run materialThe CA key or token, a relay’s upstream account; no listener

All-in-one is still the default and nothing about it changes: acme-proxy serve with no --role behaves exactly as it always did. The split is worth doing when you want privilege separation — the process parsing untrusted JWS and CSRs from the internet is then not the one holding operator sessions, and neither is the one making outbound connections to client-chosen hosts. Each can run under its own uid and its own systemd sandbox.

Only the worker holds signing material. finalize queues the signing and answers processing, a revocation is a database row or a queued job, and the CA’s certificate, its CRL and renewal information are served from ca.pem and the database. So the acme and admin processes never read ca.key, never log in to a PKCS#11 token and never use a relay’s upstream account. Make ca.key (0600) readable by the worker’s uid alone; the others need ca.pem and the database. A custom signer’s script must still be present where acme runs if it serves the CRL or renewal information, since those hooks answer a request.

Three things to get right:

  1. Initialise once, first. acme-proxy init migrates the database and generates the CA key, the upstream account and any self-signed TLS certificate. Run it as the uid that should own those files. A process that does not run worker refuses to start against a schema that is behind, naming acme-proxy migrate, and without the CA certificate, naming acme-proxy init — so starting the others before the schema or the CA exists fails loudly rather than racing, and never generates a second CA.
  2. Give each process its own metrics.bind_address. The counters are per-process memory, so three processes are three scrape targets; sharing one address means the second one to start fails to bind. Set ACME_PROXY_METRICS__BIND_ADDRESS per unit, point each at its own file with ACME_PROXY_CONFIG, or turn metrics.enabled off where you do not want it. Every series carries a role label naming the roles that process runs, so one scrape config can tell them apart.
  3. Run at least one worker. A process without it logs server_role_no_worker at startup; a deployment without one issues nothing, because challenge validation, issuance, CRL signing, relay and custom revocations, notifications and the periodic sweeps are all queued work. A worker in another process picks a row up within the job poll interval, so clients see a second or so more processing than all-in-one; a revocation for a relay or custom profile still running when its request’s deadline nears answers 503 with Retry-After, and asking again follows it.

A worked topology, one systemd unit per role:

# acme-proxy-worker.service
ExecStart=/usr/local/bin/acme-proxy serve --role worker
Environment=ACME_PROXY_METRICS__BIND_ADDRESS=127.0.0.1:3002

# acme-proxy-acme.service
ExecStart=/usr/local/bin/acme-proxy serve --role acme
Environment=ACME_PROXY_METRICS__BIND_ADDRESS=127.0.0.1:3012

# acme-proxy-admin.service
ExecStart=/usr/local/bin/acme-proxy serve --role admin
Environment=ACME_PROXY_METRICS__BIND_ADDRESS=127.0.0.1:3022

One host, one filesystem — on SQLite. SQLite across processes is fine on a local disk in WAL mode, and busy_timeout is already set — it is not safe on NFS or across nodes. Multi-node needs PostgreSQL: point database.url at a server instead of a file, run acme-proxy migrate once, and the three roles can then live on different hosts. Nothing else about the deployment changes.

An existing SQLite deployment moves across with its accounts, orders and audit trail intact — stop the server, create and migrate the target, then acme-proxy transfer --to <url>. Starting the new deployment empty instead would leave every certificate it has already issued impossible to revoke, so this is not an optional step.

Run one admin process either way; its login rate limiter is in memory, so two would each get their own budget.

Each process reloads independently on SIGHUP, so a configuration change means reloading all three.

Upgrading

Replace the binary, migrate, and restart. The schema is append-only as of 0.1.0 — a new release only ever adds migrations, never rewrites the ones your database has already applied.

Migrations no longer run as a side effect of opening the database. acme-proxy serve running the worker role (which the default does) still applies them at startup, so a single-process deployment can simply restart; anything else — an admin command, or a split deployment’s acme/admin process — checks the schema and refuses by name until acme-proxy migrate has run.

systemctl stop acme-proxy
install -m 0755 acme-proxy /usr/local/bin/acme-proxy
acme-proxy migrate          # explicit; the default `serve` would also do it
systemctl start acme-proxy
journalctl -u acme-proxy -n 50

A container is upgraded the same way, with the image tag in place of the binary: pull the new tag, migrate with it against the same /data, then recreate the container on it. With Docker Compose, after changing the image: line:

docker compose pull
docker compose stop acme-proxy
docker compose run --rm acme-proxy migrate
docker compose up -d
docker compose logs -n 50 acme-proxy

migrate after the service name replaces the image’s default serve command for that one container. A single-container deployment would also migrate as it starts; running it as a step of its own stops the upgrade on the error, instead of leaving a server in a restart loop. A split deployment must migrate before any of its acme or admin containers start on the new tag. The advice below applies unchanged, the database backup first of all.

Worth knowing before you do it:

  • Take a copy of the database first. SQLite in WAL mode means three files; copy them together with the server stopped, or use sqlite3 acme.db ".backup backup.db" on a running one. Migrations are not reversible, so a downgrade means restoring this copy.
  • db_migration_failed at startup means the process is not serving. The most likely cause is running an older binary against a database a newer one has already migrated.
  • Files outside the database are untouched. The CA key and certificate, the exported ca.crl, the upstream account key and its .kid sidecar all persist across an upgrade — back them up on the same schedule as the database, since the CA key is the one thing that cannot be regenerated without redistributing trust. See Trusting the CA.
  • Read the changelog’s Breaking section first. Before 1.0.0 the schema is the only compatibility guarantee: configuration keys, profile names, the admin JSON API, log event names and the CLI may all have moved, and every such change is listed there. See Compatibility. A renamed key is normally refused by name at startup — the server stops with an error naming the replacement rather than coming up looking configured — so acme-proxy filter show and a --help are cheap pre-restart checks.

TLS Termination

ACME strictly expects traffic over HTTPS (RFC 8555 §6.1).

While you can place acme-proxy behind a reverse proxy (like Nginx, HAProxy, or Traefik) and let it terminate TLS, acme-proxy is fully capable of terminating TLS itself via the highly secure rustls crate.

Architectural security (Slowloris protection)

Terminating TLS in async frameworks requires care to avoid denial-of-service vectors. In acme-proxy, handshakes run off the accept path. When a TCP connection arrives, a background task spawns the TLS handshake into a bounded channel. The handshake is strictly bounded by handshake_timeout_ms. If this were done inline, a single stalled client (e.g., a Slowloris attack) could block every other connection for the length of the timeout.

Client IP preservation

Under the hood, acme-proxy wraps the TLS listener in a TapIo struct. This is load-bearing. Without this wrapper, the Axum HTTP layer would lose visibility into the underlying TCP socket’s peer address once TLS is wrapped around it. By using TapIo, acme-proxy ensures that the allowed_ip and reverse_dns filters can correctly identify the true client IP, failing closed if it cannot be determined.

Configuration

To serve HTTPS directly, configure the [server.tls] section.

Note: Setting server.tls.enabled = true replaces the cleartext HTTP listener entirely. You do not get both HTTP and HTTPS simultaneously.

[server]
base_url = "https://acme.internal:3000"
bind_address = "[::]:3000"

[server.tls]
enabled = true

# The PEM encoded certificate chain (leaf first) and private key
cert_path = "server.pem"
key_path  = "server.key"

# Budget for one TLS handshake
handshake_timeout_ms = 10000

If cert_path or key_path files are missing on disk at startup, acme-proxy will automatically generate a self-signed certificate for the host specified in server.base_url and write it to disk.

The web admin listener

[admin.tls] is the same mechanism on a second socket: the same load-or-generate provisioning, the same TapIo wrapper preserving the peer address, the same 0600 on a generated key. Only the defaults differ.

[admin.tls]
enabled   = true
cert_path = "admin.pem"     # not server.pem
key_path  = "admin.key"

The paths are separate on purpose: the two listeners answer to different names (admin.base_url versus server.base_url, and a generated certificate takes its name from whichever applies), and sharing one certificate between them should be a decision, not an accident. The log lines carry a listener field ("acme" or "admin") so certificate churn on one is distinguishable from the other.

Unlike the ACME listener, TLS here is not optional once the panel leaves loopback — startup refuses that combination outright. See Web Admin.

Profiles

acme-proxy is a multi-tenant ACME server. It serves ACME entirely through Profiles.

A single process, running on a single port with a single database, can host multiple isolated ACME endpoints. Each profile is mounted under the /profile/<name>/directory namespace.

Why use profiles?

  • Serve an internal self-signed CA at /profile/local/directory for dev environments.
  • Serve a strict Let’s Encrypt relay at /profile/prod/directory for production services.
  • Apply different network filters (e.g., strict IP allowlists for production, bypass for dev) without needing to run multiple binary instances.

One process, several endpoints

Everything below the router is per profile — its own signer, filters, challenge validators and EAB policy. Everything above it is shared: one socket, one database, one process.

graph TD
    REQ["Incoming request"] --> ROOT["Root router<br/>/health, /, tracing, hardening headers"]
    ROOT -->|"/profile/dev/*"| PDEV["Profile: dev"]
    ROOT -->|"/profile/prod/*"| PPROD["Profile: prod"]
    ROOT -->|"/profile/staging/*"| PSTG["Profile: staging"]

    PDEV --> SDEV["signer: local_ca<br/>filters: none"]
    PPROD --> SPROD["signer: relay<br/>filters: ipam"]
    PSTG --> SSTG["signer: local_ca<br/>filters: none"]

    SDEV --> B1["Arc&lt;dyn SignerBackend&gt; #1"]
    SSTG --> B1
    SPROD --> B2["Arc&lt;dyn SignerBackend&gt; #2"]

    B1 --> DB[("one SQLite file<br/>rows tagged by profile")]
    B2 --> DB

Note the fan-in: dev and staging have identical [signer] sections, so they share one backend instance rather than constructing two. Two profiles sharing ca.key while differing elsewhere is a startup error: one CA key under two configurations would issue under two policies from one identity — two different crl_distribution_points for one CRL, say — and that is refused rather than left to half-work.

Hard database isolation

Profiles act as a strict isolation boundary in the SQLite database.

The accounts table uses a constraint: UNIQUE(profile, pubkey). This means if a client registers a key at /profile/default, and then uses the exact same cryptographic key to connect to /profile/le, the database creates two independent ACME accounts. This ensures endpoints cannot cross-pollinate authorizations, orders, or nonces.

Inheritance and configuration

Profiles inherit from the base (global) configuration keys. A profile only needs to override the specific keys that differ.

Eight sections can be overridden, and no others: signer, filter, ipam, challenge, eab, order, notify and meta. Everything else — the listen socket, the database, logging, audit, the admin listener — is process-wide.

graph LR
    ENV["ACME_PROXY_* env"] --> BASE
    FILE["config.toml"] --> BASE["Base configuration"]
    BASE --> MERGE{{"merge, per key"}}
    OVR["[profiles.prod]<br/>only the keys that differ"] --> MERGE
    MERGE --> EFF["Effective configuration<br/>for profile 'prod'"]

Inheritance is per key, not per section. A profile that sets only challenge.bypass keeps the global challenge.enabled rather than reverting it to the compiled default. Arrays, however, replace wholesale — they never append. Precedence is: profile key, then global key, then compiled default.

That split is why [filter] is shaped the way it is. filter.rules is an array, so a profile naming its own rules replaces the sequence outright — which is right, because order is the policy. [filter.check.<name>] and [filter.rule.<name>] are tables, so they merge per key, and a profile can dry-run one rule without restating anything:

[profiles.staging.filter.rule.inventory-owned]
mode = "warn"

A profile inherits every globally defined check and cannot remove one, which costs nothing: a check no selected rule names is never built. A global [filter] section can therefore carry a library of checks and each profile pick the subset its own filter.rules uses.

[signer]
backend = "local_ca"

[filter]
rules = ["corp-only"] # The base selection; each profile may replace it

# The library. Both checks and both rules are declared once, globally.
[filter.check.corp-net]
type  = "allowed_ip"
allow = ["10.0.0.0/8"]

[filter.check.corp-names]
type  = "identifiers"
allow = ["*.corp.example.com"]

[filter.rule.corp-only]
when = "corp-net"
then = "allow"

[filter.rule.named-and-corp]
when = "corp-net and corp-names"
then = "allow"

# Profile 1: Uses the global local_ca, and the inherited address-only rule.
# `corp-names` is named by no rule it selects, so it is never built.
[profiles.default]
enabled = true

# Profile 2: Relays to Let's Encrypt, and replaces the selection with the
# stricter rule — which is what pulls `corp-names` into existence here.
[profiles.le]
enabled = true
signer.backend = "relay"
signer.relay.directory_url = "https://acme-v02.api.letsencrypt.org/directory"
filter.rules = ["named-and-corp"]

acme-proxy filter show --profile <name> prints the built policy for one profile, which is the quickest way to confirm a profile selected what you intended. A check that no selected rule names is reported as filter_check_unused — an advisory, not an error.

At runtime, acme-proxy deduplicates the signer backends in memory so that two profiles sharing the exact same signer configuration (e.g., two profiles using the same local_ca) don’t duplicate background polling threads or memory overhead.

Signers

The signer is what actually produces a certificate, once the client has proved control of its names and the filters have allowed the request. Everything before this point is the same whichever backend you choose; everything after it is the backend’s business.

There are three, and they answer three different questions.

BackendUse it whenThe certificate is signed by
Local CAThe certificates only need to be trusted by machines you control.This server, from a CA key on disk or in a PKCS#11 token.
RelayYou need publicly trusted certificates, but your clients cannot reach a public CA — or you have one scarce upstream credential to share.A real upstream ACME CA, which this server becomes a client of.
Custom ScriptThe authority already exists and does not speak ACME.Whatever your script talks to: a legacy PKI, an internal API, an offline process.

Choosing one

Start from what has to trust the certificate:

  • Only your own machines? local_ca. You distribute the CA certificate once (see Trusting the CA) and the whole thing works offline, including revocation via the CRL.
  • Browsers, partners, anything you do not control? You need a public CA, so relay. Your internal clients keep proving control to this server — over HTTP, DNS or TLS, whichever suits them — while the upstream challenge is solved once, centrally, with a credential no client ever holds.
  • An existing corporate PKI that issues by ticket, script or API? custom. It is the escape hatch, and it is deliberately a shell contract rather than a plugin API, so anything that can be scripted can be a signer.

Nothing stops you from running more than one. [signer] is a per-profile section, so a local_ca at /profile/dev can sit beside a relay at /profile/prod in the same process.

What every backend has to provide

All three implement the same trait, and the shape of it is worth knowing because it is what the rest of the server can rely on. It comes in two halves, split by who holds the key: the backend — issue, revoke — is built only by the process running the worker role, and the read side — the CRL, the trust anchor, renewal information — is built by every process from public material, so the one parsing client requests never holds signing material:

  • issue — the only required capability. It receives the order’s identifiers and the client’s CSR and returns a chain, or refuses with badCSR.
  • revoke — must be idempotent. Revoking twice is not an error, because the server cannot always know whether a previous attempt reached the authority.
  • crl_der and renewal_info — optional, and default to “nothing to say here”. local_ca publishes a CRL; the relay passes the upstream’s renewal window through, explanationURL and all.

issue never runs inside a client’s request: finalize queues it, answers the order processing, and the worker calls the backend. A backend may itself answer processing rather than a certificate, meaning “this finishes later”. Only the relay does — signing upstream takes as long as the upstream takes — and the order simply stays processing until it has.

Configuration

[signer]
backend = "local_ca"   # or "relay", or "custom"

Each backend then reads its own table — [signer.local_ca], [signer.relay], [signer.custom] — documented on its own page.

Reference

backend (String) — Default: "local_ca" | Env: ACME_PROXY_SIGNER__BACKEND

Which backend issues certificates: local_ca, relay or custom. Any other value is a startup error.

Backends are shared by configuration, not per profile

Two profiles whose [signer] sections are identical share one backend instance rather than constructing two. Revocations and the CRL live in the database, keyed by the CA’s key, so two instances over one CA would still agree on what is revoked; what they could not agree on is everything else in the section.

Two profiles sharing ca.key while differing anywhere else in [signer] is therefore a startup error, not a race to discover later. See Profiles & Routing.

Local CA Signer

The local_ca backend uses an internal ECDSA key to act as a fully functional Certificate Authority. It is capable of generating its own self-signed root, or it can act as a subordinate (Intermediate) CA if provided with an existing key and certificate.

Reach for it for development environments, CI pipelines, and isolated internal networks — anywhere the certificates only need to be trusted by machines you control.

Security constraints

When processing a Certificate Signing Request (CSR) from a client, the local_ca is deeply distrustful of the requested extensions:

  1. Overwriting Extensions: The Local CA overwrites every extension the CSR asked for before signing — a fresh random serial, fixed key usages, and its own validity window. This matters more than it sounds: the CSR parser otherwise copies a requested basicConstraints/keyUsage straight into the signed leaf.
  2. Basic Constraints: The issued leaf is never a CA. Without the reset above, a client authorized for one name could submit a CSR carrying CA:TRUE
    • keyCertSign and receive a working intermediate CA, which it could then use to mint arbitrary trusted certificates. (In implementation terms the leaf is built with IsCa::NoCa rather than ExplicitNoCa — an explicit CA:FALSE broke certbot’s chain parser — but the security property is the same.)
  3. Subject Alternative Names (SANs): The DNS SANs in the CSR must be exactly the set of identifiers the order authorized — no more, no fewer — or issuance fails with badCSR.
  4. Non-DNS SANs are rejected, not stripped: a CSR carrying an IP, email or URI SAN is refused outright with badCSR. Do not expect local_ca to quietly drop them.
  5. The subject is emptied: the issued leaf carries no distinguished name at all, so a Common Name in the CSR cannot leak into it. (A CN that looks like a domain the order does not cover is separately rejected earlier, at finalize — see Troubleshooting.)

Certificate validity

Leaves are valid for leaf_validity_days, starting one hour in the past to absorb clock skew between the CA and its clients.

A client may narrow that window using the order’s notBefore/notAfter (RFC 8555 §7.4), but never widen it: a requested start is honoured only if it is later than the default start, and a requested end only if it is earlier than the default end. A request whose clamped window would be empty or inverted is discarded whole, the policy default is used, and a local_ca_requested_validity_discarded warning is logged.

Configuration

[signer]
backend = "local_ca"

[signer.local_ca]
cert_path = "ca.pem"
key_path = "ca.key"
key_type = "ecdsa-p256"
crl_path = "ca.crl"
leaf_validity_days = 90

Reference

cert_path (String) — Default: "ca.pem" | Env: ACME_PROXY_SIGNER__LOCAL_CA__CERT_PATH

The path where the CA certificate is stored. If this file does not exist, a new self-signed Root CA is generated on startup and saved here.

key_path (String) — Default: "ca.key" | Env: ACME_PROXY_SIGNER__LOCAL_CA__KEY_PATH

The path where the CA private key is stored. A generated key is created with mode 0600 at creation time, not chmod’ed afterwards, so there is no window in which another local user could read it.

key_type (String) — Default: "ecdsa-p256" | Env: ACME_PROXY_SIGNER__LOCAL_CA__KEY_TYPE

Algorithm for a generated CA key. "ecdsa-p256" is currently the only accepted value; anything else is a startup error. (An existing key supplied on disk is used as-is, whatever its type.)

crl_path (String) — Default: "ca.crl" | Env: ACME_PROXY_SIGNER__LOCAL_CA__CRL_PATH

Where the current Certificate Revocation List (RFC 5280) is exported as PEM, for publishing from a static web server. The revocations and the CRL GET /crl serves live in the database; this file is rewritten whenever a new CRL is stored and is never read back. Writers lock ca.json.lock beside it, which stays empty. The path also locates the JSON ledger a CA kept before the database did (ca.crl → ca.json), imported once; see Revocation.

crl_distribution_points (Array<String>) — Default: [] | Env: ACME_PROXY_SIGNER__LOCAL_CA__CRL_DISTRIBUTION_POINTS

Where a relying party can fetch that CRL. Each URL is written into every issued leaf as cRLDistributionPoints (RFC 5280 §4.2.1.13); empty — the default — emits no extension at all, which is why a certificate from this CA says nothing about revocation until you set this.

Nothing derives it, deliberately. The URL is frozen into every certificate signed while it is set, so a value read from server.base_url would silently break certificates already issued the day that changed. And this server’s own copy is served at {base_url}/profile/<name>/crl, inside the profile router and therefore behind that profile’s filter policy — an address-based rule would refuse it to exactly the relying parties the extension exists for. Name a URL you know is publicly reachable; a webroot or CDN copy of crl_path is the usual answer.

Several entries mean one CRL reachable in several places, not several different CRLs. http:// is idiomatic and gets no warning: fetching a signed CRL over TLS means validating that connection’s certificate first, which is the loop this extension exists to break. Credentials in the URL, a non-http(s) scheme, and any value the URL parser would normalize (a missing trailing /, a leading space from an environment-variable list) are each a startup error naming the key and the value.

Note that two profiles sharing one CA share this list too: they share one [signer] section, one set of revocations and one CRL, so there is one place that CRL is published. Giving them different URLs while they share ca.key is refused at startup — see Profiles & Routing.

ca_issuer_urls (Array<String>) — Default: [] | Env: ACME_PROXY_SIGNER__LOCAL_CA__CA_ISSUER_URLS

Where a relying party can fetch this CA’s own certificate, written into every issued leaf as authorityInfoAccess with the caIssuers access method (RFC 5280 §4.2.2.1). Empty — the default — emits no extension. Same reasoning, same validation and the same startup errors as crl_distribution_points above.

The sibling access method, id-ad-ocsp, is never written: this server runs no OCSP responder, and a pointer at one that does not exist is worse than none.

leaf_validity_days (Integer) — Default: 90 | Env: ACME_PROXY_SIGNER__LOCAL_CA__LEAF_VALIDITY_DAYS

The validity period (in days) for issued leaf certificates. Distinct from order.validity_seconds, which bounds the ACME order object, not the certificate.

key_source (String) — Default: "file" | Env: ACME_PROXY_SIGNER__LOCAL_CA__KEY_SOURCE

Where the issuing private key lives. "file" is everything described on this page: a PEM key at key_path, loaded or generated. "pkcs11" puts the key in a hardware token instead, and reads the [signer.local_ca.pkcs11] table — see Hardware Keys. Any other value is a startup error, and so is "pkcs11" on a binary built without --features hsm: there is deliberately no silent fallback, since falling back would hand an operator who asked for hardware a software key with no indication of it.

[signer.local_ca.subject]

The X.509 Subject (Distinguished Name) of the CA certificate acme-proxy generates for itself. It is read only on the startup that generates the CA — supplying your own cert_path means supplying your own DN, and editing this table afterwards changes nothing, because the certificate already exists. To change it, retire the CA and generate a new one.

Every key is optional and String-valued. An unset or empty value is omitted from the Subject rather than written blank; the empty string is treated as absent because the environment source cannot distinguish the two. The one exception is common_name, which falls back to "acme-proxy local CA" so the CA always carries one. Leaving the whole table out therefore yields a CommonName-only Subject — which is exactly what this backend produced before the table existed.

This is the CA’s own DN, not the leaf’s. Issued certificates deliberately carry an empty Subject and identify themselves through subjectAltName, which is what every modern client reads.

common_name — Default: "acme-proxy local CA" | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__COMMON_NAME

organization — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__ORGANIZATION

organizational_unit — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__ORGANIZATIONAL_UNIT

country — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__COUNTRY

A two-letter ISO 3166-1 country code, e.g. "US".

state — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__STATE

locality — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__LOCALITY

[signer.local_ca.subject]
common_name  = "Example Corp Internal CA"
organization = "Example Corp"
country      = "US"

Hardware-backed keys

The CA key described above is a file, and a file can be copied. If the CA is one your fleet actually trusts, the issuing key can instead live in a PKCS#11 token — a YubiKey, an enterprise HSM, or SoftHSM2 for development — where it is created once and can never be read back out.

Everything on this page still applies; only where the signature comes from changes. See Hardware Keys (PKCS#11). Note that it requires a build with --features hsm, which is not the default.

Multi-tier PKI (using an intermediate CA)

For production internal deployments, you should avoid using an auto-generated Root CA directly on the server. Instead, you can use a Multi-Tier Hierarchy: create an offline Root CA, use it to sign an Intermediate CA, and hand the Intermediate CA to acme-proxy.

Here is how you can do this in practice using OpenSSL.

Step 1 — Create the offline root CA

Generate a private key and a self-signed Root certificate (keep this key highly secure and offline):

# Generate the Root private key
openssl ecparam -genkey -name prime256v1 -out root_ca.key

# Create the self-signed Root certificate (valid for 10 years)
openssl req -x509 -new -nodes -key root_ca.key -sha256 -days 3650 \
  -out root_ca.pem \
  -subj "/CN=My Company Offline Root CA"

Step 2 — Create the intermediate CA for acme-proxy

Generate the private key and CSR for the Intermediate CA:

# Generate the Intermediate private key
openssl ecparam -genkey -name prime256v1 -out acme_intermediate.key

# Create the CSR
openssl req -new -key acme_intermediate.key -out acme_intermediate.csr \
  -subj "/CN=My Company ACME Intermediate CA"

Create an OpenSSL extension file (v3_ext.cnf) to ensure the Intermediate CA is allowed to sign other certificates:

[ v3_intermediate_ca ]
subjectKeyIdentifier = hash
authorityKeyIdentifier = keyid:always,issuer
basicConstraints = critical, CA:true, pathlen:0
keyUsage = critical, digitalSignature, cRLSign, keyCertSign

(Note: pathlen:0 ensures this intermediate can issue leaf certificates, but cannot issue further intermediate CAs).

Now, sign the Intermediate CSR using the Offline Root CA:

openssl x509 -req -in acme_intermediate.csr -CA root_ca.pem -CAkey root_ca.key \
  -CAcreateserial -out acme_intermediate.pem -days 1825 -sha256 \
  -extfile v3_ext.cnf -extensions v3_intermediate_ca

Step 3 — Configure acme-proxy

Finally, point acme-proxy to your newly minted Intermediate CA. The cert_path must contain the Intermediate certificate followed by the Root certificate (the bundle), so clients can verify the full chain.

cat acme_intermediate.pem root_ca.pem > ca_bundle.pem

Update your config.toml:

[signer.local_ca]
cert_path = "ca_bundle.pem"
key_path = "acme_intermediate.key"

Now acme-proxy issues certificates signed by the Intermediate CA, mirroring a standard enterprise PKI hierarchy.

The order inside the bundle is load-bearing. The signing issuer is parsed from the first PEM block in cert_path; the remaining blocks are only appended to the chain served to clients. Concatenating root-first instead would not fail loudly — it would sign with the root’s identity using the intermediate’s key, producing certificates nothing can verify. Always cat intermediate.pem root.pem, never the reverse.

Two further caveats:

  • Nothing checks that key_path actually corresponds to the certificate in cert_path. A mismatched pair produces unverifiable certificates rather than a startup error. (This caveat is specific to key_source = "file"; the PKCS#11 path does verify the pair at startup.)
  • The whole bundle is emitted with every issued certificate, root included. Most clients tolerate this, but if you would rather not ship the root, put only the intermediate in cert_path and distribute the root out of band — see Trusting the CA.

Hardware Keys (PKCS#11)

By default the Local CA’s issuing key is a PEM file on disk, protected by nothing but its 0600 permissions. key_source = "pkcs11" moves that key into a hardware token — a YubiKey, a network HSM, or SoftHSM2 for development — where it is created once and can never be read back out. acme-proxy sends the token the bytes to be signed and receives a signature; the private key never enters this process’s memory.

PKCS#11 rather than a vendor-specific PIV library, so a YubiKey today and an enterprise HSM tomorrow are the same configuration with a different module_path.

What this protects, and what it does not

Everything else about the Local CA is unchanged: the same CSR sanitisation, the same leaf_validity_days clamping, the same CRL and revocations. Only where the signature comes from moves.

It protects the CA issuing key — the one that, if stolen, lets an attacker mint certificates your fleet trusts. It does not protect the ACME account keys, the TLS server key (server.tls.key_path), or the database; those stay on disk.

Requirements

  • A build with the hsm feature. It is off by default, so the stock binary does not have it and key_source = "pkcs11" on one is a startup error naming the feature:

    cargo build --release --features hsm
    
  • A PKCS#11 module (.so), loaded at runtime — nothing is linked at build time.

  • An existing CA certificate, for the reason below.

Two rules that differ from the software path

Both are startup errors, so you will meet them immediately rather than in production.

The CA is never generated

With key_source = "file", a missing cert_path/key_path means “generate a CA and write it here”. With key_source = "pkcs11" there is no such thing: the private key is created inside the token by its own tooling, and this server cannot produce one that a token would then hold. So cert_path must already exist, and key_path is neither read nor written.

Both walkthroughs below cover creating that certificate.

The key and the certificate are cross-checked

At startup, the token key’s SubjectPublicKeyInfo is compared against the one in cert_path. A mismatch — almost always a wrong key_label — stops the server.

This is stricter than the file path, where (as Local CA warns) nothing checks that key_path corresponds to cert_path, and a mismatched pair simply produces certificates that verify nowhere. Here a typo is caught before the first certificate is issued rather than discovered by a client days later.

Reference

Reaching this table at all takes signer.local_ca.key_source = "pkcs11", which is documented with the rest of [signer.local_ca] in Local CA. Everything below is [signer.local_ca.pkcs11], read only when that key is set.

pkcs11.module_path (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__MODULE_PATH

The PKCS#11 module to load. Required. See each walkthrough for the usual paths.

pkcs11.token_label (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__TOKEN_LABEL

Which token to use, by label. Preferred over slot_id: slot numbers are assigned dynamically and change across reboots and re-plugs on most drivers (SoftHSM2 will hand you something like 276468771).

pkcs11.slot_id (Integer) — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__SLOT_ID

Which slot to use, for tokens with no usable label. Consulted only when token_label is empty.

pkcs11.key_label (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__KEY_LABEL

The private key’s CKA_LABEL. Required. On a YubiKey the labels are fixed by the driver, so this is something you look up rather than choose — see below.

pkcs11.key_id (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__KEY_ID

The key’s CKA_ID as hex (01, or 01:ff), to disambiguate a token holding several keys under one label. Optional; two keys sharing a label and no key_id to separate them is a startup error rather than a coin flip.

pkcs11.pin_file (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__PIN_FILE

A file holding the user PIN. Trailing whitespace is trimmed, so a PIN written with echo works. The file is checked for permissions and warns if it is world-readable, exactly as ca.key does.

pkcs11.pin (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__PIN

SENSITIVE. The PIN directly. Prefer pin_file, or set this through the environment variable; a PIN in config.toml is a long-lived secret in a file that tends to get copied around. pin_file wins when both are set, and having neither is a startup error.

A PIN is not a password. Tokens block after a small number of wrong attempts — three on a YubiKey PIV applet, after which you need the PUK. This is why acme-proxy retries a failed signature at most once, and why the trailing newline in your PIN file is worth getting right.


Walkthrough A — SoftHSM2

SoftHSM2 is a software token: no hardware needed, and the same setup the project uses in CI. Use it to try the feature before committing to hardware.

# Debian/Ubuntu
sudo apt install softhsm2
# Arch
sudo pacman -S softhsm

Step 1 — Create a token

softhsm2-util --init-token --free --label acme-ca --so-pin 3737 --pin 1234

--free takes the first uninitialised slot. Note that the token is reassigned to a new slot number afterwards — which is exactly why token_label is the selector to use, not slot_id.

Step 2 — Create the CA key and certificate

The key must exist inside the token, and cert_path must hold a certificate for it. For SoftHSM2 the simplest route is to generate both locally, import the key, and destroy the local copy:

# The CA key and its self-signed certificate
openssl genpkey -algorithm EC -pkeyopt ec_paramgen_curve:P-256 -out ca.key

openssl req -x509 -new -key ca.key -sha256 -days 3650 -out ca.pem \
  -subj "/CN=Example Corp Issuing CA/O=Example Corp" \
  -addext "basicConstraints=critical,CA:true,pathlen:0" \
  -addext "keyUsage=critical,keyCertSign,cRLSign"

# Move the key into the token, then remove it from disk
softhsm2-util --import ca.key --token acme-ca --label ca-key --id 01 --pin 1234
shred -u ca.key

For a real HSM, generate the key in the token instead so it never exists outside it — pkcs11-tool --module <module> --token-label acme-ca --login --keypairgen --key-type EC:prime256v1 --label ca-key --id 01 (from the opensc package), then certify that public key with your offline root. The import above is a development convenience, and the reason it is acceptable here is that a SoftHSM2 token is a directory of files anyway.

pathlen:0 matches what the Local CA generates for itself: it may issue leaves but no further CAs.

Step 3 — Configure

[signer]
backend = "local_ca"

[signer.local_ca]
cert_path  = "ca.pem"
crl_path   = "ca.crl"
key_source = "pkcs11"

[signer.local_ca.pkcs11]
module_path = "/usr/lib/softhsm/libsofthsm2.so"
token_label = "acme-ca"
key_label   = "ca-key"
pin_file    = "/etc/acme-proxy/hsm.pin"
printf '1234' > /etc/acme-proxy/hsm.pin
chmod 600 /etc/acme-proxy/hsm.pin

If SoftHSM2’s token store is not in its default location, SOFTHSM2_CONF must be set in the server’s environment — it is read by the module, not by acme-proxy.

Step 4 — Confirm it is really using the token

RUST_LOG=info acme-proxy serve
INFO acme_proxy::signer::local_ca::pkcs11: the local CA's issuing key is on a PKCS#11 token
  event="local_ca_pkcs11_opened" module=/usr/lib/softhsm/libsofthsm2.so
  slot=276468771 key_label=ca-key algorithm=PKCS_ECDSA_P256_SHA256
  mechanism=CKM_ECDSA_SHA256
INFO acme_proxy::signer::local_ca: event="local_ca_pkcs11_loaded" cert_path="ca.pem" key_label=ca-key

local_ca_pkcs11_opened is the line that proves it: it names the module, the slot the token actually landed in, the curve read off the key, and the mechanism chosen. If you see local_ca_loaded or local_ca_generated instead, the configuration is still on the file path.

Then issue something and check it chains:

openssl verify -CAfile ca.pem /path/to/issued/cert.pem
# cert.pem: OK

Walkthrough B — YubiKey (libykcs11)

A YubiKey 5 exposes its PIV applet through libykcs11, shipped with yubico-piv-tool.

# Debian/Ubuntu → /usr/lib/x86_64-linux-gnu/libykcs11.so
sudo apt install yubico-piv-tool
# Arch → /usr/lib/libykcs11.so
sudo pacman -S yubico-piv-tool

Both paths are in circulation; check which one you have before configuring module_path.

Step 1 — Generate the key and certificate on the device

Use slot 9c (Digital Signature). Its PIV policy requires the PIN for every private-key operation, which is the right posture for a CA key and the reason to prefer it over 9a.

# Generate the key inside the YubiKey — it never leaves
yubico-piv-tool -s 9c -a generate -A ECCP256 -o ca_pub.pem

# Self-sign a certificate for it, on the device
yubico-piv-tool -s 9c -a verify-pin -a selfsign-certificate \
  -S '/CN=Example Corp Issuing CA/O=Example Corp/' \
  --valid-days 3650 -i ca_pub.pem -o ca.pem

# Store the certificate in the slot as well (optional, but conventional)
yubico-piv-tool -s 9c -a import-certificate -i ca.pem

Copy ca.pem to wherever cert_path points.

Touch policy must be never for the CA slot. If the slot is provisioned to require a touch, every issuance blocks until somebody physically touches the key. That is correct for an offline root and catastrophic for an ACME server expected to issue unattended.

Step 2 — Find the key label

You do not choose the label on a YubiKey — libykcs11 assigns fixed ones per PIV slot. Read it off the device:

pkcs11-tool --module /usr/lib/libykcs11.so --list-objects --login

Slot 9c reports as Private key for Digital Signature; 9a as Private key for PIV Authentication. Use that string verbatim.

Step 3 — Configure

[signer.local_ca]
cert_path  = "ca.pem"
crl_path   = "ca.crl"
key_source = "pkcs11"

[signer.local_ca.pkcs11]
module_path = "/usr/lib/libykcs11.so"
token_label = "YubiKey PIV #12345678"
key_label   = "Private key for Digital Signature"
pin_file    = "/etc/acme-proxy/hsm.pin"

The PIN is the PIV PIN (factory default 123456), not the PIV management key and not the FIDO PIN.

Step 4 — Expect CKM_ECDSA

libykcs11 does not offer CKM_ECDSA_SHA256, so acme-proxy computes the SHA-256 digest itself and asks the token to sign that. The startup line reads:

mechanism=CKM_ECDSA+SHA256

This is normal and not a downgrade — the same signature, with the hashing done on this side of the USB cable.

Performance

A YubiKey signature takes roughly 50–300 ms, and signings are serialised by a mutex. That is comfortable for hundreds of certificates a day and is not a throughput solution; the signing call runs on the blocking thread pool, so it does not stall the rest of the server while it waits. For higher volumes use a networked HSM, or the Custom Script signer against a KMS.


Operations

Backup and disaster recovery

The key cannot be backed up. That is the point of the feature, and it makes recovery something to plan before you need it. Two workable approaches:

  • Two tokens, one offline root. Keep an offline root CA, use it to certify an intermediate held on each of two tokens, and hand acme-proxy one of them. A lost token is replaced by provisioning a new intermediate; clients trust the root and never notice. See Multi-Tier PKI.
  • Accept re-enrolment. For a small internal fleet, losing the CA and distributing a new one is survivable — just make sure it is a decision rather than a discovery.

Revocations live in the database like every other CA’s, so backing up the database backs them up; crl_path is only an export of the current CRL.

When the token disappears

If the session drops — the YubiKey is unplugged, a network HSM times out — acme-proxy reopens the session, logs back in and retries the signature once. The relevant log lines are local_ca_pkcs11_session_lost followed by either a successful issuance or local_ca_pkcs11_reconnect_failed.

If that fails, finalize requests return serverInternal (500) and clients retry, which is the right behaviour: the order stays valid and issuance resumes once the token is back. GET /crl keeps working throughout — the current CRL is read from the database and serving it signs nothing.

Sharing one token between profiles

Several profiles may use the same module, and even the same key. acme-proxy opens one PKCS#11 context per module for the whole process and shares it, so this works without special configuration. Two profiles naming the same token key with otherwise different signer settings is refused at startup, for the same reason two profiles sharing ca.key are: one key under two configurations would issue under two policies from one identity.

Troubleshooting

SymptomCause and fix
key_source = "pkcs11" … built withoutThe binary has no PKCS#11 support. Rebuild with --features hsm.
CKR_PIN_INCORRECT at startupUsually a stray character in pin_file. Trailing newlines are trimmed, but leading or embedded whitespace is not. Check with xxd. Do not retry blindly — see the PIN warning above.
CKR_PIN_LOCKEDToo many wrong attempts. A YubiKey PIV PIN is unblocked with the PUK (yubico-piv-tool -a unblock-pin).
is not the key certified by …The SPKI cross-check failed: key_label/key_id resolve to a different key than cert_path describes. List the token’s objects and compare.
no PKCS#11 token labelled …The label is wrong, or the token is not plugged in. The message lists the labels actually present.
N PKCS#11 private keys are labelled …Set key_id to pick one.
unsupported PKCS#11 curveOnly P-256 and P-384 are supported. The message prints the CKA_EC_PARAMS it found.
supports neither … nor CKM_ECDSAThe token cannot do ECDSA signing at all, or will not report its mechanisms.
Certificates that verify nowhereShould not happen — the SPKI cross-check catches the usual cause at startup. If it does, capture the local_ca_pkcs11_opened line and the failing certificate and open an issue.

See also Maintenance & Troubleshooting.

Relay

The relay backend keeps this server answering ACME to its own clients while a real upstream CA does the signing. It captures internal ACME requests and fulfils them through an upstream external CA (Let’s Encrypt, ZeroSSL, or a commercial CA).

How it works

There are two ACME conversations, and keeping them apart is the whole trick to reading this page. acme-proxy is the server in one and the client in the other. Each has its own account, its own account key, its own order and its own challenges, and nothing crosses between them.

sequenceDiagram
    autonumber
    participant C as Internal client<br/>(certbot, acme.sh)
    participant P as acme-proxy
    participant U as Upstream CA<br/>(Let's Encrypt)

    Note over C,P: Conversation 1 — acme-proxy is the SERVER.<br/>Client's own account key, client's own order.
    C->>P: newOrder (example.internal)
    P-->>C: 201, order pending
    C->>P: trigger challenge
    P->>C: validate control (http-01 / dns-01 / tls-alpn-01)
    P-->>C: 200, challenge valid — order ready
    C->>P: finalize (CSR)

    Note over P,U: Conversation 2 — acme-proxy is the CLIENT.<br/>Its OWN upstream account key, a SECOND order.
    P-->>C: 200, order processing
    P->>U: newOrder (same identifiers)
    U-->>P: 201, upstream order pending
    alt challenge_strategy = dns01
        P->>P: publish TXT via RFC 2136 + TSIG
    else challenge_strategy = http01
        P->>P: publish token on its own /.well-known route
    else challenge_strategy = bypass
        P->>U: trigger — the upstream already trusts this account
    end
    U->>P: validate (against the proxy's thumbprint)
    P->>U: finalize (the client's CSR, relayed unchanged)
    U-->>P: certificate

    Note over C,P: Back in conversation 1.
    C->>P: poll order
    P-->>C: 200, order valid + certificate URL

Two consequences fall straight out of the diagram:

  • The key authorization at the upstream uses the proxy’s own thumbprint, never the client’s. They are different accounts on different servers, so the client could not answer the upstream’s challenge even in principle.
  • finalize stays processing for longer. Every backend’s finalize answers processing and is signed by the worker, but here conversation 2 takes as long as the upstream takes, so the client polls for minutes rather than for the moment local_ca needs — see Core Concepts.

Challenge strategies

The strategy names the one challenge type the proxy answers upstream. A CA offers several — Let’s Encrypt currently poses four — and every type the strategy does not name is ignored, including ones this server has no implementation for at all. Nothing falls back: if the upstream authorization does not offer the type the strategy names, the relay says so and stops, rather than trying a type it could not finish.

dns01 (RFC 2136 TSIG)

This is the strategy that can prove a wildcard. The proxy intercepts internal HTTP-01 or DNS-01 challenges, but to satisfy the external CA, the proxy solves the external DNS-01 challenge itself. It does this using an rfc2136 provider powered by hickory-proto. It securely authenticates with the DNS server using TSIG (Transaction Signature) to publish the TXT record. Note: The TXT record uses the thumbprint of the proxy’s upstream account key, not the internal client’s key.

DNS alias mode

By default the record is written at _acme-challenge.<domain>, so the update key needs write access to every zone a profile issues for. Alias mode moves the record to one name in a zone set aside for it, the way acme.sh’s --challenge-alias does. Delegate each domain once with a CNAME, then point zone and challenge_alias at the alias zone:

_acme-challenge.www.example.com.  CNAME  _acme-challenge.acme-alias.net.
_acme-challenge.api.example.org.  CNAME  _acme-challenge.acme-alias.net.
[signer.relay.dns01]
challenge_alias = "acme-alias.net."

[signer.relay.dns01.rfc2136]
zone = "acme-alias.net."

The CA follows the CNAME itself; nothing changes on its side. Every domain of the profile shares the one record name, which is safe because values are added and removed one by one. The alias is configured rather than found by following the CNAME, because this server’s resolver need not see what the CA sees, and a record published at the wrong name invalidates the upstream authorization for good.

Whoever can write the alias name can pass dns-01 for every domain pointing at it. Domains that must not share that power belong in separate profiles, each with its own alias and key.

http01

The proxy answers the upstream’s http-01 challenge by serving the key authorization itself, from a route on its own root router at /.well-known/acme-challenge/<token>.

Like dns01, the value served is derived from this proxy’s upstream account thumbprint, not the internal client’s — the two are different accounts on different servers, so the client cannot answer it even in principle. Unlike dns01, the body is the key authorization verbatim rather than its SHA-256 digest (RFC 8555 §8.3 versus §8.4).

This is the opposite direction of the inbound http-01 challenge, which is this server validating its own clients. The two share the well-known path and nothing else.

Two constraints are worth knowing before choosing it:

  • It needs a forwarder. acme-proxy does not open a second listener and does not bind port 80. The upstream CA fetches http://<identifier>:80/.well-known/acme-challenge/<token>, so something must forward or redirect that path to acme-proxy — see Deploying the http-01 responder below.
  • It cannot prove a wildcard. Nothing answers HTTP on the name *.example.com. An upstream authorization for a wildcard is refused with an error naming dns01, which is the strategy that can.

There is no [signer.relay.http01] table: setting challenge_strategy = "http01" is the whole configuration.

bypass

Used when the upstream CA implicitly trusts the proxy’s account (e.g., a commercial CA with pre-validated domains). The proxy simply tells the upstream “I have validated this”, bypassing external challenges entirely.

This is the one strategy that does not name a type: it triggers whatever the authorization offers. It prefers a challenge it could have satisfied had it needed to, since an upstream that validates out of band still decides against the challenge it was pointed at.

Deploying the http-01 responder

Only relevant with challenge_strategy = "http01".

The upstream CA fetches on port 80 of each name being issued, so put the existing web server for that name in front of acme-proxy and hand it that one path. Either shape works — RFC 8555 §8.3 explicitly permits following redirects, and every real CA does, so the target need not share the name:

# nginx, on the name being certified. Proxy it:
location /.well-known/acme-challenge/ {
    proxy_pass http://acme-proxy:3000;
}

# ...or just redirect it, which needs no upstream block:
location /.well-known/acme-challenge/ {
    return 301 http://acme-proxy:3000$request_uri;
}
# Caddy
handle /.well-known/acme-challenge/* {
    reverse_proxy acme-proxy:3000
}
# Traefik, as a dynamic-configuration router
http:
  routers:
    acme-challenge:
      rule: "PathPrefix(`/.well-known/acme-challenge/`)"
      service: acme-proxy
      priority: 100
  services:
    acme-proxy:
      loadBalancer:
        servers:
          - url: "http://acme-proxy:3000"

The route is mounted on the root router, beside GET /health — it is not under a profile’s /profile/<name> prefix, carries no filter chain, mints no nonce, and answers a plain 404 rather than an ACME problem document for an unknown token. It exists only while some profile’s signer uses this strategy; with any other backend the path is not routed at all. A http_01_responder_mounted line at startup confirms it is live.

Configuration

[signer]
backend = "relay"

[signer.relay]
directory_url = "https://acme-staging-v02.api.letsencrypt.org/directory"
account_key_path = "upstream_account.key"
contact = ["mailto:admin@example.com"]
challenge_strategy = "bypass"
poll_interval_ms = 2000
poll_timeout_secs = 300

Several relaying profiles

[signer] is a per-profile section, so one server can relay to several upstreams at once — a Let’s Encrypt endpoint beside a commercial CA, or one internal CA per environment. Each profile gets its own [signer.relay], and they must not share an account_key_path: two backends over one upstream account key would overwrite each other’s registration, and startup refuses it by name.

[profiles.public.signer]
backend = "relay"
relay.directory_url = "https://acme-v02.api.letsencrypt.org/directory"
relay.account_key_path = "public_account.key"

[profiles.partner.signer]
backend = "relay"
relay.directory_url = "https://acme.commercial-ca.example/directory"
relay.account_key_path = "partner_account.key"

Two profiles whose [signer] sections are byte-for-byte identical share one backend and one upstream account; anything that differs makes them independent. See Profiles for what else a profile separates.

Reference

directory_url (String) — Default: "" | Env: ACME_PROXY_SIGNER__RELAY__DIRECTORY_URL

The upstream ACME server’s directory URL.

Starting the server registers an account. The first acme-proxy serve with this backend configured contacts directory_url and performs newAccount there, writing the assigned account URL to the .kid sidecar. There is no confirmation step and no dry-run — merely booting a configuration that names a production CA creates a real account at it, and account creation is itself rate limited (Let’s Encrypt allows 10 per IP address per 3 hours). Point directory_url at a staging endpoint (https://acme-staging-v02.api.letsencrypt.org/directory) while you are still working out a configuration, and switch to production only once it is settled. Subsequent starts reuse the .kid sidecar and do not contact the upstream.

account_key_path (String) — Default: "upstream_account.key" | Env: ACME_PROXY_SIGNER__RELAY__ACCOUNT_KEY_PATH

Path to this proxy’s own account key at the upstream CA. If the file is absent, an ECDSA P-256 key is generated on startup. The assigned kid is stored beside it with a .kid extension.

contact (Array) — Default: [] | Env: ACME_PROXY_SIGNER__RELAY__CONTACT

Optional contacts sent with newAccount to the upstream CA.

challenge_strategy (String) — Default: "bypass" | Env: ACME_PROXY_SIGNER__RELAY__CHALLENGE_STRATEGY

How the proxy satisfies the upstream’s domain-control checks: bypass (the upstream validates nothing), dns01 (publish the TXT record the upstream asks for) or http01 (serve the challenge file from this server’s own root router, which requires a reverse proxy in front of it and cannot prove a wildcard). Any other value is a startup error.

poll_interval_ms (Integer) — Default: 2000 | Env: ACME_PROXY_SIGNER__RELAY__POLL_INTERVAL_MS

How often to poll an upstream order/authorization while it resolves.

poll_timeout_secs (Integer) — Default: 300 | Env: ACME_PROXY_SIGNER__RELAY__POLL_TIMEOUT_SECS

Total budget (in seconds) for one upstream issuance before the local order is marked invalid.

[signer.relay.dns01]

Only consulted when challenge_strategy = "dns01".

provider (String) — Default: "rfc2136" | Env: ACME_PROXY_SIGNER__RELAY__DNS01__PROVIDER

DNS provider used to publish the upstream TXT record. rfc2136 is currently the only implementation.

challenge_alias (String) — Default: "" | Env: ACME_PROXY_SIGNER__RELAY__DNS01__CHALLENGE_ALIAS

A domain, e.g. acme-alias.net., under which every challenge record is published as _acme-challenge.<alias> — see DNS alias mode. Empty publishes at each domain’s own _acme-challenge name. The alias must lie inside rfc2136.zone; one outside it, a wildcard, or a value that already starts with _acme-challenge. is a startup error.

[signer.relay.dns01.rfc2136]

All default to "" and are required once the dns01 strategy is selected.

server — Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__SERVER host:port of the nameserver accepting the dynamic update, e.g. 10.0.0.53:53. This is the update target, distinct from dns.resolver, which governs lookups.

zone — Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__ZONE The zone the update is sent for, fully qualified with a trailing dot, e.g. internal.company.com..

tsig_key_name — Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_KEY_NAME Name of the TSIG key the update is signed with.

tsig_key_secret — Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_KEY_SECRET The TSIG shared secret, in standard base64 — note this differs from EAB secrets, which are base64url. A value that is not valid base64 is a startup error, not a runtime one. This key is legitimately long-lived, so unlike a one-shot EAB credential it belongs in configuration; still prefer the environment variable over a file on disk.

tsig_algorithm (String) — Default: "" (read as hmac-sha256) | Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_ALGORITHM The TSIG algorithm: hmac-sha256, hmac-sha384 or hmac-sha512. Must match the key as your nameserver defines it.

Records are added and removed by value, so other TXT values at the same name — another order’s for that name, or ones this server did not write — are left alone. The challenge name must lie inside zone; one outside it is refused before anything is sent.

Updates are sent over UDP and retried over TCP when the response is truncated — a TSIG-signed update readily exceeds 512 bytes, so the TCP path is a normal occurrence rather than an edge case. Only a UDP answer from server itself is accepted.

A successful answer counts only when it is TSIG-signed by the configured key, for this update, within the server’s time window (RFC 8945); anything else fails the update. A refusal is reported as the server sent it, with its TSIG error explained: BADKEY means the server does not know tsig_key_name under tsig_algorithm, BADSIG that tsig_key_secret does not match, and BADTIME that the two clocks disagree.

[signer.relay.dns01.propagation]

What the dns01 strategy waits for between publishing the TXT record and asking the upstream to validate it. The upstream looks once: a record it cannot see yet makes the authorization invalid for good, which fails the client’s order and, at a public CA, counts against its failed-validation limit.

mode (String) — Default: "none" | Env: ACME_PROXY_SIGNER__RELAY__DNS01__PROPAGATION__MODE

none triggers the challenge right after the update succeeds, which is right when the update server is itself what the CA asks. delay sleeps delay_secs first, for a provider that accepts an update before serving it (a DNS API behind an RFC 2136 bridge) or secondaries that lag their primary. Any other value is a startup error.

delay_secs (u64) — Default: 30 | Env: ACME_PROXY_SIGNER__RELAY__DNS01__PROPAGATION__DELAY_SECS

Seconds to wait under delay; ignored under none. Zero is refused (use none), and so is a value not less than poll_timeout_secs, which bounds the whole relay attempt. The delay runs once for each name in the order, so a multi-name order needs poll_timeout_secs above the delay times the number of names plus the time the upstream takes.

The wait is a fixed delay rather than a poll of a public resolver on purpose. A CA resolves through the zone’s authoritative nameservers itself, while a recursive resolver answers from whichever one it reached and caches a “no such record” from a lagging one for the zone’s negative TTL — and never sees an internal or split-horizon zone at all.

[signer.relay.eab]

An upstream External Account Binding credential supplied in configuration rather than through acme-proxy upstream register. Both keys are empty by default, which means “no configuration-file credential”. Read only by acme-proxy serve, and only on a startup that finds no .kid sidecar beside account_key_path — see EAB considerations below for which of the two mechanisms to prefer.

kid (String) — Default: "" | Env: ACME_PROXY_SIGNER__RELAY__EAB__KID

The key id the upstream’s operator issued alongside the secret.

hmac_key (String) — Default: "" | Env: ACME_PROXY_SIGNER__RELAY__EAB__HMAC_KEY

Sensitive. The shared secret, in base64 — url-safe, unpadded url-safe or standard are all accepted, the same three forms acme-proxy upstream register takes. A value that decodes as none of them is a startup error, as is setting either key without the other. Prefer the environment variable to a file on disk, and clear it once registration has succeeded: while it stays non-empty the server logs a signer_relay_eab_secret_in_config warning on every startup.

EAB considerations

An External Account Binding (EAB) credential is a one-time use token that authorizes a single newAccount request and is useless afterwards — registration itself only ever runs once, guarded by the .kid sidecar that ends up next to account_key_path. There are two ways to supply it:

The Admin CLI, registering the proxy with the upstream CA out of band:

acme-proxy upstream register --profile prod --eab-kid "..." --eab-hmac-key-file /path/to/secret

The secret may also be piped on stdin. It is never accepted as a command-line argument, because argv is visible to every user on the host via ps. Nothing about this credential is ever written to disk — only the resulting account kid persists.

--profile is required whenever the configuration defines more than one profile: [signer] is a per-profile section, so registering “the upstream” without saying which one would be registering nothing.

[signer.relay.eab] in configuration, read by acme-proxy serve itself on the first startup with no .kid sidecar yet:

[signer.relay.eab]
kid = "..."
hmac_key = "..."   # base64: url-safe, unpadded url-safe, or standard

This is the trade-off the CLI path exists to avoid: a bootstrap secret sitting in configuration for the life of the server, in exchange for not needing a separate imperative step — useful when config.toml is already populated by a secrets manager or a templated deployment. Once registration succeeds, serve logs a signer_relay_eab_secret_in_config warning on every startup for as long as hmac_key stays non-empty, the same treatment challenge.bypass and ipam.netbox.insecure_skip_verify get — clear it out once acme-proxy upstream show confirms a kid is stored. Setting kid without hmac_key, or vice versa, is a startup error.

If the upstream requires EAB and neither mechanism supplies a working credential, acme-proxy serve fails at startup naming both.

The outbound client validates the upstream’s TLS certificate against webpki-roots — unlike challenge.http_01, which deliberately does not validate the responder’s certificate. Here the certificate is the only thing identifying the CA being handed your CSRs.

Custom Script Signer

The custom signer backend delegates certificate issuance, revocation and metadata retrieval to an external script (Bash, Python, Go, …). Use it to integrate acme-proxy with legacy PKI systems, HSMs, or internal APIs that do not speak ACME natively.

acme-proxy still serves ACME to its own clients and still enforces domain control, filters and EAB; only the signing step is handed off.

Configuration

[signer]
backend = "custom"

[signer.custom]
script_path = "/usr/local/bin/legacy-pki-bridge.sh"
timeout_ms = 15000
args = []
supports_crl = false
supports_renewal_info = false

Reference

script_path (String) — Default: "" | Env: ACME_PROXY_SIGNER__CUSTOM__SCRIPT_PATH

Path to the executable. An empty value is a startup error once this backend is selected.

timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_SIGNER__CUSTOM__TIMEOUT_MS

Budget for one invocation. issue and revoke run in the job queue, so there it bounds one attempt. The crl and renewal_info hooks answer a request inline, so while either is enabled this must stay below server.request_timeout_ms — the server refuses to start otherwise.

args (Array) — Default: [] | Env: ACME_PROXY_SIGNER__CUSTOM__ARGS

Static arguments passed on every invocation. These are the only command-line arguments the script receives; the hook is not passed as an argument.

supports_crl (Boolean) — Default: false | Env: ACME_PROXY_SIGNER__CUSTOM__SUPPORTS_CRL

Whether the script implements the crl hook. While false, the hook is never invoked — no process is spawned at all — and GET /crl has nothing to serve.

With it on, the answer is cached for a minute and re-served, so a burst of requests to the unauthenticated GET /crl is one run of the script rather than one per request. At most four read hooks (crl and renewal_info together) run at a time, whatever the request rate.

supports_renewal_info (Boolean) — Default: false | Env: ACME_PROXY_SIGNER__CUSTOM__SUPPORTS_RENEWAL_INFO

Whether the script implements the renewal_info hook. While false, the hook is never invoked and GET /renewalInfo/{certID} falls back to the server’s own local estimate.

These two flags default to false and gate the hooks entirely. A renewal_info hook written without setting supports_renewal_info = true will simply never run, with no error to explain why.

Hooks

The hook is selected by the ACME_SIGNER_HOOK environment variable, not by a command-line argument. Every hook receives a JSON object on stdin.

HookstdinstdoutGated by
issue{"hook":"issue","order_id":…,"identifiers":[{"type":"dns","value":"…"}],"csr_der_base64":"…"}PEM certificate chain, leaf firstalways
revoke{"hook":"revoke","cert_der_base64":"…","reason":<int|null>}ignoredalways
crl{"hook":"crl"}raw DER of the CRLsupports_crl
renewal_info{"hook":"renewal_info","cert_der_base64":"…"}see belowsupports_renewal_info

csr_der_base64 and cert_der_base64 are standard base64 of the DER bytes — not PEM, and not ACME’s base64url.

issue

Exit codes are the contract:

  • 0 — stdout is the PEM chain (leaf first, issuers after). Trailing whitespace is trimmed and exactly one newline re-appended, since a strict parser needs a newline after the final -----END CERTIFICATE-----.
  • 3 — reserved: the CSR is bad. The order becomes invalid with a badCSR error, which the client reads when it polls. Do not use this exit code for backend failures.
  • anything else — an internal failure. The issuance is retried under the job queue’s attempt budget, and the order is marked invalid (terminal, but pollable) once it runs out.

The script runs in the worker role, in the signer_issue job finalize queues: the client is answered processing and polls until the certificate is there. It never runs inside a client’s request, so it may take as long as its timeout_ms without holding one open.

The order’s requested notBefore/notAfter (RFC 8555 §7.4) are not passed to the script — there is no contract for it, and inventing one would break existing scripts. Your script decides validity on its own. (The local_ca backend does honour them, clamped.)

revoke

Exit 0 means revoked. Any non-zero exit is an internal failure, and acme-proxy then leaves the order un-revoked so the operation can be retried — the CA-side action is authoritative.

Revocation must be idempotent: acme-proxy may call this hook for a certificate your PKI already considers revoked, and that must succeed rather than error.

renewal_info

stdout drives RFC 9773:

  • empty — no opinion; the server falls back to its own estimate.
  • <start> <end> — the renewal window, as epoch seconds.
  • <start> <end> <explanationURL> — additionally supplies RFC 9773 §4.2’s optional explanationURL. The URL is last and optional so an existing two-token script keeps working unchanged.

Any other token count, or a non-integer timestamp, is an internal failure.

crl

stdout is the raw DER of the CRL, served by GET /crl. Empty stdout means “no CRL”. Failures here are logged and swallowed — a broken crl hook degrades to no CRL rather than taking the endpoint down.

Environment variables

VariableSet forValue
ACME_SIGNER_HOOKevery hookissue, revoke, crl or renewal_info
ACME_SIGNER_ORDER_IDissueThe order being finalized
ACME_SIGNER_IDENTIFIERSissueComma-joined identifier values
ACME_SIGNER_REASONrevokeRFC 5280 reason code, empty when none given

There is no ACME_SIGNER_PROFILE; the signer backend is never told which profile it is serving. (Backends are shared between profiles with identical [signer] configuration, so there would not always be one answer.)

Security & process isolation

  1. Environment clearing: env_clear() is called. The script inherits a minimal PATH plus the ACME_SIGNER_* variables above — nothing else. The server’s own environment may hold the NetBox token, SMTP password or RFC 2136 TSIG key, and a signing script has no business reading them.
  2. Zombie protection: the child runs with kill_on_drop(true) under a tokio::time::timeout. A timeout alone only drops the future, so without this a hung script would outlive its deadline and leak a process per request.
  3. Failure reporting: on a non-zero exit, the first non-empty line of stdout (falling back to stderr) is used as the error detail.

See Custom Plugins Examples for a complete script.

Challenge Validation

A challenge is how a client proves to acme-proxy that it controls the identifiers it is asking for. Every authorization created by newOrder carries one challenge per enabled type; satisfying any one of them makes the authorization valid, and an order becomes ready once every authorization is.

acme-proxy implements all three challenge types RFC 8555 and RFC 8737 define:

  • http-01 — serve a token over HTTP.
  • dns-01 — publish a TXT record.
  • tls-alpn-01 — present a special certificate in a TLS handshake.

This is validation acme-proxy performs against its own clients. It is separate from how the relay signer backend satisfies an upstream CA’s challenges on your behalf; the two are configured independently and need not match.

Configuration

[challenge]
# Types offered, in this order. Empty or unknown = startup error.
enabled = ["http-01", "dns-01"]

# Skip validation entirely. Testing only.
bypass = false

# Budget for one validation attempt.
timeout_ms = 5000

[challenge] is a per-profile section: each endpoint can offer a different set of types. See Profiles & Routing.

Reference

enabled (Array) — Default: ["http-01"] | Env: ACME_PROXY_CHALLENGE__ENABLED

Which types each new authorization offers, and in what order. Clients pick one. Valid values are http-01, dns-01 and tls-alpn-01. An empty list, or an unrecognised name, is a startup error — even when bypass is on, because a bypassing server still has to advertise a challenge for the client to trigger.

bypass (Boolean) — Default: false | Env: ACME_PROXY_CHALLENGE__BYPASS

Mark a triggered challenge valid immediately, with no network check. With this on, [filter] is the only access control there is — which is why it is not the default.

timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_CHALLENGE__TIMEOUT_MS

Budget for one validation attempt, applied at the registry level whatever the type. It bounds a job attempt in the runner, not a request — see Validation runs in the job queue — and it must stay below server.request_timeout_ms, which server::profile::build_all refuses to start otherwise.

max_in_flight_per_account (Integer) — Default: 32 | Env: ACME_PROXY_CHALLENGE__MAX_IN_FLIGHT_PER_ACCOUNT

How many of one account’s challenges may be validating at once. 0 is no limit.

A validation is queued work that reaches out to an address the client named, and an account can create as many orders as it likes. Without a cap, one busy — or hostile — account fills the runner with outbound probes while signings, revocations and CRL regenerations wait behind them. A trigger over the cap is answered 429 rateLimited with a Retry-After, and the challenge is left pending: the client re-triggers it once one of its own validations has settled, and loses no order.

Raise it for a deployment that renews many certificates at once from one account; the ceiling that matters is jobs.max_concurrent, which is how many validations actually run in parallel.

Per-type keys live under [challenge.http_01] and [challenge.tls_alpn_01]; see those pages. dns-01 has no table of its own — it is governed by dns.resolver.

Bypass is not a shortcut

With challenge.bypass = true, [filter] is the only access control the server has. Anyone who can reach the endpoint can obtain a certificate for any name it will accept, without proving anything.

Bypass exists for two legitimate cases: local testing (as in the Quick Start), and a deployment where an IPAM-backed filter such as ipam is genuinely the authority on which host may hold which name, making a network round-trip redundant.

It defaulted to true early in this project’s life. That was reconsidered: an empty filter.rules plus the default bind on every interface made the combination an open CA, so the default is now false.

The two state machines

An authorization and its challenges are separate objects with separate statuses, and the edge between them is the one worth internalising: an authorization becomes valid as soon as any one of its challenges does. The siblings stay pending for ever and that is correct.

stateDiagram-v2
    direction LR
    state "authorization" as A {
        [*] --> a_pending: created with the order
        a_pending --> a_valid: any one challenge valid
        a_pending --> a_invalid: a challenge failed
        a_pending --> a_deactivated: §7.5.2
        a_valid --> a_deactivated: §7.5.2
        a_pending --> a_expired: expires passed
    }
    state "challenge" as C {
        [*] --> c_pending: created with the authorization
        c_pending --> c_valid: proof accepted
        c_pending --> c_invalid: proof refused — terminal
    }
    c_valid --> a_valid: promotes its parent

a_invalid and c_invalid are terminal: the client must create a new order, and re-triggering the same challenge will not retry it. Deactivation is the operator- or client-initiated exit — see Deactivation below.

Validation runs in the job queue

Triggering a challenge with POST /chall/{id} does not perform the check. It claims the challenge, writes a challenge_validate job and answers straight away with the challenge in the processing state, plus a Retry-After. The job runner performs the outbound check and records the verdict.

RFC 8555 has the states for exactly this. §7.1.6: challenges “transition to the processing state when the client responds to the challenge”. §8.2 pairs that with a Retry-After on the challenge resource, which is what the client polls against. certbot, acme.sh and lego all poll.

sequenceDiagram
    participant C as ACME client
    participant P as acme-proxy
    participant W as job runner
    participant T as The name being proven<br/>(port 80 / 443 / DNS)
    participant D as SQLite

    C->>P: POST /chall/{id}
    P->>D: claim: pending → processing
    P->>D: enqueue challenge_validate
    P-->>C: 200 + challenge object (processing)<br/>+ Retry-After + Link: rel="up"
    W->>D: claim the job
    rect rgb(240, 240, 240)
        Note over W,T: in the runner, under challenge.timeout_ms
        W->>T: fetch token / query TXT / TLS handshake
        T-->>W: answer, or timeout
    end
    W->>D: one transaction:<br/>challenge + authorization + order
    Note over D: "is every authorization valid?"<br/>is read INSIDE this transaction
    C->>P: POST /chall/{id} (retry — not a state change)
    P-->>C: 200 + challenge object — valid or invalid

Three consequences:

  • challenge.timeout_ms bounds a job attempt, not an HTTP request. It is therefore independent of server.request_timeout_ms, and the server no longer refuses to start when it exceeds it.
  • A client that points a name at an unreachable host no longer occupies one of server.max_concurrent_requests while the server waits for it.
  • The server still needs egress to the client. For http-01 and tls-alpn-01 it must be able to open a connection back to the machine requesting the certificate — a common source of “the order just sits at pending” in firewalled networks.

A validation is attempted once: a check that ran records its verdict, pass or fail. The job is retried only when the attempt could not happen at all — the database was unreachable, or the endpoint’s profile is not mounted by the process that picked the row up.

Both outcomes are 200

A validation failure returns 200 OK with the challenge object, its status set to invalid and an error member describing what went wrong. It is not a 4xx.

This follows RFC 8555 §7.5.1, and it is load-bearing: certbot’s acme library surfaces an HTTP error status as a transport failure, which would obscure the actual reason the challenge failed. Read the challenge object’s status, not the HTTP status.

Responses also carry a Link: rel="up" header pointing at the authorization, which that same library requires.

A challenge that reaches invalid is terminal. The client must create a new order; re-triggering the same challenge will not retry it.

Wildcards

A wildcard identifier such as *.example.com is accepted only when dns-01 is among enabled — it is the only challenge type that can prove control of a whole subtree. Otherwise newOrder refuses with rejectedIdentifier, naming dns-01.

For a wildcard identifier:

  • The authorization is created on the base name (example.com), with "wildcard": true in the authorization object.
  • It offers dns-01 alone, even if other types are enabled.

Ordering example.com and *.example.com together therefore produces two authorizations on the same base name, and the TXT record for each goes to the same _acme-challenge.example.com. acme-proxy matches any TXT record at that name, so publishing both values side by side works.

Only a single leading *. is legal. *.*.example.com and foo.*.example.com are rejected as malformed.

What happens on success

Success is committed as one transaction covering the challenge, its authorization, and the order — including the “is every authorization now valid?” read that promotes the order to ready. Doing that read inside the transaction is deliberate: two concurrent validations of one order could otherwise each read before the other’s write landed, and neither would promote the order.

Note the promotion depends on every authorization being valid, not every challenge. An authorization with three challenges needs only one of them.

Deactivation

A client can deactivate an authorization it no longer wants by POSTing {"status": "deactivated"} to the authorization URL (§7.5.2). If the order had already reached ready, it is demoted back to pending.

Deactivation is refused once the order is valid — at that point the certificate exists, and revocation, not deactivation, is what undoes it.

HTTP-01

The default challenge type. The client serves a token at a well-known path over plain HTTP, and acme-proxy fetches it.

Not to be confused with the http01 upstream strategy. This page is about acme-proxy validating its own clients. The relay signer has a challenge_strategy = "http01" that runs the same challenge type in the opposite direction — acme-proxy serving the file, to prove itself to an upstream CA. The two share the well-known path and nothing else, and are configured independently.

How it works

  1. The client is given a random token on the challenge object.
  2. It computes the key authorization: token + "." + base64url(SHA256(JWK thumbprint of its account key)).
  3. It serves that string at http://<identifier>/.well-known/acme-challenge/<token>.
  4. It triggers the challenge with POST /chall/{id}.
  5. acme-proxy resolves the identifier, fetches the URL, trims whitespace from the body, and compares it to the key authorization it computed independently.

The body is compared to the key authorization verbatim — unlike dns-01, which compares a SHA-256 digest of it. Serving the digest here is a common mistake when hand-rolling a responder.

Every ACME client implements a responder for this type, usually via a --standalone mode or a webroot.

Limitations

  • No wildcards. *.example.com cannot be proven this way; use dns-01.
  • The name must resolve, and be reachable. acme-proxy connects to the client, using dns.resolver to find it. A host behind a firewall that blocks inbound port 80 from the proxy cannot be validated.

Configuration

[challenge]
enabled = ["http-01"]

[challenge.http_01]
port = 80
https_port = 443
follow_redirects = true
max_redirects = 5
max_response_bytes = 4096

Reference

port (Integer) — Default: 80 | Env: ACME_PROXY_CHALLENGE__HTTP_01__PORT

Port the challenge is fetched from. RFC 8555 fixes this at 80 for the public Internet; it is configurable here because internal deployments frequently cannot bind low ports.

https_port (Integer) — Default: 443 | Env: ACME_PROXY_CHALLENGE__HTTP_01__HTTPS_PORT

Port used when a redirect sends the fetch to https.

follow_redirects (Boolean) — Default: true | Env: ACME_PROXY_CHALLENGE__HTTP_01__FOLLOW_REDIRECTS

Follow 3xx responses. Required by the specification, and commonly needed in practice — many hosts redirect all HTTP to HTTPS.

max_redirects (Integer) — Default: 5 | Env: ACME_PROXY_CHALLENGE__HTTP_01__MAX_REDIRECTS

Hop limit before the validation fails.

max_response_bytes (Integer) — Default: 4096 | Env: ACME_PROXY_CHALLENGE__HTTP_01__MAX_RESPONSE_BYTES

Cap on how much of the response body is read. A key authorization is under 100 bytes; this exists so a client cannot make the server read an unbounded stream.

The TLS certificate presented after a redirect to HTTPS is not validated (RFC 8555 §8.3) — at validation time the client by definition does not yet have a trusted certificate for the name.

Redirects are an SSRF surface

Following redirects means a client can steer the server’s fetch at an address of its choosing. Boulder’s usual mitigation — refusing to connect to RFC 1918 space — cannot apply here, because serving private networks is the entire point of this server.

What contains it instead:

  • Only http and https schemes are followed.
  • Only the two configured ports (port and https_port) are connected to.
  • At most max_redirects hops.
  • The shared challenge.timeout_ms bounds the whole attempt, redirects included.
  • follow_redirects = false turns the surface off entirely.
  • Nothing the fetch learned is echoed back to the client. The error a client sees names its kind — connection or unauthorized — and the identifier, and says the server’s log has the reason. The status, the body length, the socket error and any redirect target stay in challenge_validation_failed; a truncated body preview is logged at debug. This is what stops the challenge from becoming a read primitive or a port scanner against your internal network. What remains is the kind itself: a client can still tell “nothing answered” from “something answered wrongly”.

If your threat model does not tolerate this, disable redirects, or use dns-01, which makes no outbound connection to the client at all.

Troubleshooting

Look for these events in the log:

EventMeaning
challenge_http_01_loadedThe fetch succeeded; a body was read.
challenge_http_01_matchedThe body matched. Validation passed.
challenge_http_01_mismatchThe responder answered with the wrong content. Check that it is serving the key authorization, not just the token, and not the digest.
challenge_http_01_redirectA redirect was followed; the target is logged.
challenge_validation_failedThe attempt failed. The detail says whether it was a connection error, a timeout, or a mismatch, with the status or socket error — the client is told only the kind.

A connection failure usually means one of: the name does not resolve through dns.resolver; nothing is listening on port; or a firewall blocks the proxy’s egress to the client.

DNS-01

The client publishes a TXT record proving control of the domain. This is the only challenge type that can authorize a wildcard, and the only one that requires no inbound connectivity to the client at all.

How it works

  1. The client is given a random token.
  2. It computes the key authorization: token + "." + base64url(SHA256(JWK thumbprint)).
  3. It publishes a TXT record at _acme-challenge.<identifier> whose value is base64url(SHA256(keyAuthorization)).
  4. It triggers the challenge, and acme-proxy queries that name through dns.resolver.

The value is the digest, not the key authorization. This is the single most common mistake when writing a DNS responder by hand: http-01 serves the key authorization verbatim, dns-01 serves the base64url-encoded SHA-256 of it. The two are not interchangeable.

acme-proxy matches any TXT record present at the name. That is required rather than lenient: an order covering both example.com and *.example.com produces two authorizations that publish two different values to the same _acme-challenge.example.com, and both must be able to validate.

Configuration

There is no [challenge.dns_01] table — unlike the other two types, dns-01 has no per-type key. Enable it, and point dns.resolver wherever your authoritative data lives:

[challenge]
enabled = ["dns-01"]

[dns]
# host:port of the nameserver every lookup goes through.
resolver = "10.0.0.53:53"

Leaving dns.resolver unset uses the system configuration (/etc/resolv.conf).

The resolver is deliberately uncached

The shared resolver performs no caching. This matters here more than anywhere else: a client typically publishes its TXT record and triggers the challenge seconds later, and a cached negative answer — an NXDOMAIN or empty NOERROR from the moment before publication — would defeat validation for the whole negative-TTL window.

(The one component that does cache is filter.reverse_dns, which builds its own resolver. PTR lookups for an address that keeps connecting are exactly what a cache is for.)

What this does not remove is propagation delay in your own DNS infrastructure. If dns.resolver points at a recursive resolver rather than the authoritative server, the record still has to reach it. Pointing straight at the authoritative nameserver is the reliable choice for an internal deployment.

Wildcards

[challenge]
enabled = ["dns-01"]     # required for *.example.com to be accepted at all

Without dns-01 enabled, newOrder refuses a wildcard identifier with rejectedIdentifier. With it enabled:

  • the authorization is created on the base name, flagged "wildcard": true;
  • it offers dns-01 and nothing else, regardless of what else is enabled;
  • the TXT record still goes to _acme-challenge.example.com — there is no _acme-challenge.*.example.com.

Client support

Every major client implements dns-01, but each needs a provider plugin or hook script to write the record: certbot’s --manual with an auth hook or a DNS plugin, acme.sh --dns dns_<provider>, lego’s --dns <provider>.

If you would rather your clients did not hold DNS credentials, that is exactly what the relay signer backend exists for: your clients prove control to acme-proxy however you like, and acme-proxy holds the single RFC 2136 TSIG key that answers the upstream CA.

Troubleshooting

EventMeaning
challenge_dns_01_loadedTXT records were retrieved for the name.
challenge_dns_01_matchedOne of them matched. Validation passed.
challenge_validation_failedNo record matched, or the lookup failed.

If the lookup returns nothing, check in this order: the record exists at _acme-challenge.<name> (not at the name itself); its value is the digest; dns.resolver can actually see the zone; and the record has finished propagating to whatever dns.resolver points at.

A TXT record split into multiple character-strings is reassembled by concatenation before comparison, so a long value chunked by your DNS server is handled correctly.

TLS-ALPN-01

Defined by RFC 8737. The client proves control by answering a TLS handshake on port 443 with a special self-signed certificate, rather than by serving anything over HTTP.

Its appeal is operational: validation happens entirely within the TLS layer, so a host that terminates TLS but serves no plain HTTP at all can still be validated, and nothing needs to be routed on port 80.

How it works

  1. The client is given a random token and computes the key authorization as usual.
  2. It generates a self-signed certificate that contains:
    • a single dNSName SAN equal to the identifier being validated, and
    • a critical id-pe-acmeIdentifier extension (OID 1.3.6.1.5.5.7.1.31) whose value is SHA256(keyAuthorization).
  3. It arranges for that certificate to be presented when a handshake arrives with SNI set to the identifier and ALPN protocol acme-tls/1.
  4. acme-proxy performs that handshake and inspects the certificate it gets back.

The proof is verified without ever completing an application-layer exchange: the certificate presented during the handshake is the answer.

What acme-proxy checks

  • Exactly one dNSName SAN, matching the identifier. Not “at least one”.
  • The id-pe-acmeIdentifier extension is present and marked critical, as RFC 8737 requires.
  • Its value equals the expected digest, compared in constant time.

The presented certificate is otherwise untrusted by design — it is self-signed and issued by the entity being challenged, so there is no chain to validate.

Configuration

[challenge]
enabled = ["tls-alpn-01"]

[challenge.tls_alpn_01]
port = 443

Reference

port (Integer) — Default: 443 | Env: ACME_PROXY_CHALLENGE__TLS_ALPN_01__PORT

Port the validation handshake connects to. RFC 8737 fixes this at 443 for the public Internet; it is configurable here for internal deployments that cannot bind it.

Client support

Support is thinner than for the other two types:

  • lego implements a responder (--tls). This is what the project’s end-to-end suite uses.
  • certbot and acme.sh do not implement a tls-alpn-01 responder. Listing this type is harmless for them — they simply pick another enabled type — but it cannot be their only option.
  • Servers that manage their own certificates, such as Caddy and Traefik, generally do support it natively.

If tls-alpn-01 is the only entry in challenge.enabled, certbot and acme.sh clients will not be able to complete an order at all.

Limitations

  • No wildcards. Use dns-01.
  • Requires inbound connectivity from acme-proxy to the client on port, the same constraint as http-01.
  • The responder must serve the ACME certificate only for the acme-tls/1 ALPN protocol, and its normal certificate otherwise. Serving it unconditionally would break ordinary traffic to that host for the duration.

Troubleshooting

EventMeaning
challenge_tls_alpn_01_loadedThe handshake completed and a certificate was obtained.
challenge_tls_alpn_01_matchedThe certificate carried the right digest. Validation passed.
challenge_validation_failedHandshake failed, or the certificate did not satisfy the checks above.

A handshake that fails outright usually means the responder did not negotiate acme-tls/1 — many TLS servers simply fall through to their normal certificate when they do not recognise the ALPN protocol, and the certificate then has no id-pe-acmeIdentifier extension to find.

Filters

Filters provide access control for acme-proxy. They restrict which clients can reach the server and which identifiers (e.g. DNS names) those clients can request certificates for.

[filter] is a small policy engine with two halves:

  • a check is one named question about a request — “is this address in the management network?”, “does the inventory say this address owns this name?”. Each is a [filter.check.<name>] with a type saying which question it asks.
  • a rule is a boolean expression over check names plus what a match means. Each is a [filter.rule.<name>], and filter.rules lists the ones to evaluate, in order. First match wins.

Everything is a named check — custom included — and two checks of the same type are ordinary rather than impossible.

[filter]
rules = ["mgmt-bypass", "inventory-owned"]

[filter.check.mgmt-net]
type  = "allowed_ip"
allow = ["10.0.0.0/8"]

[filter.check.corp-names]
type  = "identifiers"
allow = ["*.corp.example.com", "corp.example.com"]

[filter.check.inventory]
type = "ipam"

[filter.rule.mgmt-bypass]
when = "mgmt-net"
then = "allow"

[filter.rule.inventory-owned]
when    = "corp-names and (inventory or mgmt-net)"
then    = "allow"
message = "this address owns no such name in the inventory"

See Policy: rules and conditions for the condition language and how rules are evaluated, and Checks for the types and their keys.

The two hooks

A check can act at two points, and both default to “pass” for a check that does not implement them:

  • the connection stage — runs on every request, before anything else. Refusal is 403 access_denied.
  • the identifier stage — runs at newOrder and again at finalize against the names projected out of the CSR. It runs after the account is resolved, so a check can bind names to an account as well as to an address. Refusal is 403 rejectedIdentifier at newOrder, 400 badCSR at finalize.

Checking again at finalize is what stops a client from passing a benign newOrder and then smuggling extra names into the CSR:

graph LR
    REQ["Any request"] --> CONN["connection stage<br/>rules over address/path checks"]
    CONN -->|deny| D403["403 access_denied"]
    CONN -->|allow| ROUTE{"which resource?"}
    ROUTE -->|"newOrder"| ID1["identifier stage<br/>names from the order"]
    ROUTE -->|"finalize"| ID2["identifier stage<br/>names from the CSR"]
    ROUTE -->|"anything else"| OK["handler"]
    ID1 -->|deny| DREJ["403 rejectedIdentifier"]
    ID2 -->|deny| DCSR["400 badCSR"]
    ID1 --> OK
    ID2 --> OK

Both stages must allow. They are evaluated independently, each over the subset of rules that can run there, so a connection-stage allow does not skip the identifier stage. A stage no rule applies to allows without consulting filter.default — otherwise a policy made entirely of identifiers checks would refuse every request before a name had been mentioned.

Available check types

TypeStage(s)Purpose
allowed_ipbothCIDR allow/deny on the client address
pathconnectionGlob allow/deny on the request path
reverse_dnsconnectionPTR lookup with optional forward confirmation
identifiersidentifiersGlob or regex allow/deny on requested names
eabidentifiersWhich EAB credential the account registered under
ipamidentifiersAsk an IPAM whether the client owns the names
custombothShell out to an operator-supplied script

allowed_ip reads nothing but the client address, which both stages carry, and that is what makes mgmt-net or inventory writable at all: the address half is still answerable at the point where the inventory is consulted.

Seeing what a policy does

acme-proxy filter show prints the resolved policy, with every condition re-parenthesized so you can see what the parser understood. acme-proxy filter explain evaluates it against a hypothetical request and reports each check’s verdict, which rule matched, and the HTTP answer. See the CLI reference.

Reference

rules (Array) — Default: [] | Env: ACME_PROXY_FILTER__RULES

Which [filter.rule.<name>] entries to evaluate, and in what order. Empty means no filtering at all — anyone who can reach the server can obtain a certificate if they satisfy the challenges (or, with challenge.bypass on, with no proof at all). The server logs a filter_disabled warning at startup saying so.

default (String) — Default: "deny" | Env: ACME_PROXY_FILTER__DEFAULT

What happens at a stage where a rule was applicable and none of them matched: "allow" or "deny". Never consulted at a stage no rule applies to.

trusted_proxies (Array) — Default: [] | Env: ACME_PROXY_FILTER__TRUSTED_PROXIES

CIDRs (or bare addresses) whose forwarded-for header is believed. A request from any other peer is attributed to its own address, and its forwarded header is ignored. Note this is a [filter] key — placing it under [server] silently does nothing.

forwarded_header (String) — Default: "x-forwarded-for" | Env: ACME_PROXY_FILTER__FORWARDED_HEADER

The header consulted for the forwarded client address.

Fail-closed semantics

Every check that depends on the client address fails closed when it cannot determine one — including in deny-only (blocklist) mode. An address the server cannot see is not “absent from the deny list, therefore fine”.

This makes one deployment detail load-bearing: the server must be served with connection info attached, which it is by default. Behind a reverse proxy, set trusted_proxies — otherwise every request is correctly, but unhelpfully, attributed to the proxy.

A check that cannot reach its authority — a DNS timeout, an unreachable inventory — is a third answer, neither pass nor fail, and becomes a retryable 500 rather than a refusal the client would believe permanent. What that means for a policy built out of or is worth reading in Policy.

Keys the policy engine replaced

Each of these is refused by name at startup, so a configuration written against the older shape stops the server rather than coming up looking configured and filtering nothing. The refusal is a diagnostic and nothing more: none of these keys is still read, and the errors themselves go away at 1.0.0. Before then a configuration key may be renamed in any release, with every such change listed in the changelog.

RemovedReplacement
filter.enabledDeclare each filter as a [filter.check.<name>] with a type, write a [filter.rule.<name>] naming them, list it in filter.rules.
filter.exempt_pathsA path check plus a rule — which can also combine the path with an address, and can glob.
filter.custom_enabledcustom is an ordinary check type; filter.rules already says which run and in what order.
[filter.allowed_ip], [filter.reverse_dns], [filter.identifiers], [filter.custom.<name>]The type’s keys move onto its [filter.check.<name>] entry.
[filter.netbox][ipam.netbox], read by a type = "ipam" check — see IPAM.

Policy: rules and conditions

A rule is a boolean expression over check names plus what a match means. filter.rules lists the rules to evaluate, in order, and the first match decides.

[filter]
rules   = ["mgmt-bypass", "inventory-owned"]
default = "deny"

[filter.rule.mgmt-bypass]
when = "mgmt-net"
then = "allow"

[filter.rule.inventory-owned]
when    = "corp-names and (inventory or mgmt-net)"
then    = "allow"
message = "this address owns no such name in the inventory"
mode    = "enforce"

The condition language

expr   := term ( "or" term )*
term   := factor ( "and" factor )*
factor := "not" factor | "(" expr ")" | name
name   := [a-z0-9-]+

not binds tightest, then and, then or; and and or are left-associative, so a or b and c means a or (b and c). The three keywords are matched case-insensitively and cannot be used as check names. A parse error names the column it gave up at:

filter.rule.r.when: expected a check name, `not` or `(` at column 9 in "net and )"

acme-proxy filter show re-prints every condition with the grouping made explicit, which is the quickest way to confirm the parser read what you meant.

Where a rule runs

A rule is evaluated at the intersection of the stages its checks can decide at — never the union. Evaluating a rule at a stage where one of its checks cannot run would silently treat that check as passing and change the boolean answer, so the intersection is the only composition that cannot lie.

The consequence worth knowing: a rule combining a connection-only check with an identifiers-only one has no stage at all. That is a startup error naming both sides, not a rule that quietly never fires.

filter.rule.strict combines `has-ptr` (connection only) and `corp-names`
(identifiers only), so there is no point in a request where both can be
evaluated. Give `has-ptr` stages = ["identifiers"] if it can decide there, or
split the rule in two.

Both stages must allow, and each evaluates its own applicable subset independently. A stage no rule applies to allows without consulting filter.default.

When a check cannot decide

A check has three possible answers, not two: it passed, it failed, or it could not decide. The third covers a DNS timeout, an unreachable inventory, a script that would not spawn — cases where the server learned nothing, which is not the same as learning “no”.

Conditions combine them with three-valued logic, where an unknown propagates only if it could change the answer:

andor
pass, passpasspass
pass, failfailpass
fail, failfailfail
fail, unknownfailunknown
pass, unknownunknownpass
unknown, unknownunknownunknown

The two bold rows are the point. mgmt-net or inventory keeps working through an inventory outage, because a disjunction whose other side already passed does not care what the unknown would have been. inventory on its own still becomes a retryable 500 — the or buys resilience for the addresses it names and nothing more.

The same principle applies at the rule level. A rule whose condition came back unknown is not skipped: it is remembered, and once the policy reaches an answer it is asked whether that would have mattered. If the unknown rule’s effect differs from the effect actually reached, the whole stage is a 500. If it agrees, the answer stands. This is what stops rule order from deciding whether an outage is survivable.

Warn mode

mode = "warn" makes a matching rule log filter_rule_warned and not decide — evaluation continues to the next rule. It is how a tightened policy is rolled out: deploy it in warn mode, watch for the event, and switch to enforce once no legitimate client trips it.

[filter.rule.inventory-owned]
mode = "warn"

A policy of nothing but warn rules therefore falls through to filter.default.

Because rules are a map rather than an array of tables, a profile can dry-run one rule and inherit the rest:

[profiles.staging.filter.rule.inventory-owned]
mode = "warn"

What the client is told

In order: the matching rule’s message if it has one; otherwise the first check that actually refused, in evaluation order; otherwise a generic sentence. A message is the way to say “ask the network team” instead of exposing which check bit.

A check that could not decide never reaches the client — the response is a plain 500, and the specifics stay in the logs.

Reference

filter.rule.<name>.when (String) — Required | Env: ACME_PROXY_FILTER__RULE__<NAME>__WHEN

The condition, in the language above. Empty is a startup error: a rule with no condition is what filter.default is for.

filter.rule.<name>.then (String) — Required | Env: ACME_PROXY_FILTER__RULE__<NAME>__THEN

"allow" or "deny". No default — a rule that does not say what a match means is one whose author has not finished writing it.

filter.rule.<name>.message (String) — Default: "" | Env: ACME_PROXY_FILTER__RULE__<NAME>__MESSAGE

Shown to the client verbatim in place of whichever check failed.

filter.rule.<name>.mode (String) — Default: "enforce" | Env: ACME_PROXY_FILTER__RULE__<NAME>__MODE

"enforce" or "warn". A warn rule matches, logs and does not decide.

Startup refusals

Every one of these stops the server rather than producing a policy that does not mean what it says.

ConfigurationRefusal
filter.rules names a rule with no [filter.rule.<name>]names the missing entry
when names a check with no [filter.check.<name>]names the missing check
A rule is defined but filter.rules is emptysays to list the rules to evaluate
when will not parsenames the column, and quotes the expression
then missing, or not allow/denynames the value
mode not enforce/warnnames the value
filter.default not allow/denynames the value
A rule’s checks share no stagenames both sides and suggests stages
A check is named and, or or notsays they are the language’s own words

Checks

A [filter.check.<name>] is one named question about a request. It takes a type saying which question, plus that type’s own keys:

[filter.check.mgmt-net]
type  = "allowed_ip"
allow = ["10.0.0.0/8"]

[filter.check.corp-names]
type  = "identifiers"
allow = ["*.corp.example.com", "corp.example.com"]

Two checks of the same type are ordinary — two identifier lists with different rules, an address list per network, several script hooks — which is the main thing the older filter.enabled shape could not express.

Naming

Each name must match ^[a-z0-9-]+$: lowercase letters, digits and -, the same restriction profile names have. The reason is that a name is also an environment variable segment (ACME_PROXY_FILTER__CHECK__<NAME>__…), and the config crate lowercases those, so MgmtNet in a file and MGMTNET in the environment would silently become two entries instead of one overriding the other.

and, or and not are the condition language’s own words and cannot name a check.

Only what a rule names is built

A check defined here but mentioned by no selected rule is never constructed. It opens no connection, spawns no client and validates nothing; it is reported once at startup as filter_check_unused.

That is deliberate, and it is what makes profile inheritance usable: a global [filter] section can carry a library of checks, every profile inherits all of them, and each profile’s filter.rules picks the subset it actually wants without paying for — or failing startup on — the rest.

Keys by type

type and stages are universal. Everything else belongs to one type, and setting a key that belongs to a different type is a startup error naming both — the keys are one flat namespace, so without that check script_path on an allowed_ip check would simply be read by nothing.

TypeKeysPage
allowed_ipallow, denyallowed_ip
pathallow, denypath
reverse_dnsallow, deny, allow_regex, deny_regex, require_forward_confirm, timeout_msreverse_dns
identifiersallow, deny, allow_regex, deny_regex, allowed_types, allow_wildcardsidentifiers
eaballow, deny, allow_regex, deny_regex, kids, require_activeeab
ipam(none — configured by the [ipam] section)ipam
customscript_path, timeout_ms, pass_stdin, argscustom

Defaults, where a type has one: require_forward_confirm = true, timeout_ms = 2000 for reverse_dns and 5000 for custom, pass_stdin = true, allow_wildcards = false, require_active = false, allowed_types = ["dns", "cn"].

Every list defaults to empty, and empty always means “this type’s natural default”, never “none”. That is not a style choice: an unset list environment variable arrives as an empty list rather than as absent, so stages = [] has to mean “infer” and allowed_types = [] has to mean ["dns", "cn"].

Matching: globs first, regexes on request

The name-matching checks take globs in allow/deny, where * matches one label:

  • *.example.com matches a.example.com; it does not match a.b.example.com, and it does not match example.com. List the bare name too, exactly as you would in a certificate.
  • Everything else is literal. No ?, no character classes, no **.
  • Matching is case-insensitive.

allow_regex/deny_regex take regexes instead, automatically anchored as ^(?:…)$, and are unioned with the globs — so a policy can be mostly globs with one regex where a glob will not do. Anchoring is not optional: the regex crate searches rather than matches, so an unanchored example\.com would also accept example.com.evil.net, which is precisely the bypass an allowlist exists to prevent. Write .*\.example\.com for a suffix.

On allowed_ip the two lists are CIDRs or bare addresses instead; on path they are path globs where * stops at /.

Allow and deny

One rule, shared by every check that has the pair:

  • deny is checked first and wins. Plain membership, not longest-prefix-match: a /32 in allow does not beat a /8 in deny.
  • An empty allow imposes no constraint, so a deny-only configuration is a working blocklist rather than a list that refuses everything.

Which gives three usable shapes: allow-only (a strict allowlist), deny-only (a blocklist, everything else served), or both (an allowlist with holes punched in it).

Stages

There are two hook points — the connection, and the identifiers — and each type answers at the ones it can:

TypeDefault stagesCapable of
allowed_ip, customconnection + identifiersthe same
pathconnectionconnection
reverse_dnsconnectionconnection + identifiers
identifiers, ipam, eabidentifiersidentifiers

reverse_dns is the one whose default is narrower than its capability: it could answer at the identifier stage from the same address, but a PTR plus forward-confirmation exchange at newOrder and again at finalize triples the lookups for an answer that has not changed. Opt in when you need it in an identifier-stage rule:

[filter.check.has-ptr]
type   = "reverse_dns"
stages = ["identifiers"]

Naming a stage the type cannot serve is a startup error saying why — an ipam check at the connection stage would query the inventory on every newNonce.

That allowed_ip answers at both stages is load-bearing rather than incidental: it is what makes mgmt-net or inventory a rule that can be evaluated at all, since the address half must still be answerable at the point where the names are known.

Reference

filter.check.<name>.type (String) — Required | Env: ACME_PROXY_FILTER__CHECK__<NAME>__TYPE

Which check type this instance is: allowed_ip, path, reverse_dns, identifiers, eab, ipam or custom. An unknown value is refused by name, as is the old netbox, which became ipam with its settings in [ipam.netbox].

filter.check.<name>.stages (Array) — Default: [] | Env: ACME_PROXY_FILTER__CHECK__<NAME>__STAGES

Override where this instance decides: "connection", "identifiers", or both. Empty infers from the type, per the table above.

Allowed IP Filter

The allowed_ip filter provides network-level access control based on the client’s IP address. It can operate as an allowlist, a blocklist, or both.

How it works

This filter implements standard allow/deny semantics using CIDR matching:

  • Deny wins: The deny list is checked first. If a client matches any CIDR in the deny list, the request is immediately rejected.
  • Allow list: If an allow list is provided, the client must match at least one CIDR. If the list is empty, no allow constraint is imposed (functioning purely as a blocklist).

Client IP resolution & proxies

acme-proxy resolves the client IP securely. If the server is behind a reverse proxy, the proxy must be trusted. Trusted proxies are declared under [filter], not [server]:

[filter]
trusted_proxies = ["10.0.0.0/8"]
forwarded_header = "x-forwarded-for"   # the default

Get the section right. Unknown keys are ignored rather than rejected, so trusted_proxies written under [server] is silently dropped — no warning, no startup error. The forwarded header is then never believed, and every request is attributed to the reverse proxy’s own address instead of the client’s.

When a connection originates from a trusted proxy, acme-proxy walks the forwarded-for header right-to-left, skipping trusted hops until it finds the true client IP. Requests arriving from an address that is not in trusted_proxies are attributed to their peer address, and any forwarded header they carry is ignored — which is what makes a spoofed X-Forwarded-For header useless. IP addresses are canonicalized internally (IpAddr::to_canonical()), meaning IPv4-mapped IPv6 addresses (e.g., ::ffff:192.168.1.1) are properly treated as IPv4.

Important (Fail Closed): The filter subsystem operates with strict fail-closed semantics. If the client IP cannot be determined (e.g., misconfigured reverse proxy or missing TapIo wrapper), ConnectionContext::require_client_ip() fails. The filter will deny access rather than assuming the client is safe.

Configuration

[filter]
rules = ["internal-only"]

[filter.check.internal-nets]
type = "allowed_ip"
# Deny external bad actors
deny = ["203.0.113.9", "198.51.100.0/24"]
# Allow internal networks
allow = ["192.168.1.0/24", "10.0.0.0/8", "fd00::/8"]

[filter.rule.internal-only]
when = "internal-nets"
then = "allow"

allow and deny take CIDRs or bare addresses (a bare address becoming a host route), IPv4 or IPv6. Both empty while a rule names the check is a startup error: an empty allow imposes no constraint, so the check would accept everything, and an operator who configured it did not mean to turn on something inert.

The shared allow/deny semantics, and the keys themselves, are documented under Checks.

This check answers at both stages, since it reads nothing but the client address — which is what lets a rule say internal-nets or inventory and still have the address half answerable once the requested names are known.

Path Check

type = "path" matches the request path, so a rule can be about what is being asked for as well as who is asking.

[filter.check.public-paths]
type  = "path"
allow = ["/crl"]

[filter.rule.public]
when = "public-paths"
then = "allow"

Connection stage only. By the identifier stage the path is always /newOrder or /finalize/{id}, so a rule combining this with a name check would be asking a question with a constant answer.

The /crl and /ca.pem trap

Both are served by the profile router, which means they sit behind the filter policy exactly like /newOrder does. Turn on an address-based check without accounting for them and two things break quietly:

  • every relying party outside your allowlist loses revocation checking, and relying parties are precisely not the ACME clients you allowlisted;
  • a host that has not installed the root yet cannot fetch it — which is the one moment it needs to, and the refusal looks like the CA being down.

So any address-based policy wants a companion rule:

[filter]
rules = ["public", "mgmt-only"]

[filter.check.public-paths]
type  = "path"
allow = ["/crl", "/ca.pem"]

[filter.check.mgmt-net]
type  = "allowed_ip"
allow = ["10.0.0.0/8"]

[filter.rule.public]
when = "public-paths"
then = "allow"

[filter.rule.mgmt-only]
when = "mgmt-net"
then = "allow"

Because public comes first and first match wins, both are served to anyone while everything else still requires the management network.

Server-level routes — GET /health, GET /, and the http-01 responder — are served by the root router, which no profile’s policy ever sees. They are already unfiltered and need no entry here.

Paths are profile-stripped

Matching is against the path with the /profile/<name> prefix removed, so /directory — not /profile/default/directory — is the value to list, and one check covers every endpoint the process serves.

Globs

* matches one or more characters other than /, i.e. one path segment:

GlobMatchesDoes not match
/crl/crl/crl/extra
/renewalInfo/*/renewalInfo/abc123/renewalInfo/a/b, /renewalInfo/
/*/directory, /newOrder/renewalInfo/abc

Everything else in a glob is literal, so a path containing regex syntax cannot smuggle a pattern in. allow/deny follow the shared rule described in Checks: deny wins, an empty allow imposes no constraint.

A check with both lists empty is a startup error — it would match every request, which is a check that does nothing.

Combining with other checks

The reason this is a check rather than the flat filter.exempt_paths list it replaces is that a path is rarely the whole answer:

# Only the operator network may revoke.
[filter.check.revocation]
type  = "path"
allow = ["/revokeCert"]

[filter.rule.no-remote-revocation]
when    = "revocation and not mgmt-net"
then    = "deny"
message = "revocation is restricted to the management network"

The old list could only say “skip the connection stage entirely for this exact string”, and could not express /renewalInfo/* at all.

Reverse DNS Filter

The reverse_dns filter provides advanced access control by resolving the client’s IP address back to its associated hostnames via DNS PTR records.

This allows you to write policies based on the physical infrastructure naming convention rather than static IP addresses, which is incredibly useful in environments with dynamic IP allocation.

Resolution logic

When a client connects, the filter performs the following:

  1. It queries the DNS for PTR records associated with the client’s IP.
  2. Forward Confirmation (Optional): If require_forward_confirm = true (the default), it takes the resulting hostnames and queries their A/AAAA records to ensure they point back to the original client IP. This prevents malicious actors from setting up a fake PTR record on an IP they control to spoof an internal hostname.
  3. The resulting, validated hostnames are then checked against the allow and deny regex lists.

Crucial Detail: acme-proxy applies the deny list across every PTR candidate returned by the DNS query, not just the one that would otherwise be accepted. If a client’s IP resolves to good.internal AND evil.hacker.net, and your deny list blocks evil, the connection is denied, even if good matches the allow list.

Timeout budget

DNS queries happen on the hot path — this is a connection-level filter, so it runs on every non-exempt request, newNonce included. The filter operates with a strict timeout_ms budget across all DNS queries so slow nameservers cannot tie up the ACME server.

To keep that affordable, reverse_dns builds its own, caching resolver — deliberately unlike every other DNS consumer in the server, which shares one uncached resolver. A PTR lookup for an address that keeps connecting is exactly what a cache is for, whereas the shared resolver must stay uncached so a dns-01 TXT record published moments before a challenge is triggered is not defeated by a cached negative answer. Both honour dns.resolver.

Configuration

[filter]
rules = ["known-hosts"]

[filter.check.has-ptr]
type = "reverse_dns"
# Require the PTR record to correctly forward-resolve back to the IP
require_forward_confirm = true
# Allow any host in the specific internal domain
allow = ["*.corp.example.com"]
# Deny the guest network infrastructure
deny = ["*.guest.example.com"]
timeout_ms = 2000

[filter.rule.known-hosts]
when = "has-ptr"
then = "allow"

allow/deny take globs over the resolved hostname; allow_regex/deny_regex take anchored regexes and are unioned with them. The keys and their defaults are documented under Checks.

Connection stage by default. It is capable of the identifier stage from the same address, but a PTR plus forward-confirmation exchange at newOrder and again at finalize triples the lookups for an answer that has not changed — so opt in with stages = ["identifiers"] when a rule needs it there.

Identifiers Filter

The identifiers filter controls which domains, IPs, or URIs a client is allowed to request a certificate for. This is crucial for preventing a compromised client from requesting a certificate for a sensitive internal domain.

Type flattening and cn handling

A Certificate Signing Request (CSR) can contain identifiers in multiple places (Subject Alternative Names (SANs) and the legacy Subject Common Name (CN)). acme-proxy flattens all of these into a single list of typed identifiers (dns, ip, email, uri, other, cn).

  • Deny applies everywhere: A deny rule applies to every single type. If you deny *.evil.com, a client cannot sneak it into the Subject Common Name to bypass the filter.
  • Allow skips cn: allow rules explicitly skip the cn type (SUBJECT_ONLY_TYPES). A CN is legacy metadata and often contains human labels (e.g., "rcgen self signed cert"). It is not a true identifier the certificate is for, so it is exempt from strict allow-listing.

What reaches the filter

An order’s dns identifiers are checked for shape before any rule runs: a value that is not a DNS name is malformed, and one the URL parser would read as an IPv4 or IPv6 address — 10.0.0.5, but also 2130706433 or 0x7f.1, which name 127.0.0.1 — is rejectedIdentifier. A rule written for names therefore never has to anticipate an address spelled as one.

Regex anchoring

All matching is performed via Regular Expressions (Regex).

Security Notice: acme-proxy automatically anchors all regexes as ^(?:pattern)$ and makes them case-insensitive. You do not need to manually anchor your regexes. This prevents substring bypasses (e.g., an unanchored example\.com would accidentally match example.com.evil.net).

Configuration

[filter]
rules = ["corp-names-only"]

[filter.check.corp-names]
type = "identifiers"
allowed_types = ["dns", "cn"]
allow = ["*.corp.example.com", "corp.example.com"]
deny  = ["secret.corp.example.com"]
allow_wildcards = false

[filter.rule.corp-names-only]
when = "corp-names"
then = "allow"

allow/deny take globs, where * is one label — so *.corp.example.com does not cover corp.example.com and both are listed, exactly as they would be in a certificate. allow_regex/deny_regex take anchored regexes and are unioned with the globs, for what a glob cannot express. The keys and their defaults are documented under Checks.

Two instances of this type with different lists are ordinary, which is the usual way to say “these names from this network, those names from that one”:

[filter.rule.tenant-a]
when = "tenant-a-net and tenant-a-names"
then = "allow"

[filter.rule.tenant-b]
when = "tenant-b-net and tenant-b-names"
then = "allow"

EAB Check

type = "eab" matches on the External Account Binding credential the requesting account registered under. It is the multi-tenant lever: mint one credential per tenant, bind each to its own name space, and no tenant can request another’s names.

[filter]
rules = ["tenant-a", "tenant-b"]

[filter.check.is-tenant-a]
type  = "eab"
allow = ["tenant-a"]

[filter.check.tenant-a-names]
type  = "identifiers"
allow = ["*.tenant-a.example.com"]

[filter.rule.tenant-a]
when = "is-tenant-a and tenant-a-names"
then = "allow"

Identifier stage only — at the connection stage no account has been authenticated, so there is no credential to ask about.

Why the label and not the account

The obvious handle for “which client is this” would be the account id, and it is the wrong one. An account id is a UUID v7 generated when the account is created, so a policy naming one can only be written after the fact, and you would be editing configuration in response to a client registering.

An EAB credential is the other way round: acme-proxy eab create --label tenant-a mints it before any account exists, and you choose the label. It also survives what the account does next — a key rollover keeps the same account, and a client re-registering under the same credential lands in the same tenant. Credentials are deliberately reusable, so one label naturally covers a whole team.

Blocking a single misbehaving account is a different job and already has a lever: acme-proxy account deactivate <id>.

No credential means refused

An account registered without EAB has no credential, and this check refuses it. That is the only defensible reading of a question about which credential authorised the account — “none” cannot satisfy a tenant rule.

It follows that an eab check under a profile whose eab.enabled is false could never do anything but refuse, so that combination is a startup error rather than a policy.

Matching

allow/deny are globs over the label, with the usual semantics described in Checks — deny first and winning, an empty allow imposing no constraint. allow_regex/deny_regex take anchored regexes and union with them.

kids pins credentials by their kid instead, for an operator who would rather not rely on labels being unique. It is a second allow source rather than a separate gate: either the kid being listed or the label matching is enough.

[filter.check.is-tenant-a]
type  = "eab"
allow = ["tenant-a"]
kids  = ["4f1c…"]     # this exact credential, whatever its label says

A credential minted with no label cannot match a label allowlist. That is the same rule seen from the other side, and it is why eab create is worth always giving a --label.

Making revocation retroactive

Revoking an EAB credential stops new registrations. Accounts already created under it keep issuing, for ever — which is what the credential’s role as a registration-time authorisation implies, and what accounts.eab_kid’s own migration meant by calling it an audit trail.

require_active = true changes that for this check: the credential must still be active, so acme-proxy eab revoke <kid> reaches existing accounts too.

[filter.check.live-credential]
type           = "eab"
require_active = true

A deleted credential goes further than a revoked one, with or without require_active: an account whose kid names a row that no longer exists resolves to no credential at all, and every eab check refuses it, exactly as it refuses an account registered without EAB. See Deleting a credential.

Off by default, because turning it on retroactively changes what eab revoke means for a deployment. It is the lever to reach for when a tenant’s credential leaks — and it is a usable policy on its own, with no labels at all: “any tenant, but not one whose credential we have withdrawn”.

Cost

Resolving the credential is two indexed reads, and they happen only when the policy contains an eab check. A deployment without one pays nothing.

Reference

filter.check.<name>.kids (Array) — Default: [] | Env: ACME_PROXY_FILTER__CHECK__<NAME>__KIDS

Credential kids matched exactly, beside the label globs in allow. A second allow source, not a separate gate.

filter.check.<name>.require_active (Boolean) — Default: false | Env: ACME_PROXY_FILTER__CHECK__<NAME>__REQUIRE_ACTIVE

Also require the credential to still be active, so revoking it refuses accounts already registered under it.

A check that sets none of allow, deny, kids or require_active only asks whether the account used EAB at all, which the [eab] section already guarantees. That is a startup error rather than a check that always passes.

Custom Script Filter

The custom filter lets operators write arbitrary scripts (Bash, Python, …) to decide whether a connection or a CSR should be permitted. Use it for policy that cannot be expressed with the built-in filters.

Configuration

[filter]
rules = ["scripted"]

[filter.check.check-network]
type        = "custom"
script_path = "/etc/acme-proxy/filters/check-network.sh"
timeout_ms  = 5000
pass_stdin  = true
args        = []

[filter.rule.scripted]
when = "check-network"
then = "allow"

custom is an ordinary check type: there is no separate selection list, because filter.rules already says which checks run and in what order. Several [filter.check.<name>] entries may point at the same script — each is told which one invoked it through ACME_FILTER_CHECK_NAME, so one script can serve them all and branch on it.

The keys and their defaults are documented under Checks. An empty script_path on a check some rule names is a startup error.

Hooks

The script is invoked at both filter hooks, distinguished by ACME_FILTER_HOOK:

ACME_FILTER_HOOKWhenRefusal becomes
connectionEvery non-exempt request403 access_denied
identifiersAt newOrder and at finalize403 rejectedIdentifier (newOrder) / 400 badCSR (finalize)

If your script only cares about one hook, branch on this variable and exit 0 otherwise — the script is called for both.

Data passing

Environment variables

connection hook:

VariableValue
ACME_FILTER_HOOKconnection
ACME_FILTER_CLIENT_IPThe resolved client address, canonicalized (an IPv4-mapped IPv6 address is flattened to IPv4). Empty when unknown.
ACME_FILTER_METHODHTTP method
ACME_FILTER_PATHRequest path

identifiers hook:

VariableValue
ACME_FILTER_HOOKidentifiers
ACME_FILTER_CLIENT_IPAs above
ACME_FILTER_ACCOUNT_IDThe authenticated ACME account
ACME_FILTER_STAGEnewOrder or CSR
ACME_FILTER_IDENTIFIERSComma-joined identifier values. Guaranteed free of commas and control characters — see below.

ACME_FILTER_IDENTIFIERS is safe to split on , and safe to read line by line. A request whose identifiers could not survive that join is refused before the script runs, with badCSR, so a script never has to defend against it.

That guarantee needs stating because it is not free. At the newOrder stage every identifier is a DNS name the server has already validated. At the CSR stage the list also carries the certificate request’s subject CommonName, which is arbitrary text — routinely a human label such as Example Corp Issuing CA rather than a host name, and therefore not validated as one. A CommonName holding a comma or a newline would otherwise reach the script as extra entries.

The typed JSON on stdin has no such ambiguity and is the better source when a script cares which identifier is which: each entry is its own object with a type, so a cn is distinguishable from a dns there and not here.

JSON on stdin

When pass_stdin is true (the default), a JSON object — not a bare array — is written to the script’s standard input.

connection:

{"hook":"connection","client_ip":"203.0.113.5","method":"POST","path":"/newOrder"}

identifiers:

{
  "hook": "identifiers",
  "client_ip": "203.0.113.5",
  "account_id": "…",
  "stage": "newOrder",
  "identifiers": [{"type": "dns", "value": "a.example.com"}]
}

client_ip is null rather than a string when the address is unknown.

At the CSR stage the identifiers list is the flattened projection of the whole CSR — SANs and the subject Common Name — so entries of type ip, email, uri, other and cn appear alongside dns. That is deliberate: a deny rule cannot be dodged by moving a name from a SAN into the CN.

Return codes

  • 0 — permitted.
  • Any non-zero exit — denied. This includes exit code 255 and death by signal; the check is simply “did it exit successfully”.
  • Timeout, or failure to spawn — treated as an internal error (500), so the client retries rather than seeing a permanent refusal.

On denial, the reason sent to the ACME client is the first non-empty line of stdout, falling back to the first non-empty line of stderr, and finally to a generic “script exited with status …” message. Keep it to one line, and remember it is client-visible — do not leak internal detail into it.

Execution model and security

  1. Environment clearing (env_clear): the child runs with a scrubbed environment, inheriting only a minimal PATH and the ACME_FILTER_* variables above. The server’s own environment may hold secrets — the RFC 2136 TSIG key, the NetBox token, the SMTP password — and a filter script has no business reading them.
  2. Zombie prevention (kill_on_drop): execution is wrapped in a Tokio timeout with kill_on_drop(true). A timeout alone only drops the future, so without this a hung script would outlive its deadline and leak a process per request.
  3. Cost: the connection hook runs on every non-exempt request, including newNonce. A script doing network I/O there will dominate your latency; prefer the identifiers hook when the policy only concerns names.

IPAM

An IPAM backend answers one question: which names does this address own?

acme-proxy asks it through the ipam filter, which turns the answer into a decision — may this client have a certificate for the names it is asking for? The two halves are deliberately separate. The inventory reports what it holds; the filter alone decides what to do about that, and it decides the same way whichever product answered.

Renamed at this release. This used to be a filter called netbox, with its settings under [filter.netbox]. Both moved. See Migrating from filter.netbox below — the old spelling is refused by name at startup, with an error naming all three moves, rather than silently ignored.

Backends

BackendProductPage
netboxNetBoxNetBox
phpipamphpIPAMphpIPAM
customan operator scriptCustom Script

One backend per profile. [ipam] is a per-profile section, so two endpoints served by the same process may consult different inventories.

Sources

Each backend that reads an inventory itself takes a sources list naming the places a permitted name may come from. It is validated at startup against that backend’s own vocabulary. custom has no such list at all — the script reads whatever it reads, so there is nothing to declare and no key to set.

SourceWhat it readsnetboxphpipamcustom
dns_namethe address object’s own name✓✓—
custom_fieldthe custom field on the address✓✓—
devicethe same field on the assigned device or VM✓✓—
viprole-tagged service addresses on the same device✓——
fhrpaddresses of an FHRP group the client’s interface is in✓——

Two things the list does not say, and both matter:

  • Order is meaningless. The result is a union of sets, unlike filter.rules (evaluation order) or challenge.enabled (offer order).
  • device is a fallback, not a union. It is read only when the address object itself carried no value for the custom field. A value set on the address is the more specific statement, and an operator narrowing one address of a machine must not have the machine-wide list quietly widen it again. vip and fhrp are unions.

Empty, or an unknown entry, is a startup error. An inventory trusted for nothing can never permit a name, so it is a filter that refuses everything — and a typo that silently narrows an allowlist is worse than a refusal to boot. A source that exists but belongs to another backend is refused by name too: naming fhrp under [ipam.phpipam] says a check will run that phpIPAM cannot run, and answering that with silence would leave an operator believing it does.

The default for both backends is ["dns_name", "custom_field", "device"] — exactly what the old netbox filter always did. The two service-address sources are off because both widen what a client may certify.

Matching is exact

Case-insensitive and ignoring a trailing dot, but otherwise literal. There is no suffix rule and no wildcard expansion: an entry example.com does not permit a.example.com, and a request for *.example.com requires that exact string in the inventory. Same reasoning as the anchored patterns in Identifiers — a rule that quietly covers more than it says is the bypass an allowlist exists to prevent.

An ip identifier is permitted when it is the connecting address (a machine may always certify the address it is talking from) or when it is listed like any other name. A common name (cn) is skipped, as in identifiers. Any other type is refused: an inventory has nothing to say about an email address or a URI, and a filter whose job is to confirm entitlement must refuse what it cannot confirm.

Denied versus Internal

The most consequential property of the whole subsystem.

“The inventory does not associate this name with this address” is a decision about the client and denies the request (403 rejectedIdentifier). So is “the inventory holds no record of this address at all”, worded differently so an operator can tell the two apart.

“It answered 500”, “the token was refused”, “the lookup timed out” are not decisions about anybody — the server failed to reach one. Those become a 500 the client can retry.

This is enforced by the types, not by care at each call site: the error an Ipam backend can return has no “denied” variant to reach for. An inventory outage therefore stops issuance rather than permitting everything, and never looks like a permanent refusal.

What it costs per request

The lookup runs inline in newOrder and again at finalize, so it is part of those requests’ worst case. timeout_ms is one budget covering the whole lookup however many requests the backend makes to answer it — not one per request — and it must stay below server.request_timeout_ms.

Migrating from filter.netbox

Three changes, and the server refuses the old spelling of each by name: a check still declared type = "netbox" names the first, a [filter.netbox] section the other two:

  1. A check declared with type = "netbox" becomes type = "ipam", plus ipam.backend = "netbox".
  2. [filter.netbox] becomes [ipam.netbox]. Most keys are unchanged.
  3. Three keys moved or changed shape:
WasNow
filter.netbox.timeout_msipam.timeout_ms
filter.netbox.use_dns_name = true"dns_name" in ipam.netbox.sources
filter.netbox.device_fallback = true"device" in ipam.netbox.sources

Since the default sources is ["dns_name", "custom_field", "device"], a deployment that left both booleans at their defaults needs no sources line at all. Environment variables move from ACME_PROXY_FILTER__NETBOX__* to ACME_PROXY_IPAM__NETBOX__*.

Startup fails on the old filter name rather than aliasing it. The section moved too, so a silent alias would leave [filter.netbox] read by nothing while the server came up looking configured — the same reasoning that made signer.backend = "acme_proxy" a named refusal when it became relay. Both are error messages rather than compatibility paths: nothing reads the old spelling, and the refusals go away at 1.0.0. Renames like this are expected before then and are listed in the changelog.

Configuration

[ipam]
backend = "netbox"
timeout_ms = 5000

[ipam.netbox]
url = "https://netbox.internal.example.com"
token = "..."
sources = ["dns_name", "custom_field", "device"]

[filter]
enabled = ["ipam"]

Reference

backend (String) — Default: "" | Env: ACME_PROXY_IPAM__BACKEND

Which inventory to consult: netbox, phpipam, custom, or empty for none. Anything else is a startup error rather than a silent fallback. Enabling the ipam filter while this is empty is also a startup error.

timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_IPAM__TIMEOUT_MS

Budget for one whole lookup, however many requests the backend makes to answer it — or, for custom, however long its script takes. Applied once around all of them, so a wedged inventory cannot pin a request open. Exceeding it is reported as a server error, not a denial.

NetBox

Reads what NetBox associates with the client’s address. Supports every source, including the two that resolve a shared service address.

What NetBox is asked

One lookup always happens:

GET <url>/api/ipam/ip-addresses/?address=<client ip>

Up to four more are made, each gated by a source:

QuerySource
dcim/devices/{id}/ or virtualization/virtual-machines/{id}/device
ipam/ip-addresses/?device_id=N&role=…vip
ipam/fhrp-group-assignments/?interface_type=…&interface_id=…fhrp
ipam/ip-addresses/?fhrpgroup_id=…fhrp

A read-only API token is enough.

Authenticating

NetBox has two generations of API token, and this backend sends whichever it is given: the scheme is derived from the token itself, so there is nothing to configure.

TokenSent asWhere it comes from
nbt_<key>.<secret>Authorization: Bearer nbt_<key>.<secret>v2, the default since NetBox 4.5
anything elseAuthorization: Token <token>the legacy v1 token, not accepted from NetBox 4.7

The nbt_ prefix is NetBox’s own marker for a v2 token, and the whole string — key, dot and secret — is displayed once when the token is created, so paste it verbatim. A value starting nbt_ that carries no . is the key half on its own: that is refused at startup, because NetBox would otherwise answer every lookup with a 403, which looks exactly like a token that has been revoked.

Declaring names

Two places, and either or both can be trusted:

  • dns_name on the IP address object — the ordinary case, one name per address.
  • A custom field (custom_field, by default acme_domains) for the extra names that address may request. Configure it in NetBox as a multi-select or a text field on ipam.ipaddress and — for the device source — on dcim.device and virtualization.virtualmachine. A single string is accepted as well as a list.

With the device source, an address that carries no value of its own falls back to the field on its device or virtual machine, so names can be declared once per machine rather than once per address. It is a fallback and not a union: see Sources.

Shared and service addresses

A VRRP, CARP or keepalived pair answers on an address that belongs to the pair, not to either member — but the client connects from its own member address, so without one of the sources below it is refused a certificate for the service name. Both are unions: the member’s own names and the service address’s names are true at the same time.

Which one an estate needs depends on how it models redundancy in NetBox.

vip — a role on an address of the same device

The classic modelling: the service address is created on one of the members’ interfaces and tagged with a role.

GET <url>/api/ipam/ip-addresses/?device_id=3&role=vip&role=vrrp

vip_roles says which roles count. The role is re-checked on the answer as well as sent as a filter — a filter parameter this server got wrong must never degrade into “every address on the device”, which would widen an allowlist without saying so.

fhrp — membership of an FHRP group

NetBox’s own model for first-hop redundancy: the service address is assigned to an ipam.FHRPGroup, and each member’s interface is recorded as belonging to that group.

client address ─▶ its interface ─▶ fhrp-group-assignments?interface_id=7
                                        └─▶ group ids ─▶ their addresses

The direction of that chain is the membership proof. A group is only ever reached through an assignment naming the client’s own interface. Nothing is ever looked up by group name, by the service address, or by the identifier the client asked for — so there is no query that could reach a group the client is not recorded in, and no way to turn the check into a lookup of “who owns this name?” by choosing a request carefully. An interface in no group contributes nothing and costs one request.

No role filter applies here: an address assigned to an FHRP group is the group’s service address by construction, and applying vip_roles would drop legitimately untagged VIPs.

A client connecting from the service address needs neither source — that address object comes back from the first query with its own names attached.

TLS

Unlike the challenge validators, where the certificate is deliberately not checked because the proof is what matters, NetBox’s certificate is the only thing identifying the service whose answers decide who may have a name certified. The public roots apply, plus any operator-supplied CA. Switching that off is explicit, logged on every start, and meant to be temporary.

Configuration

[ipam]
backend = "netbox"

[ipam.netbox]
url = "https://netbox.internal.example.com"
token = "your_netbox_read_only_token"
custom_field = "acme_domains"
sources = ["dns_name", "custom_field", "device"]

Turning on service addresses:

[ipam.netbox]
sources = ["dns_name", "custom_field", "device", "vip", "fhrp"]
vip_roles = ["vrrp", "carp"]

Reference

url (String) — Default: "" | Env: ACME_PROXY_IPAM__NETBOX__URL

Base URL of the NetBox instance. Any path is kept, so an instance served under a subpath works. Required when ipam.backend is netbox.

token (String) — Default: "" | Env: ACME_PROXY_IPAM__NETBOX__TOKEN

NetBox API token, of either generation — see Authenticating. A secret: prefer the environment variable.

custom_field (String) — Default: "acme_domains" | Env: ACME_PROXY_IPAM__NETBOX__CUSTOM_FIELD

Custom field holding the permitted names, on the address object and on its device or virtual machine. Only read when sources names custom_field or device.

sources (Array) — Default: ["dns_name", "custom_field", "device"] | Env: ACME_PROXY_IPAM__NETBOX__SOURCES

Where a permitted name may come from. All five sources are available here. See Sources.

vip_roles (Array) — Default: ["vip", "vrrp", "hsrp", "glbp", "carp", "anycast"] | Env: ACME_PROXY_IPAM__NETBOX__VIP_ROLES

Which NetBox address roles mark a service address. Read only when sources names vip, so this is which roles rather than whether to look at all.

ca_cert_path (String) — Default: "" | Env: ACME_PROXY_IPAM__NETBOX__CA_CERT_PATH

Extra CA certificates (PEM) to trust on top of the public roots, for a NetBox behind an internal PKI. Ignored when insecure_skip_verify is on.

insecure_skip_verify (Boolean) — Default: false | Env: ACME_PROXY_IPAM__NETBOX__INSECURE_SKIP_VERIFY

Skip verification of NetBox’s TLS certificate entirely. Meant as a temporary way out of an expired NetBox certificate. Startup logs an ipam_netbox_tls_verification_disabled warning for as long as it is set.

phpIPAM

Reads what phpIPAM associates with the client’s address. Supports dns_name, custom_field and device; phpIPAM records no address roles and no redundancy groups, so vip and fhrp are refused by name at startup.

What phpIPAM is asked

One lookup always happens:

GET <url>/api/<app_id>/addresses/search/<client ip>/

One more is made when sources names device:

GET <url>/api/<app_id>/devices/<deviceId>/

Setting up the API application

In phpIPAM, under Administration → API, create an application:

  • App id — becomes app_id, and appears in every API path.
  • App permissions — Read is enough.
  • App security — SSL with App code. The app code becomes token and is sent as a bare token header (phpIPAM’s own scheme, not Authorization).

The alternative SSL with User token scheme — user credentials exchanged for a six-hour session token — is not implemented: it needs a refresh loop and somewhere to keep the token, for no gain over a credential that can be rotated in the environment.

Declaring names

  • hostname on the address — the direct analogue of NetBox’s dns_name, read when sources names dns_name.
  • A custom column (custom_field, by default custom_acme_domains). phpIPAM prefixes custom columns with custom_, so the default carries that prefix; write whatever your column is actually called. Add it under Administration → Custom fields for IP addresses and — for the device source — for Devices.

A phpIPAM custom field is a plain text column, so several names are written comma-separated:

www.example.com, api.example.com, mail.example.com

The split is on commas and is not configurable — a comma is not legal in a DNS name, so there is no estate it could need to differ for.

With the device source, an address whose column is empty falls back to the same column on the device named by its deviceId. A fallback, not a union: see Sources.

An unknown address is a 404

The one place phpIPAM’s wire behaviour differs in a way worth knowing about. NetBox answers an unknown address with 200 and an empty result list; phpIPAM answers 404 with its own envelope:

{"code": 404, "success": false, "message": "No addresses found"}

That is read as “no such address” — a denial naming the address — rather than as a transport failure, which would turn every request from an unrecorded machine into a retryable 500. Every other non-2xx status is still a failure, so a broken or misconfigured phpIPAM stops issuance rather than reading as “this address owns no names”. See Denied versus Internal.

Configuration

[ipam]
backend = "phpipam"

[ipam.phpipam]
url = "https://ipam.internal.example.com"
app_id = "acme"
token = "your_app_code"
custom_field = "custom_acme_domains"
sources = ["dns_name", "custom_field", "device"]

Reference

url (String) — Default: "" | Env: ACME_PROXY_IPAM__PHPIPAM__URL

Base URL of the phpIPAM instance. Any path is kept, so an instance served under a subpath works. Required when ipam.backend is phpipam.

app_id (String) — Default: "acme" | Env: ACME_PROXY_IPAM__PHPIPAM__APP_ID

The API application’s identifier, which appears in every API path. One path segment: letters, digits, - and _. Anything else is a startup error, since a stray slash would silently retarget the API rather than fail.

token (String) — Default: "" | Env: ACME_PROXY_IPAM__PHPIPAM__TOKEN

The application’s app code, sent as a token header. A secret: prefer the environment variable.

custom_field (String) — Default: "custom_acme_domains" | Env: ACME_PROXY_IPAM__PHPIPAM__CUSTOM_FIELD

Column holding the permitted names, on the address and on its device. Only read when sources names custom_field or device.

sources (Array) — Default: ["dns_name", "custom_field", "device"] | Env: ACME_PROXY_IPAM__PHPIPAM__SOURCES

Where a permitted name may come from. Only these three are available here; naming vip or fhrp is a startup error.

ca_cert_path (String) — Default: "" | Env: ACME_PROXY_IPAM__PHPIPAM__CA_CERT_PATH

Extra CA certificates (PEM) to trust on top of the public roots. Ignored when insecure_skip_verify is on.

insecure_skip_verify (Boolean) — Default: false | Env: ACME_PROXY_IPAM__PHPIPAM__INSECURE_SKIP_VERIFY

Skip verification of phpIPAM’s TLS certificate entirely. Startup logs an ipam_phpipam_tls_verification_disabled warning for as long as it is set.

Custom Script

Runs an operator-supplied script to answer the one question the whole subsystem asks: which names does this address own?

Use it when the inventory of record is something this server carries no client for — a CMDB, a hosts file, an LDAP tree, a spreadsheet exported nightly, a vendor API behind a Python wrapper. It is the same escape hatch the custom filter and the custom signer are for their subsystems, and it runs under the same hardening.

It is also the backend with the least in it. There is no URL, no credential, no TLS setting and no sources list: the script is the inventory, and where it looks is its own business.

The contract

The script is told the client address twice — once in the environment, once in a JSON object on stdin — so neither a four-line shell script nor a Python one has to reach for the channel it finds awkward.

VariableValue
ACME_IPAM_HOOKAlways names_for. There is one hook.
ACME_IPAM_CLIENT_IPThe resolved client address, canonicalized (an IPv4-mapped IPv6 address is flattened to IPv4).

The same on stdin:

{"hook": "names_for", "client_ip": "203.0.113.5"}

A script that exits without reading stdin is fine and is not an error.

The answer

stdout plus an exit code.

ExitstdoutMeans
0one name per lineThe inventory holds this address, and these are its names
0emptyHeld, and entitled to nothing
3ignoredNo record of this address at all
anything elsethe reasonThe script failed — a retryable 500, never a denial

Names are compared exactly, but the script need not tidy them: each line is lowercased and stripped of a trailing dot, and a blank line is ignored. So WWW.Example.COM. and www.example.com are the same answer, and printing whatever form the inventory happens to hold is correct.

One name per line rather than a separated list because a newline is the shell idiom, and plain text rather than JSON because there is nothing here a structure would carry that a list of lines does not — a contract needing jq for what echo already does would be paid for by every script ever written against it. The custom signer hands back its certificate chain the same way.

Why 3 is reserved

“Held, and entitled to nothing” and “no record of this address” are different answers — the ipam check words a different refusal for each, so an operator reading a 403 can tell them apart — and once stdout means the names, an exit status is the only channel left to say the second one in. So it gets a code of its own, exactly as the custom signer’s badCSR does.

Every other non-zero exit stays a failure, and that direction is the one that matters. A script that breaks, is missing, or runs past ipam.timeout_ms produces a server error the client can retry, never a refusal — see Denied versus Internal. It is the same guarantee an unreachable NetBox gets, and it is enforced by the types rather than by care here: the error an IPAM backend can return has no denied variant to reach for, so a broken script cannot fail open or look permanent.

An example

#!/bin/sh
case "$ACME_IPAM_CLIENT_IP" in
  203.0.113.5)
    echo www.example.com
    echo api.example.com
    exit 0
    ;;
  203.0.113.6)
    exit 0          # held, entitled to nothing
    ;;
esac
exit 3              # no record of this address

A fuller one, reading a real source, is in Custom Plugins Examples.

No sources

sources exists because NetBox and phpIPAM each read several places and an operator has to say which are trusted. A script reads whatever it reads, so there is nothing here to list and no key to set. The vocabulary is closed and validated per backend, so a sources line under [ipam.custom] is not a source this backend “does not support” — it is a setting that does not exist.

What it costs

One forked process per lookup, inside newOrder and again at finalize. ipam.timeout_ms is the budget, and it is also what kills the child at the deadline; there is deliberately no second timeout here.

Note that acme-proxy filter explain really runs the policy, so it executes this script — as it does the custom filter’s. That is the point of the command, and it is reported in its side-effects list.

Security and process isolation

  1. Environment clearing: env_clear() is called. The script inherits a minimal PATH plus the two ACME_IPAM_* variables above — nothing else. The server’s own environment may hold the NetBox token, the SMTP password or the RFC 2136 TSIG key, and an inventory script has no business reading them.
  2. Zombie protection: the child runs with kill_on_drop(true) under a tokio::time::timeout. A timeout alone only drops the future, so without this a hung script would outlive its deadline and leak a process per request.
  3. Failure reporting: on a non-zero exit other than 3, the first non-empty line of stdout (falling back to stderr) becomes the error detail. It is logged, not sent to the client: an internal error tells the client nothing about why.

Configuration

[ipam]
backend = "custom"
timeout_ms = 5000

[ipam.custom]
script_path = "/etc/acme-proxy/ipam/lookup.sh"
args = []

Reference

script_path (String) — Default: "" | Env: ACME_PROXY_IPAM__CUSTOM__SCRIPT_PATH

Path to the executable answering the lookup. Empty while ipam.backend is custom is a startup error, the same as signer.custom.script_path.

args (Array) — Default: [] | Env: ACME_PROXY_IPAM__CUSTOM__ARGS

Fixed arguments passed before the script is told anything about the request. One script can serve several deployments by branching on them.

There is deliberately no timeout_ms here. ipam.timeout_ms is the budget the whole lookup runs under, and a second one would contradict the rule the other backends follow.

Configuration Reference

This page is the structured reference for every core configuration parameter.

For deep dives into specific subsystems (Signers, Filters, Notifications, EAB), see their dedicated chapters — linked at the bottom, and the place where those sections’ keys are documented.

How configuration is resolved

Sources are layered, lowest precedence first:

  1. Built-in defaults — everything documented below has one, so an empty configuration is valid apart from [profiles].
  2. A configuration file — config.toml in the working directory, or the path in ACME_PROXY_CONFIG (the extension may be omitted, in which case the format is inferred). A missing file is not an error.
  3. ACME_PROXY_* environment variables — __ separates nested keys, a single _ separates the prefix. server.tls.enabled is ACME_PROXY_SERVER__TLS__ENABLED.

config.toml.example in the repository is the annotated companion to this page: it lists every key with its default and its environment variable name in context.

[profiles] is mandatory. Everything else can be left at its default, but the server serves ACME only through profiles and refuses to start without at least one enabled. See [profiles] below.

Every section, and where it is documented

Six sections are large enough to have a chapter of their own, and their keys are documented there rather than restated here. This page stays the complete map: every top-level section acme-proxy reads appears below, whether or not its text lives here. A delegated section’s own sub-tables — [signer.local_ca.subject], [ipam.netbox], [filter.check.<name>] and the rest — are listed in the chapter that owns them.

Overridable marks the sections a [profiles.<name>] block may override. Everything else is process-wide — one setting for the whole server, however many endpoints it mounts.

SectionControlsOverridableDocumented
[database]The SQLite file or PostgreSQL servernobelow
[server]Listen socket, public URL, admission controlnobelow
[server.tls]HTTPS on the ACME listenernobelow
[admin]The web admin listener and its sessionsnobelow
[admin.filter]Who may reach the admin listenernobelow
[admin.notify]Operator security notificationsnobelow
[admin.tls]HTTPS on the admin listenernobelow
[nonce]Replay-nonce freshnessnobelow
[audit]Reverse lookups and retention for the trailnobelow
[jobs]The durable background-work queuenobelow
[metrics]The Prometheus listenernobelow
[dns]The resolver every outbound lookup usesnobelow
[proxy]The forward proxy outbound clients dial throughnobelow
[logging]Filter, format, targetnobelow
[order]The ACME order object’s lifetimeyesbelow
[meta]Directory meta membersyesbelow
[profiles.<name>]An ACME endpoint—below
[signer]How a certificate is obtainedyesSigners
[filter]Who may ask, and for whatyesFilters
[ipam]The inventory an ipam filter check consultsyesIPAM
[challenge]How control of a name is provenyesChallenge Validation
[notify]Outbound notificationsyesNotifications
[eab]External Account BindingyesEAB

The criterion for the last six is having a chapter, not being overridable — [order] and [meta] are overridable and documented here, because neither is large enough to be worth a page. That is the whole rule; there is nothing subtler going on.

config.toml.example in the repository carries the same list as a comment header, with every key in context.


[database]

url (String) — Default: "sqlite://sqlite.db" | Env: ACME_PROXY_DATABASE__URL

Database connection URL, for accounts, orders, certificates and everything else this server keeps. The scheme picks the backend, and any other scheme is refused by name at startup:

  • sqlite://<path> — a file, created on first use. The default, and what a single-host deployment wants. sqlite:///var/lib/acme-proxy/acme.db is an absolute path (three slashes).
  • postgres:// or postgresql:// — a server, which must already exist: creating a database is an operator’s act, not something a server does to a cluster it was pointed at. Run acme-proxy migrate once against it.

PostgreSQL is what a multi-node deployment needs. SQLite across processes is fine on one local disk, and not safe on NFS or across hosts — see Deployment. Nothing else changes with the backend: the same binary, the same configuration and the same ACME behaviour.

A PostgreSQL URL usually carries user:password@. The password is never logged: the startup line and the SIGHUP refusal both print it as postgres://acme:***@host/db. It is still a credential in a configuration file, so give that file the permissions it deserves.

This key cannot be changed by a reload — see Reload.


[server]

bind_address (String) — Default: "[::]:3000" | Env: ACME_PROXY_SERVER__BIND_ADDRESS

Network address the server binds and listens to. SIGHUP moves it without restarting, and an address that cannot be bound refuses the reload rather than taking the running socket down — see Configuration Reload.

base_url (String) — Default: "http://localhost:3000" | Env: ACME_PROXY_SERVER__BASE_URL

Public base URL advertised in the ACME directory, with no trailing slash. It is never derived from the request, and every signed request’s url field is checked against it (RFC 8555 §6.4) — so behind a reverse proxy, or with TLS enabled, this must be set to the public URL or every client is rejected.

max_concurrent_requests (Integer) — Default: 100 | Env: ACME_PROXY_SERVER__MAX_CONCURRENT_REQUESTS

How many ACME requests may be in flight at once before the server sheds load.

admission_wait_ms (Integer) — Default: 50 | Env: ACME_PROXY_SERVER__ADMISSION_WAIT_MS

How long a request may wait for a slot before it is refused. Past the limit a request waits this long and then gets 503 + Retry-After — it is shed, not queued.

request_timeout_ms (Integer) — Default: 60000 | Env: ACME_PROXY_SERVER__REQUEST_TIMEOUT_MS

Whole-request deadline. It must exceed signer.custom.timeout_ms when that backend is installed with its crl or renewal_info hook enabled, since those hooks run inline inside a request; the server refuses to start otherwise. It is deliberately independent of challenge.timeout_ms and of the script’s issue and revoke hooks, which run in the job queue — see Challenge Validation. A relay or custom revocation that a request queues is waited on for this long, less a second.

max_body_bytes (Integer) — Default: 131072 | Env: ACME_PROXY_SERVER__MAX_BODY_BYTES

Largest request body accepted (128 KiB).

These four keys govern the ACME routes only. GET /health is mounted outside all of them.

trusted_proxies and forwarded_header are not [server] keys — they live under [filter]. See Filters.

[server.tls]

Full treatment in TLS Termination.

enabled (Boolean) — Default: false | Env: ACME_PROXY_SERVER__TLS__ENABLED

Serve HTTPS on bind_address instead of cleartext — one listener, not two. base_url is not rewritten for you; set it to https://… yourself or every signed request fails the §6.4 URL check above.

cert_path (String) — Default: "server.pem" | Env: ACME_PROXY_SERVER__TLS__CERT_PATH

PEM certificate chain, leaf first. A self-signed certificate is generated and written when either this or key_path is missing.

key_path (String) — Default: "server.key" | Env: ACME_PROXY_SERVER__TLS__KEY_PATH

Path to the private key.

handshake_timeout_ms (Integer) — Default: 10000 | Env: ACME_PROXY_SERVER__TLS__HANDSHAKE_TIMEOUT_MS

Budget for one TLS handshake. Handshakes run concurrently, off the accept path, so this never delays another client.


[admin]

The web admin interface — a second listener, on its own socket, serving no ACME. Process-wide, so there is no [profiles.<name>].admin: an operator manages every endpoint this process serves. Full treatment in Web Admin.

enabled (Boolean) — Default: false | Env: ACME_PROXY_ADMIN__ENABLED

Off by default: a certificate authority should not grow a management surface because somebody upgraded it. Bootstrap it with acme-proxy admin user create; there is no sign-up page.

bind_address (String) — Default: "127.0.0.1:3001" | Env: ACME_PROXY_ADMIN__BIND_ADDRESS

Loopback on purpose. This listener has no admission control, filters nothing until [admin.filter] names a rule, and — until [admin.tls] is on — has no transport security.

Startup refuses a non-loopback bind while admin.tls.enabled is false. The session cookie is always sent Secure, and a browser silently declines to store one over plain HTTP on anything but localhost; the symptom would be “sign-in works, then I am immediately signed out”, with nothing in any log to explain it. Either enable [admin.tls], or keep the loopback bind and reach it through an SSH tunnel:

$ ssh -N -L 3001:127.0.0.1:3001 ca.example.com

It is also an error for this to equal server.bind_address.

base_url (String) — Default: "http://localhost:3001" | Env: ACME_PROXY_ADMIN__BASE_URL

The origin the panel is reached at. Load-bearing three times over: the CSRF origin check compares against it, a generated self-signed certificate takes its host, and the pages build absolute URLs from it — exactly as server.base_url does for the ACME listener. Through a tunnel this stays localhost. The resolved origin is logged at startup (admin_origin_resolved) so a mismatch is visible before the first refused request.

session_ttl_seconds (Integer) — Default: 43200 (12 h) | Env: ACME_PROXY_ADMIN__SESSION_TTL_SECONDS

Absolute session lifetime. Never extended by activity: past it, the operator signs in again.

session_idle_timeout_seconds (Integer) — Default: 3600 (1 h) | Env: ACME_PROXY_ADMIN__SESSION_IDLE_TIMEOUT_SECONDS

Idle lifetime, advanced on use — at most once a minute, so a polling page is not a stream of database writes. Whichever deadline comes first wins.

login_max_attempts (Integer) / login_window_seconds (Integer) — Defaults: 5 / 300 | Env: ACME_PROXY_ADMIN__LOGIN_MAX_ATTEMPTS, ACME_PROXY_ADMIN__LOGIN_WINDOW_SECONDS

Failed sign-ins allowed from one address per window, then 429 with a Retry-After. The password hash is deliberately expensive (PBKDF2-HMAC-SHA256 at 600 000 iterations), so this is an availability control as much as a credential one: over the limit, the hash is not computed at all.

Keyed on the client address. The forwarded-for header is believed only from a peer listed in admin.filter.trusted_proxies — never from filter.trusted_proxies, which governs the ACME listener. With no trusted proxy listed, the limiter behind a reverse proxy counts the proxy.

require_mfa (Boolean) — Default: false | Env: ACME_PROXY_ADMIN__REQUIRE_MFA

Require a second factor (TOTP) of every operator. What this changes is the operator who has none: with it on, their next sign-in lands on the enrolment page and their session stays half-authenticated until they finish. An operator who already has one is challenged whether this is set or not.

It deliberately does not refuse a password-only sign-in outright: enrolling needs a session and a session would then need a factor, so that would brick the panel including the way in to fix it.

Turning it on does not retroactively end sessions that predate it — acme-proxy admin session revoke --all is the lever that does. While it is on and some operator has no factor, every start logs admin_mfa_enrolment_pending. See Operators and sessions.

max_body_bytes (Integer) — Default: 65536 | Env: ACME_PROXY_ADMIN__MAX_BODY_BYTES

Largest admin request body. An admin body is a small JSON object or a form, never a certificate.

page_size_max (Integer) — Default: 200 | Env: ACME_PROXY_ADMIN__PAGE_SIZE_MAX

Ceiling on ?limit= for the list endpoints; the default page size is 50. A larger request is clamped, not refused.

template_dir (String) — Default: "" | Env: ACME_PROXY_ADMIN__TEMPLATE_DIR

Override individual page templates on disk, mirroring notify.template_dir. Empty means the compiled-in defaults. The override is per file: a directory holding only layout.html restyles the chrome of every page and leaves the other fifty-five at their defaults. Every template is compiled at startup, so a broken override refuses to start rather than serving a 500 later. Applies to the /ui pages only; the JSON API has nothing to template. See Customizing the Panel.

[admin.filter]

Who may reach the admin listener at all — for the deployment that cannot put a host firewall in front of the port, such as a container. The same policy engine and the same keys as the ACME listener’s [filter] (rules, default, trusted_proxies, forwarded_header, [admin.filter.check.<name>], [admin.filter.rule.<name>]), under ACME_PROXY_ADMIN__FILTER__…, with four differences:

  • It is its own section. Nothing is inherited from the global [filter], which is the ACME profiles’ base: an edit made for one listener must not open or shut the other.
  • Only the connection stage exists here, so only allowed_ip, path, reverse_dns and custom checks are accepted. An identifiers, eab or ipam check, or a rule whose checks were moved to the identifier stage with stages, is refused by name at startup and on reload.
  • Empty rules filters nothing, and logs no warning.
  • trusted_proxies also keys the login limiter (see admin.login_max_attempts): it is the one list the admin listener believes a forwarded-for header from, and it applies whether or not a rule is written.

Every request is evaluated, /health included. A refusal is 403 — the admin API’s JSON error (access_denied) under /api, the HTML error page elsewhere — and is logged as filter_request_blocked with listener = "admin". See Web Admin — Restricting who can reach it.

[admin.notify]

A whole [notify] section, process-wide, built only while [admin] is enabled. It delivers the admin_sign_in / admin_credential_changed operator security events (ASVS V6.3.5 / V6.3.7). Every key is the one documented for the per-profile [notify] — under ACME_PROXY_ADMIN__NOTIFY__… — with one difference: [admin.notify.email].to may be empty, since each event carries the affected operator’s own contact address as its recipient. See Web Admin — Security notifications.

[admin.tls]

HTTPS on admin.bind_address instead of cleartext — the same one-listener-not-two shape as [server.tls], and the same load-or-generate provisioning.

enabled (Boolean) — Default: false | Env: ACME_PROXY_ADMIN__TLS__ENABLED

Anything but http://localhost needs this on, or the browser will not store the session cookie at all.

cert_path / key_path (String) — Defaults: "admin.pem" / "admin.key" | Env: ACME_PROXY_ADMIN__TLS__CERT_PATH, ACME_PROXY_ADMIN__TLS__KEY_PATH

PEM chain (leaf first) and its private key. When either is missing, a self-signed certificate for the host of admin.base_url is generated and written at startup; the generated key is created 0600. Separate paths from [server.tls] on purpose — the two listeners answer to different names and should not share a certificate by accident.

handshake_timeout_ms (Integer) — Default: 10000 | Env: ACME_PROXY_ADMIN__TLS__HANDSHAKE_TIMEOUT_MS

As [server.tls].


[challenge]

Which challenge types each new authorization offers, whether they are validated at all, and the per-type keys under [challenge.http_01] and [challenge.tls_alpn_01].

Documented in full in Challenge Validation — this section is a per-profile subsystem with its own chapter, so its keys live there rather than being restated here.


[order]

validity_seconds (Integer) — Default: 604800 (7 days) | Env: ACME_PROXY_ORDER__VALIDITY_SECONDS

Lifetime of the ACME order object before it expires. This is housekeeping for the order resource, not the issued certificate’s validity — that is signer.local_ca.leaf_validity_days, or whatever the delegating backend decides.

max_identifiers (Integer) — Default: 100 | Env: ACME_PROXY_ORDER__MAX_IDENTIFIERS

Most identifiers a single newOrder may name. A request naming more is refused with urn:ietf:params:acme:error:malformed.

The only bound before this was server.max_body_bytes (128 KiB), which at roughly thirty bytes per identifier admits some four thousand names in one request. Every one of them becomes an authorization plus a challenge per offered challenge type, all inserted in a single transaction — and SQLite has one writer, so that transaction stalls every other write in the process while it runs. 100 is what Let’s Encrypt allows and is far above what a real client asks for; the point is that there is a ceiling.

Refused as malformed rather than rateLimited deliberately: the order is malformed for this server whenever it is sent, so rateLimited — which tells a client to come back later — would be a lie.

retention_days (Integer) — Default: 30 | Env: ACME_PROXY_ORDER__RETENTION_DAYS

Days an order is kept after it expires, before the daily order_sweep deletes it. The authorizations and challenges beneath it go with it, through the schema’s ON DELETE CASCADE. 0 keeps everything for ever.

A valid order is never swept, whatever its age. Its row is how revokeCert and the CRL resolve a certificate by serial, and what RFC 9773 renewal information is derived from — deleting one would make an issued certificate unrevokable and unrenewable, which is a far worse outcome than a large table. Only orders that ended some other way are eligible: invalid, or an abandoned pending/ready/processing order past its own expires, at which point no client can act on them either.

Nothing pruned orders before this key existed, so the table and its children grew for the life of a deployment — one order plus one authorization per identifier plus one challenge per offered type, on a default configuration where newAccount is open to anyone.

Per-profile like the rest of [order]. The sweep is a single job handler (JobRegistry refuses two handlers for one kind) that applies each mounted profile’s own value to that profile’s rows, emitting one order_reaper_swept line per profile.


[nonce]

ttl_seconds (Integer) — Default: 300 | Env: ACME_PROXY_NONCE__TTL_SECONDS

Freshness window for JWS anti-replay nonces. Expired nonces are swept on an interval for the life of the process.


[audit]

Traceability and the CA’s audit trail. Process-wide, not per-profile — the trail describes the CA, not one of its endpoints, so this section may not appear under [profiles.<name>].

There is deliberately no enabled key. The address columns on accounts and orders, and the audit_log table, are always written: recording who asked the CA to sign something is what a CA does, not a feature to switch on. The only thing here that can be turned off is the reverse lookup.

reverse_dns (Boolean) — Default: true | Env: ACME_PROXY_AUDIT__REVERSE_DNS

Resolve a PTR record for the client’s address and freeze it into the row beside the address. Turn it off on an estate with no usable reverse zone: every lookup would fail, every *_ptr column would end up NULL anyway, and all that would be left is the round trip. Lookups go through dns.resolver like every other DNS query this server makes.

reverse_dns_timeout_ms (Integer) — Default: 2000 | Env: ACME_PROXY_AUDIT__REVERSE_DNS_TIMEOUT_MS

Budget for one PTR lookup. Deliberately small: this runs inside a request that has already done its real work, so a slow nameserver costs a NULL in one column rather than latency on issuance. Every failure is a NULL, never a refused request.

retention_days (Integer) — Default: 0 | Env: ACME_PROXY_AUDIT__RETENTION_DAYS

Delete audit_log rows older than this many days. 0 keeps everything for ever, which is the right default for a trail whose value is that it is complete. Any non-zero value schedules a daily audit_sweep job beside the nonce one, running the same DELETE as acme-proxy audit cleanup --older-than <days>.

See Audit Trail.


[jobs]

The durable background-work queue: the jobs table plus the one runner that drains it. Process-wide, not per-profile — there is one queue and one runner for the process, so this section may not appear under [profiles.<name>].

There is deliberately no enabled key. The queue is how the server finishes work it has already promised a client: an order answered processing is owed a certificate. Switching it off would not disable a feature, it would strand the orders. What is tunable is how hard and how long the server tries.

Most of what the server does after answering a request runs here, so this section’s reach is wider than the name suggests:

  • Client-visible work — challenge validation (challenge_validate), issuance (signer_issue, and signer_relay_issue under the relay backend), revocation through a relay or script (signer_revoke), and a local CA’s CRL signing (local_ca_crl_regenerate).
  • Deliveries — every notification (notify_deliver) and the expiry digest (notify_expiry_digest).
  • Periodic sweeps — expired nonces, orders past order.retention_days, audit.retention_days, expired admin sessions, stale http-01 tokens, the daily CRL refresh (local_ca_crl_sweep), and this queue’s own retention_days.

A runner that is not running is a server that validates, issues, sweeps and notifies nothing — job_runner_started is the line that says it is. With role processes, only a worker process runs one.

poll_interval_ms (Integer) — Default: 1000 | Env: ACME_PROXY_JOBS__POLL_INTERVAL_MS

How often the runner looks for work nobody woke it for. This bounds only scheduled work — a backoff coming due, a periodic sweep firing. Queueing a job wakes the runner directly, so a relay queued by finalize starts immediately whatever this says, which is what keeps a polling ACME client from waiting on a tick.

max_concurrent (Integer) — Default: 8 | Env: ACME_PROXY_JOBS__MAX_CONCURRENT

How many jobs may run at once. A restart after an upstream outage that left a few thousand orders in flight would otherwise become a few thousand concurrent pollers against one CA, which is how a recoverable backlog turns into a rate-limit ban.

max_attempts (Integer) — Default: 5 | Env: ACME_PROXY_JOBS__MAX_ATTEMPTS

How many attempts a job gets before it is retired permanently. Counted when the job is claimed, so one that kills the process still exhausts its budget rather than crash-looping. With the 30-second base below and doubling, five attempts span roughly seven and a half minutes.

Unlike every other key here, this one is frozen onto each row as it is queued rather than read afresh each attempt, so raising it applies to work queued from then on and not to a backlog already waiting.

retry_base_seconds (Integer) — Default: 30 | Env: ACME_PROXY_JOBS__RETRY_BASE_SECONDS

The first retry delay; each subsequent one doubles. Longer than any single upstream round trip, so a retry is not simply the same failure again, and short enough that a blip clears inside one client poll cycle.

retry_max_seconds (Integer) — Default: 3600 | Env: ACME_PROXY_JOBS__RETRY_MAX_SECONDS

Where the doubling stops, and a real ceiling — the jitter applied to each delay only ever subtracts, so no retry is scheduled past this. An hour sits well under the default order lifetime, which keeps a job’s own deadline the binding constraint rather than this.

lease_seconds (Integer) — Default: 300 | Env: ACME_PROXY_JOBS__LEASE_SECONDS

The default budget for one attempt, and therefore how long a claim is held before another runner may take the row. A handler needing a different one says so itself: the relay signer asks for its own signer.relay.poll_timeout_secs.

retention_days (Integer) — Default: 7 | Env: ACME_PROXY_JOBS__RETENTION_DAYS

Delete settled job rows older than this many days. Unlike audit.retention_days this defaults to a non-zero value: a finished job is a receipt, not evidence, and the trail that has to be complete is audit_log’s. 0 keeps everything for ever and stops the sweep being scheduled at all.


[metrics]

The Prometheus exposition, served on a listener of its own — a third socket beside the ACME and admin ones, not a route on either. Process-wide, not per-profile: there is one counter set for the process, and the endpoint is a dimension of it rather than something each endpoint configures for itself.

The separate port is what settles the access question. A scrape carries no session and needs none, because reaching the port at all is the permission — so the control is your firewall, not a credential this server checks. Putting it on the ACME listener would have meant an unauthenticated route on a public socket; putting it on the admin listener would have meant an auth exemption on a listener whose rule is that every route but sign-in needs a session, plus coupling metrics to the panel being enabled.

There is deliberately no path key — /metrics is what every scrape configuration already assumes, and this listener serves nothing else. There is no [metrics.tls] either, unlike [server.tls] and [admin.tls]: those carry a client’s signed requests and an operator’s session cookie, while a scrape carries no credential and the exposition holds no secret.

What is exposed, and how to point Prometheus at it, is in Monitoring.

enabled (Boolean) — Default: false | Env: ACME_PROXY_METRICS__ENABLED

Bind the metrics listener and serve GET /metrics. Off by default, the same posture [admin] takes: a certificate authority should not open a new socket because somebody upgraded it. Off means no socket at all, rather than one answering 404.

bind_address (String) — Default: 127.0.0.1:3002 | Env: ACME_PROXY_METRICS__BIND_ADDRESS

The socket the exposition is served on — beside the other two, server on 3000 and admin on 3001.

Unlike admin.bind_address, a non-loopback value is neither refused nor warned about. The admin listener refuses one without TLS because its cookie is always Secure, which a browser silently declines to store over plain HTTP, so the symptom would be an unexplained sign-out loop. Nothing here has a cookie, and a port reachable from a Prometheus host on another machine is exactly the intended deployment.

A value equal to server.bind_address, or to admin.bind_address while the panel is enabled, is refused — at startup and on a reload alike — since only one of the three could then bind.

Both keys reload: SIGHUP moves this listener, or switches it on and off, without restarting. The new socket is bound before anything is published, so an address that cannot be bound refuses the reload and leaves the running one serving. See Configuration Reload.


[dns]

resolver (String) — Default: unset (system configuration, i.e. /etc/resolv.conf) | Env: ACME_PROXY_DNS__RESOLVER

host:port of the nameserver every DNS lookup this server makes goes through: the dns-01 TXT query, the connect target that http-01/tls-alpn-01 resolve before reaching out, and filter.reverse_dns’s PTR and forward lookups.

The shared resolver is deliberately uncached, so a TXT record published moments before a challenge is triggered is not defeated by a cached negative answer. filter.reverse_dns is the one exception and keeps its own cached resolver.


[proxy]

The forward proxy every outbound HTTP client dials through: the upstream CA the relay signer talks to, the IPAM inventory, the notification webhooks, and the http-01 and tls-alpn-01 challenge validators — the last through a CONNECT tunnel, since it is TLS rather than HTTP.

Not everything outbound. SMTP (notify.email) and the RFC 2136 updates signer.relay.dns01 makes are not HTTP and keep dialling directly; an estate whose egress is proxy-only needs a separate route for those two.

Every key is empty by default, which means no proxy at all. There is no enabled key — the presence of a URL is the switch.

http_url (String) — Default: "" | Env: ACME_PROXY_PROXY__HTTP_URL

Proxy for http:// targets, e.g. http://proxy.example.com:3128. Falls back to $http_proxy when empty. A cleartext target is forwarded rather than tunnelled: the request line carries the whole URL, which is what RFC 9112 §3.2.2’s absolute-form is for.

https_url (String) — Default: "" | Env: ACME_PROXY_PROXY__HTTPS_URL

Proxy for https:// targets, reached by CONNECT. Falls back to $https_proxy, then $HTTPS_PROXY.

Normally the same http://proxy.example.com:3128 as http_url: this key names the proxy used for https targets, not a proxy spoken to over https. An https:// value is a startup error rather than a second TLS layer with no trust anchor configured for it, and so is a socks5:// one.

The two keys are independent on purpose. An estate that proxies only its TLS egress is ordinary, and “it worked for http and silently did nothing for https” is the failure a single key would produce.

no_proxy (Array<String>) — Default: [] | Env: ACME_PROXY_PROXY__NO_PROXY

Targets that bypass the proxy. An entry is * (everything), a domain — which also matches everything under it — a .domain (the same thing), an address, or a CIDR block. Matching is case-insensitive and ignores a trailing root dot.

A network entry is compared only against a target that is already an address literal: a hostname is never resolved to test one, which would mean a DNS lookup on every outbound request and a race with the connect that follows.

An entry carrying a port is a startup error. Matching is on the host, and an entry that silently ignored half of itself is worse than one that is refused.

Loopback and localhost bypass unconditionally, before this list is consulted. That rule exists because of the environment fallback: an operator’s inherited shell http_proxy must not route this server’s own loopback traffic through a corporate proxy, and the failure that would cause carries no signal at all.

Environment fallback

Each key falls back to its conventional variable when left empty:

keythenthen
http_url$http_proxy—
https_url$https_proxy$HTTPS_PROXY
no_proxy$no_proxy$NO_PROXY

An empty string counts as unset in both sources, so a ${VAR:-} shell default does not become a proxy at the empty URL.

Uppercase HTTP_PROXY is deliberately not read. Under CGI a client-supplied Proxy: request header lands in the environment under exactly that name (httpoxy, CVE-2016-5385 and its siblings). This server is never a CGI process, so the vector does not reach it — but Go’s net/http dropped the variable for this reason, matching it costs nothing, and honouring a variable purely because everything else does is the kind of decision that is only ever wrong. HTTPS_PROXY has no such history and is honoured.

Credentials

Userinfo in the URL becomes a Proxy-Authorization: Basic header: http://user:password@proxy.example.com:3128. Percent-encode anything unusual — DOMAIN%5Cuser is decoded before the header is built, since a backslash or an @ in a proxy username is entirely ordinary and encoding it verbatim sends the wrong credential.

On a tunnelled connection the credential is spent on the CONNECT and is not repeated inside the tunnel, where the origin server would read it. The configured URL is never logged with its password: every log line and every error message renders it with the password replaced.

When the proxy is down

A configured but unreachable proxy is an error, every time. There is no fallback to a direct connection: dialling around a controlled egress path at exactly the moment the control fails is the opposite of what the setting is for. Errors name the proxy rather than the origin, so a refused connection points at the host that actually refused it.


[meta]

The optional meta members of the directory (§7.1.1). All are empty by default and omitted from the directory when empty — never sent as an empty value.

terms_of_service (String) — Default: "" | Env: ACME_PROXY_META__TERMS_OF_SERVICE

URL of your terms of service. This one has teeth: setting it turns on §7.3.3, so newAccount then refuses any request without termsOfServiceAgreed: true (403 userActionRequired + a Link: rel="terms-of-service" header), and account objects begin reflecting termsOfServiceAgreed.

website (String) — Default: "" | Env: ACME_PROXY_META__WEBSITE

Informational URL about the ACME server. Advertised only.

caa_identities (Array) — Default: [] | Env: ACME_PROXY_META__CAA_IDENTITIES

Hostnames this CA recognizes in CAA records. Advertised only — this server performs no CAA checking.


[logging]

All six keys are validated at startup: an unknown value is a refusal to start with a message naming the key, never a silent fallback — a CA running at a log level or to a destination its operator did not ask for is the worse failure. The same validation runs on a reload, where a bad value refuses the whole thing rather than half-swapping the log stream.

All six also reload on SIGHUP, so raising the level or switching to JSON mid-incident does not cost a restart.

What the resulting records actually contain, and what to alert on, is Monitoring & Observability.

filter (String) — Default: "acme_proxy=info" | Env: ACME_PROXY_LOGGING__FILTER

EnvFilter directive, and the last of three layers to be consulted. The precedence, highest first:

  • --log-level, typed on the command line. A flag was asked for here and now, where the other two are ambient — the same reasoning that has --color always outrank NO_COLOR. It sets acme_proxy alone, so it never turns on a dependency’s logging by accident.
  • RUST_LOG, when set to a non-empty value. It replaces the whole filter, so a bare RUST_LOG=debug also turns on debug logging for every dependency.
  • this key.

That precedence is the same on a reload as at startup, which means editing this key while either of the two outranks it changes nothing; the server logs server_logging_filter_overridden, whose source field names which one, rather than letting the edit pass for applied.

This section describes the server’s log stream. An admin command emits nothing at all unless asked — see the CLI’s --log-level.

json_format (Boolean) — Default: false | Env: ACME_PROXY_LOGGING__JSON_FORMAT

Emit JSON instead of the human-readable format. Set this in production if you ship logs to ELK, Loki, Datadog or similar: the structured fields become first-class keys rather than text to be re-parsed.

target (String) — Default: "stdout" | Env: ACME_PROXY_LOGGING__TARGET

Where records are written: stdout or stderr. Any other value is a startup error.

ansi (Boolean) — Default: true | Env: ACME_PROXY_LOGGING__ANSI

ANSI colour in the human-readable format. Turn it off when the log is piped to a file or a collector that does not strip escape sequences. Ignored when json_format is on.

span_events (String) — Default: "none" | Env: ACME_PROXY_LOGGING__SPAN_EVENTS

Span lifecycle records: none, close or full. Any other value is a startup error. close emits one record as each span ends, carrying the time spent busy and idle inside it — the closest thing to per-operation timing available without a metrics endpoint, and cheap enough to leave on. full adds new/enter/exit and is a debugging tool.

flatten_event (Boolean) — Default: false | Env: ACME_PROXY_LOGGING__FLATTEN_EVENT

JSON only: lift a record’s own fields (event, and everything beside it) to the top level instead of nesting them under fields. What most pipelines want; off by default because a field can then collide with one of the format’s own keys.


[profiles.<name>]

An ACME endpoint is a profile. The server serves ACME only through them, and at least one enabled profile is required — startup fails otherwise, with a copy-pasteable minimal configuration in the error.

# The whole minimum. `enabled` defaults to true, so naming it is enough.
[profiles.default]

enabled (Boolean) — Default: true | Env: ACME_PROXY_PROFILES__<NAME>__ENABLED

Parks a profile without deleting its configuration. It also doubles as the one key an environment-only profile needs: ACME_PROXY_PROFILES__DEFAULT__ENABLED=true defines a working profile with no configuration file at all.

Load-bearing rules:

  • The mount path is derived from the name, never configured. [profiles.le] answers at {base_url}/profile/le/directory. Names must match ^[a-z0-9-]+$. The name is public API — it appears in every kid and order URL a client stores — so renaming a profile invalidates every client’s saved account.
  • Inheritance is per key, not per section. The global [signer], [filter], [challenge], [eab], [order], [notify] and [meta] sections are the base each profile overlays. A profile that sets only challenge.bypass keeps the global challenge.enabled rather than reverting it to the compiled default. Precedence: profile key > global key > compiled default.
  • Arrays replace wholesale, never append. A profile’s filter.rules fully replaces the global one — order is the policy, so it is stated once per profile rather than accumulated from two places.
  • Profiles are a database boundary, not just a URL prefix. Accounts and orders carry a profile column, and accounts are keyed UNIQUE(profile, pubkey) — one client key used at two endpoints is two independent ACME accounts.
  • Signer backends are shared by configuration. Two profiles with identical [signer] sections share one backend instance. Two profiles sharing a local_ca key path while differing elsewhere is a startup error.
[signer]
backend = "local_ca"

[filter]
rules = ["corp-only"]

[filter.check.corp-net]
type  = "allowed_ip"
allow = ["10.0.0.0/8"]

[filter.rule.corp-only]
when = "corp-net"
then = "allow"

# Inherits everything above, but dry-runs the one rule: `[filter.rule.<name>]`
# is a table, so overriding `mode` keeps `when` and `then` from the global one.
[profiles.dev]
filter.rule.corp-only.mode = "warn"

# Overrides two keys; keeps the whole filter policy.
[profiles.prod]
signer.backend = "relay"
signer.relay.directory_url = "https://acme-v02.api.letsencrypt.org/directory"

See Profiles & Routing.


Environment variable gotchas

These bite in production and produce no error, so they are worth knowing before you configure anything through the environment.

Array-valued keys are parsed from a comma-separated string. That means a value containing a literal comma cannot be expressed. A regex such as ^host\d{2,3}\.example\.com$ is therefore file-only — through the environment, {2,3} splits into two list entries. In a file, write the array form (deny = ["^host\d{2,3}\.example\.com$"]), which is taken exactly as written; a bare string in a file is split on commas too, since by the time the value is read there is nothing left to say where it came from.

Items are trimmed, and an empty item is a startup error. a, b is ["a", "b"], and a,,b or a trailing comma is refused by name rather than silently dropped or kept as an entry that matches nothing.

An array set to the empty string is the empty list. Shell defaults like ACME_PROXY_FILTER__RULES="${RULES:-}" set the variable to "", which the configuration layer cannot distinguish from a deliberate value — so it is read as “no values”, which is what clears a list the file set. Do not expect "" to mean “fall back to the file”.

A numeric-looking value loses its leading zeros. The environment source parses 007 as a number before the list is built, so it arrives as 7. A value whose spelling matters — an argument to a custom signer, say — belongs in the file’s array form.

Unknown keys are ignored, not rejected. A misspelled key, or a key written under the wrong section, is silently dropped. The most common instance of this is trusted_proxies written under [server] when it belongs to [filter] — see Allowed IP.


Sections documented elsewhere

The six sections with a chapter of their own, expanded — see the map above for the rest.

Configuration Scenarios

This page outlines complete, practical examples of configuring acme-proxy for different real-world use cases. Each block is a whole config.toml.

All three scenarios below use the relay backend. Starting the server registers an account at directory_url — see Relay. Use a staging endpoint while you are still working the configuration out.

Let’s Encrypt relay with DNS validation

This scenario configures acme-proxy to act as an internal relay. It intercepts ACME clients locally, but ultimately relays the issuance requests to Let’s Encrypt.

The proxy takes the burden of solving Let’s Encrypt’s DNS-01 challenges on behalf of internal users by utilizing an RFC 2136 dynamic DNS provider. Internal clients never see a DNS credential; the single TSIG key lives here.

[server]
base_url = "https://acme.internal.company.com"
bind_address = "[::]:3000"

[signer]
backend = "relay"

[signer.relay]
directory_url = "https://acme-v02.api.letsencrypt.org/directory"
account_key_path = "le_upstream.key"
contact = ["mailto:admin@company.com"]
challenge_strategy = "dns01"
poll_interval_ms = 2000
poll_timeout_secs = 300

[signer.relay.dns01]
provider = "rfc2136"

[signer.relay.dns01.rfc2136]
server = "10.0.0.53:53"
zone = "internal.company.com."
tsig_key_name = "acme-update-key"
# MUST be standard base64 (not base64url). Prefer the environment variable
# ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_KEY_SECRET to a file.
# A non-base64 value here is a startup error, not a runtime one.
tsig_key_secret = "c2VjcmV0LXJlcGxhY2UtbWU="
tsig_algorithm = "hmac-sha256"

[challenge]
# Require local clients to prove control to the proxy via HTTP-01
enabled = ["http-01"]
bypass = false

[profiles.default]
enabled = true

The key above writes only inside internal.company.com.. For names spread over several zones, DNS alias mode points every domain’s challenge at one alias zone instead of needing a profile per zone.

Public CA relay with HTTP validation

The same relay as above, for an operator who has no RFC 2136 write access to the zone but does control the reverse proxy already fronting the names being issued. Instead of publishing a TXT record, acme-proxy serves the upstream’s challenge file itself.

[server]
base_url = "https://acme.internal.company.com"
bind_address = "[::]:3000"

[signer]
backend = "relay"

[signer.relay]
directory_url = "https://acme-v02.api.letsencrypt.org/directory"
account_key_path = "le_upstream.key"
contact = ["mailto:admin@company.com"]
# No [signer.relay.http01] table exists — this line is the whole
# configuration. The responder is a route on this server's own root router;
# acme-proxy does NOT open a second listener or bind port 80.
challenge_strategy = "http01"
poll_interval_ms = 2000
poll_timeout_secs = 300

[challenge]
# Local clients still prove control to the proxy independently.
enabled = ["http-01"]
bypass = false

[profiles.default]
enabled = true

This only works if the upstream CA’s fetch reaches acme-proxy. Add one location to whatever already answers on port 80 for each name being issued:

location /.well-known/acme-challenge/ {
    proxy_pass http://acme-proxy:3000;
    # A `return 301 http://acme-proxy:3000$request_uri;` works equally well —
    # RFC 8555 §8.3 permits following redirects.
}

Note this strategy cannot issue wildcards (nothing answers HTTP on the name *.example.com); those need scenario 1’s dns01. See Relay for Caddy and Traefik equivalents.

Commercial ACME CA with EAB and NetBox filtering

This scenario uses a commercial CA backend. It enforces External Account Binding (EAB) on the internal proxy so only authorized users can register. It additionally uses NetBox, through the ipam filter, to verify if a client’s IP is actually permitted to request a certificate for a specific DNS name.

[server]
base_url = "https://ca.internal.company.com"

[signer]
backend = "relay"

[signer.relay]
directory_url = "https://commercial-ca.example.com/acme/directory"
account_key_path = "commercial_upstream.key"

# The commercial CA already trusts this server's upstream account implicitly,
# so the upstream challenge is bypassed.
challenge_strategy = "bypass"

[eab]
# Require all internal clients to register with an EAB credential generated by the admin
enabled = true

[filter]
# Ask the inventory about every internal request
enabled = ["ipam"]

[ipam]
backend = "netbox"
timeout_ms = 5000

[ipam.netbox]
url = "https://netbox.internal.company.com"
# Store this in ACME_PROXY_IPAM__NETBOX__TOKEN ideally
token = "your_netbox_read_only_token"
custom_field = "acme_domains"
sources = ["dns_name", "custom_field", "device"]

[profiles.default]
enabled = true

Protocol Support

acme-proxy implements RFC 8555 in full, plus the extensions an ACME client in 2026 expects to find. This page is the conformance summary: what a client can call, which RFC section governs it, and what is deliberately not implemented.

Everything here is served per profile, under /profile/<name>/. There is no ACME at the bare root — see Profiles & Routing.

Resources

Every row is reachable at {base_url}/profile/<name><path>. The Advertised column says whether the directory object names it: a client is expected to find advertised resources by reading the directory rather than by constructing paths.

PathMethodRFCAdvertisedNotes
/directoryGET, POST§7.1.1—It is the entry point; §6.3 requires POST-as-GET to work too.
/newNonceHEAD, GET, POST§7.2yesAll three forms.
/newAccountPOST§7.3yesFind-or-create by public key: 201 for a new account, 200 for an existing one, Location either way.
/acct/{id}POST§7.3.2, §7.3.6—Contact update and deactivation. kid-authenticated. There is deliberately no unauthenticated GET.
/acct/{id}/ordersPOST§7.1.2.1—The account’s order list, filtered as §7.1.2.1 requires.
/keyChangePOST§7.3.5yesKey Rollover.
/newOrderPOST§7.4yesAccepts notBefore/notAfter and RFC 9773’s replaces.
/order/{id}POST§7.1.3—POST-as-GET.
/order/{id}/finalizePOST§7.4—Takes the CSR; hands it to the configured signer.
/authz/{id}POST§7.5, §7.5.2—One URL serves both the read and the deactivation, told apart by whether a payload arrived.
/chall/{id}POST§7.5.1—Triggers validation. Both outcomes are 200.
/certificate/{id}POST§7.4.2—POST-as-GET; PEM chain.
/revokeCertPOST§7.6yesRevocation & CRL.
/renewalInfo/{certID}GETRFC 9773 §4.1yesUnauthenticated. Advertised without the id — §4.1 has the client append it.
/crlGETRFC 5280noRouted but not advertised: a CRL is CA infrastructure, not an ACME resource.
/ca.pemGET—noThe profile’s trust anchor, so installing it is one curl. 404 unless the backend has one of its own. Not advertised, for the same reason as /crl. See Trusting the CA.

Two paths sit outside every profile, on the root router:

PathPurpose
/healthLiveness. Outside the filter chain, the admission limiter and the nonce middleware — see Monitoring.
/.well-known/acme-challenge/{token}Mounted only when a signer backend has an http-01 token store to publish, i.e. signer.relay.challenge_strategy = "http01". See Relay.

Two further paths are not on this listener at all. GET /metrics has a socket of its own (off by default, [metrics]), so firewalling that port is what controls who can read it — see Monitoring. The web admin’s /ui and /api likewise have their own listener; see Web Admin.

Protocol behaviour worth knowing

Every signed request is checked the same way. The media type (application/jose+json), any crit header, the JWS url against the route actually reached, and the nonce are all verified before a handler runs, so no resource can forget one. jwk and kid are mutually exclusive per §6.2, and the verification algorithm never rests on the client’s alg alone. See Architecture.

Refusals are problem documents. Every rejection is application/problem+json with an RFC 8555 URN type. A multi-identifier order rejected on several names comes back as one compound problem with a subproblems entry per identifier (§6.7.1); a single rejection stays its own type, unwrapped.

A POST always carries a fresh Replay-Nonce, errors included (§6.5). GET /directory, GET /crl, GET /ca.pem and GET /renewalInfo do not mint one — nothing asks them to, and each nonce is a committed database write.

Extensions

ExtensionRFCDefaultPage
External Account Binding§7.3.4offEAB
Account key rollover§7.3.5always onKey Rollover
Renewal Information (ARI)RFC 9773always onARI
Terms of service§7.3.3offSet meta.terms_of_service and newAccount starts enforcing it.
Wildcard identifiers§7.1.3requires dns-01Challenge Validation

Directory metadata

The directory’s meta object carries only what is configured. An unset member is omitted, never sent empty — "website": "" says less than saying nothing.

  • externalAccountRequired appears as true when EAB is on for that profile.
  • termsOfService, website and caaIdentities come from [meta].

meta.terms_of_service is the one with teeth: setting it turns on §7.3.3, so newAccount then refuses a request without termsOfServiceAgreed: true (403 userActionRequired plus a Link: rel="terms-of-service" header), and the account object starts reflecting the field.

Not implemented

Stated explicitly, because each is something a reader may reasonably expect:

  • CAA checking. meta.caa_identities is advertised to clients; this server does no CAA lookup of its own. Where the relay backend is in use, the upstream CA performs its own.
  • OCSP. Revocation is published as a CRL at GET /crl. There is no OCSP responder, and no authorityInfoAccess OCSP pointer is ever written into an issued certificate. The local CA does write the caIssuers half of that extension, and a cRLDistributionPoints pointer, once an operator names the URLs — see Local CA.
  • Identifier types other than dns. newOrder accepts DNS names, including wildcards; ip identifiers (RFC 8738) and permanent-identifier are not supported.
  • Pre-authorization (§7.4.1). The directory does not advertise newAuthz, and it is not routed. Authorizations exist only as part of an order.
  • POST to /renewalInfo — RFC 9773 §4.3’s optional client-side renewal signal. The GET half is implemented.

External Account Binding (EAB)

acme-proxy supports enforcing External Account Binding (RFC 8555 §7.3.4). When enabled, any client attempting to register a new account must provide an EAB credential that was minted out-of-band by the server operator.

This is a highly effective security mechanism for internal CA endpoints: it restricts account creation to authorized entities without relying purely on IP filtering.

Solving commercial CA EAB limits

Beyond security, EAB support in acme-proxy solves a significant operational hurdle when using Commercial CAs (like ZeroSSL, Sectigo, or GlobalSign).

When working with external or commercial CAs, organizations are sometimes restricted to a single account or a limited number of EAB credentials validated for specific domains. It is impractical to distribute these scarce upstream credentials to hundreds of individual internal servers.

By placing acme-proxy in front of the commercial CA:

  1. The proxy consumes a single upstream EAB credential to register its own master account with the Commercial CA.
  2. The proxy then issues its own unlimited local EAB credentials to your internal servers.
  3. This effectively multiplexes the upstream account, allowing thousands of internal clients to securely acquire certificates without exhausting your upstream quotas or spreading sensitive upstream secrets across your infrastructure.

Configuration

[eab]
enabled = true

Reference

enabled (Boolean) — Default: false | Env: ACME_PROXY_EAB__ENABLED

When enabled, the newAccount endpoint will refuse any request that doesn’t carry a valid, unused EAB payload. Standard onlyReturnExisting lookups are exempt because they only query existing accounts and never create new ones.

Minting credentials via CLI

You manage EAB credentials with the Admin CLI, because the secret is sensitive and is shown exactly once.

To create a new credential for a client:

acme-proxy eab create --label "DevOps Team"

# Bind it to a single endpoint in a multi-profile deployment:
acme-proxy eab create --label "DevOps Team" --profile prod

The CLI prints a kid (Key Identifier) and an HMAC secret. The secret is printed only once. It is stored but never displayed again, so a lost secret is replaced, not recovered.

--profile matters in a multi-tenant deployment: omitted, the credential is accepted at every profile, which is what an unscoped credential means. Bind it unless you intend that.

A client then uses the credential at registration:

certbot register \
  --server https://acme.internal/profile/default/directory \
  --eab-kid "the-kid" \
  --eab-hmac-key "the-secret"

Credentials are reusable, not single-use

A credential is not consumed by the account it creates. The same kid can bind any number of accounts and keeps working until you explicitly revoke it. There is no used state — only active and revoked.

This is worth being deliberate about. Handing one credential to a team means any number of hosts can register with it, and a leaked credential stays valid until someone notices. If you want one-client-one-credential, mint one per client and revoke it once that client has registered.

(This differs from the upstream credential consumed by acme-proxy upstream register — or, alternatively, signer.relay.eab in configuration. Commercial CAs typically issue single-use EAB credentials, which is precisely why that one is consumed once by registration and then no longer needed at all, unlike the credentials described on this page.)

Revocation

Revoke a credential to prevent any further use:

acme-proxy eab revoke <kid>

This takes effect immediately, with no restart — credentials are read from the live database on every newAccount. Revocation is idempotent.

Revoking does not disturb accounts that already registered with the credential; it only stops new ones from being bound. To shut out an existing account, deactivate it: acme-proxy account deactivate <id>.

A revoked credential keeps its row, so the accounts it registered still resolve to it — its label keeps matching eab filter checks, and require_active can refuse them.

Deleting a credential

Deleting removes the row. What happens to the accounts registered with it is your choice:

acme-proxy account list --eab-kid <kid>             # who registered with it
acme-proxy eab delete <kid>                         # keep the accounts
acme-proxy eab delete <kid> --deactivate-accounts   # retire them
acme-proxy eab delete <kid> --delete-accounts       # remove them
  • Keep (the default) leaves the accounts as they are. They still name the deleted kid but resolve to no credential, so every eab filter check refuses them from then on.
  • --deactivate-accounts moves them to deactivated: they can request nothing more, and their orders are kept, so every certificate they hold can still be revoked and still appears in the expiry digest until it expires. This is the way to retire a tenant.
  • --delete-accounts deletes them with their orders, authorizations and challenges. It is refused while any of those orders holds a live certificate (issued, not revoked, not expired): the order is the only record of the certificate, and without it the certificate could never be revoked. Revoke those certificates or wait for them to expire, or deactivate instead. A refusal changes nothing, not even the credential.

No choice revokes a certificate. The panel’s credential page offers the same three, and the JSON API takes them as DELETE /api/eab/{kid}?accounts=keep|deactivate|delete. Each writes an eab_deleted audit row, plus one account_deactivated or account_deleted row per account it changed.

Inspecting credentials

acme-proxy eab list --json
acme-proxy eab show <kid>

Neither ever renders the secret. eab list is newest first and paged (Admin CLI → Paging); it reads the same query GET /api/eab and the panel’s /ui/eab do, so the three cannot come to describe the credential set differently.

Key Rollover

acme-proxy fully supports Account Key Rollover (RFC 8555 §7.3.5) via the POST /keyChange endpoint.

This allows a client to proactively rotate the cryptographic key-pair associated with their ACME account without needing to register a new account or abandon their existing authorizations and orders.

Cryptographic verification

The key rollover process is highly secure and requires the client to prove possession of both the old and the new key simultaneously to prevent hijacking.

  1. Inner JWS: The client constructs a JSON Web Signature (JWS) whose payload names the account and the new key, signed by the new key itself.
  2. Outer JWS: That inner JWS becomes the payload of an outer JWS, signed by the old key currently on the account and carrying its kid.
  3. acme-proxy unwraps this nested structure and verifies both signatures (ES256 or RS256, via ring). The outer signature proves the request comes from the current account holder; the inner one proves possession of the new key, so a key the requester does not control cannot be installed.
  4. Collision check: the proxy checks that the new key does not already belong to another account. If it does, the request is rejected with 409 Conflict plus a Location header naming the account that already holds the key.
  5. Update: the account’s stored public key is replaced in a single statement.

Because accounts are keyed UNIQUE(profile, pubkey), the collision check is scoped per profile — the same key legitimately belonging to a different account at a different profile is not a conflict.

Client support

Many modern clients support key rollover natively. For instance, using lego:

lego --server https://acme.internal/profile/default/directory \
     --email "admin@example.com" \
     accounts keyrollover

(Note: Ensure you are using a recent version of the client, as older clients may lack support for the keyChange endpoint).

Renewal Information (ARI)

acme-proxy implements the ACME Renewal Information extension (RFC 9773). ARI allows a CA to signal to ACME clients when they should ideally renew their certificates. This prevents the “stampede” effect where thousands of certificates expire simultaneously, and allows the CA to orchestrate graceful mass-revocations by shortening the renewal window.

GET /renewalInfo/{certID}

This is an unauthenticated endpoint clients use to ask “when should I renew this certificate?”. The response provides a time window (Start and End).

The window is resolved in this order:

  1. Revoked certificates win. If the certificate is known to be revoked locally (looked up by its serial), the answer is a window entirely in the past, prompting a compliant client to renew immediately. This is checked before the signer backend is consulted, so a locally revoked certificate is never talked out of renewing by an upstream CA that has not noticed yet.
  2. The backend’s own opinion, if it has one. The relay backend relays the upstream CA’s answer verbatim, including RFC 9773 §4.2’s optional explanationURL — so if Let’s Encrypt signals early renewal, that reaches your internal clients unchanged. The custom backend can do the same through its renewal_info hook, but only when signer.custom.supports_renewal_info = true; otherwise the hook is never invoked.
  3. A local estimate, when the backend has no opinion — always the case for local_ca. The certificate’s notBefore and notAfter are read by parsing the certificate itself, and the suggested window runs from ⅔ to ¾ of the validity period. For a 90-day certificate that is roughly day 60 to day 67, leaving a comfortable margin before expiry rather than running up to it.

Spreading clients across that window is what prevents the “stampede” effect where a whole fleet renews at the same moment.

The replaces field

During a newOrder request, an ARI-aware client can include a replaces field containing the certID of the certificate it intends to replace.

acme-proxy validates it against all three of RFC 9773 §5’s correspondence rules. The certID is decoded into an Authority Key Identifier and a serial number, the predecessor order is looked up, and then:

  1. The AKI must match the predecessor’s issuer — the serial alone is not enough, since serials are only unique per issuer. (A certificate stored before this check existed may have no recorded AKI; that case falls back to matching on serial alone, for backwards compatibility.)
  2. The predecessor must belong to the requesting account. You cannot claim to replace someone else’s certificate.
  3. The new order must share at least one identifier with the predecessor. A renewal that covers none of the same names is not a replacement.

Failing any of them is malformed; an unknown certID likewise.

Concurrency guard: a client cannot have two orders replacing the same predecessor. A second newOrder naming a replaces value already claimed gets 409 alreadyReplaced. An order that later becomes invalid releases its claim, so a failed attempt does not permanently block a retry.

The accepted replaces value is stored and reflected back on the 201 response and on every subsequent poll of the order.

Reloading the configuration

acme-proxy reloads its configuration on SIGHUP, without restarting:

sudo systemctl reload acme-proxy
# or, directly:
kill -HUP "$(pidof acme-proxy)"

Nothing is dropped. Connections stay open, in-flight ACME orders keep their state, and queued background work keeps its place. What changes is what the server does with the next request — and, if you moved a listener, on which port it answers it.

A reload is a rebuild and a swap, not a patch. Both routers, every profile’s filter chain and challenge registry, the notification backends, the job registry and both TLS acceptors are constructed fresh from the file on disk — and only once every one of them has succeeded is any of it published.

All or nothing

A reload either applies completely or changes nothing at all.

Three things stop one, and each says so in the log:

Log eventWhat happened
server_config_reload_refusedA key that cannot change while the process runs did change.
server_config_reload_failedThe file did not load, or what it asks for could not be built.
server_config_reloadedIt applied.

In the first two cases the server carries on with exactly the configuration it already had. That matters more than it sounds: a reload that applied the half it understood would leave a running server that no file on disk describes, and “what is this thing actually running?” would stop having an answer.

Every success carries a generation — 1 for the configuration the process started with, and one higher for each reload that landed. It is the quickest way to answer “did my SIGHUP take?”:

journalctl -u acme-proxy | grep server_config_reloaded

What a reload cannot change

One key. A refusal names it, what the server is running, and what the file now says:

`database.url` cannot be changed while the server is running
(running with `sqlite://sqlite.db`, the file now says `sqlite://other.db`):
restart to apply it
KeyWhy a restart is needed
database.urlThe connection pool is open, and the accounts and orders this CA has issued against it do not follow it elsewhere. A different database is a different CA.

Everything else reloads. That was not always true, and the four keys that came off this table most recently are worth knowing about because they are the ones an operator is most likely to remember as frozen: the set of enabled profiles, each profile’s [signer] section, dns.resolver and [proxy]. See Adding and removing endpoints below.

What a reload does change

Everything else, including the things operators reach for most:

  • Access policy — [filter] rules and checks, and the [ipam] inventory they consult. Tightening a rule takes effect on the next request.
  • The set of endpoints, and each one’s [signer] — see below.
  • Egress — dns.resolver and [proxy]. Every outbound client is rebuilt, including the signer backends, which used to be the reason these two were frozen.
  • Notification backends — a new webhook, a changed template directory, a different events list.
  • Challenge settings, order and account policy, [meta], [eab].
  • TLS certificates — server.tls.cert_path and key_path (and the admin listener’s own pair) are re-read, so a renewed certificate is served to the next connection while established ones are undisturbed.
  • The listeners themselves — see below.
  • The panel’s templates — admin.template_dir is recompiled. A template that does not parse fails the reload, so a mistake never reaches a browser.
  • Retention — audit.retention_days and jobs.retention_days. A sweep already scheduled keeps its current time and picks the new cutoff up on its next run.
  • Background work — every [jobs] key, so slowing a retry storm, widening a lease or raising concurrency mid-incident costs nothing. See below.
  • Logging — every [logging] key, so raising the level or switching to JSON mid-incident costs nothing. The swap is the first thing a reload publishes, so the server_config_reloaded line that confirms it is already under the new settings. One caveat: RUST_LOG still outranks logging.filter, exactly as it does at startup, so with it set an edited filter changes nothing — the server says so with server_logging_filter_overridden.

For what each key means, see the configuration reference.

Adding and removing endpoints

Add a [profiles.<name>] section and signal, and that endpoint starts serving:

server_config_reloaded generation=2 profiles=["le", "staging"]
profile_mounted profile=staging directory=https://ca.example/profile/staging/directory

Remove it and signal again, and it stops. Endpoints that were already running are not disturbed either way — no connection is dropped and no in-flight order loses its state, which is the whole reason this is worth doing without a restart.

Three things to know.

profile_mounted fires only for an endpoint that was not there before. It is a lifecycle event, delivered to whichever [notify] backends are configured, so re-firing it for every endpoint on every SIGHUP would make the notification surface noisiest in exactly the config-managed deployments that would least want it.

Unmounting keeps the data. The accounts and orders belonging to that endpoint stay in the database and come back exactly as they were if you mount it again — the profile name is what they are keyed on. Setting enabled = false is the same thing as removing the section.

Drain a relay endpoint before removing it. Unmounting the last profile a relaying backend serves takes its job handler with it, so any issuance still waiting on the upstream has nothing left to finish it. Those rows sit in the queue until the endpoint is mounted again, and the orders behind them expire. Nothing else is affected, and no other backend has background work to lose.

Editing a signer

A profile’s [signer] section reloads, including the one an endpoint is actively issuing with. The obvious worry — that a rebuilt local CA would forget what it had revoked — does not arise: revocations and the CRL live in the database, which the running CA and its replacement both read, so a revocation that lands during the reload is not lost either, and GET /crl answers identically across the swap. A relay serving http-01 keeps its published key authorizations in the database the same way, so an upstream CA fetching one mid-reload still gets it.

A backend whose section did not move is not rebuilt at all. That matters most with key_source = "pkcs11": an ordinary reload does not log in to the token again.

Two edges:

  • Changing where the CA material lives is a new CA. Point crl_path at a different file and the endpoint starts from that file’s revocation history, not the old one’s. That is the intended reading of the key, but it is worth saying, because the certificates already issued do not move with it.
  • [dns] and [proxy] rebuild every backend. They are not [signer] keys, but an outbound client caches them, so a change to either has to reach the signers to mean anything. Nothing is lost — the same handover applies.

Moving a listener

All seven keys that decide where a socket is, or whether there is one, reload: server.bind_address, server.tls.enabled, admin.enabled, admin.bind_address, admin.tls.enabled, metrics.enabled and metrics.bind_address. A reload that moved one says which:

server_config_reloaded generation=2 listeners_rebound=["acme"]

Three things are worth knowing before you use it.

A bad address refuses the reload; it does not take the socket down. Every new socket is bound before anything is published, so a port already in use, a name that does not resolve or a privileged port you no longer have the capability for is a server_config_reload_failed — with the listener that is already running still answering on the address it always had. Fix the file and signal again.

Established connections are not disturbed. A rebind replaces the socket new connections arrive on; a request already in flight finishes, and a keep-alive connection opened before the move stays usable until its client closes it. The old socket stops accepting immediately, so nothing new arrives there.

Turning TLS on or off does not move the socket at all. The mode is decided per connection, like the certificate: the next client to connect speaks the new protocol on the same port, and listeners_rebound stays empty. Remember to move server.base_url with it, or every signed request fails RFC 8555 §6.4’s URL check — the server warns tls_base_url_mismatch when it can see the two disagree.

One caveat on the panel. Switching admin.enabled off releases the socket and empties its router, so nothing answers on it — but it does not sign anybody out: sessions live in the database and are waiting when you switch it back on. Use acme-proxy admin session revoke --all if that is what you meant. Switching it off and on again does clear the login-attempt lockout, since the limiter goes with the panel.

Retuning the job runner

All seven [jobs] keys reload. These are the knobs you reach for while something is going wrong — an upstream CA rate-limiting you, a backlog draining too slowly — so a restart to apply them would have dropped exactly the in-flight orders you were trying to save.

The runner does not restart; it picks the new values up on its next pass and says so:

job_runner_retuned poll_interval_ms=250 lease_seconds=120 max_concurrent=16

That line is the confirmation worth grepping for. server_config_reloaded means a generation was published; this means the runner is actually running under it. It is only emitted when the pacing really moved, so reloads that touch other sections stay quiet.

Each key lands at its own grain, and the differences are all in the same direction — nothing already in flight is disturbed:

  • poll_interval_ms takes effect immediately, without waiting out the old interval first.
  • max_concurrent widens at once when raised. Lowered, it takes back the slots that are free and reaches the new figure as running jobs finish; no job is cancelled to get there sooner.
  • lease_seconds, retry_base_seconds and retry_max_seconds apply to the next job claimed. One already running keeps the budget and the backoff it started under.
  • max_attempts is frozen onto each job when it is queued, so a change applies to work queued from then on. Raising it is not a way to rescue a backlog that is about to give up — those rows keep the budget they were queued with.
  • retention_days rebuilds the sweep, including registering it when it goes from 0 to a real value.

Systemd

Add ExecReload to the unit so systemctl reload works:

[Service]
ExecReload=/bin/kill -HUP $MAINPID

See Deployment for the rest of the unit.

Two things to expect

Neither is a problem, but both look odd if you are not expecting them.

Startup warnings repeat. A reload re-emits the advisories that describe the configuration it just applied — challenge_validation_bypassed, filter_disabled, tls_disabled and the rest. That is deliberate: they are written to stay visible for as long as the condition holds.

profile_mounted does not repeat. It is a lifecycle notification meaning “this endpoint came up”, delivered to whichever [notify] backends are configured. Firing it on every reload would make the notification surface noisiest in exactly the config-managed deployments that would least want it.

Reloads a restart still handles better

Two edges, both brief and both identical to what a restart does:

  • A request that was already in flight when the reload landed finishes under the old configuration. That includes the notification it may queue.
  • For the moment it takes in-flight requests to drain, the old and new admission limits are both in force, so concurrency can briefly reach twice server.max_concurrent_requests.

If a change is important enough that neither is acceptable, restart.

Admin CLI

acme-proxy embeds an administrative command-line interface in the same binary, so a full deployment never needs a separate tool to manage its state: accounts, orders, the audit trail, the job queue, nonces, EAB credentials, the upstream account, the web admin’s operators and sessions, and the database itself (migrate, init, transfer). profile and filter read the configuration back as the server would build it.

Invoking it

serve is the default subcommand, so a bare acme-proxy starts the server. You reach the admin CLI by naming a subcommand:

acme-proxy account list

Every command reads the same configuration as the server — config.toml in the working directory, or ACME_PROXY_CONFIG, plus ACME_PROXY_* environment overrides — so it must be run where the configuration points at the same database. There is no --config flag.

Commands operate directly on the database, SQLite or PostgreSQL. Running them against a live server is safe (SQLite runs in WAL mode; PostgreSQL is built for concurrent writers), but they act immediately and are not transactional across the server’s own in-flight requests.

The schema is applied explicitly

Opening the database does not migrate it. acme-proxy migrate applies any migrations that have not run yet, and a serve running the worker role does the same at startup — so a default single-process acme-proxy serve against a fresh database still just works.

Everything else checks the schema and refuses by name:

database error: the schema is 3 migration(s) behind; run `acme-proxy migrate` first

This used to be a side effect of opening the database, which made every subcommand an upgrade step — acme-proxy audit list from a newer binary silently rewrote the schema — and let two processes starting together race the migration runner, SQLite offering sqlx no lock to serialise them.

acme-proxy migrate is idempotent and safe to run repeatedly; it prints how many migrations it applied, or says the schema is already up to date.

acme-proxy init migrates and then generates whatever first-run material the configuration calls for — the local CA key and certificate, an upstream account for a relay profile, a self-signed TLS certificate. It is the one command that creates key material, so a split deployment runs it once, as the uid that should own those files, before starting anything.

Moving between backends

acme-proxy transfer --to <url> copies every row of the configured database into another one. The scheme of each URL picks its backend, so this is how a SQLite deployment becomes a PostgreSQL one — and the reverse is the same command with the two swapped.

$ acme-proxy transfer --to postgres://acme@db.internal/acme
Copy 14203 row(s) from sqlite://acme.db to postgres://acme:***@db.internal/acme?
The source server must be stopped, or the copy is a torn snapshot.
Continue? [y/N] y
  accounts                 412
  orders                  9881
  audit_log               3910
  …
Copied 14203 row(s) into 15 table(s).

Stop the server first. Nothing can check it: a worker that issues a certificate while the copy is running writes rows the copy has already walked past, and the result looks exactly like a good one. That is the only part of this an operator has to get right unaided.

The target must already exist, be migrated and be empty. Create the database, run acme-proxy migrate against it (this command will not — applying a schema belongs to migrate, init and the worker role, and nothing else), then transfer. A target that already holds rows is refused by name, listing them: a transfer is a copy, not a merge, and there is no flag that makes it one.

Row ids, certificate serials and audit ids all survive, because all three are things something outside the database still refers to — a kid a client stored, a serial the CRL carries, an id an operator typed. A certificate issued before the move is revocable after it.

What does not travel is the schema’s own history: each backend keeps its own migration set and checksums. And --json answers {"tables": [{"table", "rows"}], "total"} — a report of a copy, not a listing, so it has its own shape rather than the paged envelope.

Roles

serve takes --role, a comma-separated list of acme, admin and worker. With no --role it runs all three in one process, which is the default and what every deployment before the flag existed did.

RoleDoes
acmeServes ACME to certificate clients: the ACME listener and the root router. Enqueues work, runs none.
adminServes the web admin, /ui and /api. Enqueues work, runs none.
workerDrains the job queue, and owns the schema and the first-run material.

Splitting them puts the process that parses untrusted JWS and CSRs, the process that holds operator sessions, and the process that reaches out to client-chosen hosts in three different places, each able to run under its own uid. See Deployment.

An unknown role is refused with usage before anything is read:

error: invalid value 'wroker' for '--role <ROLES>': unknown role `wroker`
(expected one of: acme, admin, worker)

A process running no worker logs the advisory server_role_no_worker at startup: nothing there drains the queue, so challenge validation, notifications and the periodic sweeps all wait for a process that does.

Global flags

-y, --yes — skip the interactive “Are you sure?” prompt on destructive commands. It is a global flag, so it may be given anywhere on the line. account delete, order delete, eab delete, jobs cancel, audit cleanup, nonce cleanup, transfer, admin user delete and admin user totp reset prompt; nothing else is gated by it.

--json — where supported, emit JSON instead of the human-readable line format. Single-item commands print one JSON object. Every list command prints the same envelope the admin JSON API returns, so a script does not learn one shape for the shell and another for the API:

{ "items": [ … ], "total": 137, "limit": 50, "offset": 0 }

total is what the same filters match unpaged, which is the difference between having read the table and having read a page of it. See Paging. (It is not newline-delimited JSON.)

--color <auto|always|never> — when to colour the human-readable output. Also global. The default is auto: colour when the stream is a terminal and NO_COLOR is unset or empty, which means a piped or redirected run is plain without your having to say so.

  • always colours regardless of the stream and regardless of NO_COLOR — it was typed on this command line, so it outranks both. That is what makes acme-proxy audit list --color always | less -R work.
  • never never colours, whatever the terminal is.
  • An unrecognised value is refused rather than treated as auto.

Colour is decided separately for stdout (the output) and stderr (error messages), since the two are redirected independently. It is semantic, never decorative: statuses (valid, pending, revoked…), audit events that name a refusal, filter explain’s per-check verdicts, and the standing warnings such as eab create’s “shown only this once”. Labels, timestamps and identifiers are never coloured.

--json output never carries colour, at any setting — it is the same bytes a script parses today. So is every human-readable line under --color never.

Note this is not the same switch as logging.ansi, which colours the server’s log stream and is a configuration key rather than a flag; the CLI’s colour is not configurable, on purpose, since the right answer depends on the terminal in front of you rather than on the deployment.

--log-level <off|error|warn|info|debug|trace> — emit log records for this run. Also global.

An admin command prints only its own output by default: no log records at all, on either stream. That is what makes acme-proxy account list --json | jq work — before this flag existed, a db_migration_completed record landed on stdout ahead of the JSON on every invocation, because [logging] was installed for every subcommand and its target defaults to stdout.

  • Records go to stderr, whatever logging.target says. stdout is the answer; a diagnostic does not belong in it.
  • The level covers acme-proxy alone, so --log-level debug does not also turn on sqlx and hyper. Set RUST_LOG for a directive that reaches further — a non-empty RUST_LOG turns records on by itself, with no flag.
  • --log-level off is silence stated explicitly, which is what a script wants when the environment it runs in may carry a RUST_LOG.

It is worth reaching for on filter show, which builds the policy exactly as startup does and so is the cheapest pre-restart check: the refusals are printed either way, but the advisories are log records, so filter show --log-level warn is how you see filter_disabled and filter_check_unused before a restart rather than after it.

On serve the flag outranks both RUST_LOG and logging.filter — a flag was typed where the other two are ambient — and [logging]’s other five keys still decide the format, the target and the rest. It survives a SIGHUP, so an operator who started the server at --log-level debug keeps it across a reload; an edit to logging.filter is then a no-op the server warns about (server_logging_filter_overridden).

--version — print the build’s own version and exit. It is the first thing a bug report asks for, and the answer a checkout cannot give on a host where the binary was copied in. --help is its counterpart and works at every level: acme-proxy audit --help lists that group’s subcommands.

Exit codes

An admin command exits with one of these. A script can branch on the code without parsing stderr; the deciding question between 1 and 3 is whether re-running the identical command is worth it.

CodeMeaningExamples
0Success — the command did what was asked.
1The host could not carry out the request. Worth retrying, or fixing the host and retrying.A database that will not open, a signer or CA error, an unreadable --password-file, an unreachable upstream, a broken [dns]/[proxy] section, invalid configuration.
2The command line itself was rejected. Emitted by the argument parser.An unknown flag or subcommand, a missing argument.
3The request cannot be satisfied as written. Re-running the identical command will not help.No object with that id (no such order …); an object in the wrong state (… is already revoked, only ready or failed jobs can be cancelled); an unknown --status/--event/--outcome value, or --role on admin user create; contradictory flags (--hide-superseded without --expiring-in); nothing supplied on stdin where a password or an EAB key was asked for.

serve exits 1 for any startup failure and otherwise runs until it is signalled.

Shell completions

acme-proxy completions <shell> prints a completion script on stdout, for bash, elvish, fish, powershell or zsh. It is generated from the same command tree clap parses, so it covers every subcommand and flag, four levels deep — acme-proxy admin user totp completes to status, reset and recovery-codes. Flag values complete only where the flag has a fixed set clap knows about, which today is --color, --log-level and completions’ own <shell>; --status and --outcome take a string the command refuses by name, so a shell has nothing to offer for them.

The command reads neither the configuration nor the database, so it works anywhere, including in a shell startup file and before a deployment exists.

# bash — system-wide, or ~/.local/share/bash-completion/completions/acme-proxy
acme-proxy completions bash | sudo tee /etc/bash_completion.d/acme-proxy

# zsh — any directory on $fpath; the file must be named _acme-proxy
acme-proxy completions zsh > ~/.zfunc/_acme-proxy

# fish
acme-proxy completions fish > ~/.config/fish/completions/acme-proxy.fish

Regenerate after upgrading: before 1.0.0 the CLI is not frozen, so a script kept from an older binary can go on offering a subcommand that no longer exists. The policy is stated in the changelog.

One limitation is the generator’s rather than this CLI’s: the fish script stops completing at three levels, so acme-proxy admin user totp offers nothing past totp. The other four shells complete the whole tree.

Man page

acme-proxy man prints the roff source of acme-proxy.1 on stdout. Like completions, it reads nothing and is generated from the command tree:

acme-proxy man | sudo tee /usr/share/man/man1/acme-proxy.1 > /dev/null
man acme-proxy

It is one page for the top-level command — the options, every subcommand with its one-line purpose, the environment variables, the configuration file, and a pointer back to this book. The per-flag detail of each subcommand lives in the tables below rather than in the page, which is why SEE ALSO names the book. Read it without installing anything with acme-proxy man | man -l -.

Account management

CommandFlags
account list--profile <name>, --eab-kid <kid>, --limit <n>, --offset <n>, --json
account show <id>--json
account update-contact <id>--contact <uri> (repeatable)
account deactivate <id>—
account delete <id>(prompts)
  • account list --profile restricts the listing to one ACME endpoint. Without it, accounts from every profile are listed — the admin CLI is deliberately unscoped by default, unlike the request path, which always scopes by profile. The listing is newest first and paged; see Paging below.
  • account list --eab-kid lists the accounts one EAB credential registered — what to look at before eab delete.
  • account list shows, per account, the address its key was last seen from and that address’s reverse name (ip (ptr), the address alone when no name resolved, - when neither was recorded). account show prints one field per line and adds where the account was registered from. Nothing in the server ever compares against these — pinning an identity to an address breaks CGNAT and mobile — and none of them reaches an ACME object.
  • account deactivate prevents the account from making any further requests. It is the operator-side equivalent of a client deactivating itself.
  • account delete cascades: every order, authorization and challenge belonging to the account is destroyed with it. The prompt names what will go.
  • account delete and order delete are refused while a live certificate would go with them — one issued, not revoked, and not known to have expired. An order row is the only record of its certificate: without it, revokeCert and order revoke cannot find the certificate, and it drops out of the expiry digest and renewal information. Revoke it first (order revoke), or wait for it to expire. A certificate whose expiry was never recorded counts as live. There is no flag to override this.

Order management

CommandFlags
order list--profile <name>, --account-id <id>, --status <status>, --identifier <name>, --identifier-contains <text>, --cert-serial <hex>, --expiring-in <days>, --hide-superseded, --limit <n>, --offset <n>, --json
order show <id>--json
order chain <id>—
order delete <id>(prompts)
order revoke <id>--reason <n>, --wait <seconds> (default 30)
  • --status is one of pending, ready, processing, valid and invalid, and is refused by name otherwise.

  • --identifier <name> finds the orders that name that identifier exactly (case-insensitive): the answer to “which order covers web.corp.example.com”. It is an exact match on purpose — --identifier example.com will not surface evil-example.com, which is the wrong thing to hand somebody hunting a misissuance. --identifier-contains <text> is the substring form for when only a fragment of the name is remembered; the two are mutually exclusive.

    A wildcard order stores the wildcard form, so --identifier matches *.example.com and not host.example.com — “which order named this?” is not “which certificate covers this?”, and only the first is a question an exact match can answer. --identifier-contains example.com spans both.

  • --cert-serial <hex> finds the order whose issued certificate carries that serial — the value an abuse report hands you, and the same one audit list --cert-serial filters on. Case and separators do not matter: what openssl x509 -serial prints (upper case) and what a report quotes (often colon-separated) are both folded to the form the column holds.

  • order list --expiring-in <days> asks a different question over a different query: the certificates this CA issued that reach their notAfter inside the window, soonest first, each annotated with whatever has already replaced it. It is the same listing the [notify.expiry] digest mails and the panel shows at /ui/expiring, so the three cannot come to disagree about what “expiring” or “already replaced” means. Add --hide-superseded to drop the rows that have a successor and leave only the ones to act on. Under --json the envelope carries two more members, as GET /api/expiring does: hidden, the rows --hide-superseded dropped from this page, and days, the window asked for.

  • --status, --account-id, --identifier, --identifier-contains and --cert-serial are refused with --expiring-in, by name. The expiry listing is issued, unrevoked certificates by definition, has no account predicate and is ordered by expiry, so any of them would silently mean something other than it does elsewhere — the rule --status and audit list --event already follow.

  • order show prints one field per line, omitting every field that was not recorded rather than rendering it empty — the shape audit show and account show have. It covers the certificate’s serial and the leaf’s own notAfter beside the requested notAfter the client asked for, and the revocation timestamp and reason, which are deliberately absent from the ACME JSON a client sees — revocation state is admin-visible only.

  • order show and order show --json describe the same order, field for field, with one deliberate exception: the issued chain. --json carries it as certificatePem and the web panel offers it as a download; the text rendering does not, a command run to get one’s bearings being the wrong place for several kilobytes of PEM. The three URL members (authorizations, finalize and the ACME certificate URL, which is reachable only by signed POST-as-GET) are likewise --json’s alone — the indented authorization tree is what a terminal reads instead.

  • order chain <id> prints that chain and nothing else, so it pipes:

    $ acme-proxy order chain 0198f3b1-... > web.example.com.pem
    

    It is the terminal’s spelling of the panel’s GET /ui/orders/{id}/chain.pem download, and it keeps that route’s rule: an order that never reached issuance is an error, not an empty file, because zero bytes named .pem read as a broken certificate rather than an absent one. There is no --json — the PEM is the output.

  • order revoke is the operator-side equivalent of POST /revokeCert, for an out-of-band compromise report a client cannot or will not act on. It never loads a CA key or contacts an upstream itself; the running server does the signing, as Revocation & CRL describes. It is not confirm-gated, because revocation only ever tightens trust.

    --reason is the RFC 5280 reason code: 0–6 or 8–10, since 7 is unused; any other value is refused. Omitted, the revocation carries no reason. For a relay or custom profile the revocation is queued for the running server, and --wait is how many seconds the command waits for it to land before returning (default 30); --wait 0 queues it and returns at once. A local CA’s revocation is recorded immediately and does not wait.

Job queue

The background queue (relayed issuance, notification delivery, the periodic sweeps) is what keeps an order processing through a transient upstream failure instead of failing it. When an order is stuck, this is where to look.

CommandFlags
jobs list--kind <k>, --status <s>, --limit <n>, --offset <n>, --json
jobs show <id>--json
jobs cancel <id>(prompts)
jobs run-now <id>—
  • jobs list is paged like the other listings, newest first. --status is one of ready, running, done, failed, cancelled and is refused by name — passed to SQL an unknown value answers “no rows”, which reads as “nothing is in that state”. --kind is not refused: a job kind is an open set, so a typo simply matches nothing.

  • jobs show prints one field per line, and for a relay issuance job (kind = signer_relay_issue) it appends the upstream order it drives — the upstream URLs and the upstream’s own error text. The stored CSR is never shown.

  • jobs cancel retires a job (status = cancelled) and is confirm-gated. Two things it will not do:

    • A running job is refused — a runner owns it, and its lease will expire or it will settle. Wait it out, then cancel the resulting ready/failed row.
    • Cancelling a periodic sweep job (nonce_sweep, audit_sweep, order_sweep, …) stops that sweep until the server restarts. The prompt says so.

    Cancelling an in-flight signer_relay_issue job also abandons the ACME order: the local order is marked invalid (so the client stops polling), the upstream mapping is abandoned (so a restart does not resume it), and a certificate_issue_failed audit row is written naming you. This is the operator-side way to stop a relayed issuance that will never complete.

  • jobs run-now makes a job eligible immediately — it is picked up within jobs.poll_interval_ms (it does not wake the runner). On a ready job it just pulls run_at forward; on a failed job it grants exactly one more attempt (attempts is set to max_attempts - 1), not a fresh budget, because a full reset is what turns a permanently failing job into an infinite retry loop. It is not confirm-gated.

The web admin has the same surface at /ui/jobs (GET /api/jobs), with cancel and run-now behind an operator-or-higher session.

Audit trail

CommandFlags
audit list--profile <name>, --account-id <id>, --order-id <id>, --cert-serial <hex>, --event <e>, --outcome success|failure, --since-days <n>, --limit <n>, --offset <n>, --json
audit show <id>--json
audit cleanup--older-than <days> (prompts)
  • audit list is paged like every listing; see Paging.
  • An unknown --event or --outcome is refused by name, listing the values this build knows. Passed through to SQL it would answer “no rows”, which reads exactly like “nothing happened”.
  • audit show prints one field per line, omitting every field that was not recorded rather than rendering it empty.
  • audit cleanup is the only command in this binary that destroys audit history, so it is confirm-gated and its prompt names the row count. audit.retention_days runs the same sweep daily.

The web admin can read this trail but not prune it — see Audit Trail.

Paging

Every listing in this binary is paged. account list, order list (both of its queries), audit list, jobs list, eab list, upstream order list, admin user list and admin session list all take --limit <n> and --offset <n>, defaulting to 50 rows. orders and audit_log each grow a row per issuance for the life of the deployment, so there is deliberately no “everything” spelling and --limit 0 is not a way around it: on a year-old CA that is a terminal full of scrollback and a table loaded into memory. A nonsense window is corrected rather than refused — a --limit 0 becomes one row, a negative --offset becomes zero.

eab list, admin user list and admin session list used to answer a bare JSON array with no window, on the argument that an operator mints those rows by hand a few at a time. That was true of how the tables fill and said nothing about how long they have been filling — and it made a script learn one shape for the shell and another for /api.

Every paged listing ends with a count, always and not only when the page is short:

$ acme-proxy order list --limit 2
...
2 of 1877 row(s).

“42 of 1877” is the difference between having read the table and having read a page of it. Page with --offset; the listings are ordered newest first, tie-broken on the row id, so a row cannot swap between pages and go unseen.

admin user list is the one exception, and is oldest first. The bootstrap operator — created before there was a panel to sign in to — is precisely the row whose position should not move as colleagues are added.

order list --expiring-in adds a third number when --hide-superseded drops rows, because supersession is decided per row and cannot become part of the query — so the total counts the window, not the rows printed under it:

$ acme-proxy order list --expiring-in 30 --hide-superseded --limit 20
...
6 of 8 row(s), 2 superseded hidden.

Under --json every one of them answers the same envelope the admin API returns, member for member (order list --expiring-in adds hidden and days, as its API twin does):

$ acme-proxy eab list --limit 2 --json
{"items":[…],"total":37,"limit":2,"offset":0}

The window is not clamped to admin.page_size_max. That key is a ceiling on what an HTTP caller may ask the server for; this front end already answers to a shell on the host.

Access policy

CommandFlags
filter show--profile <name>, --json
filter explain--profile <name>, --client-ip <ip>, --identifier <name> (repeatable), --path <p> (default /newOrder), --account-id <id> (default explain), --json

--profile may be omitted only when exactly one profile exists, the same rule upstream show follows: [filter] is per-profile, so acting on “the policy” without saying which one would be acting on nothing.

filter show prints the resolved policy — every check with its type and the stages it decides at, then every rule in evaluation order with its condition re-parenthesized. That last part is the point: an operator who wrote a or b and c sees a or (b and c) printed back and has their answer about precedence without reading the grammar.

Both commands build the policy rather than reading the file back, so every startup refusal reaches you here too. filter show is therefore the cheapest way to check a policy before restarting the server. Like every command but completions and man, they open the configured database first, so on a fresh host run acme-proxy migrate before them:

$ acme-proxy filter show
profile: default
default: deny (when a rule was applicable and none matched)

checks
  inventory            ipam         identifiers only
  mgmt-net             allowed_ip   connection and identifiers

rules (first match wins)
  mgmt-bypass          mgmt-net -> allow
                         evaluated at: connection and identifiers
  inventory-owned      inventory or mgmt-net -> allow
                         evaluated at: identifiers only

--json prints the same policy as a document, which is what a configuration check in CI reads — and is the shape the web panel renders, so the two front ends cannot come to describe one policy differently:

{
  "profile": "le",
  "active": true,
  "defaultEffect": "deny",
  "warning": null,
  "checks": [
    { "name": "mgmt-net", "type": "allowed_ip", "stages": "connection and identifiers" }
  ],
  "rules": [
    { "name": "inventory-owned", "when": "corp-names and (inventory or mgmt-net)",
      "then": "allow", "mode": "enforce", "stages": "identifiers only" }
  ]
}

checks is name-sorted and rules is evaluation order. An endpoint with no rules answers "active": false and the warning above, with defaultEffect null rather than the configured word: filter.default is consulted only where some rule was applicable, so with no rules it is not a fact about that endpoint at all.

filter explain evaluates it against a hypothetical request and reports all three stages — connection, newOrder and CSR — because every stage must allow, and that is the thing most easily misread. For each it prints every check’s verdict with its reason, which rule matched, and the HTTP answer that stage would produce.

$ acme-proxy filter explain --client-ip 10.0.0.5 --identifier web.corp.example.com

Checks the evaluation never reached are listed as skipped: a short-circuited operand and a passing one look identical in the outcome, so this is the only way the output can answer “why did my inventory check not run”.

This really runs the policy. filter explain executes your custom scripts and issues real IPAM and DNS requests, exactly as a request would, because a stubbed answer would be worse than nothing the first time it disagreed with production. It writes nothing, and it names the checks that reached outside the process at the end of its output (sideEffects under --json).

That is also why explain is a host-only command with no web-admin equivalent: the address and names are chosen by the caller, so behind a session it would be script execution and outbound requests driven from one stolen cookie. show is the opposite case and the panel does serve it — see Web Admin — because it reads an already-built policy and reaches nothing outside the process.

Nonce housekeeping

CommandFlags
nonce count--json
nonce cleanup--ttl-seconds <n> (prompts)

nonce cleanup deletes expired nonces. The server already sweeps them on an interval for the life of the process, so this is mainly a debugging tool. --ttl-seconds defaults to the configured nonce.ttl_seconds.

nonce count reports the table size and the window a nonce is fresh for — the terminal’s spelling of GET /api/nonces, answering the same {count, ttlSeconds} under --json:

$ acme-proxy nonce count
1284 nonce(s), ttl 300s.

The two numbers are only meaningful together. The count should sit near the request rate times the TTL; one far above that says the reaper is not running. Values are never listed, on either surface: a nonce is a bearer credential until it is consumed, so a listing would put live ones on a screen.

Profiles

CommandFlags
profile list--json

The ACME endpoints this configuration mounts, name-sorted, each with the two facts that decide whether it is safe to expose and then its directory URL — last, because it is the only field here with no bounded width, and any column after it would be ragged:

$ acme-proxy profile list
internal              challenges=bypassed   eab=off  https://ca.example.com/profile/internal/directory
le                    challenges=validated  eab=on   https://ca.example.com/profile/le/directory

challenges=bypassed is painted as a warning, because it is one: an endpoint that marks a challenge valid without checking anything is an open CA wherever [filter] is empty. A profile parked with enabled = false is absent, which is the honest answer to “what does this mount”.

Like filter show, this builds the answer the way startup does rather than reading a file back, so a configuration serve would refuse is refused here too — which makes it a pre-restart check as well as a listing. The panel’s GET /api/profiles and /ui/profiles render the identical document from the mounted profiles instead, so between an edit and its SIGHUP the two legitimately disagree, and only this one can be pointed at a configuration the server would not start on.

Upstream account management

Only relevant with signer.backend = "relay".

CommandFlags
upstream show--profile <name>, --json
upstream register--profile <name>, --eab-kid <kid>, --eab-hmac-key-file <path>
upstream order list--profile <name>, --status <s>, --limit <n>, --offset <n>, --json
upstream order show <local-order-id>--json

--profile is required for show/register whenever the configuration defines more than one profile. [signer] is a per-profile section, so acting on “the upstream” without saying which one would be acting on nothing. It may be omitted only when exactly one profile exists. upstream order list takes --profile as an ordinary filter and is cross-profile without it.

upstream register performs this proxy’s own newAccount at the upstream CA and stores the resulting account URL beside account_key_path with a .kid extension. It is the only time the account is registered: serve reads the stored kid and never needs the EAB credential again.

Security note: the EAB HMAC secret is read from --eab-hmac-key-file, or prompted on stdin. It is deliberately not accepted as a command-line argument, because argv is visible to every user on the host via ps. Omit --eab-kid entirely when the upstream requires no External Account Binding.

upstream order list|show reads the upstream_orders table — one row per local order this proxy relayed to its upstream. --status is processing, valid or invalid, refused by name. Each row carries the upstream order/finalize/ certificate URLs, the upstream CA’s own error text for a failed relay, and the finalize request’s request_id; show takes the local order id and cross-links to the relay job. It is read-only — to stop an in-flight relay, use jobs cancel on the signer_relay_issue job (see Job queue), which abandons the local order too. The panel twins are GET /api/upstream-orders and /ui/upstream-orders.

External Account Binding (EAB)

CommandFlags
eab create--label <text>, --profile <name>, --json
eab list--limit <n>, --offset <n>, --json
eab show <kid>--json
eab revoke <kid>—
eab delete <kid>--deactivate-accounts or --delete-accounts (prompts)
  • eab create prints the generated HMAC secret once. It is stored but never shown again, so a lost secret is replaced, not recovered.
  • --profile binds the credential to one endpoint. Omitted, the credential is accepted at every profile — which is what an unscoped credential means, and is usually not what you want in a multi-tenant deployment.
  • eab revoke takes effect immediately, with no restart: credentials are read from the live database on every newAccount.
  • eab delete removes the credential. By default its accounts are kept, but they no longer resolve to any credential, so every eab filter check refuses them. --deactivate-accounts also deactivates them and keeps their orders, so their certificates can still be revoked. --delete-accounts also deletes them and everything under them; like account delete, it is refused while any of their orders holds a live certificate, and then nothing changes. The prompt names how many accounts and orders are involved.
  • eab list is newest first and paged; see Paging. It reads the same query GET /api/eab and /ui/eab do, so the three cannot come to describe the credential set differently.

See External Account Binding for the protocol side.

Web admin operators and sessions

The web admin has no sign-up page: the first operator is created here. These commands work whether or not [admin] is enabled, and whether or not the server is running.

CommandFlags
admin user create <username>--password-file <path>, --role admin|operator|viewer (default admin), --contact <address>
admin user list--limit <n>, --offset <n>, --json
admin user show <username>--json
admin user passwd <username>--password-file <path>
admin user role <username> <admin|operator|viewer>revokes the operator’s sessions
admin user contact <username>--contact <address> (omit, or pass empty, to clear)
admin user delete <username>confirm-gated; -y skips
admin user disable|enable <username>—
admin user totp status <username>--json
admin user totp reset <username>confirm-gated; -y skips
admin user totp recovery-codes <username>prints them once
admin session list--user <u>, --limit <n>, --offset <n>, --json
admin session revoke--user <u> (optionally --session <id>) or --all
$ printf '%s' "$PASSWORD" | acme-proxy admin user create alice
Created admin user alice (bac6a47e-711b-4e8e-858e-417da905dab9), role admin.
  • The password never goes in argv. There is no --password flag and clap rejects one: argv is visible via ps and lands in shell history. Supply it on stdin or with --password-file (which strips one trailing newline). Typing it interactively works but echoes, and the command says so.
  • Minimum 12 characters. Stored as PBKDF2-HMAC-SHA256 at 600 000 iterations, and not recoverable — a lost password is replaced with admin user passwd.
  • admin user passwd and admin user disable both revoke every session that user holds. A password changed because it may have leaked, that left the leaked session alive, would be a change in name only.
  • --role / admin user role set the operator’s privilege tier, which scopes what their web sessions may do — viewer reads only, operator adds every CA action, admin adds managing other operators. It does not restrict the CLI: the host is the trusted plane. An unknown value is refused by name, and admin user role revokes the operator’s sessions so a demotion takes effect at once. A row that predates the feature reads as admin. See Web Admin — Roles.
  • Usernames are stored lowercased, so Alice and alice cannot become two logins that read as one in a log line.
  • There is deliberately no admin user totp enrol. Enrolling happens in the panel, which shows the setup key once behind Cache-Control: no-store; there is no way to do it from a terminal that does not put that key into scrollback and shell history — the same reasoning that keeps a password out of argv. What the shell is for is the case the panel cannot serve: totp reset is how an operator who has lost their authenticator gets back in. It asks first, because it removes a security control rather than tightening one, and it takes the recovery codes and every live session with it.
  • admin user show is the detail beside the listing, and adds the two things a row cannot carry: whether enrolment was started and never confirmed, and how many recovery codes are left. That first one matters because “enrolment pending” and “no factor” behave identically at the login prompt — an operator who believes they enrolled has no other way to find out. admin user totp status says the same thing about the factor alone. It also shows the operator’s contact address and the recent login addresses that raise a “new address” notification.
  • --contact / admin user contact set the address a web-admin operator receives security notifications at (a sign-in from an unfamiliar address, a refused second factor, a lockout, a credential change — see Web Admin — Security notifications). The address is validated as a mailbox; a bad one is refused. admin user contact with no --contact, or an empty one, clears it. Notifications are delivered only when admin.enabled and [admin.notify] are configured. A change made from this CLI — a contact address, passwd, totp reset, recovery-codes — is notified like one made in the panel: the CLI queues the message and the running server’s worker sends it.
  • admin session list shows a fingerprint of the stored token hash, never the hash itself. That fingerprint is the <id> admin session revoke --user <u> --session <id> takes to end one session rather than all of an operator’s — the granularity the panel’s Operators and Account pages already have. --session needs --user, since the fingerprint only names a row within one operator’s sessions. Both listings are paged; see Paging, which also has the reason admin user list is the one listing ordered oldest first.

See Web Admin — Users & Sessions for the full treatment.

Web Admin

A browser- and script-facing management interface for the server: the accounts it has registered, the orders it has issued, the EAB credentials it honours, and the nonce table.

It is a second listener, on its own socket, serving no ACME — and it is off by default. A certificate authority should not grow a management surface because somebody upgraded it.

It has two faces over the same operations: HTML pages at /ui for a browser, and a JSON API at /api for a script. Neither is built on the other; both are thin layers over the same crates/admin/src/admin/ operations the CLI calls.

[admin]
enabled = true

There is no sign-up page and never will be. The first operator is created from a shell on the host — see Users & Sessions.

Why a second listener

The ACME listener is public, unauthenticated, and frequently internet-facing. Everything about its defaults follows from that: admission control that sheds load, a filter chain, a small body limit.

The admin listener has the opposite shape. It defaults to loopback, requires a session on every route but sign-in, and deliberately carries no admission control (the availability concern here is credential brute force, which the login rate limiter handles and admission control would not touch). It does not inherit the profiles’ [filter] either: it has a policy of its own, [admin.filter].

Access control on this listener is the bind address, TLS, [admin.filter], and the session.

Putting both on one socket would have meant one set of defaults for two very different threat models.

Exposing it

The default binds loopback only:

[admin]
bind_address = "127.0.0.1:3001"
base_url     = "http://localhost:3001"

The recommended way to reach it from elsewhere is an SSH tunnel, which needs no configuration change at all:

$ ssh -N -L 3001:127.0.0.1:3001 ca.example.com

Then open http://localhost:3001 locally. admin.base_url stays as it is, because from the browser’s point of view the panel really is on localhost.

Binding to a real interface

Startup refuses a non-loopback bind_address while admin.tls.enabled is false:

admin.bind_address `0.0.0.0:3001` is not loopback while admin.tls.enabled is
false: the session cookie is sent `Secure`, which a browser will not store over
plain HTTP on anything but localhost, so signing in would appear to succeed and
then fail silently. Set admin.tls.enabled = true, or bind 127.0.0.1 and reach
it through an SSH tunnel

This is an error rather than a warning on purpose. The session cookie is always sent Secure — making that conditional is how a session cookie leaks — and browsers accept a Secure cookie on http://localhost but silently refuse it on http://192.0.2.10:3001. The operator would see “sign-in works, then I am immediately signed out” with nothing in any log to explain it.

So a public bind means TLS:

[admin]
enabled      = true
bind_address = "0.0.0.0:3001"
base_url     = "https://admin.example.com:3001"

[admin.tls]
enabled = true

As with [server.tls], a self-signed certificate for the host of admin.base_url is generated on first start when the files are missing, and the key is created 0600. The paths default to admin.pem/admin.key, separate from the ACME listener’s — the two answer to different names and should not share a certificate by accident.

Give the panel its own host name

Every response carries Strict-Transport-Security: max-age=31536000; includeSubDomains. Once a browser has seen that over HTTPS it will refuse plain HTTP for the whole host, and for every name under it, for a year — which is the point when the panel has a host of its own, and a trap when it does not.

The scope is the host in admin.base_url, so base_url = "https://admin.example.com:3001" commits admin.example.com and anything below it, and nothing else. But base_url = "https://example.com:3001" commits every subdomain of example.com, including services that have nothing to do with this server and may not speak HTTPS at all. HSTS is not scoped by port, so running on :3001 does not narrow it.

Give the panel a name of its own. There is no configuration key for the header: weakening it for everyone is the wrong trade when a dedicated host name costs a DNS record.

Nothing to worry about while admin.tls.enabled is false and the bind is loopback — a browser ignores the header entirely over plain HTTP (RFC 6797 §7.2), which is why it is emitted unconditionally rather than gated.

Restricting who can reach it

Where a host firewall can guard the port, use it. Where it cannot — a container runtime owns the host’s netfilter tables, so an nftables rule on the host is not yours to write — [admin.filter] puts the same policy engine the ACME listener uses in front of every admin request, /health included:

[admin]
enabled      = true
bind_address = "0.0.0.0:3001"
base_url     = "https://admin.example.com:3001"

[admin.tls]
enabled = true

[admin.filter]
rules = ["mgmt"]

[admin.filter.check.mgmt-net]
type  = "allowed_ip"
allow = ["10.20.0.0/24", "127.0.0.1/32"]

[admin.filter.rule.mgmt]
when = "mgmt-net"
then = "allow"

A caller outside 10.20.0.0/24 gets 403 before any handler runs: the admin API’s access_denied error under /api, the HTML error page elsewhere. Only checks that read the request itself make sense here — allowed_ip, path, reverse_dns, custom — and startup refuses the others by name. The check and rule syntax is the one Filters documents; the section’s own rules are in the Configuration Reference.

Two things are worth knowing in a container:

  • Container networking rewrites addresses. Published ports usually keep the client’s address, but some setups (rootless Docker’s port forwarder, a userland proxy) make every connection appear to come from the bridge gateway. Check the client_ip on the admin access line before relying on an allowlist.
  • A reverse proxy in front (Traefik, Caddy, nginx) is every peer. List its network in admin.filter.trusted_proxies, and the policy, the access line and the sign-in rate limiter all see the real client from its X-Forwarded-For. The loopback-or-TLS rule above still applies to the bind: a proxy in another container reaches the panel over the network, so keep [admin.tls] on and have the proxy speak HTTPS to it.

A health check run inside the container (Docker’s HEALTHCHECK) arrives from loopback, so allow 127.0.0.1/32 as above. An orchestrator probing from outside needs a type = "path" check for /health and a rule allowing it.

Authentication

Sign-in exchanges a username and password for an opaque session token:

  • The token is 256 random bits, and only its SHA-256 is stored. A database read — a backup, a .dump, an injection — yields nothing replayable.
  • It travels in a __Host-acme_admin_session cookie, HttpOnly; Secure; SameSite=Strict; Path=/. The __Host- prefix is browser-enforced: it requires those attributes, so an edit that drops one breaks visibly.
  • Sessions have both an absolute lifetime (session_ttl_seconds, never extended by activity) and an idle timeout (session_idle_timeout_seconds).
  • Passwords are hashed with PBKDF2-HMAC-SHA256 at 600 000 iterations. A failed sign-in costs the same whether the username exists or not, so the endpoint cannot be used to enumerate operators — and every failure returns the same invalid_credentials, whatever the real cause. The log says which.

CSRF

Every request with an unsafe method must carry an X-CSRF-Token header matching the session’s csrfToken, which sign-in and GET /api/session both return.

SameSite=Strict is set as well, but it is not sufficient on its own here, and the reason is specific: SameSite is scoped to the registrable domain, not the origin — different ports of the same host are same-site. A panel on :3001 beside anything else on :8080 of the same box is exactly what it does not cover.

There is also an origin gate: a request whose Origin does not match admin.base_url, or whose Sec-Fetch-Site says cross-site, is refused. That is what covers sign-in itself, which by definition carries no session token yet. Both headers are checked only when present, so a script or curl is unaffected.

Roles

A session also carries the operator’s role (Users → Roles). A mutating request is refused with 403 insufficient_role unless the role permits it: a CA action (revoke, delete an account or order, EAB, nonce sweep) needs operator or admin; the /operators/* and /ui/operators/* colleague-management routes (including a colleague’s role and notification address) need admin; the own-account routes (password, notification address, own sessions, own second factor, sign-out) need only a live session. An operator whose row predates the feature is admin.

Reads are open to every role with one exception: the /operators surface needs admin to read as well as to act, on both front ends. A tier that gated only the writes would be a control over what a colleague can do and not over what they can learn — and that surface carries every operator’s role and contact address, the addresses each recently signed in from, and every live session’s fingerprint. So GET /api/operators, /api/operators/{username} and /api/operators/{username}/sessions, and the two /ui/operators pages, answer 403 insufficient_role to an operator or a viewer.

SameSite=Strict also means clicking a link into the panel from another site will not carry your session. For an admin panel that is a feature, but it surprises people.

The panel

Open admin.base_url in a browser: / redirects to /ui/, and anything under it that needs a session bounces to /ui/login.

PageWhat is on it
/ui/What needs attention — failed jobs, certificates expiring within seven days, refusals in the last day, and any endpoint bypassing validation — each linking to its list; then totals for accounts, orders, EAB credentials and nonces, and the mounted endpoints
/ui/accountsEvery account, filterable by profile and EAB credential, listed with the address its key was last seen from; a detail page carries both recorded addresses, the contact editor, deactivate and delete
/ui/ordersEvery order, filterable by profile, status, account, identifier (exact or substring) and certificate serial; a detail page shows the authorizations and challenges, offers the issued chain for download, and revokes or deletes
/ui/eabCredentials, paged; minting (the secret is shown once); a detail page counts the accounts it registered, links to them, and revokes or deletes — keeping, deactivating or deleting those accounts
/ui/expiringCertificates lapsing inside a window, soonest first, each annotated with whatever has already replaced it; filterable by profile and window, with a control to hide the replaced ones. Read-only
/ui/jobsThe background queue — relayed issuance, notification delivery, the periodic sweeps — filterable by kind and status. A detail page cross-links a relay job to its upstream order and carries Cancel and Run now (see below)
/ui/upstream-ordersOne row per local order relayed to an upstream CA: the upstream URLs, the upstream’s own error text, the local order status. Read-only — stop an in-flight relay from /ui/jobs instead
/ui/auditThe CA’s audit trail — every issuance and every refusal, filterable, with a detail page per row. Read-only: there is no route here that prunes it
/ui/noncesThe table size, and a manual sweep
/ui/profilesThe endpoints this process serves, and a warning for any that bypass validation
/ui/profiles/{name}/filterOne endpoint’s resolved access policy: every check, and every rule in evaluation order with its condition re-parenthesized. Read-only

The account pages surface the CA’s account-side traceability columns: where newAccount was called from and the reverse name that address had at the time (Created from), and when and where the key last authenticated a request (Last seen, Last seen from). Each is shown only when it was recorded — a reverse lookup that found nothing leaves the address alone, and an estate where it can never succeed leaves both names blank rather than every row saying “unknown”. Nothing in the server ever compares against them; they answer “who asked for this certificate, and from where”, not “may this request proceed”.

How it is built

[htmx], vendored into the binary — no npm, no build step, no CDN. The templates are [minijinja] and can be overridden on disk without rebuilding.

Each list and detail route serves two representations of one URL: a whole document for a normal navigation, and the bare fragment htmx is going to swap when the request carries HX-Request. So /ui/orders?status=valid is a real, bookmarkable URL whether you got there by clicking a filter or by typing it.

The CSRF token reaches the browser as an hx-headers attribute on <body> and comes back as the same X-CSRF-Token header the API uses — the pages needed no second CSRF mechanism, which is the main reason htmx was chosen over plain forms.

Sign-in is the exception: a plain HTML form, no JavaScript, protected by the origin gate rather than a token (there is no session to have one yet). It works with scripting disabled.

The API

Mounted at /api, unversioned. Every response is application/json with Cache-Control: no-store.

A filter left blank is the same as leaving it out: ?profile= is every profile, not a profile whose name is the empty string. That matters because the panel’s own controls are <select> elements inside a submitted form, which always send their name — so every profile and any status reach the API as blanks.

MethodPath
POST/api/sessionsign in — {username, password}
GET/api/sessionwho am I, and my csrfToken
DELETE/api/session[?all=true]sign out (of this browser, or all)
GET/api/session/mfawhat a half-authenticated cookie still owes — {step}
POST/api/session/mfafinish the sign-in — {code}, a TOTP or a recovery code
GET/api/mfa{totpEnabled, enrolmentPending, recoveryCodesRemaining}
POST/api/mfa/totpbegin an enrolment — returns the secret, once
POST/api/mfa/totp/confirm{code} — returns the recovery codes, once
DELETE/api/mfa/totpturn it off; 409 while admin.require_mfa is on
POST/api/mfa/recovery-codesreissue — returns them once
GET/api/accounts?profile=&eabKid=&limit=&offset=eabKid: the accounts one credential registered
GET/api/accounts/{id}
GET/api/accounts/{id}/orders?limit=&offset=
PATCH/api/accounts/{id}{contact: [...]}
POST/api/accounts/{id}/deactivate
DELETE/api/accounts/{id}cascades to the account’s orders; 409 live_certificates while one holds a live certificate
GET/api/orders?profile=&accountId=&status=&identifier=&identifierContains=&certSerial=&limit=&offset=identifier is exact, identifierContains a substring, the two mutually exclusive; certSerial is the issued leaf’s serial
GET/api/orders/{id}order + authorizations + challenges, plus certificatePem once issued
POST/api/orders/{id}/revoke{reason} optional
DELETE/api/orders/{id}409 live_certificates while its certificate is live
GET/api/eab?limit=&offset=never shows a secret
POST/api/eab{label, profile} — returns the secret, once
GET/api/eab/{kid}
POST/api/eab/{kid}/revokethe row survives, moved to revoked
DELETE/api/eab/{kid}?accounts=keep|deactivate|deletethe row goes; 409 live_certificates for delete while an account holds a live certificate
GET/api/expiring?profile=&days=&superseded=&limit=&offset=read-only; superseded=hide drops the replaced rows
GET/api/audit?profile=&accountId=&orderId=&certSerial=&event=&outcome=&limit=&offset=read-only
GET/api/audit/{id}one row
GET/api/nonces{count, ttlSeconds}, the shape nonce count --json prints
POST/api/nonces/cleanup{ttlSeconds} optional
GET/api/profilesthe endpoints actually mounted; profile list’s document
GET/api/profiles/{name}/filterone endpoint’s resolved access policy; read-only
GET/healthunauthenticated, no database access

Lists return an envelope, not a bare array — total is what the same filters match unpaged, which is what a page control needs:

{ "items": [ … ], "total": 137, "limit": 50, "offset": 0 }

limit defaults to 50 and is clamped to admin.page_size_max rather than refused. Every listing on the CLI answers the same four members (Admin CLI → Paging), so a script learns one shape rather than two.

POST /api/session for an operator with a second factor answers 200 with {"mfaRequired": true, "step": "verify"} and no user member — a half-authenticated session must not read operator metadata. It is not a 401: the password was right, and a script has to be able to tell those apart. Send the code to POST /api/session/mfa, which answers the ordinary session body and a new cookie; the pending one is dead by then.

The two …/session/mfa routes are the only mutating endpoints that do not take X-CSRF-Token. The sign-in page they serve is a plain form with no token to send, exactly as POST /api/session has none, and the origin check covers both.

The issued chain

An order’s detail carries certificatePem, the chain as it was issued, and the order page renders it with a GET /ui/orders/{id}/chain.pem download (application/pem-certificate-chain). The ACME certificate member beside it is a URL, and one a browser cannot follow — RFC 8555 §7.4.2 serves it by signed POST-as-GET only, so it is there for completeness rather than as a link.

certificatePem is on the detail shape only, never on a listing: a page of fifty orders would otherwise carry fifty chains for a column no list shows. An order that never reached issuance has neither the field nor the download, and the route answers 404 rather than an empty file.

On a host holding the database, acme-proxy order chain <id> prints the same bytes on stdout under the same rule — see Admin CLI → Order management.

The card names the leaf two more ways, and both are on the listing shape as well since each is one short string: certSerial, which is what an abuse report quotes and what GET /api/audit?certSerial= filters on, and certNotAfter, the certificate’s own expiry — a different date from the Not after above it, which is the §7.4 window the client requested. Both are omitted on an order that never issued rather than sent empty.

Revocation

POST /api/orders/{id}/revoke resolves the revocation route from the order’s own profile. Two profiles can hold two different CAs, and revoking against the wrong one would record nothing useful and leave the real CRL untouched. An order belonging to a profile this process no longer mounts answers 409 profile_not_mounted rather than guessing.

The panel holds no signing backend — only the worker role does — so it revokes the way the CLI does. A local_ca revocation is recorded at once and the worker signs it into the CRL. A relay or custom revocation is queued for the worker and waited on; if it is still running when the wait ends, the API answers 202 {"status": "queued", "job": "<id>"} and the page says so. Follow it on the Jobs page.

The job queue

/api/jobs and /ui/jobs list the background queue and expose two mutations, both requiring an operator-or-higher session (a viewer is refused):

  • POST /api/jobs/{id}/cancel — retires the job (status = cancelled). A running job is refused with 409 job_not_cancellable; wait out its lease and cancel the resulting ready/failed row. Cancelling a signer_relay_issue job also marks its ACME order invalid and abandons the upstream mapping, with a certificate_issue_failed audit row naming the operator — the way to stop a relayed issuance that will never complete. Cancelling a periodic sweep job stops that sweep until the server restarts.
  • POST /api/jobs/{id}/run — makes the job eligible immediately (picked up within jobs.poll_interval_ms; the runner is not woken). On a ready job it pulls run_at forward; on a failed one it grants exactly one more attempt — attempts is set to max_attempts - 1, not reset, so repeatedly clicking it is the only way to loop a permanently failing job.

On the page, a 409 renders as a banner beside the still-present card; a 5xx replaces the page. The list rows carry no controls — the two buttons live on the detail card, and only when the job is ready or failed.

/api/upstream-orders and /ui/upstream-orders are read-only, like the audit trail below but for the queue’s reason: abandoning an in-flight relay is POST /api/jobs/{id}/cancel, not a route here. They list one row per relayed local order — the upstream URLs, the upstream CA’s own error text, the finalize request’s request_id — joined to the local order for its identifiers and status, and never carry the stored CSR. Both contribute no entry to the CSRF test table.

The audit trail is read-only here

/api/audit and /ui/audit list, filter, page and resolve a row by id, and that is all they do. There is no route on this listener that deletes audit history, and that is a deliberate limit rather than a missing feature: the first thing a stolen session would do is erase what it had just done, and a trail the watched thing can erase proves nothing. Pruning happens on the host with acme-proxy audit cleanup, or on a schedule via audit.retention_days.

This is also why /api/audit contributes no entry to the CSRF test table — with no mutating verb, there is nothing to protect. Every other verb on those paths is unroutable. See Audit Trail.

The expiry list is read-only too, for a different reason

/api/expiring and /ui/expiring answer the digest’s question on demand: what lapses soon, and has anything replaced it? They share the query, the ordering and the supersession rule with [notify.expiry] and with order list --expiring-in, so a page and a mail never disagree about what is about to expire.

Neither has a mutating route, and the reason is not the audit trail’s: renewal is the client’s action, driven by its own ACME flow against a key this server does not hold. There is simply nothing here for a button to do. Both therefore contribute no entry to the CSRF test table.

Two members of the answer need reading together. total counts the rows the window matches; hidden counts the ones this page dropped as already replaced. They are separate because supersession is computed per row rather than in SQL, so the count beside the page cannot follow the filter down — and a pager whose arithmetic quietly disagrees with the rows under it would be worse than saying so. The page says it in a line above the table.

The default window is [notify.expiry] lead_days wherever the digest is on, and 30 days where it is off: an operator who has chosen a lead time gets that one back.

The policy is shown, never explained

/api/profiles/{name}/filter and /ui/profiles/{name}/filter answer what acme-proxy filter show prints: the default effect, every check with its type and the stages it decides at, and every rule in evaluation order with its condition re-parenthesized, so an operator who wrote a or b and c reads a or (b and c) back. It is the missing half of the warning on /ui/profiles: an endpoint with challenge.bypass on has [filter] and nothing else between it and its clients, and this is where that policy can be read.

filter explain has no equivalent here and is not getting one. It really runs the policy — it executes the operator’s custom scripts and issues real IPAM and DNS requests, against an address and a list of names the caller chose. Behind a session that is script execution plus SSRF from one stolen cookie, on a listener that deliberately carries no filter chain and no admission control. show is the opposite case: it reads an already-built policy through four accessors, runs no check, and reaches nothing outside the process. Neither surface has a mutating verb — a policy is configuration, and configuration is edited in config.toml and reloaded — so both contribute no entry to the CSRF test table, for the same structural reason /api/audit does not.

One difference from the CLI is worth knowing, because it is the useful kind. The panel reads the live policy, the one this process is enforcing right now; acme-proxy filter show rebuilds one from configuration, which is what makes it the cheapest pre-restart check. [filter] reloads on SIGHUP and the whole policy is swapped, so between an edit and its reload the two legitimately disagree — and a configuration that would be refused is reported by the CLI while the panel goes on serving the last good policy. Both answers are correct; comparing them is how an operator finds out which state they are in.

An endpoint with no rules is a state, not an error: the answer carries "active": false and says the endpoint filters nothing, rather than a 404. That is the one policy an operator most needs to be told about.

Security notifications

Every web-admin sign-in and every second-factor change is logged, and — when [admin.notify] is configured — the operator it happened to is also told (ASVS V6.3.5 / V6.3.7). Two events fire:

  • admin_sign_in — a completed sign-in from an address not among the operator’s recent ones (admin_users.known_login_ips keeps the last five distinct addresses; a first-ever sign-in has no baseline and is silent), a correct password followed by a refused second factor, or a per-session second-factor lockout.
  • admin_credential_changed — the operator’s password changed, a second factor was enrolled or removed, recovery codes were regenerated, or another administrator reset this operator’s second factor (by_self = false). Also when the operator’s notification address changed (change = "contact_address") — and that one message goes to the address that was replaced, not the new one, since whoever made the change controls the new address. A first-ever address replaced nothing and tells nobody.

[admin.notify] has exactly the shape of the per-profile [notify] section — enabled, email / webhook / custom backends, template_dir, per-backend events — but is process-wide and built only while [admin] is enabled (ACME_PROXY_ADMIN__NOTIFY__ENABLED). Email delivery goes to each operator’s own address rather than to notify.email.to, stored in admin_users.contact_email. An operator sets their own from Your account (it asks for the current password), an admin sets a colleague’s from the Operators page, and the host sets anybody’s with acme-proxy admin user create --contact <address> or admin user contact <username> --contact <address>. See Users → Notification address. An operator with no address on file gets no message (the event is still logged); [admin.notify.email].to, if set, is the fallback for that case.

Changes made from the host CLI (admin user passwd, contact, totp reset, totp recovery-codes) notify too. The CLI queues the delivery and exits, and the running server’s job runner sends it, so a change made while no server runs is delivered when one starts. Nothing is queued while admin.enabled is off, since there is then no [admin.notify] to deliver through.

Errors

Not ACME problem documents. Every Problem type in this server is a hardcoded urn:ietf:params:acme:error:* URN, and nothing on this listener is an ACME error:

{ "error": "not_found", "message": "no such account: acct-1" }

error is a stable snake_case code you may branch on; message is for a human and may change. The codes:

  • Request: bad_request, invalid_status, invalid_contact, conflicting_identifier_filter, not_found, method_not_allowed.
  • Session and access: session_invalid, session_expired, session_idle, invalid_credentials, csrf_failed, rate_limited, access_denied, insufficient_role, mfa_required, mfa_not_enabled.
  • The row’s state: order_not_issued, already_revoked, live_certificates, last_admin, job_not_cancellable, job_not_runnable, profile_not_mounted.
  • Server: signer_failed, internal.

The pages answer the same failures as HTML carrying the same code, with one deliberate split: a refusal that is about the row’s state — 409 already_revoked, order_not_issued — comes back as a banner beside the button you pressed, with the record still on screen, while a server problem replaces the page. A missing session is neither: a browser gets 303 to /ui/login, and an htmx request gets the HX-Redirect header, because a 303 is followed by fetch before htmx ever sees it and the sign-in page would be swapped into whatever you clicked.

Driving it with curl

$ curl -sc jar -X POST http://127.0.0.1:3001/api/session \
       -H 'content-type: application/json' \
       -d '{"username":"alice","password":"…"}'
{"csrfToken":"…","expiresAt":"…","user":{…}}

$ curl -sb jar 'http://127.0.0.1:3001/api/orders?limit=5'

$ curl -sb jar -X POST http://127.0.0.1:3001/api/eab \
       -H "x-csrf-token: $CSRF" -H 'content-type: application/json' \
       -d '{"label":"team-a"}'

Security notes

  • This is a second attack surface on a certificate authority. It is off by default, binds loopback, and needs a session — but it has no admission control, and filters nothing until [admin.filter] names a rule. That is worth knowing rather than discovering.
  • Sign-in is protected by a fixed-window rate limiter (admin.login_max_attempts per admin.login_window_seconds, keyed on the client address — an IPv6 client by its /64, which one subscriber can rotate through at will). An attempt counts from the moment it starts, so a parallel burst gets no more guesses than a sequence. Over the limit, the password hash is not computed at all — 600 000 iterations is a denial-of-service lever otherwise — and the hash that does run is on the blocking pool, off the workers that serve requests.
  • A forwarded-for header is believed only from admin.filter.trusted_proxies, never from the ACME listener’s filter.trusted_proxies; honouring it from anyone else would let a caller spoof the key the rate limiter counts on. Behind a reverse proxy that is not listed there, the limiter counts the proxy.
  • Every response carries Content-Security-Policy: default-src 'none'; script-src 'self'; style-src 'self'; img-src 'self' data:; connect-src 'self'; form-action 'self'; frame-ancestors 'none'; base-uri 'none'. No unsafe-inline, no unsafe-eval — affordable because htmx is served from this origin and drives everything through hx-* attributes rather than inline handlers. Alongside it: no-store, nosniff, X-Frame-Options: DENY, Referrer-Policy: same-origin and HSTS.
  • htmx.min.js is a vendored third-party file, and cargo deny audits the crate graph and cannot see it. Its version, source URL, SHA-256 and licence are recorded in crates/admin/src/webadmin/static/README.md, which is the only provenance record there is — check it when you update.
  • A second factor (TOTP) is available per operator, and admin.require_mfa makes it compulsory — see Operators and sessions. It is off by default, so the loopback bind plus an SSH tunnel remains the baseline posture and not a substitute for one.
  • WebAuthn is not implemented. webauthn-rs 0.5 hard-depends on openssl/openssl-sys and is MPL-2.0, neither of which this tree carries; the design does not preclude it later (another factor is another branch in the same state machine, not a change to it).

Configuration

See the Configuration Reference for every [admin], [admin.filter] and [admin.tls] key, and Customizing the Panel for admin.template_dir.

[htmx]: https://htmx.org [minijinja]: https://docs.rs/minijinja

Web Admin — Users & Sessions

The web admin has no sign-up page. Operators are created from a shell on the host, with acme-proxy admin.

Bootstrapping the first operator

$ printf '%s' "$PASSWORD" | acme-proxy admin user create alice
Created admin user alice (bac6a47e-711b-4e8e-858e-417da905dab9).

That is enough to sign in — the panel itself does not need to be running, and does not need to be enabled yet.

If the panel is enabled with no operators, startup says so:

WARN event="admin_no_users" the web admin is enabled but has no operators:
     create one with `acme-proxy admin user create <username>`

The password never goes in argv

There is deliberately no --password flag, and clap refuses one:

$ acme-proxy admin user create alice --password hunter2
error: unexpected argument '--password' found

argv is visible to every process on the host via ps, and shells routinely write it to history. The same reasoning already applies to the upstream EAB secret (acme-proxy upstream register).

Two ways in, both of which keep it out of the process table:

$ printf '%s' "$PASSWORD" | acme-proxy admin user create alice   # stdin
$ acme-proxy admin user create alice --password-file /run/secrets/pw

--password-file strips a single trailing newline, so it does not matter whether the file was written with printf '%s' or printf '%s\n'.

Typing it interactively works too, but the password will echo — there is no rpassword here, because echo suppression needs a real TTY and that would break the injectable-reader design the whole admin layer is testable through. The command warns when it notices a terminal.

Password policy

Three rules, checked in this order, each of which ends the check — one refusal names one reason:

  1. Length. Minimum 12 characters, maximum 1024 bytes.
  2. It must not name this deployment — see the word list below.
  3. It must not be a commonly used password — see the corpus below.

There are deliberately no composition rules. “One digit, one symbol” measurably pushes people towards weaker, more guessable passwords; the two checks above refuse the passwords composition rules were reaching for, without dictating shape.

Length is counted in characters, so a 12-character passphrase in a non-Latin script is 12, not its byte count. The maximum is in bytes, and is a denial-of-service control rather than a security one: without it a sign-in could hand 600 000 iterations a multi-megabyte input.

All three run on admin user create and admin user passwd, and none of them runs on sign-in. A password that predates a rule change must still work, and a corpus refresh must never lock an operator out of the panel they would need to be signed in to fix.

The context-specific word list

A password may not contain any word that names this deployment, compared case-insensitively. Words shorter than four characters are ignored: a subject holding CA, or a host label io, would refuse a large share of every password anyone typed and buy nothing.

The list is derived from your own configuration, not fixed:

SourceWhat is taken
Alwaysacme, proxy
The operator’s usernameSplit on anything that is not a letter or a digit
server.base_url, admin.base_urlThe host only, split on . and -
[signer.local_ca.subject] — global and each profile’s owncommon_name, organization, organizational_unit, state, locality, each split into words
Each [profiles.<name>] nameThe name, which appears in every kid and order URL that endpoint issues

So a CA at https://ca.example.com with a CommonName of “Example Corp Issuing CA”, managed by operator, refuses any password containing acme, proxy, example, corp, issuing or operator — acmeproxy2026! among them.

Two limits are deliberate:

  • A CA already on disk is not described here. [signer.local_ca.subject] is read only when this server generates a CA, so an adopted ca.pem carries a subject the configuration never saw. Add those words to the subject section if you want them barred.
  • country is not read, being two characters and so below the floor whatever it holds.

If no profile resolves — which is every admin user create run against a configuration that has none yet — the rest of the list still applies. A missing section is never a reason to refuse a password change.

The common-password corpus

A password may not be (whole-string, case-insensitively) one of 13 918 known common passwords compiled into the binary. It is not a substring test: a-long-enough-password contains password and is accepted.

The corpus is the top 700 000 entries of SecLists’ xato-net-10-million-passwords-1000000.txt, filtered to entries of at least 12 characters. That filter is the whole reason a compiled-in list is affordable: password, qwerty and 123456 are refused by the length rule before this check is reached, so carrying them would cost every deployment bytes — including the ones running with admin.enabled = false — in exchange for nothing. Filtering turns 8.5 MB into 195 KB.

Provenance, the exact rank cut, the size budget it was derived from and the refresh command live in crates/admin/src/admin/corpus/README.md.

Passwords are stored as PBKDF2-HMAC-SHA256 at 600 000 iterations (OWASP’s current recommendation for the non-Argon2 case), in a self-describing format:

pbkdf2-sha256$600000$<salt-b64url>$<hash-b64url>

Self-describing so the cost can be raised, or the algorithm swapped, without a migration: a row written under older parameters is re-hashed in place on its owner’s next successful sign-in. A lost password is replaced, never recovered — nothing can read one back.

Second factor (TOTP)

Off by default and per-operator. Once enrolled, signing in is two requests: the password mints a half-authenticated session that reaches nothing but the page finishing the login, and a code from an authenticator app turns it into a real one.

stateDiagram-v2
    [*] --> anonymous
    anonymous --> anonymous: wrong password<br/>(counts against login_max_attempts)
    anonymous --> locked_out: limiter tripped, by peer address
    locked_out --> anonymous: login_window_seconds elapses

    anonymous --> active: password ok, no factor enrolled<br/>and require_mfa is off
    anonymous --> pending_mfa: password ok, factor enrolled
    anonymous --> enrolling: password ok, no factor<br/>and require_mfa is on

    pending_mfa --> pending_mfa: wrong code<br/>(counts against mfa_attempts)
    pending_mfa --> [*]: mfa_attempts exceeded —<br/>the pending row is DELETED
    pending_mfa --> active: correct TOTP code
    pending_mfa --> active: unused recovery code
    enrolling --> active: enrolment confirmed

    active --> [*]: sign out, expiry, idle timeout,<br/>or a revoked session

Three things the diagram is making explicit:

  • Promotion mints a new session. pending_mfa → active is an insert plus a delete, not an UPDATE — a new cookie and a new CSRF token — because the pending token crossed the wire before authentication had finished.
  • Guessing is bounded twice, and the two bounds are not redundant. The limiter is keyed on the peer address; mfa_attempts is keyed on the session, because a pending_mfa cookie is deliberately valid from any address and one IPv6 /64 supplies 2⁶⁴ fresh addresses.
  • enrolling is only reachable with no factor. A session that owes a code can never reach an enrolment route, or the factor would be bypassable by enrolling a new one over it.

Enrolling

From the panel, not from a shell. Sign in, click your username in the top-right corner, and press Set one up:

  1. The page shows a base32 setup key and an otpauth:// URI. Type the key into an authenticator app, or open the URI on the device holding it.
  2. Type the code the app shows and press Confirm. Nothing is enabled until you do — an enrolment begun and abandoned leaves you exactly where you were.
  3. Ten recovery codes appear. Store them now; they are shown once.

There is no QR code, and the panel’s Content-Security-Policy is not why — it already permits a data: image. A QR renderer would be a dependency for a convenience rather than a capability, and every authenticator has “enter a setup key manually”.

There is deliberately no acme-proxy admin user totp enrol. There is no way to enrol from a terminal that does not put the setup key into scrollback and the shell’s own history — the same reasoning that keeps a password out of argv.

The codes are HMAC-SHA-1, six digits, thirty seconds, which is RFC 6238’s default. Not a lapse: Google Authenticator ignores the algorithm= parameter of an otpauth:// URI and always computes SHA-1, so anything else would produce an entry that yields wrong codes forever with no diagnosis from either side. SHA-1’s collision attacks do not weaken HMAC-SHA1.

Recovery codes

Ten, each usable once, hashed the same one-way as a password — so a lost set is replaced, never recovered:

$ acme-proxy admin user totp recovery-codes alice
New recovery codes for alice — the previous set no longer works.
Store these now; they are not recoverable.

  K7QF2-3BXTM
  …

Type one into the code box at sign-in exactly as you would a six-digit code: the server tells them apart by shape, so there is no mode to choose. Case and the - do not matter.

When somebody is locked out

A lost phone with no recovery codes left is a shell command on the host — the same place the first operator was created:

$ acme-proxy admin user totp status alice
alice                 totp=enabled  recovery-codes=3

$ acme-proxy admin user totp reset alice
Remove the second factor and every recovery code for alice, and revoke their
sessions? [y/N] y
Removed the second factor for alice. Their sessions were revoked; they can sign
in with a password alone until they enrol again.

reset asks first, unlike most admin commands, because it removes a security control rather than tightening one. It takes the recovery codes and every live session with it.

status distinguishes three states, and the middle one matters: pending means an enrolment was started and never confirmed, which behaves exactly like off at the sign-in prompt. An operator who believes they enrolled has no other way to find out they did not.

Requiring it of everybody

[admin]
require_mfa = true

This governs the operator who has no factor: their next sign-in lands on the enrolment page and their session stays half-authenticated until they finish. An operator who already has one is challenged whether this is set or not.

It deliberately does not refuse a password-only sign-in. Enrolling needs a session and a session would then need a factor, so refusing would brick the panel — including the way in to fix it.

Two consequences worth knowing:

  • It does not retroactively end sessions that predate it. The lever that does is acme-proxy admin session revoke --all.
  • Turning the factor off is refused while it is set, since the operator would simply be made to enrol again on their next sign-in.

While it is on and somebody still has no factor, every start says so:

WARN event="admin_mfa_enrolment_pending" count=2
     admin.require_mfa is on and some operators have no second factor

What it costs an attacker

A half-authenticated session lives five minutes and no longer — it is a password that has been accepted and nothing more. Code attempts share the login rate limiter (admin.login_max_attempts per admin.login_window_seconds, per address), so five wrong codes also lock the password out from that address: one address, one budget. A code accepted once cannot be replayed inside its own thirty-second window.

The password asked for again when an operator replaces or removes a live factor shares that same budget, and is checked before the hash rather than after it. So a stolen session cookie cannot be used to grind the account password — which matters, because a correct guess there would let the thief enrol their own authenticator, end every other session and void the recovery codes. Past the budget those routes answer 429 with Retry-After, and the account card shows it as a banner. What the budget does not bound is an attacker with many source addresses: the cookie is deliberately valid from anywhere, so each address buys its own admin.login_max_attempts. Rotate the password and run acme-proxy admin session revoke --all if you believe a cookie has been taken.

Changing your own password

Also from the panel, not only from the host. Sign in, open your username in the top-right corner, and use the Password card: current password, then the new one.

The current password is asked for unconditionally — unlike the second-factor controls above, this runs whether or not you have TOTP enrolled. ASVS 5.0 V6.2.3 is why: a live cookie is not proof you still know the password, only that you did at sign-in. Guessing it is bounded the same way and shares the same budget as What it costs an attacker — five wrong attempts lock the address out. The new one still has to satisfy Password policy.

$ acme-proxy admin user passwd alice --password-file /run/secrets/new
Password changed for alice. Every session they held was revoked.

That command, run from the host, still ends every session — there was no request to preserve. The panel’s own change is different in exactly one way: the session that submitted it stays signed in, and every other session of that operator is revoked. Rotating a credential from inside a session you are already trusted on need not sign you out of the tab that did it.

Your notification address

The address your security notifications are delivered to — a sign-in from a new address, a change to your password or second factor. Set it on Your account under Notification address, or from the API:

$ curl -X POST https://admin.example.com/api/account/contact \
    -H 'Cookie: __Host-acme_admin_session=…' -H 'X-CSRF-Token: …' \
    -d '{"current_password": "…", "contact": "alice@example.com"}'

An empty or absent contact clears it. It asks for your current password although an address is not a credential, and revokes no session: it is where the alarms go, so a stolen cookie that could change it silently would switch off the one signal that the cookie was stolen. For the same reason, a change is reported to the address it replaced — the new one is controlled by whoever made the change. Setting the address you already have does nothing and sends nothing. An address that does not parse as a mailbox is refused (400 invalid_contact).

Roles

Every operator has one of three roles, stored in admin_users.role and governing what that operator’s web sessions may do. The CLI is unaffected — a shell on the host runs every subcommand whatever the row says, the same way it already ignores disabled for read commands.

RoleCan do
viewerRead every page and API route except the Operators surface, which is admin-only to read as well as to act. Act on their own account only: change their password, revoke their own sessions, manage their own second factor, sign out.
operatorEverything a viewer can, plus every CA action: revoke a certificate, deactivate or delete an ACME account, delete an order, mint or revoke an EAB credential, run a nonce sweep.
adminEverything an operator can, plus the Operators surface — reading it as well as acting on it: disable, enable, reset a colleague’s second factor, revoke one of their sessions.

A web session that tries something above its role gets 403 insufficient_role (an error document in the panel, a JSON error from the API); the request never reaches the operation.

The panel does not offer what a role cannot reach: a lower tier is shown no Operators nav entry, and the per-row Delete/Revoke/Deactivate controls are hidden from a viewer. That is presentation, not authorization — the write extractors decide either way — but a control that always answers 403 is a worse page than no control.

Demoting the last admin is refused, since the Operators surface is admin-only on both front ends and a deployment with none could not manage operators from the panel at all. Promote somebody else first. Nothing is unrecoverable: this host can always set a role back.

A username must match ^[a-z0-9._-]+$. It is a URL segment on the panel, so a name holding /, ?, # or a space would produce links and routes that never match. Existing operators are unaffected — the rule is checked when one is created, never when one signs in.

An operator whose row predates this feature has role admin. The column is NULL for them, which reads as admin, so an upgrade changes nobody’s authority and the first operator — the only way into the panel — is never locked out of it.

Set the role when you create the operator, or change it later:

$ acme-proxy admin user create noc --role viewer --password-file /run/secrets/pw
Created admin user noc (…), role viewer.

$ acme-proxy admin user role noc operator
Role of noc set to operator. Every session they held was revoked.

--role defaults to admin and an unknown value is refused by name. admin user role revokes the operator’s sessions, the same as passwd and disable — a demotion that left a live admin session alive would take effect only when that cookie expired.

An admin can also change a colleague’s role from their page on the Operators surface (From the panel), with the same revocation. Not their own: that surface refuses to target the caller, which is also why the last-admin refusal cannot be reached from the panel — an admin changing somebody else always leaves at least themselves.

Managing operators

$ acme-proxy admin user list
alice                 active    admin     totp=on   2026-08-08T13:21:18Z  2026-08-08T15:07:17Z
1 of 1 row(s).

$ acme-proxy admin user list --json

$ acme-proxy admin user show alice
id             bac6a47e-711b-4e8e-858e-417da905dab9
username       alice
status         active
role           admin
totp           enabled
recovery_codes 7
created        2026-08-08T13:21:18Z
updated        2026-08-08T15:07:17Z
last_login     2026-08-08T15:07:17Z

$ acme-proxy admin user passwd alice --password-file /run/secrets/new
Password changed for alice. Every session they held was revoked.

$ acme-proxy admin user disable alice
Disabled alice. Their sessions were revoked.

$ acme-proxy admin user enable alice

$ acme-proxy admin user delete alice          # asks first; -y skips

admin user list is paged like every other listing (Admin CLI → Paging) and is the one that is oldest first: the bootstrap operator is the row whose position should not move as colleagues are added. admin user show adds the two things a row cannot carry — whether enrolment was started and never confirmed, and how many recovery codes are left. The first matters because “pending” and “no factor” behave identically at the login prompt.

Usernames are stored lowercased, so Alice and alice cannot become two logins that read as one in a log line.

A password change revokes every session that user held. A password changed because it may have leaked, that left the leaked session alive, would be a change in name only. Disabling does the same.

From the panel

Also from the Operators page, not only from the host — for the operations that do not mint a credential. The page and its actions need the admin role (Roles); an operator or viewer session is refused it. Sign in, open Operators, and pick a colleague:

bob                   active    off    2026-08-08T13:21:18Z  2026-08-08T15:07:17Z

Their page shows the same status, role and second-factor summary admin user show does, their notification address and the addresses they recently signed in from, plus their own live sessions (see Sessions below). It offers three buttons — Disable, Reset second factor, and, per session, Revoke — and two forms: Change role (which revokes every session they hold) and Save address (reported to the address it replaces; empty clears it). Every one of them asks for your own password again first, whether or not you have a second factor. Disabling a colleague’s account or ending one of their sessions is a much larger blast radius than anything on your own account page, and a stolen cookie alone should not be sufficient authority for it — which is true of an operator who has enrolled no factor exactly as it is of one who has. (The second-factor routes on your own account page make the opposite trade, and deliberately: a first enrolment protects nothing, and a password prompt there would stand in front of the admin.require_mfa bootstrap.)

$ curl -X POST https://admin.example.com/api/operators/bob/disable \
    -H 'Cookie: __Host-acme_admin_session=…' -H 'X-CSRF-Token: …' \
    -d '{"password": "…"}'

$ curl -X POST https://admin.example.com/api/operators/bob/role \
    -H 'Cookie: __Host-acme_admin_session=…' -H 'X-CSRF-Token: …' \
    -d '{"password": "…", "role": "viewer"}'

$ curl -X POST https://admin.example.com/api/operators/bob/contact \
    -H 'Cookie: __Host-acme_admin_session=…' -H 'X-CSRF-Token: …' \
    -d '{"password": "…", "contact": "bob@example.com"}'

An unknown role is 400 naming the three that exist; an address that is not a mailbox is 400 invalid_contact. In the panel both are a banner on the card.

Two things this surface deliberately does not do, both already settled above: create and passwd stay on the host. Minting a credential is where “no sign-up page” already draws the line: the Operators page can disable an account, reset its factor, end a session, move it between roles or change where its notifications go, but never set a password or bring an account into being. And an operator can never target themself here — GET /ui/operators/{your own username} redirects straight to Your account, which already owns every one of those actions for yourself, so there is exactly one page an operator manages their own account from.

Sessions

$ acme-proxy admin session list
01234567  bac6a47e-…  active  2026-08-08T15:07:17Z  expires=2026-08-09T03:07:17Z  192.0.2.1
1 of 1 row(s).

$ acme-proxy admin session list --user alice --json

$ acme-proxy admin session revoke --user alice
Revoked 2 session(s) for alice.

$ acme-proxy admin session revoke --user alice --session 01234567
Revoked session 01234567 for alice.

$ acme-proxy admin session revoke --all

The id shown is a fingerprint of the stored token hash, not the hash itself — printing the hash would put every live session’s lookup key on a terminal. It is what revoke --session <id> takes to end one session; --session needs --user, since the fingerprint only names a row within one operator’s sessions.

Two deadlines apply, and whichever comes first wins:

KeyDefault
admin.session_ttl_seconds43200 (12 h)absolute; never extended
admin.session_idle_timeout_seconds3600 (1 h)advanced on use

The idle deadline is advanced at most once a minute, so a page polling every few seconds is not a stream of database writes.

An admin_session_sweep job removes expired and idle rows for the life of the process, starting with one pass at startup. Unlike nonces, sessions outlive a restart, so a startup-only sweep would leak every session an operator never explicitly signed out of.

created_ip and user_agent are recorded for forensics and are never compared against the live request: pinning a session to an address breaks every mobile and CGNAT operator, and pinning it to a User-Agent breaks on the next browser update.

From the panel

admin session revoke takes a whole operator (--user), one of their sessions (--user with --session), or the whole server (--all). The panel has the single-session form too, at two different trust levels:

Your own sessions. Sign in, open your username in the top-right corner, and scroll to Sessions: every browser currently signed in as you, the one answering this request labelled, and a Revoke beside each of the others. No password re-entry — this is the same trust level as Sign out everywhere, which sits right below it and is now a button rather than only an API route nothing linked to.

$ curl https://admin.example.com/api/account/sessions \
    -H 'Cookie: __Host-acme_admin_session=…'

Revoking the session making the request behaves exactly like signing out of just this browser: the cookie is cleared and you land back on the sign-in page. Revoking another one of your own ends it immediately, wherever it is signed in.

Another operator’s sessions. Reached from their page under Operators — the “colleague’s laptop went missing” answer that used to require SSH. Listing and revoking there behaves the same way, with one difference: it asks for your own password first, the same gate described above. A session id is a fingerprint of the stored token hash either way — printing the hash would put every live session’s lookup key on a terminal — and it only ever resolves within the one operator it was listed under, so an id copied from one operator’s page can never revoke another’s session by accident or by guessing.

What the log says

A failed sign-in returns one invalid_credentials whatever went wrong, so the endpoint cannot be used to enumerate operators. The log keeps the distinction:

WARN event="admin_login_failed" username="alice" client_ip=… reason="wrong_password"
WARN event="admin_login_failed" username="ghost" client_ip=… reason="unknown_user"
WARN event="admin_login_failed" username="bob"   client_ip=… reason="account_disabled"
WARN event="admin_login_failed" username="alice" client_ip=… reason="rate_limited"
INFO event="admin_login_succeeded" username="alice" client_ip=…

The second step keeps the same shape — one refusal to the client, the reason in the log — and note that admin_login_succeeded is emitted at promotion, not when the password is accepted:

INFO event="admin_login_mfa_pending" username="alice" client_ip=… step="verify"
WARN event="admin_mfa_failed" username="alice" client_ip=… reason="wrong_code"
WARN event="admin_mfa_failed" username="alice" client_ip=… reason="replayed"
INFO event="admin_mfa_verified" username="alice" method="totp"
WARN event="admin_mfa_recovery_code_used" username="alice" remaining=6
INFO event="admin_mfa_enabled"  username="alice" recovery_codes=10
INFO event="admin_mfa_disabled" username="alice"

reason="replayed" is the one to look at twice: it means a correct code arrived a second time inside its own window, which is what somebody replaying an observed code looks like.

One more worth knowing, because nobody will guess it from a 401:

WARN event="admin_password_hash_unreadable" username="alice"
     stored password hash could not be decoded; run
     `acme-proxy admin user passwd` to rewrite it

A corrupt admin_users row refuses the sign-in rather than erroring the endpoint, and the account stays unusable until the password is rewritten.

Revoking a single session, and every mutation on the Operators page, each leave their own line — surface says which front end it came through and target_username is absent on the self-service one, since there is nothing to distinguish it from:

INFO event="admin_session_revoked"         surface="ui"  username="alice"
INFO event="admin_operator_disabled"       surface="api" username="alice" target_username="bob"
INFO event="admin_operator_enabled"        surface="ui"  username="alice" target_username="bob"
INFO event="admin_operator_totp_reset"     surface="api" username="alice" target_username="bob"
INFO event="admin_operator_session_revoked" surface="ui" username="alice" target_username="bob"

Customizing the Panel

The pages at /ui are minijinja templates compiled into the binary. Any one of them can be replaced on disk without rebuilding, the same way notification templates work — an operator who has already overridden a notification should not have to learn a second scheme.

[admin]
template_dir = "/etc/acme-proxy/admin-templates"

Each name is looked for in that directory first and falls back to the compiled-in default. The override is per file, not per directory: a directory holding only layout.html restyles the chrome of every page and leaves the other fifty-five exactly as shipped.

Every template is compiled at startup. A broken override refuses to start, naming the file and the parse error, rather than serving a 500 the first time somebody opens that page.

The files

Paths are relative to template_dir, and are also how the templates refer to each other in {% extends %} and {% include %}.

FileWhat it is
layout.htmlThe chrome every full page extends: <head>, navigation, <body>
login.htmlSign-in. Standalone — extends nothing, and uses no JavaScript
mfa/challenge.htmlThe second sign-in step. Standalone and JavaScript-free for the same reason; branches on step between proving a code and setting one up
mfa/_setup.htmlThe setup key and the otpauth:// URI. Included by both the sign-in flow and the account page, so it renders no <form> of its own
mfa/_codes.htmlA fresh recovery set, the one time it exists in the clear
mfa/enrolled.htmlWhere a forced enrolment lands: the codes, then a link into the panel
account/index.html, account/_mfa.htmlThe operator’s own page, and the fragment every mutation on it swaps
account/_card.htmlThe second-factor card itself, with no id — so _codes.html can wrap it without nesting two elements carrying one
account/_enrol.html, account/_codes.htmlThe enrolment step, and the codes plus the refreshed card
account/_password.html, account/_password_card.htmlThe password-change swap target, and the form inside it
account/_contact.htmlWhere this operator’s own security notifications go
account/_sessions.htmlThis operator’s own live sessions, the swap target of revoking one
index.htmlThe overview: four counts and the endpoint list
partials/_flash.htmlThe inline banner every mutation’s answer renders
partials/_pager.htmlThe previous/next controls under a list
partials/_filter_meta.htmlThe tail of every list’s filter form: the loading indicator and the way back to the unfiltered list
partials/_sessions_table.htmlA table of live sessions, shared by the account page and an operator’s card
accounts/list.html, accounts/_table.htmlThe account list, and the table htmx swaps
accounts/detail.html, accounts/_card.htmlOne account, and the card every account mutation returns
orders/list.html, orders/_table.htmlThe order list
orders/detail.html, orders/_card.htmlOne order with its authorizations and challenges
expiring/list.html, expiring/_table.htmlThe expiry list, its window and profile filters, and the hidden-count line
eab/list.html, eab/_table.htmlThe credential list and the create form
eab/detail.html, eab/_card.htmlOne credential
eab/_created.htmlThe one-time HMAC secret
nonces/index.html, nonces/_panel.htmlThe nonce count and the sweep control
profiles/list.html, profiles/_table.htmlThe mounted endpoints
profiles/filter.htmlOne endpoint’s resolved access policy. No fragment: nothing on it swaps
audit/list.html, audit/_table.htmlThe audit trail
audit/detail.html, audit/_card.htmlOne audit row
jobs/list.html, jobs/_table.htmlThe background job queue
jobs/detail.html, jobs/_card.htmlOne job, and the card its cancel and run-now actions return
upstream_orders/list.html, upstream_orders/_table.htmlThe relay backend’s upstream orders. Read-only
upstream_orders/detail.html, upstream_orders/_card.htmlOne upstream order, cross-linked to its job
operators/list.html, operators/_table.htmlThe web admin’s operators
operators/detail.html, operators/_card.htmlOne operator and their live sessions, re-rendered by every mutation on the page

A file whose name starts with _ is a fragment: htmx swaps it on its own, so it must not contain <html> or <body>, and it must keep the id on its root element — that id is what the page’s hx-target points at.

Two things not to break

The extension is a security control

Every page template is named .html, and that is deliberate. minijinja decides auto-escaping from the template name, and the notify templates are named .j2 precisely so that escaping is off for them (an email body is not markup). Renaming a page template to .j2 — or adding a new one under a name minijinja does not recognise as HTML — turns an account contact or an EAB label into stored XSS.

The CSRF token has to stay on <body>

layout.html carries:

<body hx-headers='{"X-CSRF-Token": "{{ csrf_token }}"}'>

That attribute is the only route by which the token reaches a mutating request. A layout that drops it loses every write at once — which is the intended failure mode; a partial loss would be far harder to notice.

The same file sets three htmx options the other templates depend on:

<meta name="htmx-config"
      content='{"includeIndicatorStyles":false,"defaultSwapStyle":"outerHTML","responseHandling":[...]}'>

includeIndicatorStyles: false stops htmx injecting an inline <style> element that the Content-Security-Policy’s style-src 'self' would block — the rules it would have injected live in admin.css instead. responseHandling makes htmx swap non-2xx responses, without which a 409 conflict would fail silently instead of showing the operator a banner.

defaultSwapStyle: "outerHTML" means a response replaces the element it targets. Every fragment is written for that: its root element carries the same id as the swap target (accounts/_table.html is <div id="accounts-table">, the target of the accounts filter form). An override of a fragment must keep that root and its id, or the next swap on the page finds no target.

Context

Every full page gets csrf_token, user, can_write, nav (the active navigation item) and title, plus its own data; the fragment a mutation answers with gets the first three. Gate a control on {% if can_write %} — true for an operator or admin session — which is also false when absent, so a control never appears by accident. Timestamps are RFC 3339 strings, and the ago filter renders one as 3 h ago or in 5 d, or as nothing when it is not a timestamp:

{{ order.createdAt }} ({{ order.createdAt | ago }})
``` That data is the **same JSON the API returns** —
`render_account_json`, `render_order_detail_json`, `render_eab_json` — so `GET
/api/accounts/{id}` is an accurate description of what `account` holds in
`accounts/_card.html`. Lists additionally get `page` (`{items, total}`),
`pager`, `filters` and `profiles`.

A quick way to see a context in full is to render it:

```jinja
<pre>{{ account | tojson(indent=2) }}</pre>

Starting from the shipped version

The defaults are in the source tree under crates/admin/src/webadmin/templates/. Copy the one you want to change:

$ mkdir -p /etc/acme-proxy/admin-templates
$ cp crates/admin/src/webadmin/templates/layout.html /etc/acme-proxy/admin-templates/

Then send SIGHUP — see Reloading the Configuration. Templates are compiled up front, so a mistake fails the reload and the panel goes on serving the last set that worked; it never reaches a browser. The same compile happens at startup, where a mistake stops the process instead.

Stylesheet and scripts

admin.css and htmx.min.js are served from /ui/static/ and are not covered by template_dir — they are embedded assets, not templates. To restyle beyond what CSS variables allow, override layout.html and point its <link> at your own file. Note that the Content-Security-Policy is default-src 'none' with style-src 'self': a stylesheet must be served from this origin, and an inline <style> block or style= attribute will be blocked.

Revocation & CRL

acme-proxy implements certificate revocation per RFC 8555 §7.6, and — with the local_ca backend — publishes the resulting Certificate Revocation List.

POST /revokeCert

Revocation is available to a client through the standard ACME endpoint, advertised in the directory. The request payload carries the base64url DER of the certificate and an optional reason code.

Two ways to authorize it

RFC 8555 §7.6 names three, and acme-proxy accepts two:

  1. The order’s account, signing with its kid as usual.
  2. The certificate’s own key pair, signing with an embedded jwk and no account at all. This is the RFC’s accountless case, and it is what lets the holder of a compromised key revoke it even if the ACME account is gone.

The third — an account holding valid authorizations for every identifier in the certificate, without being the one that ordered it — is not supported. It would let one account revoke another’s certificate on the strength of authorizations obtained later, and an operator who needs that has acme-proxy order revoke and the panel, which are attributable.

Because of the second form, this endpoint resolves authorization itself rather than going through the usual account lookup, and it is deliberately not gated on the account’s status — a deactivated account can still revoke its certificates.

How the certificate is identified

The submitted DER is decoded, the order is looked up by the certificate’s serial number, and the stored leaf is then compared to the submitted bytes for an exact DER match. A serial-only lookup would not be enough on its own; the byte comparison is the safety net.

One subtlety worth stating, because getting it wrong is a vulnerability: the key checked against the account is the one stored with the order, never a key re-derived from the submitted certificate. Re-deriving it would let anyone who merely observed the certificate on the wire revoke it, since the certificate contains its own public key.

Responses

  • 200 OK — revoked. For a local_ca profile the order reads revoked at once and the CRL follows as soon as the worker has signed it.
  • 503 serverInternal + Retry-After — relay or custom only: the revocation was queued for the worker and had not completed by the request’s deadline. Nothing failed, and it carries on; asking again waits on the same queued revocation.
  • 400 alreadyRevoked — the certificate was already revoked. This is checked after authorization, so an unauthorized caller cannot use the endpoint to probe whether a certificate has been revoked.
  • 400 badRevocationReason — the reason code is out of range.
  • 401 unauthorized — the signer is neither the order’s account nor the certificate’s key.

Reason codes

Reason codes are RFC 5280 §5.3.1 values. Codes 7 and 11 are not valid CRL reasons, and out-of-range values are meaningless; all three are refused with 400 badRevocationReason, naming the code, so a client that meant a real reason can send it rather than have one silently dropped. Omitting reason entirely is always accepted, and records the revocation with no reason.

The first revocation of a certificate is the one that counts. Two at once — a client and an operator, or acme-proxy order revoke beside a running server — are one withdrawal of trust: whichever writes first sets the reason and the time, on the order and on the CRL alike, and the other is answered alreadyRevoked. The same holds for the queued path: a second request while the first is still queued joins that job instead of queueing another, so it inherits the first request’s reason.

Who revokes: never the request

Withdrawing trust needs what only the worker role holds — the CA key, a token login, a relay’s upstream account — so no request, from a client or from the panel, calls a backend itself:

  • local_ca: the revocation is a database row and the order’s stamp, in one transaction, and the worker signs it into the CRL (local_ca_crl_regenerate). No key is needed to record it.
  • relay / custom: the revocation is a signer_revoke job for the worker, and the request waits on it.

For those two, the backend’s own revoke is called before the order is marked revoked. The CA-side action is authoritative, so if the backend fails, the order is deliberately left un-revoked and the job retries. A backend’s revoke must therefore be idempotent — it may legitimately be called again for a certificate it has already revoked.

Revocation is orthogonal to the order state machine

RFC 8555 defines no “revoked” order status, so a revoked order’s status stays valid. The revocation timestamp and reason are stored in separate columns and are not exposed in the ACME JSON a client polls.

They are visible through the admin CLI:

acme-proxy order show <id>          # prints revoked, reason and the serial
acme-proxy order show <id> --json   # the same as revokedAt / revocationReason

Revoking as an operator

For an out-of-band compromise report that the certificate holder cannot or will not act on:

acme-proxy order revoke <order-id> --reason 1

It is not confirm-gated — unlike order delete — because revocation only ever tightens trust; there is no destructive outcome to protect against. It runs the same checks and writes the same audit row and certificate_revoked notification as the ACME endpoint, and like it never loads the CA key or talks to an upstream itself: whatever holds the signing material is the running server’s worker.

  • local_ca: the command records the revocation in the database and marks the order revoked in one transaction, then queues a local_ca_crl_regenerate job and prints its id. The order reads revoked at once. The server’s job runner signs the new CRL, usually within jobs.poll_interval_ms, and with no server running the CRL catches up when one starts. A CA that no server has started with yet is refused: its revocation state is imported the first time a server meets it.
  • relay / custom: revoking means asking the upstream CA or running the operator’s script, which is the server’s job. The command queues a signer_revoke job and waits for its answer, up to --wait seconds (default 30; 0 returns at once). A job that has not run by then is not an error: the command exits 0 naming it, and acme-proxy jobs show <id> follows it. A job that failed exits 1 with its error.

See Admin CLI.

GET /crl

With the local_ca backend, the CRL (RFC 5280) is served unauthenticated at {base_url}/profile/<name>/crl, with content type application/pkix-crl.

  • It is routed but deliberately not advertised in the ACME directory. A CRL is CA infrastructure, not an ACME resource, so it has no directory entry.
  • A valid, correctly signed empty CRL exists from the moment a worker has started with the CA, before anything has ever been revoked. Clients fetching it do not have to special-case “no revocations yet”. Every role serves the stored CRL; none signs one but the worker.
  • It is signed again for every revocation — immediately where the process holding the key records it, and otherwise by the local_ca_crl_regenerate job that revocation queues — and by a daily refresh; see below.

The revocations and the current signed CRL live in the database, keyed by the CA’s key, so every process over one database — the server, a reloaded configuration, a second server — serves the same CRL. Backing up the database backs up the revocations; there is no separate file to keep in step with it.

crl_path still holds the current CRL, as PEM, rewritten every time a new one is stored. It is an export for operators who publish the CRL from a static web server: nothing ever reads it back, and losing it loses nothing. A second file, ca.json.lock, sits beside it and is held while the export is written, so two processes cannot leave an older CRL in place of a newer one. It is always empty and needs no backup.

Upgrading from a JSON ledger

Before revocations moved into the database, they were kept in a JSON sidecar beside crl_path — the same path with the extension swapped to .json, so ca.crl was accompanied by ca.json. The first time a CA meets a database with no CRL for it, the sidecar is imported: every entry becomes a revocation, and the CRL number resumes above the last one the sidecar published. The server logs local_ca_ledger_imported when it happens, which is at startup through the daily refresh’s first pass, or at the first revocation or CRL fetch if that comes sooner.

The import happens once. The sidecar is never read again nor written, so edits to it after that change nothing. Keep it with an old backup if you like, or delete it.

A sidecar that cannot be imported — a hand-edited serial that is not hex, say — is logged as local_ca_crl_initialization_failed, naming the entry. Until it is fixed, that CA’s revocations and GET /crl fail rather than proceed without the history the sidecar holds; the import is tried again on the next attempt.

Expired entries are dropped

A revocation entry is not kept for ever. RFC 5280 §3.3 lets one go once it has appeared on a CRL issued after the certificate expired — nothing can present the certificate any more, and every relying party has had a CRL saying so — and this is what stops the CRL growing for the life of the deployment. The prune runs daily, starting shortly after startup.

The same daily pass re-signs a CRL that has less than half of its seven-day validity left, even when nothing was revoked or pruned, so a quiet CA’s CRL never lapses. When neither applies, nothing is signed.

Two rules are worth knowing:

  • An entry is dropped an hour after the certificate’s own notAfter, not at it. A relying party whose clock is behind yours still considers the certificate valid for a moment, and that moment is exactly when it would otherwise accept one you revoked.
  • An entry is dropped only once the stored CRL was issued past that notAfter, which is §3.3’s actual condition. A certificate that expires between two signings therefore stays listed for one more pass: the newest CRL a relying party can fetch predates the expiry, and dropping the entry would leave that CRL as the last word on a certificate it still lists as valid to anyone whose clock disagrees.
  • An entry whose expiry is unknown is never dropped. That is any entry recorded before this server started tracking expiries (see below), and an unknown expiry is not an expired one.

Each CRL carries a crlNumber that only ever increases, including across a restart, across a prune that shortens the list, and across several processes signing for one CA. A client that meets a lower number than it has cached keeps its cached CRL, so the number is stored with the CRL and a new one is only ever stored over the one it was numbered after.

Sidecars written by 0.1.0 are a bare JSON array with no expiries and no number. They are imported as-is, with the number resuming above anything that format could have published. Their entries have no known expiry, so they stay on the CRL.

Telling clients where it is

A certificate does not point at this CRL until you say where to fetch it. signer.local_ca.crl_distribution_points is that URL: set it, and every leaf issued from then on carries a cRLDistributionPoints extension naming it (and ca_issuer_urls does the same for the CA’s own certificate). Both are empty by default, so out of the box the CRL above is reachable only by somebody who already knows this server exists.

Two things follow from a URL being signed into a certificate:

  • Certificates already issued keep the URL they were signed with, for their whole validity. Changing the key changes nothing that is already out there, which is why the value is yours to choose rather than something derived from server.base_url.
  • The URL above is not automatically a good answer. /crl is served by the profile router, so it sits behind that profile’s filter policy; an address-based rule will refuse it to relying parties outside the allowlist. Either add a path check permitting /crl (see The path check) or publish a copy of crl_path somewhere unconditionally reachable and name that.

See Local CA for both keys.

Other backends

  • custom — the CRL comes from the script’s crl hook, and only when signer.custom.supports_crl = true. Otherwise there is nothing to serve. See Custom Script Signer.
  • relay — the upstream CA publishes its own CRL or OCSP responder; this server does not republish it.

Interaction with renewal information

A certificate acme-proxy knows to be revoked is reported through ARI with a renewal window entirely in the past, prompting a compliant client to renew immediately.

That check happens before the signer backend is consulted, so a locally revoked certificate is never talked out of renewing by an upstream CA that has not yet noticed.

Notifications

The notify subsystem alerts operators on lifecycle events within the ACME server, and web-admin operators on security events on their own account.

Supported events

EventFired when
profile_mountedA profile is initialized at startup.
account_createdA client registers a new account.
account_deactivatedAn account is deactivated.
certificate_issuedAn order is finalized and a certificate is minted.
certificate_revokedA certificate is revoked, via the ACME API or the admin CLI.
challenge_failedA domain-control validation attempt fails.
certificates_expiringThe periodic expiry digest, one per profile.
admin_sign_inA web-admin sign-in from an unfamiliar address, a refused second factor after a correct password, or a lockout.
admin_credential_changedA web-admin operator’s password or second factor changed.

These nine names are the only valid values wherever a backend’s events list is configured. An unrecognised name is a startup error, not a silently ignored entry.

certificates_expiring is the one that is not a thing that just happened. Every other event describes a single subject at the moment it changed; this one is a digest sent on a schedule, listing the certificates on one profile that expire inside a configured window. It sends nothing until notify.expiry.lead_days is set, so leaving it in an events list costs nothing.

admin_sign_in and admin_credential_changed are the web-admin operator security events (ASVS V6.3.5 / V6.3.7). They fire only on the process-wide [admin.notify] dispatcher — never on a profile’s [notify] — so listing either in a per-profile backend’s events costs nothing. Their email goes to the affected operator’s own address rather than to notify.email.to.

Backends

  • Email — SMTP, via lettre.
  • Webhook — any HTTP endpoint, with the URL, method, headers and body all configured. This is how Slack, Mattermost, Teams, Telegram and Matrix are reached: they differ in those four values and nothing else, so each is a configuration entry rather than a backend of its own.
  • Custom Script — shell out to a local script, for a channel that is not an HTTP request at all.

Email and webhook render their messages with MiniJinja templates you can override; see Customizing Templates.

Configuration

[notify]
# Which backends are active. Empty (the default) means no notifications at all.
enabled = ["email", "webhook"]

# Which [notify.webhook.<name>] entries to POST to, when "webhook" is listed
# above.
webhook_enabled = ["slack"]

# Which [notify.custom.<name>] entries to run, when "custom" is listed above.
custom_enabled = []

# Optional directory of template overrides, checked per template file before
# falling back to the compiled-in default.
template_dir = "/etc/acme-proxy/templates"

Reference

enabled (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__ENABLED

Active backends: any of email, webhook, custom. Empty means the subsystem is off. "mattermost" was removed in favour of webhook and is refused by name.

webhook_enabled (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__WEBHOOK_ENABLED

Which entries under [notify.webhook.<name>] to POST to, and in what order. Listing "webhook" in enabled while leaving this empty is a startup error, as is naming an entry that has no table.

custom_enabled (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__CUSTOM_ENABLED

Which entries under [notify.custom.<name>] to run, and in what order. The same two startup errors apply.

template_dir (String) — Default: "" | Env: ACME_PROXY_NOTIFY__TEMPLATE_DIR

Directory searched for template overrides. Lookup is per file, so overriding one message (say email/certificate_issued.body.j2) leaves every other message at its compiled-in default. Empty means defaults only.

Each backend additionally takes its own events list and timeout_ms; see the backend pages.

Expiry digest

A certificate approaching its notAfter is the one thing the events above cannot report: nothing happens when a certificate is a fortnight from expiring. The [notify.expiry] table adds a periodic sweep that looks, and sends one message per profile listing what it found.

Deliberately not one message per certificate. A renewal is a new order, so the certificate it replaced still reaches its own expiry on schedule — a per-certificate reminder therefore fires for every certificate the CA has ever issued, on its way out, in exactly the deployments where the automation is working. Instead, each entry in the digest says whether something has already taken its place, and the entries where nothing has are the ones worth acting on.

That annotation is drawn from two signals, and the message says which was used: the successor order’s own replaces field (RFC 9773 §5 — exact, but only from clients that send one), or a later, unrevoked certificate of the same account covering all of the same names. Both are deliberately narrow. A certificate wrongly marked as already renewed is one an operator skips over while it lapses; one wrongly left unmarked is a line of noise.

expiry.lead_days (Integer) — Default: 0 | Env: ACME_PROXY_NOTIFY__EXPIRY__LEAD_DAYS

How far ahead to look. 0 is off — the sweep is never scheduled at all, the same shape audit.retention_days and jobs.retention_days use.

expiry.interval_days (Integer) — Default: 7 | Env: ACME_PROXY_NOTIFY__EXPIRY__INTERVAL_DAYS

How often the digest is sent. There is deliberately no per-certificate rate limit beside it: the digest is the rate limit. A digest with nothing to report is not sent, so the absence of a message is what “everything is renewed” looks like.

expiry.max_entries (Integer) — Default: 50 | Env: ACME_PROXY_NOTIFY__EXPIRY__MAX_ENTRIES

The most certificates one message lists. The number that matched is carried whole regardless, so a truncated digest still says how many it did not name.

The schedule is a row in the durable job queue rather than a timer, so it survives a restart: a server restarting more often than interval_days still sends its digest on time instead of resetting the clock each start.

Delivery semantics

Dispatch is fire-and-forget: the event is written to the durable job queue and the ACME response proceeds immediately. A notification backend can never delay or fail the request that triggered it.

Delivery itself is a notify_deliver job, one row per backend per event, so one flaky webhook is retried without re-sending through an email backend that already succeeded. Two consequences worth planning around:

  • A row outlives the process that wrote it. A notification generated moments before a restart is delivered by whoever starts next, rather than lost. There is no drain at shutdown to configure or wait for.
  • A failure is retried, unless it never could have worked. A refused SMTP connection, a timeout, a 429 or a 5xx from a webhook goes back in the queue under jobs.max_attempts and the shared backoff. A template that does not render, a url that does not parse and any other 4xx are refused on the first attempt — retrying would reach the same answer four more times and delay the log line saying so.

Every attempt logs notify_delivered or notify_delivery_failed. When the attempts run out, one notify_delivery_abandoned says the notification is genuinely lost — that is the line to alert on. Because the custom backend’s contract is an exit code with no way to say “never retry”, every failure of a custom script is treated as retryable.

Customizing Templates

acme-proxy uses the MiniJinja templating engine to render notification payloads. The server embeds sensible default templates inside the binary, but you can override any of them by pointing the server to a custom template directory.

Enabling custom templates

In your config.toml, define a template_dir:

[notify]
enabled = ["email", "webhook"]
template_dir = "/etc/acme-proxy/templates"

The server will look in this directory before falling back to its embedded defaults. You only need to create the files you want to override.

Template file structure

Templates are grouped by backend and event name. Email requires separate files for the subject line and body.

/etc/acme-proxy/templates/
├── email/
│   ├── account_created.subject.j2
│   ├── account_created.body.j2
│   ├── certificate_issued.subject.j2
│   └── certificate_issued.body.j2
└── webhook/
    ├── certificate_issued.j2
    └── challenge_failed.j2

(Available event names: profile_mounted, account_created, account_deactivated, certificate_issued, certificate_revoked, challenge_failed, certificates_expiring, admin_sign_in, admin_credential_changed).

A webhook/<event>.j2 renders the message, not the payload: every [notify.webhook.<name>] entry then wraps it in its own body template. So a file here restyles the text for every webhook target at once, and an entry’s body restructures one target’s request without touching the text. See Webhook.

Context variables

When rendering a template, acme-proxy passes a context object containing the event’s data. All events include profile (the name of the profile triggered) and most include client_ip (the IP address of the ACME client that initiated the request).

certificate_issued

Triggered when an order is finalized and the signer mints a certificate.

  • profile (String)
  • order_id (String)
  • account_id (String)
  • cert_serial (String) - Hex-encoded serial number
  • identifiers (List of Strings) - The SANs/Domains requested
  • client_ip (Option<String>)

Example (webhook/certificate_issued.j2):

✅ **Certificate Issued** on profile `{{ profile }}`
**Domains:** {{ identifiers | join(", ") }}
**Serial:** `{{ cert_serial }}`
**Requested By IP:** `{{ client_ip | default("system") }}`

challenge_failed

Triggered when an HTTP-01 or DNS-01 validation attempt fails.

  • profile (String)
  • order_id (String)
  • account_id (String)
  • authz_id (String)
  • challenge_id (String)
  • challenge_type (String) - e.g. “http-01”
  • identifier (String) - The domain that failed
  • error (String) - The detailed error from the validation attempt
  • client_ip (Option<String>)

account_created / account_deactivated

Triggered on account lifecycle events.

  • profile (String)
  • account_id (String)
  • contact (List of Strings) - e.g. ["mailto:admin@example.com"] (only on created)
  • client_ip (Option<String>)

certificate_revoked

Triggered via the ACME API or Admin CLI.

  • profile (String)
  • order_id (String)
  • account_id (String)
  • cert_serial (String)
  • reason (Option<Integer>) - RFC 5280 revocation reason code
  • client_ip (Option<String>)

profile_mounted

Triggered during server startup when a profile is successfully initialized.

  • profile (String)

certificates_expiring

The periodic expiry digest — the one event whose subject is a list, so a template here loops where every other one interpolates.

  • profile (String)
  • generated_at (Integer) - Epoch seconds; the point days_remaining counts from, so a message read days later is still self-describing
  • lead_days (Integer) - The window this digest covers
  • total (Integer) - How many certificates matched, which may be more than certificates holds: notify.expiry.max_entries bounds the list, and the count is what lets a truncated message say how many it did not name
  • certificates (List) - each with order_id, account_id, cert_serial, identifiers (List of Strings), not_after (Integer, epoch seconds), days_remaining (Integer, floored) and superseded_by
  • superseded_by (Option) - absent when nothing has replaced this certificate; otherwise order_id, cert_serial, not_after and via, where via is "replaces" (the client said so, RFC 9773 §5) or "identifiers" (a later certificate of the same account covers the same names)

There is no client_ip: a digest is generated by a sweep, with no request anywhere in scope.

Example (webhook/certificates_expiring.j2):

⏰ **{{ total }} expiring** on `{{ profile }}`
{%- for cert in certificates %}
{{ cert.identifiers | join(", ") }} — {{ cert.days_remaining }}d
{{- " (already replaced)" if cert.superseded_by }}
{%- endfor %}

The superseded_by test is what makes a digest readable: an operator scans for the entries without it. Rendering every row identically would bury the handful that nobody has renewed among the many that are already taken care of.

admin_sign_in

A web-admin operator sign-in worth flagging.

  • profile (String) — always __admin__
  • username (String) — the operator
  • recipient (Option) — the operator’s own contact address; the email backend sends there rather than to notify.email.to
  • outcome (String) — succeeded_from_new_address, second_factor_refused or locked_out
  • client_ip (Option), user_agent (Option)
  • at (Integer) — epoch seconds

admin_credential_changed

A web-admin operator’s password or second factor changed.

  • profile (String) — always __admin__
  • username, recipient — as above
  • change (String) — password, second_factor_enabled, second_factor_disabled or recovery_codes_regenerated
  • by_self (Boolean) — false when another administrator made the change
  • client_ip (Option), user_agent (Option), at (Integer)

Email Notifications

The email notification backend sends alerts via SMTP, using the lettre crate to format and dispatch multipart messages.

Delivery semantics

Notifications in acme-proxy are designed to be entirely non-blocking.

  1. When an event occurs (e.g., a certificate is issued or revoked), the notification subsystem spawns a background Tokio task.
  2. The task attempts to connect to the SMTP server and send the email.
  3. Fire and Forget: If the delivery fails (e.g., the SMTP server is unreachable), acme-proxy logs an error at the warn level but does not retry. Notification failure does not cause the ACME transaction (like finalize) to fail.

Template customization

By default, acme-proxy sends plain, functional emails rendered from templates compiled into the binary. You can override any of them by setting notify.template_dir.

The engine is MiniJinja and the files are .j2. Email needs two files per event — one for the subject line and one for the body — under an email/ subdirectory of template_dir:

/etc/acme-proxy/templates/
└── email/
    ├── certificate_issued.subject.j2
    └── certificate_issued.body.j2

Lookup is per file, so overriding one message leaves the rest at their defaults. See Customizing Templates for the full layout and the context variables available to each event.

Configuration

[notify]
enabled = ["email"]
# Note: the path is the templates root, NOT the email/ directory inside it —
# "email/" is appended by the loader.
template_dir = "/etc/acme-proxy/templates"

[notify.email]
smtp_host = "smtp.internal.corp"
smtp_port = 587
smtp_username = "acme_bot"
smtp_password = "super_secret_password"
smtp_security = "starttls"
from = "acme-proxy@internal.corp"
to = ["pki-admins@internal.corp", "security@internal.corp"]
events = ["certificate_issued", "certificate_revoked"]
timeout_ms = 5000

Reference

notify.enabled and notify.template_dir are subsystem-wide keys documented in Notifications. The keys below are specific to this backend.

smtp_host (String) — Default: "" | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_HOST

SMTP server hostname.

smtp_port (Integer) — Default: 587 | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_PORT

SMTP server port.

smtp_username (String) — Default: "" | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_USERNAME

SMTP authentication username.

smtp_password (String) — Default: "" | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_PASSWORD

SMTP authentication password. Sensitive — prefer the environment variable to a file on disk, as with every other secret in the configuration.

smtp_security (String) — Default: "starttls" | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_SECURITY

TLS requirement: "starttls", "tls", or "none".

from (String) — Default: "" | Env: ACME_PROXY_NOTIFY__EMAIL__FROM

Sender email address.

to (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__EMAIL__TO

List of recipient email addresses. Required for a per-profile [notify] backend. Under [admin.notify] it is the fallback: the web-admin security events (admin_sign_in, admin_credential_changed) are delivered to the affected operator’s own contact address, and to is used only when that operator has none — so [admin.notify.email].to may be left empty.

events (Array) — Default: every event | Env: ACME_PROXY_NOTIFY__EMAIL__EVENTS

Lifecycle events this backend reacts to.

timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_NOTIFY__EMAIL__TIMEOUT_MS

Timeout budget for the SMTP exchange.

Webhook Notifications

The webhook backend makes one HTTP request per event, with the URL, the method, the headers and the body all stated in configuration.

That is the whole design. Slack, Mattermost, Microsoft Teams, Telegram and Matrix differ in those four values and in nothing else — the transport, the timeout, the TLS stack, the proxy and the retry rules are the same for all of them — so a chat provider is an entry in a table here rather than a backend in the binary. See Provider recipes below for a copy-pasteable entry per provider.

Configuration

Entries are named, selected and ordered by webhook_enabled, exactly like [notify.custom]. Two entries are two independent deliveries: a retry never re-sends through one that already succeeded.

[notify]
enabled = ["webhook"]
webhook_enabled = ["slack", "oncall"]

[notify.webhook.slack]
url = "https://hooks.slack.com/services/T00000000/B00000000/XXXXXXXX"
body = '{"text": {{ message | tojson }}}'

[notify.webhook.oncall]
url = "https://chat.internal.corp/api/v1/rooms/pki/messages"
events = ["certificate_revoked", "challenge_failed"]

[notify.webhook.oncall.headers]
Authorization = "Bearer s3cret"

An entry name must match ^[a-z0-9-]+$: it is also an environment-variable segment, which the configuration loader lowercases, so anything else could name one entry in a file and a silently different one through the environment.

How a body is rendered

Rendering has two stages, and the split is what lets you restyle every message without restructuring any payload, or the reverse:

  1. The message. webhook/<event>.j2 renders human-readable text. These are embedded in the binary and overridable file by file through notify.template_dir — override webhook/certificate_issued.j2 and every other message stays at its default.
  2. The payload. The entry’s own body is a template with message, hook (the event name) and every field of the event in scope — the same fields Customizing Templates lists.

| tojson is not decoration. A .j2 template has auto-escaping off, on purpose, so a message holding a quote or a newline — a challenge_failed quotes what the validator saw — would otherwise render a payload the provider answers 400 to. That answer is permanent, so the delivery is refused on the first attempt, for exactly the events you most wanted to hear about.

Provider recipes

Everything below is the whole entry: no other key is needed. message is the rendered text from stage 1.

Providermethodbodyheaders
SlackPOST{"text": {{ message | tojson }}}—
MattermostPOST{"text": {{ message | tojson }}}—
Microsoft TeamsPOST{"text": {{ message | tojson }}}—
Google ChatPOST{"text": {{ message | tojson }}}—
TelegramPOST{"chat_id": "-1001234567890", "text": {{ message | tojson }}}—
MatrixPUT{"msgtype": "m.text", "body": {{ message | tojson }}}Authorization: Bearer <token>

Notes on the two that are not simply a URL:

  • Telegram’s URL is https://api.telegram.org/bot<token>/sendMessage, and the destination is the chat_id in the body rather than anything in the URL.
  • Matrix’s URL is the room’s send endpoint, /_matrix/client/v3/rooms/<room>/send/m.room.message/<txn>, on your own homeserver. The transaction id is meant to change per message; a fixed one makes the homeserver deduplicate, which is a deliberate choice worth knowing you are making.

Mattermost’s channel and username overrides, which this backend replaced, are two more members of the same object:

body = '{"channel": "pki-alerts", "username": "ACME Proxy", "text": {{ message | tojson }}}'

Delivery semantics

Nothing here is specific to this backend — see Notifications for the queue, the retries and the one log line that means a notification was genuinely lost. Two points that bite webhooks in particular:

  • A 4xx is permanent. Every 4xx except 408 and 429 is the provider stating a reason — a webhook that has been deleted, a payload it will not accept — so it is refused on the first attempt rather than retried four more times. 5xx, 429, 408, a timeout and a refused connection are retried. A refusal’s response body is quoted in the log line, truncated: Slack’s invalid_payload is usually the entire diagnosis.
  • Everything unusable is refused at startup, not at delivery time: an empty or unparseable url, a scheme that is not http/https, a method outside POST/PUT/PATCH, a header a wire format will not carry, and a body that does not compile as a template.

The URL’s path and every header value routinely carry the credential, so nothing this server logs or reports ever renders more than the host and the header names.

Reference

url (String) — Default: "" | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__URL

The endpoint to call. Required once the entry is listed in webhook_enabled.

method (String) — Default: "POST" | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__METHOD

POST, PUT or PATCH, in any case. Anything else is a startup error: a webhook is a write, and a body on GET would be a request no provider answers.

headers (Table) — Default: {} | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__HEADERS__<HEADER>

Extra request headers. Applied after the defaults, so an entry may override content-type (application/json) or user-agent (acme-proxy). Header names arrive lowercased from the environment, which HTTP does not care about.

body (String) — Default: {"text": {{ message | tojson }}} | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__BODY

The request body, as a template. The default is the payload Slack, Mattermost, Teams and Google Chat all accept, so those four need a url and nothing else.

events (Array) — Default: every event | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__EVENTS

Lifecycle events this entry reacts to.

timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__TIMEOUT_MS

Budget for one delivery attempt, connect and handshake included.

Custom Script Notifications

The custom notification backend runs a local script (Bash, Python, Go, …) when an ACME event occurs. Use it to integrate with internal ticketing systems, custom logging infrastructure, or alerting pipelines that a plain HTTP webhook cannot satisfy.

Configuration

custom is a named map, like [filter.check]. Two keys switch it on: notify.enabled activates the backend, and notify.custom_enabled selects which scripts run, and in what order.

[notify]
enabled = ["custom"]
custom_enabled = ["ticket-creator"]

[notify.custom.ticket-creator]
script_path = "/etc/acme-proxy/scripts/ticket-creator.sh"
timeout_ms = 10000
args = []
events = ["certificate_issued", "certificate_revoked"]

Three ways to get this wrong, all of which fail at startup rather than at delivery time:

  • Listing "custom" in notify.enabled while custom_enabled is empty.
  • Naming an entry in custom_enabled that has no [notify.custom.<name>] table.
  • Using an entry name outside ^[a-z0-9-]+$ — ticket_creator (underscore) is rejected; ticket-creator is fine.

One process is spawned per enabled script per event.

Reference

script_path (String) — Default: "" | Env: ACME_PROXY_NOTIFY__CUSTOM__<NAME>__SCRIPT_PATH

Path to the executable. Required.

timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_NOTIFY__CUSTOM__<NAME>__TIMEOUT_MS

Maximum execution time. A script still running when this expires is killed.

args (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__CUSTOM__<NAME>__ARGS

Static arguments passed to the script on every invocation.

events (Array) — Default: every event | Env: ACME_PROXY_NOTIFY__CUSTOM__<NAME>__EVENTS

Which events this script reacts to. Valid names are profile_mounted, account_created, account_deactivated, certificate_issued, certificate_revoked, challenge_failed, certificates_expiring. An unrecognised name is a startup error.

Execution model and security

  1. Environment clearing (env_clear): the child runs with a scrubbed environment, inheriting only a minimal PATH and the injected ACME_NOTIFY_* variables. The server’s own environment may hold secrets — notify.email.smtp_password, the RFC 2136 TSIG key, the NetBox token — and a notification script has no business reading them.
  2. Zombie prevention (kill_on_drop): the script runs under a Tokio timeout with kill_on_drop(true). A tokio::time::timeout only drops the future, so without this a timed-out script would outlive its deadline and leak a process per event.
  3. Fire and forget: delivery runs in a background task. A failing or hanging script can never delay or fail the ACME response that triggered it; failures are logged (event = "notify_delivery_failed") and dropped, with no retry.

Data passing

The script receives context both ways.

Environment variables

All seven are always set. Ones that do not apply to the event are set to the empty string rather than omitted, so a script can read them unconditionally.

VariableValue
ACME_NOTIFY_HOOKThe event name, e.g. certificate_issued.
ACME_NOTIFY_PROFILEThe profile the event occurred in.
ACME_NOTIFY_CLIENT_IPThe ACME client’s address; empty when no request was in scope (e.g. profile_mounted, or an asynchronous relay completion).
ACME_NOTIFY_ACCOUNT_IDThe account, when the event has one.
ACME_NOTIFY_ORDER_IDThe order, when the event has one.
ACME_NOTIFY_CERT_SERIALThe certificate serial, on certificate_issued / certificate_revoked.
ACME_NOTIFY_IDENTIFIERSComma-joined identifier values. Only populated for certificate_issued.

On certificates_expiring every variable but ACME_NOTIFY_HOOK and ACME_NOTIFY_PROFILE is empty, and that is not an omission: a digest is about a list of certificates spanning however many accounts, so there is no one account, order or serial for a variable to hold. Read the list from the JSON on stdin, which is the channel that carries structure.

On admin_sign_in / admin_credential_changed the only populated variables are ACME_NOTIFY_HOOK, ACME_NOTIFY_PROFILE (__admin__) and ACME_NOTIFY_CLIENT_IP (the web-admin client’s address). The username, recipient, outcome / change, by_self and user_agent are on the stdin JSON — an operator is not an ACME subject, so it has no account or order id.

There is no ACME_NOTIFY_EVENT; the event name is ACME_NOTIFY_HOOK.

JSON on stdin

A JSON object is always written to the script’s standard input. It is the event’s own fields plus a "hook" key naming the event — the same data the templating backends render from. For example, on certificate_issued:

{
  "hook": "certificate_issued",
  "profile": "default",
  "order_id": "…",
  "account_id": "…",
  "cert_serial": "…",
  "identifiers": ["a.example.com", "b.example.com"],
  "client_ip": "203.0.113.5"
}

The exact fields per event are listed in Customizing Templates — the template context and the stdin payload carry the same values.

The issued certificate itself is never passed to a notification script, on stdin or otherwise. Only its serial and the identifiers it covers are available. A script that needs the PEM must fetch it out of band.

Example

#!/bin/bash
# /etc/acme-proxy/scripts/ticket-creator.sh
set -euo pipefail

payload=$(cat)   # the JSON described above

case "$ACME_NOTIFY_HOOK" in
  certificate_issued)
    echo "issued ${ACME_NOTIFY_CERT_SERIAL} for ${ACME_NOTIFY_IDENTIFIERS}" \
      >> /var/log/acme-issuance.log
    ;;
  certificate_revoked)
    # ACME_NOTIFY_IDENTIFIERS is empty for this event — read the payload
    # if you need more than the serial.
    reason=$(echo "$payload" | jq -r '.reason // "unspecified"')
    curl -sS -X POST https://tickets.internal/api/incidents \
      -H 'Content-Type: application/json' \
      -d "{\"serial\":\"$ACME_NOTIFY_CERT_SERIAL\",\"reason\":\"$reason\"}"
    ;;
esac

See Custom Plugins Examples for more.

Audit Trail

A certificate authority’s most important record is not what it holds but what it did: who asked it to sign or withdraw a certificate, who administered it, from where, and what it answered. acme-proxy writes that down in one append-only table and surfaces it in three places — the CLI, the JSON API and the panel.

There is deliberately no audit.enabled. Recording who asked the CA to sign something, and who changed its stored state, is not a feature of this server, it is what a CA does. The only thing an operator can switch off is the reverse-DNS lookup, because that one costs a network round trip and there are estates where it can never succeed.

What gets recorded

Two different things, in two different places.

Traceability columns, on the rows themselves:

TableColumns
accountscreated_ip / created_ptr — where newAccount was called from, frozen at creation; last_seen_at / last_seen_ip / last_seen_ptr — where the key last authenticated a request
orderscreated_ip / created_ptr — where the order was opened from

The audit log, one row per CA action and per refusal, plus one per successful administrative action:

EventWritten when
certificate_issuedthe signer returned a certificate
certificate_issue_failedfinalize was refused after the order was ready — a bad CSR, a CSR/order mismatch, a filter denial, or the backend’s own failure
certificate_revokeda revocation succeeded, whether through ACME, the CLI or the panel
certificate_revoke_faileda revocation was refused — including alreadyRevoked, an unauthorized caller, and an unknown certificate

The refusals are the point. A stream of certificate_revoke_failed rows naming certificates that do not exist is somebody enumerating serials, which is exactly the question a trail exists to answer.

Administrative actions

An operator (through the panel) or the host CLI changing stored state also leaves a row. These are recorded only on success — a not-found or refused admin operation is the operator being told the state of things, not the CA turning a remote party away — and each is attributed to actor_kind = "admin" (the operator’s username and resolved address) or "cli" (the host).

EventWritten when
account_deactivated / account_contact_updated / account_deletedan ACME account was deactivated, had its contact list rewritten, or was hard-deleted (the row names the account and its profile; account_deleted carries the cascade count)
order_deletedan order row was hard-deleted (names the order, account and identifiers)
eab_created / eab_revoked / eab_deletedan External Account Binding credential was minted, revoked or deleted (the detail names the kid; the secret is never recorded). eab_deleted also says whether its accounts were kept, deactivated or deleted, and each account changed gets its own account_deactivated or account_deleted row
operator_created / operator_deleteda web-admin operator was added or removed
operator_role_changed / operator_disabled / operator_enabledan operator’s privilege tier or sign-in status changed
operator_password_changed / operator_contact_updatedan operator’s password or notification address changed (self-service or by an admin)
operator_totp_enrolled / operator_totp_disabled / operator_recovery_codes_regeneratedan operator’s second factor was enrolled, removed (including an admin reset), or its recovery codes reissued
session_revokedone session, all of an operator’s sessions, or every session on the server was revoked (the detail says which)
job_cancelled / job_advanceda background job was cancelled or nudged/revived (run-now). A cancelled relay issuance instead writes certificate_issue_failed — it abandons a certificate order
nonce_cleanup_completed / audit_prunedthe nonce table or the audit log itself was swept by hand (audit_pruned records its own action, so a manual prune always leaves the one row that says it happened)
database_transferredevery row was copied into another backend (acme-proxy transfer). Written to the source, which is the database that holds the trail leading up to the move; the copy has already read past this row, so the target’s own trail begins at the transfer it arrived in

The vocabulary is defined by acme_proxy_core::audit::AuditEvent, not by a database constraint, so a newer server writing a name an older one does not know still loads on the older one — it simply shows the raw string. The web audit surface stays read-only: a stolen session that could erase the trail would make the trail prove nothing, so pruning is audit cleanup on the host or audit.retention_days and nothing else.

Two boundaries worth knowing:

  • Nothing is ever compared against any of this. Pinning an identity to an address breaks CGNAT and mobile clients, so these columns answer “who asked for this certificate, and from where”, never “may this request proceed”.
  • None of it reaches an ACME object. The wire format is RFC 8555’s and stays that way. The trail is visible through the admin surfaces only.

What is not recorded

Refusals that never reached a CA action. A request rejected for a bad signature, a replayed nonce or an unready order is protocol bookkeeping — nothing was signed, nothing was withdrawn, and recording it would bury the rows that matter. For the same reason, a revokeCert payload that cannot be parsed at all writes nothing: a row naming no subject is noise.

Reading it from the CLI

acme-proxy audit list
acme-proxy audit show 4213
CommandFlags
audit list--profile, --account-id, --order-id, --cert-serial, --event, --outcome, --since-days <n>, --limit <n>, --offset <n>, --json
audit show <id>--json
audit cleanup--older-than <days> (prompts)

audit list is paged, defaulting to 50 rows, as account list and order list are. This table grows a row per issuance for the life of the deployment, so it always prints N of M row(s) — a page must never be mistaken for the whole trail. There is no “everything” spelling on purpose; on a year-old CA that is a terminal full of scrollback and a table loaded into memory. Page with --offset, and see Paging for the window every listing shares and the --json envelope it answers with.

$ acme-proxy audit list --outcome failure --since-days 7
4213      2026-08-09T18:22:04Z  certificate_issue_failed    prod          acme:acct-9f2c              10.4.1.19 (web7.corp.example)              api.corp.example  reason=badCSR
1 of 1 row(s).

An unknown --event or --outcome is refused by name, listing the values this build knows. Passed through to SQL it would answer “no rows”, which looks exactly like “nothing happened” — the single most misleading answer an audit tool can give.

audit show prints one field per line, omitting every field that was not recorded rather than showing it as empty:

$ acme-proxy audit show 4213
id           4213
created      2026-08-09T18:22:04Z
event        certificate_issue_failed
outcome      failure
profile      prod
actor        acme:acct-9f2c
account      acct-9f2c
order        ord-71ab
client_ip    10.4.1.19
client_ptr   web7.corp.example
user_agent   certbot/2.9.0
request_id   01J9F2K7Q4
reason       badCSR
identifiers  api.corp.example

The actor

Every row names who acted, as a kind plus an optional id:

KindMeaning
acmean ACME client, identified by its account id
admina web admin operator, identified by username
clisomeone on the host — there is no request and no address, and the row says so
systemthe server itself, e.g. a relayed order settling in the background

An administrative revocation is attributed to the operator, not the certificate’s owner. Recording the client there would say the opposite of what happened.

Retention

audit.retention_days = 0 — the default — keeps everything for ever, which is the right default for a trail whose value is that it is complete. Setting it non-zero spawns a daily sweep running the identical DELETE as:

acme-proxy audit cleanup --older-than 365

This is the only command in the binary that destroys audit history, so it is confirm-gated and its prompt names the number of rows it is about to remove. Use -y to skip the prompt in a cron job.

The web admin cannot prune the trail. /api/audit and /ui/audit are read-only, and there is no route to list, because the first thing a stolen session would do is erase what it had done — and a trail that can be erased by the thing it is watching proves nothing. Pruning happens on the host, or on a schedule set in configuration.

Reading it from the web admin

Surface
GET /api/audit?profile=&accountId=&orderId=&certSerial=&event=&outcome=&limit=&offset=the paged envelope every list endpoint returns
GET /api/audit/{id}one row
/ui/auditthe list, filterable, with a detail page per row

The API omits absent fields rather than sending them as null, so a client can test for presence directly. There is no since filter on this surface: a browser filters by picking a page, and a date parser here would be a second definition of “how far back” that the CLI already has.

Rows carry remote text. A User-Agent and a reverse DNS name are both written by whoever is on the other end — the PTR for a client’s address is controlled by whoever runs that address’s reverse zone. The panel escapes them like any other untrusted value, and tests/admin_pages.rs pins it.

Why it survives deletion

audit_log is the one table in the schema with no foreign keys, and that is the design rather than an oversight. An audit row has to outlive account delete and order delete; a CASCADE would destroy the evidence along with its subject. So account_id and order_id are plain columns naming a row that may be gone, and the identifiers are frozen into the row rather than joined back to an order that no longer exists.

The rest follows from the same rule: rows are only ever inserted — there is no setter and no UPDATE anywhere in the crate — and the primary key is AUTOINCREMENT, so SQLite cannot reuse the rowid of a purged row.

Configuration

See [audit] for the three keys. The section is process-wide rather than per-profile: the trail describes the CA, not one of its endpoints.

Monitoring & Observability

Running an ACME server in production requires visibility into its health, request volume, and error rates. acme-proxy offers three surfaces for that: a health endpoint, structured logging, and a Prometheus exporter.

Health checks

  • Endpoint: GET /health
  • Response: 200 OK with the body {"healthy": true}

The handler takes no application state: it does not query the database, and it does not consult the signer. It answers 200 whenever the process is alive and able to accept a connection, and it has no other failure mode. Treat it as a liveness probe, not a readiness probe — it will not tell you that the disk holding sqlite.db filled up.

What makes it worth polling is where it is mounted. /health lives on the root router, not inside a profile, which means it is deliberately outside:

  • the admission-control layer, so a saturated server still answers the probe (inside the limit, the probe was starved exactly when it mattered, and a load balancer would go on reporting the node healthy right up to the point where the probe could no longer get a slot);
  • every profile’s filter chain, so an IP allowlist never has to be widened for your load balancer;
  • the Replay-Nonce middleware, so probing costs no database write; and
  • the Link: rel="index" middleware.

GET / redirects to /health.

Metrics

GET /metrics serves the OpenMetrics text format (application/openmetrics-text; version=1.0.0), which Prometheus reads natively. It is off by default and lives on a listener of its own — a third socket beside the ACME and admin ones, configured by [metrics]:

[metrics]
enabled      = true
bind_address = "127.0.0.1:3002"

The separate port is the access control. A scrape carries no credential and none is checked: reaching the port at all is the permission, so restrict it the way you restrict any other internal service — a firewall rule, a network policy, or leaving it on loopback and scraping from the same host. The endpoint is absent from the ACME and admin sockets entirely, so exposing the ACME listener to the internet does not expose this.

$ curl -s localhost:3002/metrics
# HELP acme_proxy_requests Requests served, by endpoint, matched route and response status.
# TYPE acme_proxy_requests counter
acme_proxy_requests_total{role="acme,admin,worker",profile="default",route="/newOrder",status="201"} 42
acme_proxy_requests_total{role="acme,admin,worker",profile="none",route="/health",status="200"} 8613
# HELP acme_proxy_request_duration_seconds Time to answer a request, by endpoint and matched route.
# TYPE acme_proxy_request_duration_seconds histogram
# UNIT acme_proxy_request_duration_seconds seconds
acme_proxy_request_duration_seconds_sum{role="acme,admin,worker",profile="default",route="/newOrder"} 0.84
acme_proxy_request_duration_seconds_count{role="acme,admin,worker",profile="default",route="/newOrder"} 42
acme_proxy_request_duration_seconds_bucket{role="acme,admin,worker",le="0.005",profile="default",route="/newOrder"} 3
…
acme_proxy_request_duration_seconds_bucket{role="acme,admin,worker",le="+Inf",profile="default",route="/newOrder"} 42
# HELP acme_proxy_certificates_issued Certificates signed, by endpoint.
# TYPE acme_proxy_certificates_issued counter
acme_proxy_certificates_issued_total{role="acme,admin,worker",profile="default"} 41
# HELP acme_proxy_certificate_issue_failures Issuance attempts the CA refused, by endpoint and ACME problem type.
# TYPE acme_proxy_certificate_issue_failures counter
acme_proxy_certificate_issue_failures_total{role="acme,admin,worker",profile="default",reason="badCSR"} 1
# HELP acme_proxy_certificate_issue_duration_seconds Time from an accepted finalize to a stored certificate, by endpoint.
# TYPE acme_proxy_certificate_issue_duration_seconds histogram
# UNIT acme_proxy_certificate_issue_duration_seconds seconds
acme_proxy_certificate_issue_duration_seconds_sum{role="acme,admin,worker",profile="default"} 45.0
acme_proxy_certificate_issue_duration_seconds_count{role="acme,admin,worker",profile="default"} 41
…
# HELP acme_proxy_database_pool_connections Connections in the database pool.
# TYPE acme_proxy_database_pool_connections gauge
acme_proxy_database_pool_connections{role="acme,admin,worker",state="idle"} 4
acme_proxy_database_pool_connections{role="acme,admin,worker",state="busy"} 1
# EOF

A counter’s # TYPE line names its family without _total, as OpenMetrics requires; its series, which is what a query names, keep the suffix.

MetricTypeLabels
acme_proxy_requests_totalcounterrole, profile, route, status
acme_proxy_request_duration_secondshistogramrole, profile, route
acme_proxy_certificates_issued_totalcounterrole, profile
acme_proxy_certificate_issue_failures_totalcounterrole, profile, reason
acme_proxy_certificate_issue_duration_secondshistogramrole, profile
acme_proxy_database_pool_connectionsgaugerole, state

Six things are worth knowing about the numbers.

role names the roles that process runs, comma-separated — acme,admin,worker for an all-in-one deployment, which is the default. The counters are per-process memory by design, so a split deployment is several scrape targets reporting the same family names, and this is what tells them apart. Each process also needs its own metrics.bind_address.

route is the matched route pattern, not the URI. A request for /profile/le/order/9f3c… is counted under route="/order/{id}", and the endpoint it reached becomes the profile label. Anything that matched no route at all — a scanner, a typo — collapses into a single route="<unmatched>" series rather than one series per path tried. Root-router requests such as /health carry profile="none".

reason is the ACME problem type the CA refused with (badCSR, serverInternal), the same vocabulary acme-proxy audit list --event certificate_issue_failed prints. Both are rendered from one record, so the metric and the trail cannot disagree about what happened.

The two histograms time different things. acme_proxy_request_duration_seconds runs from the request reaching the router to its response head, with buckets from 5 ms to 10 s; it has no status label, because a histogram multiplies every label by its bucket count. acme_proxy_certificate_issue_duration_seconds runs from the finalize request being accepted to the certificate being stored, whichever process signs it, with buckets from 1 s to 1 h. It is measured from job timestamps, which are whole seconds, and for a relayed order it includes the upstream CA’s own validation. It is observed by the process that signs, so in a split deployment it is on the worker’s scrape target.

The counters survive a reload but not a restart. SIGHUP rebuilds the routers and keeps the registry, so a configuration change does not read as a counter reset; a restart genuinely is a new process and starts from zero, which is what rate() expects.

The pool gauge is read from the pool. state="idle" is connections the pool holds but has not checked out and state="busy" those in use, so busy is in-flight database work rather than a number this server maintains in parallel.

A minimal scrape configuration:

scrape_configs:
  - job_name: acme-proxy
    static_configs:
      - targets: ['ca.internal:3002']

A dashboard over all six families ships in the repository — see Grafana Dashboard.

Logging

acme-proxy uses the tracing crate, configured by the six keys in [logging] — filter, JSON or human-readable output, stdout or stderr, ANSI colour, span timing, and whether JSON fields sit at the top level. Every one of them is validated at startup: an unknown value is a refusal to start with a message naming the key, never a silent fallback.

Two of them are worth setting deliberately in production:

[logging]
# Structured fields become first-class keys rather than text to be re-parsed.
json_format = true
# ... and at the top level, which is what most pipelines want.
flatten_event = true

The rest of this page is about what those records contain.

Log levels

RUST_LOG=acme_proxy=info   # operational logging (recommended for production)
RUST_LOG=acme_proxy=debug  # per-request detail, challenge validation steps,
                           # database writes, truncated response previews
                           # from http-01

There is no trace level to reach for: the crate emits nothing below debug.

RUST_LOG replaces the whole filter, so a bare RUST_LOG=debug also turns on debug logging for every dependency. acme-proxy does not configure SQL statement logging itself; to see sqlx’s own statement logs you must ask for them explicitly, e.g. RUST_LOG=acme_proxy=info,sqlx=debug.

Everything on this page is the server’s log stream. The admin commands emit none of it unless asked, with --log-level or a non-empty RUST_LOG, and what they then emit goes to stderr rather than into the output a script is parsing.

Request correlation

Every request passes through one server-wide middleware that reads an incoming x-request-id header or generates a UUID v7 when absent, opens the request span, and echoes the id back on the response. Every log line emitted while handling that request is nested under that span and carries the id, so one client’s failing renewal can be pulled out of a busy log in a single query — and a reverse proxy that already assigns request ids will have its value preserved rather than replaced. A header whose bytes are not valid ASCII cannot be echoed back, so it is replaced by a generated id and a request_id_header_invalid line at debug says so.

The span carries method, uri, version, request_id, client_ip and — for a request that reached an ACME endpoint rather than /health — profile.

client_ip is the peer address of the connection, replaced by the address resolved from filter.forwarded_header where filter.trusted_proxies says the peer is a reverse proxy. Both are per-profile settings, so a request that never reaches a profile — /health, the http-01 responder, anything admission control sheds — always shows the peer. A request arriving over a socket with no peer address at all leaves the field absent rather than reporting a placeholder.

The access line

Each request closes with exactly one request_completed record carrying status and latency_ms. It is emitted under this crate’s own target, so it is visible at the default filter. Its level varies:

CaseLeveloutcome
Response status is 5xxwarnfailure
GET/HEAD of /health or /debugsuccess
Everything elseinfosuccess

This is the only event name emitted at more than one level, and the only one whose outcome comes from the response rather than from its own name. Liveness probes are at debug deliberately: at one probe per second per node they would otherwise be most of the log. Set RUST_LOG=acme_proxy=debug to see them.

Structured events

Every log record carries two fields you can build alerting on, and their shape is enforced by a test rather than by convention (tests/logging_convention.rs).

event is a stable, greppable name, shaped <subsystem>_<object>_<outcome> — always a literal in the source, so a name in a log greps straight back to the line that wrote it. The subsystem prefix comes from a closed list, so a family grep finds the whole family: event = "certificate_revoke reaches every revocation refusal, db_ reaches the storage layer, challenge_http_01_ reaches one validator. Challenge types are always spelled with separators (http_01, dns_01, tls_alpn_01).

outcome is one of four values, and it exists so that “show me everything that broke” is an exact match instead of a suffix glob. Failure is spelled a dozen ways across the several hundred event names (_failed, but also _invalid, _mismatch, _missing, _unauthorized, _rejected), so globbing on the name silently misses most of it:

outcomeMeaning
successThe operation completed.
failureThe operation did not. This is the field to alert on.
progressAn operation began or was asked for; its result is not yet known. Always paired with a _started or _requested name.
advisoryNothing failed, but the operator should keep seeing it — a configuration posture like tls_disabled or challenge_validation_bypassed.

Every error-level record is outcome = "failure"; an advisory sits at warn.

The events worth building alerts on:

EventLevelMeaning
request_completedinfo / warn / debugOne per request. See “The access line” above for the level.
server_listening, profile_mountedinfoStartup completed; one profile_mounted per enabled profile. server_listening is repeated by a reload that moved this socket, carrying the address it moved to. A reload emits profile_mounted only for an endpoint that was not previously served.
profile_unmountedwarnA reload stopped serving an endpoint. Its accounts and orders stay in the database; an issuance still waiting on an upstream has no handler left to finish it.
server_fatal_error, profile_init_failederrorThe process is not serving.
server_socket_bind_failed, admin_socket_bind_failed, metrics_socket_bind_failederrorThe ACME, admin or metrics socket could not be bound. At startup the process does not serve; on a reload nothing is applied and whatever was already listening keeps listening.
metrics_listeninginfoThe metrics listener started. Its message is the reminder that the endpoint is unauthenticated by design, so the port must be firewalled.
metrics_config_invaliderrormetrics.bind_address names the same socket as the ACME or admin listener. The process did not start.
server_listener_stoppedinfoA listener was switched off by a reload (admin.enabled or metrics.enabled). The socket is released; established connections finish.
server_config_reloadedinfoA SIGHUP applied. Carries generation, which counts from 1 and rises by one per reload — the quickest check that one landed — and listeners_rebound, naming any socket that moved.
server_config_reload_refusedwarnThe new file changes a key that cannot change while the process runs. Nothing was applied; the error field names the key.
server_config_reload_failederrorThe new file did not load, or what it asks for could not be built. Nothing was applied.
server_logging_filter_overriddenwarnA reload changed logging.filter while something outranked it, so the edit had no effect. source names which: "flag" for a --log-level on the server’s own command line, "env" for RUST_LOG. Both win on a reload exactly as they do at startup; drop the one named and reload again.
db_migration_failederrorStartup aborted before serving.
request_shedwarnA request was refused with 503 + Retry-After: 5 because server.max_concurrent_requests was saturated for longer than admission_wait_ms. Sustained occurrences mean the limit is too low, or something is retrying hot.
request_deadline_exceededwarnA request exceeded server.request_timeout_ms.
request_handler_panickederrorA route handler panicked. The client got a 500 in the listener’s normal error shape rather than a dropped connection, and the process is unaffected — but a handler panic is always a bug. Alert on it. The listener (acme / admin) and, for the panel, surface (api / ui) fields say where; the panic message is in error.
filter_request_blocked, filter_deniedwarnThe filter policy refused a request. Expected in normal operation; a spike is either an attack or a policy change that broke a legitimate client. listener = "admin" on filter_request_blocked marks an [admin.filter] refusal.
filter_rule_warnedwarnA mode = "warn" rule matched and did not decide. This is the line a dry-run rollout is watched on: when it stops appearing for legitimate clients, the rule is safe to switch to enforce.
challenge_validation_failed, challenge_failed, challenge_validation_timeoutwarnDomain-control validation did not pass. The most useful signal that clients are misconfigured — or that egress to them is blocked. The timeout variant means nothing answered within challenge.timeout_ms at all.
challenge_http_01_mismatchwarnThe responder answered, but with the wrong key authorization. The body itself is never logged here; a truncated preview goes to challenge_http_01_mismatch_body at debug.
challenge_validation_abandonedwarnThe queue gave up on a validation — the attempts ran out, or the authorization expired under it. The challenge, its authorization and its order are marked invalid so the client stops polling. Distinct from challenge_failed, which means the check ran and the client’s setup did not satisfy it; this one means the check never reached a verdict.
challenge_validation_enqueue_failederrorA challenge was claimed but its validation job could not be written. The claim is released, so the client may trigger again; the request answers 500. Alert on it — it means the database refused a write on the ACME path.
server_role_no_workerwarnThis process runs no worker role, so nothing in it drains the job queue and it holds no signing backend. Expected in a split deployment; a deployment where no process runs one issues nothing at all, because challenge validation, issuance, CRL signing, notifications and the sweeps are all queued work.
server_schema_behinderrorA process that does not own the schema found migrations unapplied and refused to serve. Run acme-proxy migrate, or start the worker role, before the others.
nonce_replayedwarnA JWS carried a nonce that was unknown, already consumed or expired. Routine in small numbers (a client racing itself); a flood is a client stuck in a retry loop, or a replay attempt.
key_change_rejectedwarnPOST /keyChange refused. reason = bad_signature means the inner JWS did not verify — somebody attempted a rollover they could not prove possession for.
order_finalize_queuedinfofinalize accepted a CSR, claimed the order and queued its signing. Logged by the process serving ACME; the outcome is logged by the worker that signs.
local_ca_leaf_issued, order_finalizedinfoA certificate was issued. Logged by the worker.
order_finalize_bad_csr, order_finalize_issuance_failedwarn / errorThe signer backend rejected the CSR (the order goes invalid with badCSR), or failed to sign (the job retries).
server_reload_supervisor_goneerrorA SIGHUP arrived but nothing is left to act on it — the process is shutting down, or the reload supervisor died. The configuration on disk is not in effect; restart to apply it.
order_finalize_authority_withdrawnwarnThe account, or one of the order’s authorizations, was deactivated after finalize queued the signing. The order goes invalid with unauthorized and nothing is signed.
order_finalize_abandonedwarnThe queue gave up on an issuance — the attempts ran out, or the order expired under it — and marked the order invalid so the client stops polling. Carries the last reason. A steady stream means the backend is down; alert on it.
certificate_revoked, certificate_revoke_signer_failedinfo / errorRevocation succeeded, or the signer refused it — in which case the order is left un-revoked for a retry.
local_ca_crl_republishedinfoA revocation acme-proxy order revoke recorded without the CA key was signed into the CRL by the local_ca_crl_regenerate job. Carries the CA’s issuer id.
certificate_revoke_queuedinfoA revocation for a relay or custom backend was queued for the worker by a process that holds no backend — revokeCert, the panel or the CLI — which then waits on the job. Carries order_id and job_id.
certificate_revoke_abandonederrorA queued revocation for a relay or custom backend was given up after its attempts ran out: the certificate is still trusted. Carries order_id and the last reason. Alert on it.
local_ca_crl_not_storederrorGET /crl found no stored CRL for a CA. The worker stores one before it serves, so this means no worker has run with this CA yet — start one.
local_ca_crl_prunedinfoRevocation entries whose certificates had expired were dropped from the CRL (RFC 5280 §3.3). Carries rows_removed and the CA’s issuer id. Silent when nothing expired, which is most days.
local_ca_crl_prune_failederrorThe daily CRL refresh failed for one CA — the database, or a CRL that could not be signed. Nothing is lost and the refresh tries again tomorrow, but the refresh is also what re-signs a CRL before its nextUpdate, so two days of this in a row deserve a look.
local_ca_ledger_importedinfoA CA met the database for the first time and imported the JSON ledger it kept beside crl_path before. Carries rows_imported and the new crl_number. Logged once per CA; the sidecar is not read again.
local_ca_crl_initialization_failederrorA CA could not store its first CRL — usually a sidecar that will not import, named in error. That CA’s revocations and GET /crl fail until it is fixed, rather than proceeding without the history the sidecar holds.
local_ca_crl_export_failederrorThe CRL was stored and is served, but writing its copy to crl_path failed. Only a static publication of that file is stale; the daily refresh writes it again.
http_01_token_store_failederrorThe relay could not publish, retract or look up an http-01 key authorization in the database. A failed publish retries the relay attempt; a failed lookup answers the upstream’s fetch 500.
upstream_relay_succeeded, upstream_relay_failedinfo / warnOutcome of one relayed issuance under the relay signer backend.
job_run_completedinfoOne background job finished. Carries job_kind, the attempt it succeeded on, and duration_ms.
job_run_retriedwarnA job failed in a way that may not recur and went back in the queue. Carries the reason and the next run_at. Routine in ones; a steady stream of the same job_kind means whatever it talks to is unwell.
job_run_abandonederrorA job was retired permanently — the handler refused it, the attempts ran out, or its deadline passed. For a signer_relay_issue job this is the moment the client’s order goes invalid, so it is the line to alert on.
job_run_panickederrorA job handler panicked. The job is retried straight away rather than holding its lease, but this is always a bug — alert on it.
db_job_leases_reclaimedwarnRows whose runner died holding the lease were returned to the queue. Expected once after an unclean shutdown; recurring means the runner is being killed mid-job.
job_lease_lostwarnA job finished after its lease had already been reclaimed, so its result was discarded and another runner will repeat the work. Means an attempt is overrunning jobs.lease_seconds.
job_deadline_passedwarnA job was claimed after its own deadline and retired without running. For a relay that means the local order had already expired.
job_runner_retunedinfoA configuration reload moved the runner’s pacing, and it is now running under the new values — which server_config_reloaded alone does not tell you. Carries all five: poll_interval_ms, lease_seconds, retry_base_seconds, retry_max_seconds and max_concurrent. Silent when a reload leaves [jobs] alone.
job_runner_started, job_runner_stoppedinfoThe queue runner’s lifecycle. job_runner_stopped carries how many leases it released on the way out; a missing one after a restart is why work waits out a lease instead of resuming immediately. Note the six table sweeps and every notification run through this runner, so a runner that is not started is a server that is not sweeping or notifying either.
upstream_bad_nonce_retrydebugNormal ACME churn against the upstream; only interesting in bulk.
notify_delivery_failedwarnOne delivery attempt did not land. Never affects the ACME response, and no longer the end of the story: it carries retryable, and a true there means the delivery went back in the queue.
notify_delivery_target_missingwarnA queued delivery names a profile or backend this process does not have. It is retried rather than dropped, since another process over the same database may be on a newer configuration; a target that really is gone ends in notify_delivery_abandoned once the attempts run out.
notify_delivery_abandonedwarnA notification was given up on — the attempts ran out, or a backend reported a failure that could never succeed. This is the line that means an operator was not told something. Alert on it; notify_delivery_failed on its own is usually just a bad minute.
notify_deliveredinfoOne delivery landed. Carries the backend, the event kind and the attempt it succeeded on.
notify_delivery_queuedinfoA delivery was written to the queue, one line per backend that wanted the event. Carries a delivery_id shared by that event’s rows, which is how they are correlated.
tls_handshake_timeout, tls_handshake_faileddebugOnly with server.tls.enabled. Deliberately below the default filter: on a public listener these are scanner background noise, and one warn per failed handshake is a flood, not a signal.
nonce_reaper_sweptdebugThe periodic nonce cleanup ran. It is a nonce_sweep job, so its absence over a long window means the job runner is unwell — check job_runner_started.
audit_write_failedwarnAn audit row could not be written. The failure is swallowed deliberately — a certificate the CA has already signed must not become a 500 the client retries into a second issuance — so this line is the only evidence that the trail has a hole in it. Alert on it.
audit_reverse_dns_failed, audit_reverse_dns_timeoutdebugA PTR lookup for a client address found nothing in time. Costs a NULL in one column, never a refused request. Routine where no reverse zone exists; turn audit.reverse_dns off there.
job_retention_sweep_failedwarnThe daily sweep of finished jobs past jobs.retention_days failed. Nothing is lost; the table grows until the next one succeeds.
notify_expiry_digest_failederrorThe [notify.expiry] digest for one profile could not be built — usually the database. Carries profile. The digest is rescheduled at its interval regardless, so one failure costs one digest, not the schedule.
audit_reaper_swept, audit_reaper_faileddebug / warnThe daily retention sweep, only with a non-zero audit.retention_days. audit_reaper_swept carries the rows removed and the cutoff. Runs as the audit_sweep job.
order_reaper_swept, order_reaper_failedinfo / errorThe daily order-retention sweep, one line per profile, carrying that profile’s rows removed and cutoff. Runs as the order_sweep job, and only for profiles with a non-zero order.retention_days. It deletes expired, non-valid orders and cascades to their authorizations and challenges; a valid order is never swept, so revocation and renewal information stay available for every certificate actually issued.
ipam_netbox_tls_verification_disabledwarnEmitted on every start while insecure_skip_verify is set, deliberately not once-only.
proxy_configuredinfoEmitted once at startup when [proxy] resolves to anything, and not at all otherwise. Carries source (config, environment or both), the two proxy URLs with any password redacted, and the no_proxy rule count. Worth reading on a first start: an inherited shell https_proxy is otherwise an invisible reason for every outbound call to behave differently.

Only with [admin] enabled — see Web Admin:

EventLevelMeaning
admin_listening, admin_origin_resolvedinfoThe web admin started. admin_origin_resolved carries the origin the CSRF check will compare against, and the resolved bind address — check these agree with how you actually reach the panel.
admin_config_invaliderror[admin] cannot work; the process did not start. The message names the two keys that disagree.
admin_tls_init_failed, admin_filter_init_failederror[admin.tls] or [admin.filter] could not be built — an unreadable certificate, a policy that does not parse. At startup the process does not serve; on a reload nothing is applied.
admin_no_userswarnThe panel is enabled but has no operators — a running service with no way in. Fix with acme-proxy admin user create.
admin_login_succeededinfoCarries username and client_ip.
admin_login_failedwarnCarries reason: wrong_password, unknown_user, account_disabled or rate_limited. The client is told none of this — every failure returns one invalid_credentials — so this log line is the only place the distinction exists. A run of unknown_user from one address is somebody guessing usernames; a run of rate_limited is a brute-force attempt, or an operator locked out by their own retries.
admin_logoutinfoCarries scope = "one" or "all", and surface = "api" or "ui".
admin_mfa_enrolledinfoAn operator confirmed a second factor; carries username.
admin_mfa_recovery_codes_regeneratedinfoA fresh set of recovery codes was minted, from the panel or admin user totp recovery-codes. Carries username and minted.
admin_mfa_step_up_refusedwarnA sensitive change asked for the operator’s password again and did not get it. reason is wrong_password or rate_limited; a run of them on a signed-in session is somebody holding a session who does not know its password.
admin_password_hash_unreadablewarnA stored hash could not be decoded. The account is unusable until acme-proxy admin user passwd rewrites it, and nothing else will tell you.
db_admin_user_created, db_admin_user_deleted, db_admin_user_password_changed, db_admin_user_status_changed, db_admin_user_role_changed, db_admin_user_contact_changedinfoThe operator audit trail. db_admin_user_contact_changed carries cleared (whether the notification address was removed rather than set).
db_admin_sessions_revoked, db_admin_session_deletedinfoSessions ended, by a password change, a disable, or an explicit revoke. db_admin_sessions_revoked carries scope: user, user_except_current (what a password change does, so the operator making it is not logged out by their own action) or all.
admin_eab_created, admin_eab_revoked, admin_eab_deletedinfoCarries the kid and the operator who did it — never the secret. admin_eab_deleted carries accounts: keep, deactivate or delete.
db_account_delete_blocked, db_order_delete_blocked, db_eab_delete_blockedinfoAn operator delete was refused because it would have removed the only record of a live certificate. Carries live_certificates. Nothing was changed.
admin_order_revoked, admin_order_revoke_queued, admin_order_deleted, admin_account_contact_updated, admin_account_deactivated, admin_account_deleted, admin_nonces_cleaned, admin_session_revokedinfoAdmin writes, each naming the operator. Each carries surface = "api" or "ui": the JSON API and the HTML panel run one shared action, so the same write logs the same line whichever surface it came through.
admin_operator_role_changed, admin_operator_contact_updatedinfoAn operator’s privilege tier (carries role) or notification address (carries contact_set) was changed in the panel. username is who made the change and target_username whose it was — the same name when an operator edits their own address. Carries surface.
admin_job_cancelled, admin_job_advanced, admin_job_revivedinfoAn operator acting on the background queue through /api/jobs or /ui/jobs. (The acme-proxy jobs subcommand writes the matching audit rows but emits no log line: a non-serve invocation installs no subscriber.) Each carries surface and the operator. admin_job_cancelled on a signer_relay_issue job also carries order_abandoned = true and the order_id — the ACME order was marked invalid and the upstream mapping abandoned. admin_job_revived carries attempts (set to max_attempts - 1).
admin_revoke_signer_failederrorThe CA-side revocation failed, so the order is left un-revoked for a retry. Answered as 502 rather than 500.
admin_db_errorerrorA database failure on an admin route. The sqlx message is here and deliberately not in the response body, which says only “internal error”.
admin_session_reaper_swept, admin_session_reaper_faileddebug / errorThe periodic session sweep ran, or failed, as the admin_session_sweep job. A failure leaves expired sessions in the table, but they are refused on use regardless.
admin_session_orphanedwarnA session outlived its user despite the FK cascade. Should be impossible; the session is deleted and refused.

The full set is much broader than this table — several hundred names, of which the ones above are the curated subset worth an alert. Because the subsystem prefix is drawn from a closed list, an ad-hoc investigation can grep a whole family: event = "certificate_revoke for every revocation refusal, event = "replaces_ for RFC 9773 correspondence, event = "db_ for the storage layer.

Suggested alerts

Two sources, and they are complementary rather than alternatives. The metrics are cheap to alert on and answer “how much”; the log stream answers “which one, and why”, and covers everything the metrics do not.

With [metrics] enabled:

# Issuance is failing, whatever the reason.
rate(acme_proxy_certificate_issue_failures_total[15m]) > 0

# Nothing has been issued over a window longer than your shortest renewal
# interval -- silent breakage, and the one nothing else will tell you.
increase(acme_proxy_certificates_issued_total[6h]) == 0

# The server is shedding load: undersized, or a client is in a retry loop.
rate(acme_proxy_requests_total{status="503"}[5m]) > 0

# 5xx as a share of everything: this server failing, not clients being refused.
sum(rate(acme_proxy_requests_total{status=~"5.."}[5m]))
  / sum(rate(acme_proxy_requests_total[5m])) > 0.01

# The pool is saturated, so requests are queueing on a connection.
acme_proxy_database_pool_connections{state="idle"} == 0

# Requests are slow: the 95th percentile over a second, on any route.
histogram_quantile(0.95,
  sum(rate(acme_proxy_request_duration_seconds_bucket[5m])) by (le, route)) > 1

The up metric Prometheus synthesises per target also gives a liveness alert for free, though /health remains the better probe for a load balancer since it is on the ACME listener itself.

From the log stream, which reaches further. The first is the general case and subsumes most of the rest; the others are worth splitting out because each has a different response:

  • outcome = "failure" above its usual rate → something broke. This is the one query that catches every refusal, whatever the event is called, and it is the reason the field exists.
  • request_shed above zero for more than a few minutes → the server is undersized, or a client is in a retry loop.
  • challenge_validation_failed rising against a flat local_ca_leaf_issued rate → clients are asking and failing; check egress and the responder ports.
  • Any upstream_relay_failed → issuance is broken for every client behind the relay, not just one.
  • No local_ca_leaf_issued at all over a window longer than your shortest renewal interval → silent breakage.
  • Any audit_write_failed → the audit trail is silently incomplete, and nothing else will tell you. See Audit Trail.
  • server_fatal_error, db_migration_failed, profile_init_failed → page immediately.

Grafana Dashboard

A dashboard over the six metric families the [metrics] listener exposes ships in the repository at dashboards/acme-proxy.json. It is a starting point rather than a fixed artifact — import it, then change whatever your deployment needs.

Enable the metrics listener first; the dashboard has nothing to draw otherwise. See Monitoring for the endpoint and a scrape configuration.

Importing it

Download it from the repository:

curl -O https://raw.githubusercontent.com/acme-proxy/acme-proxy/main/dashboards/acme-proxy.json

Then Dashboards → New → Import in Grafana, upload the file, and pick your Prometheus data source when asked.

For a provisioned Grafana, drop it in the dashboards directory your provider config points at:

# /etc/grafana/provisioning/dashboards/acme-proxy.yaml
apiVersion: 1
providers:
  - name: acme-proxy
    type: file
    options:
      path: /var/lib/grafana/dashboards

The data source is a template variable, deliberately, rather than the __inputs block Grafana’s “export for sharing externally” produces: that block is never substituted under file provisioning, so a dashboard carrying one works when a human imports it and silently breaks when a machine does.

What is on it

Fourteen panels in three rows.

Issuance — certificates signed and refused over the dashboard’s time range, the success ratio between them, the issuance rate split by endpoint, refusals broken down by ACME problem type, and the p50 and p95 issuance latency by endpoint. That last panel is the one worth knowing about: reason says why the CA refused, in the same vocabulary acme-proxy audit list prints, because both are rendered from one record.

Requests — request rate by route and by status, the 5xx share, requests shed by admission control, unmatched paths, and the p50 and p95 latency by route.

Database — the database pool by state, and its saturation.

The latency panels are histogram_quantile over the buckets, so a percentile is only as fine as the buckets around it; see Monitoring for the bucket bounds and exactly what each histogram times. Issuance latency is drawn from the process that signs, which in a split deployment is the worker.

Two things to know before editing it

route is safe to group by. It is the matched route pattern, so /order/{id} is one series however many orders exist. Grouping by a raw URI would be a series per order, per account and per challenge — which is memory in Prometheus for as long as it retains them. The same reason every unmatched request collapses into a single <unmatched> series rather than one per path a scanner tries.

The pool panels must not be filtered by $profile. Everything else on the dashboard is scoped to the profile variable, so reaching for it on the two database panels is the natural next edit — but acme_proxy_database_pool_connections is process-wide and carries no profile label at all. Filtering it matches nothing, and the panel goes blank with no error to explain why. A test in the repository (tests/grafana_dashboard.rs) fails if the shipped dashboard ever grows that filter.

Empty panels are not always a fault

Four cases read as “no data” and are all correct:

  • Refusals by reason is empty on a CA that has refused nothing. Only refusals the CA itself made are counted — an order rejected at newOrder, for a wildcard on an endpoint offering no dns-01 for instance, is protocol bookkeeping rather than a CA action, and appears as a 403 on the request panels instead.
  • Success ratio and 5xx share show NaN when nothing has happened at all, because the ratio is genuinely undefined. They read 0 only when something did happen and all of it failed.
  • Issuance latency is empty on an acme-only or admin-only scrape target: only the worker signs, so only it observes the histogram.
  • Everything is empty for the first scrape or two after a restart. Counters do not survive a restart, though they do survive a SIGHUP reload.

Keeping it honest

tests/grafana_dashboard.rs runs in CI and asserts that every metric the dashboard queries is one this build actually emits — and the converse, that no emitted family is missing from it. A histogram counts as shown when any of its _bucket, _sum or _count series is. The names come from rendering a real, empty registry rather than from a list somebody maintains by hand, so a renamed metric fails the build instead of quietly blanking a panel that nobody looks at until an incident.

Troubleshooting

This section covers common issues when operating acme-proxy in production.

The server refuses to start

Several configuration mistakes are deliberately fatal at startup rather than silently degrading at runtime. The log line names the problem in each case.

MessageCauseFix
no enabled [profiles]No profile is defined, or all are enabled = false. The server serves ACME only through profiles.Add [profiles.default], or set ACME_PROXY_PROFILES__DEFAULT__ENABLED=true.
Unknown or empty challenge.enabledThe list is empty or names a type that does not exist. This is fatal even with challenge.bypass = true — a bypassing server still has to advertise a challenge type.Use one or more of http-01, dns-01, tls-alpn-01.
filter.check.<name> has neither allow nor deny entriesAn allowed_ip or path check with both lists empty permits everything, so it is a no-op and almost certainly a mistake.Populate filter.check.<name>.allow or .deny, or delete the check and any rule naming it.
request_timeout_ms too smallIt must exceed signer.custom.timeout_ms when that backend is configured with its crl or renewal_info hook on — those hooks run inline inside a request, so a smaller whole-request deadline would cut them off every time. challenge.timeout_ms and the script’s issue hook are not checked against it: both run in the job queue.Raise server.request_timeout_ms.
Two profiles sharing ca.key but differing elsewhereOne CA key under two [signer] configurations would issue under two policies from one identity.Make the [signer] sections identical (they then share one backend), or give each profile its own CA files.
An entry name with an underscore[filter.check.*], [filter.rule.*], [notify.custom.*], [notify.webhook.*] and profile names must match ^[a-z0-9-]+$. The name is also an environment-variable segment, which the config crate lowercases.Use hyphens: threat-intel, not threat_intel.
A custom notify backend enabled with no entries"custom" is listed in notify.enabled while notify.custom_enabled is empty.Populate notify.custom_enabled, or drop "custom".
filter.enabled is no longer a settingA 0.1.x [filter] section. The flat all-must-pass chain was replaced wholesale by named checks and rules; filter.exempt_paths and filter.custom_enabled are refused the same way.Declare each filter as a [filter.check.<name>] with a type, write a [filter.rule.<name>] whose when names them, and list the rules in filter.rules. See Filters.
[filter.rule.<name>] is configured but filter.rules is emptyA rule is declared but nothing selects it, so the policy it describes would never run. Inherited rules count: a profile that leaves filter.rules empty meets this too.List the rule in filter.rules, or delete it.
Upstream requires EAB, no .kid sidecarThe relay backend has never registered with the upstream.Run acme-proxy upstream register --profile <name> --eab-kid …, or set signer.relay.eab.kid/hmac_key in configuration.
key_source = "pkcs11" … built withoutsigner.local_ca.key_source = "pkcs11" on a binary with no PKCS#11 support. Deliberately fatal rather than falling back to the file key, which would silently leave the CA key on disk.Rebuild with cargo build --release --features hsm.
A PKCS#11 key that is not the certified oneThe token key’s public key does not match cert_path — almost always a wrong key_label.See Hardware Keys.
admin.bind_address … is not loopback while admin.tls.enabled is falseThe session cookie is always sent Secure, which a browser will not store over plain HTTP on anything but localhost. Refused rather than warned about, because the symptom is otherwise “signing in succeeds and then immediately signs you out” with nothing in any log to explain it.Set admin.tls.enabled = true, or bind 127.0.0.1 and reach the panel through an SSH tunnel.
metrics.bind_address and … are bothThe metrics endpoint is a listener of its own and cannot share a socket with the ACME or admin one. Checked at startup and on every reload.Give it its own port.
notify.webhook.<entry>.body is not a valid template (likewise .url, .method, .headers)Every webhook entry is compiled and validated at startup: the body template, the URL, a verb outside POST/PUT/PATCH, and a header name or value a wire format will not carry. Deliberately fatal — a broken entry would otherwise become a permanent delivery failure discovered on the first event that mattered.Fix the entry. See Webhook.
admin.template_dir … is not a directory, or a template that does not compileThe panel’s per-file overrides are compiled at startup for the same reason, so a broken one refuses to start rather than serving a 500 later.Point it at a directory, or leave it empty for the compiled-in defaults. See Customizing the Panel.
A [proxy] URL that is not http://, or a no_proxy entry with a portproxy.http_url / proxy.https_url name the proxy to dial, which is spelled http://host:port even for HTTPS targets; there is no SOCKS support. A no_proxy entry matches a host or a network, never a port.Drop the scheme to http://, and the port from the no_proxy entry.
ipam.phpipam.sources naming vip or fhrpphpIPAM records neither roles nor redundancy groups, so those two sources exist only for NetBox. Refused by name rather than silently returning nothing.Use dns_name, custom_field or device. See phpIPAM.

Failures specific to a hardware CA key — PIN, token, slot and mechanism problems — have their own table in Hardware Keys (PKCS#11).

A wildcard order is rejected

  • Symptoms — newOrder for *.example.com returns rejectedIdentifier naming dns-01.
  • Cause — Wildcards can only be proven with a DNS challenge, so acme-proxy accepts them only when dns-01 is among challenge.enabled.
  • Fix — Add dns-01 to challenge.enabled. See Challenge Validation.

A client suddenly cannot order anything

  • Symptoms — Every order-side request from one client returns unauthorized.
  • Cause — The account has been deactivated — by the client itself, or by acme-proxy account deactivate. Deactivation is permanent and blocks all issuance.
  • Fix — The client must register a new account.

An order sits in processing and never finishes

  • Symptoms — A client polls an order that reached processing and stays there. Only the relay backend defers like this; local_ca and custom answer finalize inline.
  • Cause — The issuance is a background job, and it is either waiting out a backoff after a retryable upstream failure or has been retired for good. The order object itself will not say which.
  • Fix — Ask the queue:
# What is the job doing, and how many attempts has it spent?
acme-proxy jobs list --kind signer_relay_issue --limit 20

# The upstream's own error text, and the URLs it was talking to.
acme-proxy jobs show <job-id>

A job still ready is waiting for its next attempt at the run_at it prints — acme-proxy jobs run-now <id> pulls that forward. A failed one has spent its budget: run-now grants exactly one more attempt, and acme-proxy jobs cancel <id> abandons the ACME order so the client stops polling and can order again. See Job queue, and job_run_abandoned in Monitoring for the log line that says it happened.

SQLite database locks

  • Symptoms — The server logs show database is locked errors during high concurrency order creation.
  • Cause — The database may not be utilizing Write-Ahead Logging (WAL) or your filesystem does not support proper locking mechanisms (e.g., NFS).
  • Fix — Ensure acme-proxy is running on a local filesystem and journal_mode = WAL is applied. (The server automatically attempts to enable WAL on startup).

Reading the database directly

When the CLI cannot answer a question — usually “what does the row actually say?” — the database is readable while the server runs, because WAL allows a reader alongside the writer.

# Which migrations have run, and did they all succeed?
sqlite3 sqlite.db "SELECT version, description, success FROM _sqlx_migrations;"

# Everything this server refused in the last day, and why.
sqlite3 sqlite.db "SELECT datetime(created_at,'unixepoch'), event, profile, client_ip, reason
                     FROM audit_log WHERE outcome = 'failure'
                      AND created_at > strftime('%s','now','-1 day');"

# An order that will not progress: its status and its authorizations'.
sqlite3 sqlite.db "SELECT o.status, a.identifier, a.status, c.type, c.status, c.error
                     FROM orders o JOIN authorizations a ON a.order_id = o.id
                     JOIN challenges c ON c.authz_id = a.id WHERE o.id = '…';"

Read, do not write. The CHECK constraints will catch an impossible status, but nothing re-syncs the in-memory state a running handler is holding. Use the Admin CLI to change anything.

Backups must include the WAL. Copying sqlite.db on its own gives you a database missing every recent write. Either take all three files (sqlite.db, -wal, -shm) with the server stopped, or run sqlite3 sqlite.db ".backup backup.db", which is consistent by construction and safe against a running server.

The schema itself — every table, constraint and index, and why each is shaped the way it is — is documented in Database Schema.

Upstream Let’s Encrypt rate limits

  • Symptoms — Order finalization fails with HTTP 429 Too Many Requests from the upstream CA.
  • Cause — When using the relay signer backend, all internal clients share a single external ACME account. Let’s Encrypt applies rate limits (e.g., 50 certificates per registered domain per week).
  • Fix — Request a rate limit increase for the root domain you are relaying, or implement careful caching mechanisms on your internal servers to prevent excessive renewals.

Network challenge blockages

  • Symptoms — Order stays in pending state, or challenge verification fails.
  • Cause — If challenge.bypass = false, the server must reach the internal client over HTTP (port 80) or DNS to verify domain ownership. Firewalls might be blocking this internal callback.
  • Fix — Ensure the host running acme-proxy has egress network access to reach the internal servers requesting certificates.

EAB registration fails

  • Symptoms — Client receives an error stating External Account Binding is required.
  • Cause — The client is connecting to a profile that requires EAB, but did not provide the kid and hmac credentials.
  • Fix — Create EAB credentials using the acme-proxy eab create CLI command and configure the client to use them.

Order finalization fails (413 payload too large)

  • Symptoms — Finalizing the order returns a 413 error.
  • Cause — The encoded Certificate Signing Request (CSR) exceeds the server.max_body_bytes limit (default 128 KiB).
  • Fix — Ensure your client is not generating excessively large CSRs or increase the limit in your configuration.

Order finalization fails (badCSR: identifier mismatch)

  • Symptoms — Finalizing the order returns 400 badCSR complaining about the requested identifiers.
  • Cause — The DNS Subject Alternative Names in your CSR are not exactly the set of names the order authorized. The comparison is set equality, so an extra name fails just as an omitted one does.
  • Fix — Configure your ACME client to put every ordered domain — and nothing else — into the CSR’s SANs.

Order finalization fails (badCSR: common name)

  • Symptoms — 400 badCSR, “CSR common name is a domain the order does not cover”.
  • Cause — The CSR’s Subject Common Name looks like a DNS name that the order does not authorize. acme-proxy does not ignore the CN: a domain-shaped CN must be covered by the order, precisely so a name cannot be smuggled past the identifier filters by moving it out of the SANs.
  • Fix — Either add that name to the order, or drop the CN. A CN that is not domain-shaped — a human label like rcgen self signed cert — is tolerated and ignored. Note that the local_ca signer empties the subject entirely on the issued certificate, so a CN is never carried through to the leaf.

Order finalization fails (badCSR: non-DNS SAN)

  • Symptoms — 400 badCSR from the local_ca signer for a CSR containing an IP, email or URI SAN.
  • Cause — local_ca accepts DNS SANs only, and rejects anything else outright rather than stripping it.
  • Fix — Remove the non-DNS SANs from the CSR.

Connection refused or 503 service unavailable

  • Symptoms — High-throughput ACME clients receive HTTP 503 errors during bursts.
  • Cause — The server’s load shedder activated because the number of concurrent in-flight requests exceeded server.max_concurrent_requests.
  • Fix — Increase max_concurrent_requests and admission_wait_ms, or configure your client to retry with exponential backoff.

Security Model

acme-proxy is a certificate authority, or the thing standing in front of one. Anything that can make it sign gets a certificate your infrastructure will trust. This page states the trust boundaries and what each secret protects, and links to the page that owns each detail.

It is a map, not a manual. Every claim here is explained in full somewhere else; if you need to act rather than orient, go to the hardening checklist.

What has to be true for a certificate to be issued

Four gates stand between a packet arriving and a certificate coming back. They are independent, and each covers something the others do not.

GateAnswersDefaultOwned by
Connection filtersMay this address talk to this endpoint at all?off (filter.rules = [])Filters
EABIs this client allowed to register an account here?offEAB
Challenge validationDoes the client actually control the name?onChallenge Validation
Identifier filtersIs this client allowed these particular names?offFilters

Two of the four are off by default, which makes the third load-bearing:

With challenge.bypass = true and an empty filter.rules, every client that can reach the socket can obtain a certificate for every name it asks for. That combination is why validation is on by default.

The identifier gate runs twice — at newOrder and again at finalize against the names actually in the CSR. That is not belt-and-braces: without the second run, a name could be smuggled past the policy by keeping it out of the order and putting it in the CSR. See Filters.

What each secret protects

SecretCompromise gets an attackerStored
The CA private key (signer.local_ca.key_path)The ability to mint any certificate your fleet trusts, silently, for as long as the CA is trusted. There is no audit row for a signature made outside this server.On disk at 0600, created with create_new rather than chmod’ed after the fact. Can live in a PKCS#11 token instead.
An EAB HMAC secretThe ability to register accounts at that profile, subject to every other gate. Revocable without a restart.Retrievable bytes — HMAC verification needs the same secret back each time. See Database Schema.
The upstream ACME account key (signer.relay.account_key_path)Control of your account at the upstream CA, including revoking what it issued.On disk at 0600, beside a .kid sidecar naming the account.
An RFC 2136 TSIG keyThe ability to write records in the zone it is scoped to. This is the credential the relay exists to not distribute.Configuration, or ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_KEY_SECRET.
A web admin passwordNothing on its own once a second factor is enrolled. Otherwise: an operator session.One-way KDF, unreadable.
A web admin session cookieAn operator session until it expires — but not the ability to change the second factor, which takes the password again.Only hex(SHA-256(token)) is stored.
A NetBox API tokenRead access to your IPAM.Configuration; belongs in the environment variable.

The CA key is the one whose loss is not recoverable by rotation: every certificate it signed stays trusted until the CA itself is distrusted everywhere. That is the argument for hardware-backed keys, and for keeping an offline root and giving acme-proxy only an intermediate — see Local CA.

How often to roll each of these, and how, is Secret Rotation.

Two listeners, two exposure surfaces

The ACME listener and the web admin listener are separate sockets with separate TLS configuration, separate authentication and separate defaults. They are not two paths on one server, and they should not sit on one interface.

  • The ACME listener answers unauthenticated clients by design. It carries the filter chain, the admission limiter and the nonce middleware.
  • The admin listener is off by default, binds 127.0.0.1 by default, and carries no filter chain and no admission control — its availability concern is credential brute force, which the login limiter handles. Access control is the bind address, TLS, and the session.

Startup refuses a non-loopback admin.bind_address while admin.tls.enabled is false. The session cookie is always Secure, which browsers silently decline over plain HTTP off localhost; the symptom would be “login succeeds, then immediately logs out” with nothing in any log.

See Deployment for where each socket belongs relative to a firewall.

Trusting a forwarded address

Every IP-based decision — filters, the login limiter, what lands in the audit trail — rests on which address the server believes the client has.

Behind a reverse proxy that address arrives in a header, and a header is written by whoever is talking to you. filter.trusted_proxies is the allowlist of hops whose forwarded-for header is believed; with it empty, the peer address is used and the header is ignored. Setting the header name without setting trusted_proxies does not make the header trusted.

The admin listener does no forwarded-header handling at all, deliberately: trusting one without an allowlist would let any caller choose its own rate-limiter key.

See Allowed IP.

The audit trail is the record, and it is not a control

Every issuance and every refusal is written to audit_log with the actor, the address, that address’s reverse name, the identifiers, the User-Agent and the request id. Nothing in the server ever compares a live request against any of it — address pinning breaks CGNAT and mobile clients, and that is a deliberate non-feature.

Two properties make it evidence rather than logging:

  • It survives deletion of its subject. audit_log has no foreign keys, so deleting an account or an order does not take its history with it. See Database Schema.
  • Nothing in the web admin can erase it. The panel’s audit surface is read-only; pruning is audit cleanup on the host, or the audit.retention_days sweep. A stolen session that could erase the trail would make the trail prove nothing.

See Audit Trail.

Where this server can be made to talk to something else

Three subsystems make outbound connections on behalf of a client’s request, which makes each one a request-forgery surface worth knowing about:

  • http-01 validation follows redirects, because RFC 8555 requires it. Boulder’s mitigation — blocking RFC 1918 targets — does not apply here, since serving private networks is the entire point. What contains it instead: only http/https, only the two configured ports, at most max_redirects hops, a shared timeout, an off switch, and the fetched body is never echoed into the client-visible error. See HTTP-01.
  • The custom hooks (signer, filter, notify) execute an operator-supplied script with request-derived data in its environment and on stdin. They run with env_clear(), a minimal PATH, a timeout and kill_on_drop — but the script is yours, and quoting its inputs is your job.
  • The relay backend talks to an upstream ACME server and, with dns01, writes DNS records. Unlike http-01 validation, it validates the upstream’s TLS certificate against webpki-roots: there, the certificate is the only thing identifying the CA being handed your CSRs.

What is out of scope

  • Availability of the ACME listener beyond the admission limiter. There is no per-account quota and no rate limiting by identifier.
  • Confidentiality of issued certificates. They are public objects; the audit trail records who received one.
  • CAA. This server performs no CAA lookup — see Protocol Support.
  • Protecting the database file from a local root. The SQLite file holds EAB secrets and TOTP secrets in a form the server can read back, which means so can anyone who can read the file. File modes are the boundary.

Hardening Checklist

Run through this before an acme-proxy deployment issues a certificate anything depends on. Every item links to the page that explains it; nothing here is explained only here.

The defaults are already the safe end of most of these. The items that need a decision from you are marked decide.

Before it serves anything

  • challenge.bypass is false. It is the default. With it on, [filter] is the only thing between a client and a certificate for any name it names. → Challenge Validation
  • At least one gate is configured — filters, EAB, or both. Validation alone proves the client controls the name; it does not say the client is allowed to have a certificate from you. → Filters, EAB
  • decide — server.bind_address is the interface you meant. The default [::]:3000 is every interface. → Configuration Reference
  • server.base_url matches how clients actually reach the server, including the scheme. It is checked against the JWS url of every signed request, so a mismatch fails every request rather than degrading. → Configuration Reference
  • ACME is served over HTTPS — either server.tls.enabled = true or a reverse proxy in front. RFC 8555 §6.1 expects it. → TLS Termination
  • acme-proxy filter explain agrees with what you meant, for both a client that should be served and one that should not. A policy is easier to get subtly wrong than a list. → CLI
  • /crl and /ca.pem are still reachable if any check is address-based. Both are served by the profile router, so an allowlist covers them too — and neither the relying parties that fetch the CRL nor the hosts that have yet to install the root are the ACME clients you allowlisted. → Path Check

An or is a hole you opened deliberately

A check that cannot reach its authority answers “unknown” rather than “no”, and pass or unknown is pass. That is the point — it is what keeps an inventory outage from locking every client out — but it means an or weakens the fail-closed property to whatever its other side says.

when = "mgmt-net or inventory"

reads as “the inventory decides, unless the address is already trusted”. If mgmt-net is wide, the inventory is decorative for everything inside it. That may be exactly what you want; what you must not do is write it believing both checks apply.

The rule of thumb: an or over an address check is a bypass for that address range, so keep the range as small as the outage you are insuring against. and has no such property — fail and unknown is fail, so a conjunction never becomes more permissive because something broke.

The CA key

  • decide — the issuing key is an intermediate, not a root. An offline root means a compromise is recoverable by re-issuing the intermediate rather than re-trusting every endpoint. → Local CA
  • decide — the key lives in a PKCS#11 token if the deployment justifies it. The key then cannot be copied, only used. → Hardware Keys
  • ca.key is 0600 and owned by the service user. acme-proxy creates it that way; a key restored from a backup may not be.
  • The CRL is reachable by everything that validates your certificates, and the database is backed up — the revocations recorded there are the authoritative record, not the CRL file. → Revocation & CRL

Behind a proxy

  • filter.trusted_proxies names the hops you trust, or the forwarded header is ignored entirely. Setting filter.forwarded_header alone does not make it trusted. → Allowed IP
  • The proxy does not forward /health if you do not want it public — it is mounted outside the filter chain on purpose. → Monitoring
  • With signer.relay.challenge_strategy = "http01", port 80 of every name being issued forwards or redirects /.well-known/acme-challenge/ here. Nothing in the process can do this for you. → Relay

The web admin

Skip this section entirely if admin.enabled is false, which is the default.

  • admin.bind_address is loopback, or admin.tls.enabled is true — startup refuses the other combination. → Web Admin
  • decide — reach it over an SSH tunnel or a VPN rather than exposing the socket. It has no filter chain and no admission control. → Deployment
  • admin.base_url is the origin operators actually type. It is load-bearing four ways — the CSRF origin check, the generated certificate’s host, and the label an authenticator app shows.
  • Every operator has a second factor, and admin.require_mfa = true so the next one does too. It does not retroactively end sessions that predate it; admin session revoke --all does. → Users & Sessions
  • Recovery codes are stored somewhere that is not the panel. Without them, a lost authenticator needs admin user totp reset on the host. → Users & Sessions
  • No --password anywhere in your provisioning. There is no such flag; the password arrives on stdin or via --password-file. → Users & Sessions

Secrets

  • Tokens and TSIG keys come from the environment, not the file where the option exists — ipam.netbox.token, signer.relay.dns01.rfc2136.tsig_key_secret.
  • [signer.relay.eab] is emptied after the first registration. It is a bootstrap credential that authorizes exactly one newAccount; the server warns on every startup for as long as it stays set. → Relay
  • insecure_skip_verify is unset. It warns on every startup by design, so it stays visible for exactly as long as it is needed. → NetBox
  • The database file is 0600. It holds EAB secrets and TOTP secrets in a form the server reads back, so anyone who can read the file can too. → Database Schema

Ongoing

  • decide — a rotation interval for each long-lived secret. Nothing in the server expires these on a timer; the page recommends one per secret and names the events that force a rotation early. → Secret Rotation
  • Backups copy the WAL. sqlite.db alone is missing every recent write; use .backup, or take all three files. → Database Schema
  • decide — audit.retention_days. 0, the default, keeps everything for ever, which is the right default for a trail whose value is that it is complete. → Audit Trail
  • Something watches the logs for refusals — certificate_issue_failed and certificate_revoke_failed rows, and a run of unknown-certificate revocation attempts, which is somebody enumerating serials. → Monitoring
  • Startup warnings are read, not filtered out. Three of them repeat on every start precisely so they cannot become background noise: challenge_validation_bypassed, ipam_netbox_tls_verification_disabled, signer_relay_eab_secret_in_config. → Monitoring

Secret Rotation

The Hardening Checklist is run once, before a deployment issues a certificate anything depends on. This page is the other half: what to do on a schedule after that. Every secret in the Security Model’s inventory is here, with a recommended interval and the events that force a rotation early.

The intervals are a starting point, not a protocol requirement — adjust them to what the deployment is worth to an attacker. Nothing in acme-proxy expires these on a timer: rotation is an operator action, except for the session cookie, which ages out on its own.

The schedule

SecretRotate everyRotate now whenHow
The CA issuing keynot on a timer — see belowthe root or intermediate key may have been exposed; someone who could read it leavesre-issue the intermediate from the offline root — Local CA
An EAB HMAC secret12 months, per credentialthe client system is rebuilt; the person who held it leavesacme-proxy eab revoke <kid>, then eab create for the replacement — CLI
The upstream ACME account keynot on a timerthe disk may have been read; the relay profile is being decommissionedregister a fresh upstream account — new account_key_path, delete the .kid sidecar, restart — CLI
An RFC 2136 TSIG key12 months, with whoever runs the zoneanyone with the zone’s write path leaves; an update you did not make appears in the nameserver logadd the new key on the nameserver, update signer.relay.dns01.rfc2136.tsig_key_secret in the environment, SIGHUP — Relay
A web admin passwordnot on a timer, by designit was shared, phished, or typed into the wrong window; the operator leavesacme-proxy admin user passwd <user> — it also ends every session that user holds — Users & Sessions
A web admin session cookierotates itself — admin.session_ttl_seconds (12 h absolute), and the idle timeouta laptop is lost; a session is suspected stolenacme-proxy admin session revoke --user <u>, or --all — Users & Sessions
A TOTP secretnot on a timerthe authenticator device is lost or replacedacme-proxy admin user totp reset <user> — it takes the recovery codes and the sessions with it — Users & Sessions
Recovery codesregenerate when few remainone has been used in anger; the printed or stored copy is exposedacme-proxy admin user totp recovery-codes <user> — Users & Sessions
An IPAM API token12 months, or whatever the IPAM’s own policy saysthe token appears in a log or a ticket; an operator with IPAM access leavesissue a new read-only token, update ipam.netbox.token (or ipam.phpipam.token) in the environment, SIGHUP — NetBox

Only the session cookie expires on its own — an absolute lifetime and an idle timeout, whichever comes first. The password deliberately has no forced periodic change: an operator’s stored hash is re-encoded on their next login when the KDF cost rises, so there is nothing a scheduled reset would achieve. Everything else is a manual cadence because nothing revokes it for you.

The CA key is the exception

Losing the CA key is not something rotation recovers from. Every certificate it signed stays trusted until the CA itself is distrusted everywhere, and there is no audit row for a signature made outside this server. So the practice for this one key is structural rather than scheduled.

  • Give acme-proxy an intermediate, not a root, and keep the root offline. A compromise of the online key is then recoverable by re-issuing the intermediate rather than by re-trusting every endpoint in the fleet. See Local CA and its security constraints.
  • Re-issue the intermediate from the root before it expires. Doing it on a calendar, well ahead of the notAfter, keeps the recovery path exercised rather than theoretical.
  • Rotating the root is a multi-year event. Distribute the replacement long before it is needed, run both roots in the trust store, and remove the old one only once nothing is signed by it. See Trusting the CA.
  • An actual key compromise is a distrust-and-reissue event, not a rotation. Publish the revocation, pull the root, and re-issue what mattered under the new one. See Trusting the CA.

Rotating a secret that lives in the environment

The TSIG key and the IPAM token belong in environment variables rather than in config.toml, and so does the relay’s bootstrap EAB secret until the first registration clears it — the Hardening Checklist says which. Rotating one of these is the same three steps every time.

  • Stage the new value on the system that backs it — a second TSIG key on the nameserver, a fresh token in NetBox or phpIPAM — so that both the old and the new one work for a moment.
  • Update the environment and send SIGHUP (or systemctl reload). The [signer] and [ipam] sections both reload with no restart and no dropped request. See Reloading the configuration.
  • Confirm from the logs, or from a test issuance, that the new credential is the one in use, then retire the old value on the far side.

The EAB HMAC secrets and the TOTP secrets are held in the database in a form the server reads back, so a rotation does not remove the need to protect the file itself — file mode is the boundary. See Database Schema.

ASVS 5.0 Assessment

A self-assessment of acme-proxy against OWASP ASVS 5.0, whose text is vendored under rfc/asvs-5.0/ in this repository.

  • Assessed at: Level 2. Every L1 and L2 requirement in scope is given a status; L3 requirements are listed too, as information rather than as a bar being claimed.
  • Assessed against: the tree at release 0.6.0.
  • Method: source review. The evidence column names a file, not a promise — for a control requirement, the documentation on this site is context and the code is the evidence.

This is a self-assessment and not a certification. ASVS is explicit that a verification claim means an assessor performed the work; nobody outside the project has performed it here. Read this page as the maintainers’ own answer to “which recognised controls does this meet, and where does it fall short”, which is a useful thing to have written down and a different thing from an audit.

What was assessed

Three surfaces, kept apart because their controls genuinely differ:

SurfaceWhat it isWhere
ACME listenerUnauthenticated by design, authenticated per request by JWS. Carries the filter chain, admission control and the nonce middleware.crates/server/src/, crates/protocol/src/handlers/, crates/protocol/src/extractors/, crates/protocol/src/middlewares/
Web adminThe only session-based, browser-facing surface. Off by default, loopback by default.crates/admin/src/webadmin/, crates/admin/src/admin/
CLI and processAnswers to a shell on the host and holds no session.src/cli/, src/main.rs, crates/core/src/config/

Most of V3, V6 and V7 apply only to the web admin. When it is disabled — which is the default — those chapters have no surface to apply to at all.

Chapter disposition

ChapterIn scopeNote
V1 Encoding and Sanitizationyes
V2 Validation and Business Logicyes
V3 Web Frontend SecurityyesWeb admin only
V4 API and Web ServiceyesGraphQL and WebSocket sections are n/a
V5 File HandlingpartlyThere is no upload feature; the sections that assume one are n/a
V6 AuthenticationyesWeb admin and the CLI credential lifecycle
V7 Session ManagementyesWeb admin only
V8 Authorizationyes
V9 Self-contained TokensyesThe self-contained token here is the ACME JWS, not a session JWT
V10 OAuth and OIDCnoNo OAuth, no OIDC, no external identity provider, no token endpoint. Nothing in the chapter has a subject
V11 Cryptographyyes
V12 Secure Communicationyes
V13 Configurationyes
V14 Data Protectionyes
V15 Secure Coding and Architectureyes
V16 Security Logging and Error Handlingyes
V17 WebRTCnoNo WebRTC, no media, no signalling

Summary

Counts are of the requirements enumerated in the per-chapter tables below; every requirement of every in-scope chapter is present, so the totals match ASVS 5.0’s own counts. L1+L2 is the bar being assessed; the L3 column is reported for information.

ChapterL1+L2 metpartialgapn/aL3 (met / short / n/a)
V1 Encoding and Sanitization171092 / 0 / 1
V2 Validation and Business Logic110000 / 1 / 1
V3 Web Frontend Security161026 / 4 / 2
V4 API and Web Service40066 / 0 / 0
V5 File Handling40050 / 0 / 4
V6 Authentication261088 / 1 / 3
V7 Session Management151020 / 1 / 0
V8 Authorization70004 / 2 / 0
V9 Self-contained Tokens70000 / 0 / 0
V11 Cryptography112015 / 2 / 3
V12 Secure Communication61020 / 2 / 1
V13 Configuration94005 / 3 / 0
V14 Data Protection90002 / 1 / 1
V15 Secure Coding and Architecture111018 / 0 / 0
V16 Security Logging and Error Handling151001 / 0 / 0
Total1681303647 / 17 / 16

The short version. There is no L1 or L2 gap. The four password-policy requirements that used to sit here — V6.2.4 at L1, and V6.1.2 / V6.2.11 / V6.2.12 at L2 — were one missing control seen from four angles, and closed as one: check_password_policy now refuses a password that names this deployment, or that appears in a compiled-in corpus of common passwords. V6.2.2 and V6.2.3, both L1, closed the same way: one self-service password-change route on the account page, gated the way the second-factor routes already were, rather than two separate fixes.

Everything else that falls short of met is either a partial — a control that exists but does not reach everywhere the requirement asks — or a documented deviation, where the project has knowingly chosen otherwise and argued the choice already. Both have their own sections below.

Two chapters deserve a note on their shape. V6 Authentication carries the most n/a rows because the web admin has one authentication pathway and no out-of-band, biometric or federated factors — most of the chapter has no subject here. V16 Security Logging is the only chapter with no L1 requirements at all and is met almost entirely, which is what you would hope for in a certificate authority: the audit trail is the product.

V1 Encoding and Sanitization

#RequirementLStatusEvidence
1.1.1Decode into canonical form once, before processing2metThe JWS protected header and payload are base64url-decoded exactly once in crates/protocol/src/extractors/acme.rs, before any check reads them; a DNS identifier passes normalize_dns_name (crates/protocol/src/acme/rules.rs) once, before storage and before the filter sees it
1.1.2Output encoding as the final step, or by the interpreter2metminijinja escapes at render time. The rule is per template name: .html auto-escapes, .j2 does not — see crates/core/src/templating.rs
1.2.1Context-correct output encoding for HTTP/HTML1metEvery panel template is .html and therefore auto-escaped; auto_escaping_is_on_for_pages_and_off_for_notify pins both directions
1.2.2Encode untrusted data in dynamically built URLs; safe protocols only1metPanel URLs are built from server-side ids. The one place an untrusted URL is followed — an http-01 redirect — is checked against a scheme allowlist in Http01Validator::redirect_allowed (crates/net/src/challenge/http_01.rs)
1.2.3Encode when building JavaScript or JSON1metAll JSON is produced by serde_json; no template writes into a <script> block
1.2.4Parameterized database queries1metEvery statement in crates/store/src/ is a runtime sqlx::query with .bind(). No query is assembled with format!
1.2.5Protection against OS command injection1metScriptHook::run uses Command::new(path) with an argv vector and no shell (crates/core/src/script_hook.rs); payloads go to stdin as JSON
1.2.6LDAP injection2n/aNo LDAP client
1.2.7XPath injection2n/aNo XPath
1.2.8LaTeX injection2n/aNo LaTeX
1.2.9Escape special characters in regular expressions2metcompile_anchored and the glob translation both run regex::escape over everything that is not the wildcard (crates/policy/src/filter/mod.rs)
1.2.10CSV and formula injection3n/aNo CSV or spreadsheet export; the CLI emits text or JSON
1.3.1Sanitize untrusted HTML from editors1n/aNo rich-text input anywhere
1.3.2Avoid eval() and dynamic code execution1metNo dynamic code execution. The one place operator-supplied code runs is a custom script hook, which is a configured executable, not evaluated input
1.3.3Sanitize before a dangerous context; trim over-long input2metContacts reject control characters, User-Agent is truncated to 256 characters before storage (crates/core/src/audit/mod.rs), identifier lists are capped by order.max_identifiers
1.3.4Sanitize user-supplied SVG2n/aNo user-supplied images
1.3.5Sanitize user-supplied scriptable or template content2n/atemplate_dir overrides are operator-supplied files on the host, not user input
1.3.6SSRF protection by allowlist of protocols, domains, paths, ports2partialScheme, port and hop count are allowlisted for http-01 redirects; destination addresses deliberately are not. See Documented deviations
1.3.7No templates built from untrusted input2metTemplate sources come only from the embedded table or template_dir; untrusted values are only ever bound as context
1.3.8JNDI injection2n/aNo JNDI
1.3.9Sanitize before memcache2n/aNo memcache
1.3.10Sanitize format strings2metRust format strings are compile-time literals; a runtime string can never become one
1.3.11Sanitize before mail systems (SMTP/IMAP injection)2metcontact_shape_error rejects control characters, hfields and multiple addresses before a contact can reach a notify template (crates/protocol/src/acme/rules.rs)
1.3.12Regular expressions free from exponential backtracking3metThe regex crate has no backtracking and guarantees linear time; patterns are operator configuration, not request input
1.4.1Memory-safe strings and copies2metSafe Rust. The unsafe blocks in the tree are std::env::set_var in tests and one PKCS#11 Send impl (crates/signer/src/local_ca/pkcs11.rs)
1.4.2Prevent integer overflow2metTime and TTL arithmetic uses saturating_add/saturating_sub throughout crates/store/src/; release builds are not built with overflow checks disabled beyond the default
1.4.3Release memory and resources; no dangling pointers2metOwnership and Drop. Script hooks additionally set kill_on_drop so a timed-out child is reaped
1.5.1Restrictive XML parser configuration (XXE)1n/aNo XML parser in the dependency graph
1.5.2Safe deserialization of untrusted data2metserde into concrete structs. No polymorphic or client-chosen types; crates/core/src/jws/mod.rs deliberately does not use deny_unknown_fields because RFC 8555 §6.2 allows extra header parameters, and every field it acts on is named
1.5.3Consistent parsers for one data type3metOne JSON parser (serde_json) and one URL parser (url) in the tree

V2 Validation and Business Logic

#RequirementLStatusEvidence
2.1.1Documented input validation rules1metIdentifier syntax and normalisation are specified in Filters; the ACME wire formats are RFC 8555’s and the deviations are listed in Protocol Support
2.1.2Documented rules for combined data items2metThe order/CSR consistency rule — the CSR may name only what the order named — is stated in Filters and enforced at finalize
2.1.3Documented business logic limits2metorder.max_identifiers, server.max_body_bytes, nonce TTL, order and authorization lifetimes and the admission limiter are all in Configuration Reference with their defaults
2.2.1Validate input against expectations1metIdentifiers are normalised and type-checked, wildcards are refused when dns-01 is off, contacts are shape-checked, the CSR is parsed and compared against the order
2.2.2Validation enforced at a trusted service layer1metEvery check is server-side. The panel’s only client-side code is htmx attribute dispatch
2.2.3Combinations of related data items are reasonable2meta_wildcard_identifier_is_rejected_when_dns_01_is_disabled and a_csr_requesting_ca_powers_yields_a_leaf_without_them in tests/security.rs are two of the pinned cases
2.3.1Business logic flows only in the expected step order1metThe order state machine refuses out-of-sequence transitions: an_order_missing_an_authorization_never_becomes_ready, an_expired_order_cannot_be_finalized, a_deactivated_account_cannot_finalize_a_ready_order (tests/security.rs)
2.3.2Business logic limits implemented as documented2metan_order_naming_more_identifiers_than_the_limit_is_refused (tests/security.rs)
2.3.3Transactions succeed in full or roll back2metMulti-row writes run inside pool.begin()/commit() — order creation with its authorizations (crates/protocol/src/handlers/order.rs), challenge validation (crates/protocol/src/handlers/authz.rs), session promotion (crates/store/src/admin_session.rs)
2.3.4Locking prevents double-booking of limited resources2metSingle-use resources are claimed by UPDATE … WHERE … AND <unused> and decided by rows_affected == 1, never by read-then-write: nonces, recovery codes (crates/store/src/admin_recovery_code.rs), the TOTP replay step (crates/store/src/admin_user.rs) and session promotion
2.3.5Multi-user approval for high-value flows3gapIssuance and revocation are single-actor operations. There is no second-operator approval, and no plan to add one — a CA that needs a quorum to sign is a different product
2.4.1Anti-automation on expensive functions2metThe ACME listener carries an admission limiter with a queue budget and a request deadline (crates/protocol/src/middlewares/admission.rs); the admin login path is rate-limited per address — an IPv6 client per /64 — counted from the moment an attempt starts, and the KDF runs on the blocking pool (LoginLimiter, crates/admin/src/webadmin/session.rs)
2.4.2Business flows require realistic human timing3n/aEvery consumer of the ACME API is a machine; timing gates would break the protocol

V3 Web Frontend Security

Applies to the web admin. With admin.enabled = false, the default, none of this is exposed at all.

#RequirementLStatusEvidence
3.1.1Documented expected browser security features3partialWeb Admin states the cookie and CSRF requirements and the Secure-over-plain-HTTP failure mode; there is no statement of what the panel does when a browser lacks a feature
3.2.1Prevent content being rendered in the wrong context1metdefault-src 'none' plus X-Content-Type-Options: nosniff on every response; the API is nested under /api with its own JSON fallback so a page path never returns an API body
3.2.2Text rendered as text, not HTML1metAuto-escaping templates; no innerHTML outside htmx’s own fragment swap of server-rendered HTML
3.2.3Avoid DOM clobbering3metThe panel ships no application JavaScript — htmx is the only script, and everything is driven by hx-* attributes
3.3.1Secure attribute and a __Secure-/__Host- prefix1met__Host-acme_admin_session, which browsers accept only with Secure, Path=/ and no Domain (crates/admin/src/webadmin/session.rs)
3.3.2SameSite set according to purpose2metSameSite=Strict on both the session cookie and its clearing form
3.3.3__Host- prefix unless shared with other hosts2metSame as 3.3.1
3.3.4HttpOnly for values scripts must not read2metHttpOnly is set; the CSRF token travels in the page and the x-csrf-token request header, never in a readable cookie
3.3.5Cookie name and value under 4096 bytes3metA 32-byte token, base64url-encoded, plus a fixed name
3.4.1HSTS on all responses, ≥ 1 year, includeSubDomains for L21metmax-age=31536000; includeSubDomains, applied by the shared security_headers() constructor to both listeners (crates/protocol/src/router.rs). Emitted unconditionally: a browser ignores it over plain HTTP (RFC 6797 §7.2), so gating it on TLS would remove only a header that is already inert. includeSubDomains makes the host in admin.base_url load-bearing — see give the panel its own host name
3.4.2CORS Access-Control-Allow-Origin fixed or allowlisted1metNo CORS layer exists on either listener, so no Access-Control-Allow-Origin is ever emitted
3.4.3CSP with object-src 'none' and base-uri 'none'2metdefault-src 'none'; script-src 'self'; style-src 'self'; img-src 'self' data:; connect-src 'self'; form-action 'self'; frame-ancestors 'none'; base-uri 'none' — no unsafe-inline, no unsafe-eval (crates/admin/src/webadmin/mod.rs). object-src falls back to default-src 'none'
3.4.4X-Content-Type-Options: nosniff2metsecurity_headers()
3.4.5Referrer policy2metReferrer-Policy: same-origin on the admin listener
3.4.6CSP frame-ancestors on every response2metframe-ancestors 'none', with X-Frame-Options: DENY alongside for older clients
3.4.7CSP reports a violation-reporting location3gapNo report-to/report-uri. For a single-origin panel with no inline script, the report channel would have no consumer
3.4.8Cross-Origin-Opener-Policy on document responses3gapNot set. frame-ancestors 'none' covers framing but not shared Window access from a popup
3.5.1Anti-forgery tokens or non-safelisted header fields1metA per-session CSRF token in the x-csrf-token header on every unsafe method, plus an Origin check against admin.base_url — the module doc in crates/admin/src/webadmin/session.rs explains why SameSite=Strict alone is not enough here
3.5.2Functionality cannot be called without a preflight1n/aThe panel does not rely on CORS preflight; it uses the token in 3.5.1
3.5.3Sensitive functionality uses unsafe HTTP methods1metEvery mutating route is POST/DELETE; GET routes are read-only. mutating_endpoints() and mutating_page_endpoints() are the lists that make this checkable
3.5.4Separate applications on different hostnames2partialThe ACME and admin surfaces are separate sockets with separate TLS and separate defaults, and the admin binds loopback unless TLS is on. They are usually two ports on one host, and cookies are not port-scoped — see Documented deviations
3.5.5Validate postMessage origins2n/aNo postMessage
3.5.6No JSONP3metNone
3.5.7No authorized data in script resources3metThe only script served is a static, unauthenticated copy of htmx
3.5.8Authenticated resources embeddable only when intended3metSec-Fetch is not inspected, but frame-ancestors 'none', SameSite=Strict and the CSRF token together mean no cross-origin embed carries the session
3.6.1SRI for externally hosted client assets3metNothing is externally hosted. htmx is vendored under crates/admin/src/webadmin/static/ and served from the same origin
3.7.1Only supported, secure client-side technologies2metHTML, CSS and htmx. No plugins
3.7.2Automatic redirects only to allowlisted hosts2metThe panel’s redirects are fixed relative paths (/ui/, the sign-in page); there is no next= parameter and no open-redirect surface
3.7.3Notify before redirecting outside the application3n/aThe panel never redirects off-origin
3.7.4HSTS preload3n/aThe panel is an internal service with an operator-chosen hostname; preloading is a decision for the operator’s domain, not this software
3.7.5Documented behaviour on browsers lacking security features3gapSee 3.1.1

V4 API and Web Service

#RequirementLStatusEvidence
4.1.1Accurate Content-Type with charset1metACME responses are application/json or application/problem+json; panel pages are text/html; charset=utf-8 via axum::response::Html; the two static assets set their own type with an explicit charset (crates/admin/src/webadmin/pages/assets.rs)
4.1.2Only user-facing endpoints redirect HTTP to HTTPS2metNeither listener redirects. server.tls.enabled makes the socket speak TLS instead of cleartext, not alongside it (crates/net/src/tls.rs)
4.1.3Intermediary-set header fields cannot be overridden by the user2metA forwarded-for header is believed only from a hop in filter.trusted_proxies; with the list empty the header is ignored entirely and the peer address is used (crates/core/src/client.rs). The admin listener has its own list, admin.filter.trusted_proxies, empty by default (crates/admin/src/webadmin/filter.rs)
4.1.4Only supported HTTP methods are usable3metaxum routes declare their methods and answer 405 otherwise; both routers carry an explicit method_not_allowed_fallback so the refusal is a proper problem document rather than an empty body
4.1.5Per-message digital signatures for highly sensitive requests3metEvery state-changing ACME request is a JWS signed by the account key, verified against a nonce and the request URL (RFC 8555 §6.2) — this is the protocol’s own design, not an addition
4.2.1Correct HTTP message framing (request smuggling)2methyper performs the framing and rejects conflicting Content-Length/Transfer-Encoding; the application never parses framing itself
4.2.2Generated Content-Length matches the body3metResponse bodies are axum types; the length is computed, never asserted
4.2.3No connection-specific header fields over HTTP/2 or HTTP/33methyper enforces this. No handler sets Transfer-Encoding
4.2.4Reject CR/LF in HTTP/2 and HTTP/3 header fields3methttp::HeaderValue rejects control bytes on construction, in both directions
4.2.5Avoid generating over-long URIs or header fields3metOutbound URLs are built from configuration plus a bounded token or id; the http-01 validator additionally caps redirect hops
4.3.1GraphQL query cost limiting2n/aNo GraphQL
4.3.2GraphQL introspection disabled2n/aNo GraphQL
4.4.1WebSocket over TLS1n/aNo WebSocket
4.4.2WebSocket handshake Origin check2n/aNo WebSocket
4.4.3Dedicated WebSocket session tokens2n/aNo WebSocket
4.4.4WebSocket tokens obtained through the authenticated session2n/aNo WebSocket

V5 File Handling

There is no file upload feature. The only client-supplied structured input is a CSR inside a signed JWS, which is assessed under V2 and V11 rather than here. The two paths that touch the filesystem on a request’s behalf are the embedded static-asset allowlist and the certificate-chain download.

#RequirementLStatusEvidence
5.1.1Documented permitted file types, extensions and sizes2n/aNo upload feature to document
5.2.1Only accept files of a processable size1metserver.max_body_bytes (128 KiB) and admin.max_body_bytes (64 KiB) bound every request body, applied as DefaultBodyLimit at each router root
5.2.2Extension matches content1n/aNo uploads
5.2.3Compressed-file limits2n/aNothing is decompressed
5.2.4Per-user file quota3n/aNo uploads
5.2.5Reject symlinks in archives3n/aNo archives
5.2.6Reject over-large images3n/aNo images
5.3.1Untrusted files in a public folder are not executed1n/aNothing untrusted is written to a served directory
5.3.2File paths built from trusted data, not user filenames1metGET /ui/static/{file} is a two-arm match, not a filesystem lookup — tower-http’s fs feature is deliberately off (crates/admin/src/webadmin/pages/assets.rs). The http-01 responder looks a token up in an in-memory store and touches no path
5.3.3Ignore user path information when decompressing3n/aNothing is decompressed
5.4.1Validate or ignore user filenames; set Content-Disposition2metThe chain download names the file from the stored order id, not the path segment, and sets attachment; filename="…" (crates/admin/src/webadmin/pages/orders.rs)
5.4.2Served filenames are encoded or sanitized2metSame: a generated identifier, so there is nothing to encode
5.4.3Antivirus scanning of files from untrusted sources2n/aNo files are accepted from untrusted sources

V6 Authentication

Applies to the web admin and to the CLI commands that mint and rotate operator credentials. The ACME listener authenticates keys, not people; that is assessed under V9.

#RequirementLStatusEvidence
6.1.1Documented anti-automation and lockout behaviour1metWeb Admin and the login_max_attempts / login_window_seconds entries in Configuration Reference. The limiter is keyed on the peer address, never the username, so no attacker can lock an operator out by guessing at them
6.1.2Documented list of context-specific words barred from passwords2metDerived and documented: Password policy
6.1.3Multiple authentication pathways documented together2metThere are two — a panel session and a shell on the host — and Users & Sessions states which operations belong to which and why create and passwd stay on the host
6.2.1Passwords at least 8 characters1metMIN_PASSWORD_LEN = 12, counted in characters rather than bytes (crates/admin/src/admin/password.rs)
6.2.2Users can change their password1metThe panel’s own account page carries a password card (POST /ui/account/password, POST /api/account/password) beside acme-proxy admin user passwd, so an operator with no shell can rotate their own → Users & Sessions
6.2.3Password change requires current and new password1methandlers::mfa::verify_current_password (crates/admin/src/webadmin/handlers/mfa.rs) checks the current password before admin::users::change_own_password writes a new one — unconditionally, unlike the second-factor step-up it was split out of, since this requirement has no “nothing yet to protect” exemption. admin user passwd still takes only the new one, answering as it does to a process that can already rewrite the row
6.2.4Check against the top 3000 passwords1met13 918 entries compiled in (crates/admin/src/admin/corpus/); every shorter entry is already refused on length
6.2.5No composition rules1metDeliberately none. The three rules are about length, this deployment’s own words and known-common passwords — none dictates shape (crates/admin/src/admin/password.rs)
6.2.6Password fields use type=password1mettemplates/login.html, the step-up fields in templates/account/_card.html, templates/account/_contact.html and templates/operators/_card.html, and the two fields in templates/account/_password_card.html
6.2.7Paste and password managers permitted1metStandard inputs with autocomplete="username" / "current-password"; nothing blocks paste
6.2.8Password verified exactly as received1metverify_password hashes the bytes as received: no trimming, no case folding, and an over-long password is rejected rather than truncated. The policy check folds a copy to compare against the corpus and the word list, and never touches what is stored
6.2.9Passwords of at least 64 characters permitted2metMAX_PASSWORD_LEN = 1024 bytes, a denial-of-service bound rather than a policy
6.2.10No forced periodic rotation2metNothing expires a password. The stored form is self-describing, so raising the KDF cost re-encodes a row on its owner’s next login instead of forcing a change
6.2.11Context-specific word list used2metPasswordContext in crates/admin/src/admin/password.rs, matched as a substring
6.2.12Check against breached passwords2metSame corpus: breach-derived (xato-net), filtered to the reachable length range
6.3.1Credential-stuffing and brute-force controls1metLoginLimiter refuses over the limit before the 600 000-iteration KDF runs, which makes it an availability control as much as a credential one; an attempt counts from the moment it starts, so a parallel burst cannot outrun it (crates/admin/src/webadmin/session.rs)
6.3.2No default accounts1metThe admin_users migration seeds no rows and there is no sign-up page; the first operator is created by admin user create on the host
6.3.3MFA or a combination of single factors2partialTOTP with recovery codes is implemented and admin.require_mfa enforces it for every operator — but it defaults to false, so a stock deployment is single-factor. Hardening tells operators to turn it on. For L3 this would need a hardware factor; see Documented deviations
6.3.4No undocumented pathways; consistent strength2metThe panel and API share one session layer, and every mutating route passes through AuthenticatedWrite, PageSessionWrite or EnrolWrite. The host CLI is the second pathway and is documented as such
6.3.5Notify users of suspicious authentication attempts3metA completed sign-in from an address not among the operator’s recent ones (admin_users.known_login_ips, last five), a correct password then a refused second factor, and a per-session second-factor lockout each send an admin_sign_in notification to the operator’s own contact_email, through [admin.notify] (crates/admin/src/webadmin/handlers/session.rs, crates/jobs/src/notify/)
6.3.6Email not used as an authentication factor3metIt is not
6.3.7Notify after changes to authentication details3metA password change, a second-factor enrol/disable, a recovery-code regeneration, a notification-address change and a colleague-admin second-factor reset send an admin_credential_changed notification — from the panel and from the host CLI alike (admin user passwd, contact, totp reset, totp recovery-codes), the CLI queuing the delivery for the running server’s worker
6.3.8Valid users not deducible from failed challenges3metAn unknown username still pays the KDF, against password::dummy_hash(), and every failure returns one invalid_credentials whatever the real cause (crates/admin/src/admin/users.rs)
6.4.1Initial passwords and activation codes are random, policy-compliant and short-lived1n/aNothing generates an initial password; the operator supplies one on stdin or in --password-file
6.4.2No password hints or secret questions1metNeither exists
6.4.3Secure forgotten-password reset that does not bypass MFA2metReset is admin user passwd on the host. It revokes every session the operator held and leaves the enrolled factor untouched, so the next sign-in still needs it
6.4.4Lost MFA factor requires enrolment-level identity proofing2metEither a single-use recovery code, or admin user totp reset on the host — the second being a strictly stronger proof than the live session plus password that enrolment took
6.4.5Renewal reminders before an authenticator expires3n/aNo authentication factor expires
6.4.6Administrators can reset but not choose a user’s password3gapadmin user passwd sets the password, so whoever runs it knows it. This is a host-root operation on a machine that already holds the hashes
6.5.1Lookup secrets and TOTPs usable only once2metAdminUser::claim_totp_step is an UPDATE … WHERE totp_last_step IS NULL OR totp_last_step < ? decided by rows_affected, so a code resubmitted inside its own 30-second window is refused (crates/store/src/admin_user.rs); recovery codes are consumed by UPDATE … WHERE id = ? AND used_at IS NULL
6.5.2Sub-112-bit lookup secrets hashed with an approved KDF and a 32-bit salt2metRecovery codes carry 50 bits and are stored through admin::password — PBKDF2-HMAC-SHA256, 600 000 iterations, a 128-bit per-row salt
6.5.3Seeds and codes from a CSPRNG2metring::rand::SystemRandom for the TOTP secret (crates/admin/src/admin/totp.rs) and every recovery code (crates/admin/src/admin/recovery.rs)
6.5.4Lookup secrets have at least 20 bits of entropy2metTen characters from a 32-symbol alphabet: 50 bits, with zero modulo bias because 256 is a multiple of 32
6.5.5Defined lifetime for codes and TOTPs2met30-second step with RFC 6238 §5.2’s one step of permitted skew either way (SKEW_STEPS = 1). A half-authenticated session additionally dies after PENDING_MFA_TTL, five minutes
6.5.6Any factor can be revoked3metadmin user totp reset, admin user disable, admin session revoke [--user <u> [--session <id>] | --all], and recovery codes are superseded as a set on re-enrolment
6.5.7Biometrics only as a secondary factor3n/aNo biometrics
6.5.8TOTP checked against a trusted time source3metThe server’s own clock; no client-supplied time reaches totp::verify
6.6.1PSTN OTP restrictions2n/aNo SMS or voice factor
6.6.2Out-of-band codes bound to their originating request2n/aNo out-of-band factor. The equivalent binding for TOTP is that the code is only accepted against the pending_mfa session that the password created
6.6.3Rate-limit code-based out-of-band mechanisms2n/aNo out-of-band factor. TOTP guessing is bounded twice — mfa_attempts on the pending row and the five-minute PENDING_MFA_TTL
6.6.4Rate-limit push notifications3n/aNo push factor
6.7.1Certificates verifying authentication assertions protected from modification3metAccount public keys live in accounts under the database’s file mode; a modified key is a key that no longer verifies its own account’s requests
6.7.2Challenge nonce at least 64 bits and unique3met256 bits from ring::rand::SystemRandom, unique by primary key and single-use by rows_affected (crates/store/src/nonce.rs)
6.8.1Identity cannot be spoofed across identity providers2n/aNo identity provider
6.8.2Signatures on authentication assertions validated2n/aNo external assertions. The equivalent for ACME JWS is V9.1.1
6.8.3SAML assertions processed once2n/aNo SAML
6.8.4Authentication strength verified from the IdP2n/aNo identity provider

V7 Session Management

Applies to the web admin. The ACME listener holds no sessions: every request carries its own signature and its own nonce.

#RequirementLStatusEvidence
7.1.1Documented inactivity timeout and absolute lifetime2metsession_ttl_seconds (12 h, never extended by activity) and session_idle_timeout_seconds (1 h) in Configuration Reference, restated in Web Admin
7.1.2Documented concurrent-session policy2partialThe behaviour is definite — sessions are unlimited per operator, and admin session revoke (--all, one operator’s, or one session with --session <id>) is the lever — but no page states the limit as a policy
7.1.3Federated session coordination documented2n/aNo federation
7.2.1Session verification at a trusted backend1metEvery request resolves hex(SHA-256(token)) against admin_sessions and re-checks state, expiry, idleness and the owner’s status (crates/admin/src/webadmin/session.rs)
7.2.2Dynamically generated tokens, not static secrets1metmint_token per sign-in; there are no API keys on this listener
7.2.3Reference tokens unique, CSPRNG, ≥ 128 bits1met256 bits from ring::rand::SystemRandom, base64url-encoded
7.2.4New token on authentication, old one terminated1metSign-in deletes whatever session the request carried; completing MFA is a rotation — AdminSession::promote deletes the pending_mfa row and inserts a new one with a new token and a new CSRF token, in one transaction (crates/store/src/admin_session.rs)
7.3.1Inactivity timeout2metsession_idle_timeout_seconds, checked per request and swept by the reaper
7.3.2Absolute maximum session lifetime2metexpires_at is set at creation and never advanced
7.4.1Terminated sessions cannot be reused1metSessions are reference tokens in a table; sign-out deletes the row
7.4.2All sessions terminated when an account is disabled or deleted1metset_status("disabled") and set_password both call AdminSession::delete_for_user; the liveness check also refuses a session whose owner is no longer active
7.4.3Option to terminate other sessions after a factor changes2metconfirm_totp_enrolment and disable_totp both call revoke_other_sessions; a password change revokes every session unconditionally
7.4.4Visible logout on every authenticated page2metA “Sign out” control in templates/layout.html, which every page extends
7.4.5Administrators can terminate sessions individually or globally2metadmin session list/revoke on the host terminates globally (--all), one operator’s (--user <u>), or one session (--user <u> --session <id>, the id being the fingerprint the listing prints); the panel’s Operators page does the individual form over HTTP — GET /ui/operators/{username} lists another operator’s sessions and POST /ui/operators/{username}/sessions/{id}/revoke ends one, gated by verify_current_password
7.5.1Full re-authentication before changing authentication attributes2metcheck_step_up demands the password again before any change to an existing second factor, and the module doc explains the blast radius that makes it necessary (crates/admin/src/webadmin/handlers/mfa.rs)
7.5.2Users can view and terminate their own sessions2metThe account page’s Sessions card (GET /api/account/sessions, /ui/account) lists every one of the caller’s own live sessions and terminates one individually (POST /api/account/sessions/{id}/revoke) or all at once (“Sign out everywhere”) — closing the gap between nothing and everything the panel used to leave → Sessions
7.5.3Further authentication before highly sensitive operations3partialSecond-factor changes are gated by check_step_up, and the whole /operators colleague-management surface by verify_current_password, which re-prompts even for an operator with no factor. Certificate revocation and account deletion require at least the operator role (admin_users.role), but for an operator holding it a live session is still sufficient authority — no password re-prompt on the CA mutations
7.6.1Federated re-authentication behaviour2n/aNo federation
7.6.2Session creation requires explicit user action2metA session exists only after a submitted sign-in form

V8 Authorization

#RequirementLStatusEvidence
8.1.1Documented function-level and data-specific rules1metSecurity Model states the four issuance gates; Filters specifies the policy engine; Web Admin states that every route needs a session
8.1.2Documented field-level rules2metThe ACME object shapes are RFC 8555’s, and Audit Trail states which fields are recorded and that none of them reach an ACME object
8.1.3Documented environmental and contextual attributes3metAddress, reverse name, IPAM ownership, request path and EAB identity are each documented under Filters, and the trust placed in a forwarded address under Allowed IP
8.1.4Documented use of contextual factors in decisions3metThe policy expression language, including how an or over an address check weakens a conjunction, is written out in Hardening
8.2.1Function-level access restricted to explicit permissions1metACME: POST-as-GET with a kid resolving to the owning account. Admin: three extractors every mutating route passes through
8.2.2Data-specific access restricted (IDOR/BOLA)1metOrder, authorization and certificate reads check the requesting account owns the object; accounts and orders are additionally isolated per profile, so a kid naming another profile does not resolve (crates/protocol/src/extractors/acme.rs)
8.2.3Field-level access restricted (BOPLA)2metResponses are built from explicit serializer functions, never by serializing a row
8.2.4Adaptive controls from contextual attributes3metThe filter chain evaluates per request, not per session, so a change of address is re-evaluated on the next call
8.3.1Authorization enforced at a trusted service layer1metExtractors and middleware, server-side. No decision depends on anything the client sends unsigned
8.3.2Authorization changes applied immediately3metSessions are reference tokens read from the database each request, so a disabled operator or a revoked session stops working on the next call. filter reload applies policy without a restart
8.3.3Access based on the originating subject3partialWith the relay signer, one upstream account is deliberately multiplexed across every local client — that is the feature. The local gates decide, and the upstream sees only this server. See Documented deviations
8.4.1Cross-tenant controls2metProfiles are the tenancy boundary: accounts, orders, nonces and EAB credentials are scoped to one, and a kid from another profile fails the prefix check
8.4.2Administrative access uses more than network location3partialPassword plus optional TOTP plus a session, with the bind address and TLS as further layers, and a per-operator role (admin/operator/viewer) scoping what a session may do. There is no device posture assessment and no contextual risk analysis

V9 Self-contained Tokens

The self-contained token in this system is the ACME JWS on every state-changing request (RFC 8555 §6.2), not a session JWT — the admin session is a reference token and is assessed under V7.

#RequirementLStatusEvidence
9.1.1Signature validated before the contents are accepted1metverify_jws verifies the signature over the protected header and payload before any handler sees the body (crates/protocol/src/extractors/acme.rs, crates/core/src/jws/signature.rs)
9.1.2Algorithm allowlist, no none1metExactly ES256 on P-256 and RS256 are accepted; anything else is Unsupported algorithm. The alg must additionally agree with the key type and the named curve, so alg alone never selects the verifier (crates/core/src/jws/signature.rs)
9.1.3Key material from trusted pre-configured sources1metA kid resolves to a stored account key whose URL prefix must match this profile’s base_url; a jwk is the key being registered and is only ever trusted for newAccount/revokeCert as RFC 8555 §6.2 defines. jwk and kid together are refused, and a crit header is refused outright
9.2.1Validity time span honoured1metThe equivalent is the nonce: single-use, and refused past nonce.ttl_seconds. Unknown, consumed and expired are made indistinguishable on purpose (crates/store/src/nonce.rs)
9.2.2Token type checked against the intended purpose2metThe protected header must carry exactly the fields RFC 8555 §6.2 defines for the request kind; newAccount requires a jwk, revokeCert takes either, and everything else a kid — an embedded jwk elsewhere is 400 malformed
9.2.3Audience restriction2metThe JWS url must equal profile.base_url plus the request path, byte for byte (RFC 8555 §6.4). A signature captured from one profile does not verify against another
9.2.4Same key across audiences carries an audience restriction2metSame mechanism: the audience is in the signed url, and the kid prefix pins the profile

V11 Cryptography

#RequirementLStatusEvidence
11.1.1Documented key management policy and lifecycle2metSecurity Model names every secret, what its compromise buys and how it is stored; Secret Rotation is the lifecycle half — a recommended interval and the early-rotation triggers for each — and Hardening covers the CA key specifically
11.1.2Cryptographic inventory maintained2metThe table in Security Model, plus the per-algorithm rationale carried in the module docs of crates/admin/src/admin/password.rs, crates/admin/src/admin/totp.rs and crates/core/src/eab.rs
11.1.3Cryptographic discovery mechanisms3metOne backend: ring, plus rustls for TLS and rcgen for certificate construction. grep -rn ring src/ is the discovery mechanism, and cargo deny fails the build on an unlisted crypto dependency
11.1.4Inventory includes a post-quantum migration path3gapNo PQC migration plan. The ACME wire algorithms are RFC 8555’s to change first
11.2.1Industry-validated implementations2metring (BoringSSL-derived) for hashing, HMAC, PBKDF2, signature verification and the RNG; rustls for TLS. Nothing hand-rolls a primitive — crates/admin/src/admin/totp.rs composes ring::hmac per RFC 4226 and is checked against the RFC’s published test vectors
11.2.2Crypto agility2metPassword hashes are stored self-describing (pbkdf2-sha256$600000$…), so the algorithm or cost can change with a new branch in verify_password and needs_rehash re-encodes each row at its owner’s next login — no migration. The signer backend, the CA key type and the key source (file or PKCS#11) are all configuration
11.2.3Minimum 128 bits of security2partialECDSA P-256, SHA-256, HMAC-SHA-256 and 256-bit secrets are all at or above the bar. RSA is accepted from 2048 bits (RSA_PKCS1_2048_8192_SHA256), which is about 112 — see Documented deviations
11.2.4Constant-time cryptographic operations3metring::constant_time::verify_slices_are_equal under the hood, and subtle::ConstantTimeEq for the TOTP comparison (crates/admin/src/admin/totp.rs)
11.2.5Cryptographic modules fail securely3metA verification failure is a refusal, never a fallback. A corrupt stored password hash is deliberately not folded into “wrong password” — it refuses and logs admin_password_hash_unreadable (crates/admin/src/admin/users.rs)
11.3.1No insecure block modes or weak padding1partialNothing in the tree encrypts. RS256 is RSASSA-PKCS1-v1_5, which RFC 8555 requires — a signature scheme, not the padding oracle this requirement targets. See Documented deviations
11.3.2Only approved ciphers and modes1metTransport encryption is rustls with safe defaults; the application encrypts nothing itself
11.3.3Encrypted data protected against modification2n/aNo application-layer encryption
11.3.4Single-use numbers not reused across key/data pairs3n/aNo application-layer encryption. ACME nonces are single-use by construction
11.3.5Encrypt-then-MAC3n/aNo application-layer encryption
11.4.1Approved hash functions1metSHA-256 throughout. The one SHA-1 is HMAC_SHA1_FOR_LEGACY_USE_ONLY inside TOTP, which RFC 6238 §1.2 specifies and which every authenticator app assumes — the module doc in crates/admin/src/admin/totp.rs is the argument for not “fixing” it
11.4.2Passwords stored with an approved, expensive KDF2metPBKDF2-HMAC-SHA256 at 600 000 iterations with a 128-bit per-row salt — OWASP’s current recommendation for the non-Argon2 case. See Documented deviations for why not Argon2id
11.4.3Collision-resistant hashes of adequate length in signatures2metSHA-256 for every signature and every integrity use. HMAC-SHA-1’s security rests on the PRF property, not collision resistance
11.4.4Approved KDF with key-stretching for password-derived keys2metSame PBKDF2 parameters; recovery codes go through the identical path
11.5.1Non-guessable values from a CSPRNG with ≥ 128 bits2metSession tokens, CSRF tokens, EAB secrets, challenge tokens and ACME replay nonces are all 256 bits from ring::rand::SystemRandom, base64url-encoded, through the one crates/core/src/random.rs. The nonce was a UUID v4 until 0.2.0 — 122 bits, and a form this requirement names explicitly
11.5.2RNG works securely under heavy demand3metSystemRandom draws from the OS CSPRNG; there is no userspace pool to exhaust
11.6.1Approved algorithms for key generation and signatures2metrcgen generates ECDSA P-256 by default; the accepted account-key algorithms are the two RFC 8555 defines. Key generation can be delegated to a PKCS#11 token, where the key never leaves the device
11.6.2Approved key exchange with secure parameters3metrustls with with_safe_default_protocol_versions(): TLS 1.2 and 1.3 only, and only its own vetted groups
11.7.1Full memory encryption for data in use3n/aA property of the host, not of this process
11.7.2Data minimization during processing3partialThe CA key can live in a PKCS#11 token and never enter this process at all. EAB and TOTP secrets are necessarily readable, because both are verified by recomputing an HMAC — file mode is the boundary, and that is stated in Security Model

V12 Secure Communication

#RequirementLStatusEvidence
12.1.1Only current TLS versions, newest preferred1metwith_safe_default_protocol_versions() on the rustls ring provider — TLS 1.3 and 1.2 only (crates/net/src/tls.rs)
12.1.2Recommended cipher suites, forward secrecy for L32metrustls ships no suite without forward secrecy and none that is not current; there is no knob to weaken it
12.1.3mTLS client certificates validated before use2n/aNo mTLS. The one place a client certificate is inspected is tls-alpn-01 validation, where the certificate is the challenge response and is checked for the RFC 8737 acmeIdentifier extension rather than for trust
12.1.4Certificate revocation such as OCSP stapling3partialAs a CA, the server publishes a CRL signed over the revocations recorded in its database (Revocation & CRL). As a TLS server it does not staple
12.1.5Encrypted Client Hello3gapNot offered by rustls in a form this could adopt today
12.2.1TLS for all client connectivity, no fallback1metWith server.tls.enabled the socket speaks TLS instead of cleartext; there is no downgrade path. HTTPS is on the hardening checklist for deployments that terminate elsewhere
12.2.2Publicly trusted certificates on external services1n/aThis is an internal service by design; its clients trust the CA the operator installed
12.3.1Encrypted protocols for all inbound and outbound connections2partialThe relay upstream, webhooks and IPAM are HTTPS. http-01 validation is HTTP because RFC 8555 §8.3 defines it that way, SQLite is a local file, not a connection, and a PostgreSQL connection is TLS only when database.url asks for it with sslmode — the server does not enforce it
12.3.2TLS clients validate certificates2metThe relay client validates against webpki-roots — there, the certificate is the only thing identifying the CA being handed your CSRs. The IPAM clients validate too; insecure_skip_verify exists, defaults off, and warns on every startup while on
12.3.3TLS between internal HTTP services2metSame set. The http-01 exception above is the protocol’s
12.3.4Internal TLS uses trusted certificates2metThe IPAM clients take a ca_bundle so a NetBox behind an internal PKI is trusted specifically rather than by disabling verification (crates/core/src/config/types/ipam.rs)
12.3.5Strong mutual authentication between internal services3n/aSingle process; there are no intra-service hops

V13 Configuration

#RequirementLStatusEvidence
13.1.1All communication needs documented, including user-supplied destinations2metSecurity Model names all three outbound surfaces and says which of them a client can steer
13.1.2Documented connection limits and behaviour at the limit3metThe database pool size, the admission limiter’s slots, its queue budget and its deadline are all in Configuration Reference, and shedding at the limit is a 503 problem document
13.1.3Documented resource-management strategy per external system3partialTimeouts are documented per subsystem and every outbound call has one. Retry policy is documented for the job runner but not stated as a policy for the IPAM and webhook clients
13.1.4Documented critical secrets and a rotation schedule3metThe secrets are named and classified in Security Model; Secret Rotation gives a recommended interval and the early-rotation triggers for each, with the CA key called out as structural rather than scheduled
13.2.1Authenticated backend communication with non-shared credentials2partialThe relay upstream authenticates by account key and the IPAM clients by API token, both per-deployment, and PostgreSQL by the role and password in database.url. SQLite is a local file governed by file mode, not by a credential
13.2.2Least privilege for backend accounts2metcustom hooks run with env_clear(), a minimal PATH, a timeout and kill_on_drop (crates/core/src/script_hook.rs); the systemd unit in Deployment runs as a dedicated acme-proxy user and the repository Containerfile runs as a non-root acme-proxy user (uid 1000) owning only /data; the IPAM token needs read access only
13.2.3No default service credentials2metNothing ships with a credential. Every secret is either operator-supplied or generated on first start
13.2.4Allowlist of external systems the application may contact2partialThe relay upstream, the IPAM host and the webhook URL are each a single configured destination — an allowlist of one. The http-01 validator is the exception, and deliberately so
13.2.5Server-level allowlist of destinations2partialSame. The containment for http-01 is scheme, port and hop count rather than destination
13.2.6Documented per-connection configuration followed3metEach client is built from its own configuration block at startup, so a broken setting stops the server rather than failing every later call
13.3.1A secrets management solution; no secrets in source or artifacts2partialNo secret is in the source tree or the image. Every secret can come from the environment rather than the file, and the CA key can live in a PKCS#11 token — which is the L3 hardware-backed form. There is no vault integration, and the database necessarily holds EAB and TOTP secrets in retrievable form
13.3.2Least privilege for secret access2metKeys are created 0600 with create_new rather than chmod’ed afterwards (crates/core/src/pemfile.rs); the database file mode is the documented boundary
13.3.3Cryptographic operations inside an isolated security module3partialAvailable but not required: --features hsm puts the issuing key in a PKCS#11 token, where it can be used and not copied (Hardware Keys)
13.3.4Secrets expire and rotate as documented3partialRotation is now documented per secret in Secret Rotation. Sessions expire on their own and EAB credentials are revocable live; the CA key, the TSIG key and the API tokens rotate on an operator-run cadence, not a timer
13.4.1No source-control metadata deployed1met.dockerignore is an allowlist — * then !Cargo.toml, !Cargo.lock, !src/, !crates/store/migrations/ — so .git never enters the build context, and the final stage copies only the compiled binary
13.4.2Debug modes disabled in production2metLog level is configuration and defaults to info; there is no debug endpoint and no development mode. challenge.bypass, the one setting that genuinely weakens the server, is off by default and warns on every startup while on
13.4.3No directory listings2metNothing is served from a directory. tower-http’s fs feature is off and static assets are a two-arm match
13.4.4HTTP TRACE unsupported2metNever routed; axum answers 405
13.4.5Documentation and monitoring endpoints not exposed unless intended2met/metrics is a separate listener, off by default; /health is deliberately outside the filter chain and the hardening checklist tells operators not to forward it (Monitoring)
13.4.6No detailed version information exposed3metNo Server header is set and no version appears in any response body
13.4.7Web tier serves only specific extensions3metThe static allowlist is two filenames; everything else is a 404

V14 Data Protection

#RequirementLStatusEvidence
14.1.1Sensitive data identified and classified2metSecurity Model classifies every secret by what its compromise buys; Database Schema classifies each by the form it is stored in — one-way, retrievable, or never stored
14.1.2Documented protection requirements per level2metThe same two pages, plus Audit Trail for retention
14.2.1No sensitive data in URLs or query strings1metThe session token is in a cookie, the CSRF token in a header, the EAB secret in a response body shown once. No credential is ever a path or query parameter
14.2.2Sensitive data not cached in server components2metCache-Control: no-store on every admin response — account contacts and a freshly minted EAB secret must not sit in a disk cache after the tab closes (crates/admin/src/webadmin/mod.rs)
14.2.3Sensitive data not sent to untrusted parties2metThe only outbound payloads are the relay’s own ACME traffic, a webhook to an operator-configured URL and IPAM lookups. No analytics, no third-party asset, no CDN
14.2.4Documented controls implemented2metNonces and session tokens reach logs only as fingerprints (crates/store/src/nonce.rs, crates/store/src/admin_session.rs); proxy URLs are redacted before they are logged or Debug-formatted, pinned by neither_debug_nor_redacted_leaks_the_password (crates/net/src/proxy.rs); the database URL’s password is redacted wherever it is printed (redact_url, crates/core/src/logfields.rs), and an EAB row’s Debug omits its HMAC secret (crates/store/src/eab.rs)
14.2.5Caching only for expected content types (web cache deception)3metno-store on the whole admin listener, and an unknown path returns a 404, never a different valid file
14.2.6Return the minimum sensitive data3metAn EAB secret is shown exactly once at creation; a session is displayed by the fingerprint of its token hash, never by the hash; render_admin_session_json is the one serializer
14.2.7Retention classification and scheduled deletion3partialaudit.retention_days sweeps the trail and the job runner reaps nonces, expired sessions and stale orders. The default is 0 — keep everything — which is the right default for a trail whose value is that it is complete, and is a decision the operator is asked to make
14.2.8Strip metadata from user-submitted files3n/aNo file uploads
14.3.1Authenticated data cleared from client storage on termination1metThe panel keeps nothing in localStorage or sessionStorage; sign-out clears the cookie with a Max-Age=0 Set-Cookie carrying the same attributes
14.3.2Anti-caching response header fields2metCache-Control: no-store
14.3.3No sensitive data in browser storage beyond session tokens2metOnly the session cookie exists

V15 Secure Coding and Architecture

#RequirementLStatusEvidence
15.1.1Documented remediation time frames for vulnerable components1partialSecurity Policy states that fixes land on main and in the next release, and cargo deny runs advisories on every CI run and on a schedule. No numeric time frame is committed to
15.1.2An SBOM or equivalent inventory is maintained2metsbom.cdx.json is a committed CycloneDX 1.5 inventory of the shipped dependency closure (--all-features --target all, dev-dependencies excluded), regenerated and diffed by the sbom CI job and carried in the published crate; cargo deny check gates that same closure
15.1.3Documented resource-demanding functionality2metThe expensive paths are named and bounded: http-01 and dns-01 validation have timeouts, the PBKDF2 cost is documented as a denial-of-service lever with the limiter placed before it, and the admission limiter’s shed-versus-queue reasoning is written out in crates/protocol/src/middlewares/admission.rs
15.1.4Risky third-party libraries highlighted3metdeny.toml is the allow list, run with all-features = true, and the rationale for refusing dependencies is recorded where the refusal was made — crates/admin/src/admin/password.rs on Argon2id, issue #5 on webauthn-rs
15.1.5Dangerous functionality highlighted3metSecurity Model names the three request-forgery surfaces, and Security Policy lists the behaviour that looks alarming and is deliberate
15.2.1No components past the documented remediation window1metThe Advisories, licenses & sources CI job fails the build on a RUSTSEC advisory
15.2.2Implemented defenses against availability loss2metAdmission limiter with a queue budget and a request deadline, body limits on both listeners, a login limiter ahead of the KDF, per-call timeouts on every outbound subsystem, and kill_on_drop on script hooks
15.2.3Production contains no test or development functionality2metTest helpers are #[cfg(test)] or behind testutil; there is no sample data, no seeded account and no development route
15.2.4Dependencies from expected repositories3metcargo deny check sources restricts registries, and Cargo.lock pins every transitive dependency by hash
15.2.5Extra protection around dangerous functionality3metScript hooks are the dangerous surface and run in a cleared environment with a minimal PATH, a deadline and kill_on_drop; the CA key can be moved into a PKCS#11 token; the Containerfile is the network-isolation story
15.3.1Return only the required subset of fields1metExplicit serializers per resource; no row is serialized wholesale
15.3.2Do not follow redirects unless intended2metIntended, bounded and switchable: follow_redirects and max_redirects on the http-01 validator, with scheme and port checked on every hop (crates/net/src/challenge/http_01.rs)
15.3.3Countermeasures against mass assignment2metRequest bodies deserialize into per-route structs holding only the fields that route accepts; nothing constructs a row from client JSON
15.3.4Original client IP transferred correctly and used for decisions2metfilter.trusted_proxies is the allowlist of hops whose forwarded header is believed; empty means the header is ignored. The admin listener has its own list, admin.filter.trusted_proxies, and with it empty — the default — no caller can choose its own rate-limiter key (crates/core/src/client.rs, crates/admin/src/webadmin/filter.rs, crates/admin/src/webadmin/session.rs)
15.3.5Explicit types and strict comparisons2metRust’s type system; there is no coercion to juggle
15.3.6JavaScript written to prevent prototype pollution2n/aThe panel ships no application JavaScript
15.3.7Defenses against HTTP parameter pollution2metaxum extractors read from one named source per parameter — a path segment, a typed query struct, or a JSON body — never from a merged bag
15.4.1Thread-safe access to shared objects3metSend/Sync are checked at compile time; shared mutable state is behind Mutex or Semaphore (LoginLimiter, Admission)
15.4.2State checks and dependent actions are atomic3metThe single-use idiom is one statement: UPDATE … WHERE <still unused> decided by rows_affected, never a read followed by a write. Key files are created with create_new, which is the atomic form of “exists?” then “create”
15.4.3Consistent locking, contained in the owning code3metLocks are held inside the type that owns the resource and never across an await
15.4.4Resource allocation prevents starvation3metThe admission limiter refuses past its queue budget rather than queueing without bound — the reasoning for not using GlobalConcurrencyLimitLayer is written out in crates/protocol/src/middlewares/admission.rs

V16 Security Logging and Error Handling

#RequirementLStatusEvidence
16.1.1A logging inventory exists2metMonitoring enumerates every event = "…" name; Audit Trail states what the trail records, where it lives, who can read it and how retention works
16.2.1Log entries carry when, where, who, what2metThe access line carries method, URI, status, latency, client address, profile and a request id (crates/protocol/src/middlewares/access.rs); an audit row carries actor, address, reverse name, identifiers, User-Agent and the same request id
16.2.2Synchronized time sources; UTC or explicit offset2metTimestamps come from the host clock as Unix seconds in the database and RFC 3339 in the log; host time sync is the operator’s
16.2.3Logs only go to documented destinations2metOne tracing subscriber built in one place — prepare_logging in crates/server/src/logging.rs — with logging.target naming the sink, so a reload cannot drift from startup
16.2.4Logs readable by the log processor2metlogging.json_format produces one JSON object per line, with flatten_event for pipelines that want fields at the top level
16.2.5Sensitive data logged according to its protection level2metNonces and session tokens appear only as fingerprints; proxy credentials and the database URL’s password are redacted; a password never enters a log or argv — admin user passwd reads from stdin or --password-file
16.3.1All authentication operations logged2metadmin_login_*, admin_mfa_verified, admin_mfa_attempts_exhausted, admin_logout and admin_password_hash_unreadable, each with the outcome and the method used
16.3.2Failed authorization attempts logged2metFilter denials, jws_url_mismatch, jws_jwk_and_kid_both_present, nonce_replayed and the *_failed audit rows. certificate_revoke_failed is written specifically so a run of them is visible as somebody enumerating serials
16.3.3Security events and control-bypass attempts logged2metchallenge_validation_bypassed and two other weakened-configuration warnings repeat on every startup so they cannot become background noise (Hardening)
16.3.4Unexpected errors and control failures logged2metBackend, signer, DNS and IPAM failures each log with outcome = "failure" and their own event name
16.4.1Logging components encode data to prevent log injection2metJSON mode escapes structurally. In text mode the only client-controlled fields are the request URI, which http::Uri renders percent-encoded, and header values, which HeaderValue::to_str accepts only as visible ASCII — so a User-Agent carrying a control byte is dropped before it can be stored, let alone printed
16.4.2Logs protected from unauthorized access and modification2metThe audit trail has no foreign keys, so deleting an account does not take its history; nothing in the panel can erase it — the audit surface is read-only and pruning is a host command (Audit Trail). The log stream itself is the operator’s to protect
16.4.3Logs transmitted to a logically separate system2partialThe server writes to stdout or a file in a format built for shipping, and Monitoring shows the pipeline — but shipping them is the deployment’s job, not this process’s
16.5.1Generic message to the consumer on unexpected errors2metEvery ACME refusal is an RFC 8555 problem document with a fixed type; internal detail goes to the log and not the body. The http-01 validator’s fetched body is never echoed into a client-visible error, precisely because it is attacker-chosen
16.5.2Secure operation when external resources fail2metA check that cannot reach its authority answers undecided rather than “allow”, so an IPAM outage degrades to a retryable 500 instead of failing open — the property Filters is built around
16.5.3Fail gracefully and securely; no fail-open2metStartup refuses rather than degrading: a non-loopback admin.bind_address without TLS, an unknown challenge type, a deadline below signer.custom.timeout_ms. tests/security.rs is the regression set for the request-path equivalents
16.5.4A last-resort handler for unhandled exceptions3metA CatchPanicLayer on each listener catches a handler panic, logs request_handler_panicked, and answers with that listener’s own error shape — an ACME problem document, or the admin JSON (/api) / HTML (/ui) error — instead of the aborted connection it used to be. panic = "abort" stays unset so the layer can unwind; the panic message goes to the log only

Documented deviations

Places where this project has knowingly chosen differently from what ASVS asks. Each was argued before this assessment existed; the assessment’s job is to surface them against the standard, not to reverse them.

http-01 validation does not block private addresses — V1.3.6, V13.2.4, V13.2.5. RFC 8555 requires following redirects, and Boulder’s mitigation — refusing RFC 1918 targets — cannot apply to a server whose entire purpose is serving private networks. What contains it instead: only http and https, only the two configured ports, at most max_redirects hops, a shared timeout, an off switch, and the fetched body never being echoed into a client-visible error. → HTTP-01

RS256 and RSA from 2048 bits — V11.2.3, V11.3.1. RFC 8555 §6.2 names RS256 as an algorithm an ACME server must accept, and RS256 is RSASSA-PKCS1-v1_5. Refusing it would refuse conforming clients. Two things soften it: this is a signature scheme, not the encryption padding V11.3.1 targets, and the key in question is a client’s own account key, whose compromise costs that client its account rather than costing the CA anything. Raising the accepted floor to 3072 bits is a protocol-compatibility decision, not a code change.

PBKDF2-HMAC-SHA256 rather than Argon2id — V11.4.2. Argon2id is the stronger primitive. Adopting it would add four crates to a certificate authority’s dependency graph — all audited on every cargo deny check, which runs with all-features = true — for a subsystem that is disabled by default and whose password is the bootstrap credential in a design that ends in a second factor. 600 000 iterations is OWASP’s current recommendation for the non-Argon2 case, and the stored form is self-describing so the trade can be revisited without a migration. The argument is in the module doc of crates/admin/src/admin/password.rs.

admin.require_mfa defaults to false — V6.3.3. Defaulting it on would brick a panel whose first operator has not enrolled yet, with no way in to fix it. The hardening checklist tells operators to turn it on, and the panel supports a bootstrap flow where enrolment is the only thing a session can do. An L2 claim for the web admin depends on the operator setting it. → Hardening

No hardware-based authentication factor — V6.3.3 at L3. WebAuthn was investigated and deferred, and both blocking checks were actually run: webauthn-rs 0.5.5 is MPL-2.0, which deny.toml’s allow list does not carry, and webauthn-rs-core hard-depends on openssl, which this tree has avoided at every turn. Nothing in the design precludes it — another factor kind is another MfaStep variant, not a change to the state machine. It stays open as issue #5.

The relay backend multiplexes one upstream account — V8.3.3. One upstream ACME account, and one centrally held RFC 2136 TSIG key, standing in for every local client. That is the whole point of the backend: not distributing a scarce credential is what it exists to do. Every authorization decision is made locally, before the upstream is ever asked. → Relay

Two listeners, usually two ports on one host — V3.5.4. They are separate sockets with separate TLS, separate authentication and separate defaults, and the admin one binds loopback unless TLS is on. They are not separate hostnames, and cookies are not port-scoped — which is exactly why the panel does not rely on SameSite for CSRF and carries a per-session token plus an Origin check instead. → Web Admin

The audit trail is a record, not a control — V8.2.4. Nothing in the server compares a live request against the trail. Pinning an identity to an address breaks CGNAT and mobile clients, and that is a deliberate non-feature. It is stated as such in the Security Model.

Gaps

Open shortfalls, worst first. What remains is all L3, recorded only here.

Lower-priority L3 items, recorded here only and with no issue open: no CSP violation-report endpoint (V3.4.7), no Cross-Origin-Opener-Policy (V3.4.8), no documented behaviour for browsers lacking security features (V3.1.1, V3.7.5), no OCSP stapling as a TLS server (V12.1.4), no Encrypted Client Hello (V12.1.5), no post-quantum migration plan (V11.1.4), no multi-user approval for issuance (V2.3.5), and admin user passwd letting the resetter learn the password (V6.4.6).

Re-running this

The requirement text is vendored at rfc/asvs-5.0/, so this page can be re-derived against a later ASVS release by diffing the chapter files and revisiting only the rows whose requirement text moved. The per-chapter tables enumerate every in-scope requirement rather than only the failures for exactly that reason: a list of gaps alone cannot be compared against anything.

Architecture

This page is about how the pieces fit, and about the handful of decisions that are load-bearing enough that changing them would break something non-obvious. The organising ideas are three: RFC 8555’s checks are hoisted into an extractor so no route can forget them, an ACME endpoint is a profile, and signing, filtering and notifying are each a trait with several implementations.

The workspace

One binary over a Cargo workspace. The root package, acme-proxy, holds the clap command tree (src/cli/), main.rs and the integration tests; nine library crates under crates/ hold everything else, each naming only the crates beneath it:

CrateWhat it holdsDepends on
acme-proxy-coreconfiguration, the ACME wire types, certificate parsing, the audit vocabulary—
acme-proxy-storethe storage layer over SQLite or PostgreSQL, one module per table, and both migration setscore
acme-proxy-netDNS, outbound HTTP and proxies, TLS, listeners, the challenge validatorscore
acme-proxy-policythe filter engine and the IPAM inventoriescore, net
acme-proxy-jobsthe job queue, notifications, the audit writer, metricscore, net, store
acme-proxy-signerthe signing backends and their read sidecore, jobs, net, store
acme-proxy-protocolthe ACME services, extractors, handlers and routersall of the above
acme-proxy-adminthe operation layer and the web admin panelprotocol and below
acme-proxy-serverthe runtime: roles, listeners, reload, loggingall of the above

The crate edges are the layering: a handler cannot reach the runtime that serves it, and a job queue cannot reach the filters, because the compiler refuses the import. tests/layering.rs pins each crate’s dependencies to this table, so an edge across a layer — which Cargo would accept — is a deliberate change rather than a drive-by one.

Request flow and extractors

Nearly every ACME endpoint is signed by the client using JSON Web Signatures (JWS). Rather than parsing this manually in each handler, the server leverages an AcmeRequest<T> Extractor.

The JWS extractor core

Eight checks run in a fixed order, and five of them have their own way out. The shape matters more than the list: the jwk/kid branch in the middle is where the two security properties below live, and a linear numbered list hides it.

graph TD
    REQ["Signed POST"] --> CT{"Content-Type is<br/>application/jose+json?"}
    CT -->|no| E415["415 — body never read,<br/>so no nonce is burned"]
    CT -->|yes| DEC["Decode the flattened JWS<br/>and its protected header"]
    DEC -->|unparsable| EMAL["malformed (400)"]
    DEC --> CRIT{"crit header present?"}
    CRIT -->|"yes — this server<br/>implements none"| EMAL
    CRIT -->|no| URL{"JWS url equals<br/>the route reached?"}
    URL -->|"no — §6.4"| EMAL
    URL -->|yes| AUTH{"jwk or kid?"}
    AUTH -->|"both, or neither"| EMAL
    AUTH -->|jwk| JWK["Re-encode the key as DER SPKI"]
    AUTH -->|kid| KID["Load the account, then check the<br/>stored SPKI's own OID against alg"]
    KID -->|"unknown kid"| EACC["accountDoesNotExist (400)"]
    KID -->|"OID does not match alg"| EALG["badSignatureAlgorithm (400)"]
    JWK --> SIG{"Signature verifies?<br/>ES256 or RS256, via ring"}
    KID --> SIG
    SIG -->|no| E401["unauthorized (401)"]
    SIG -->|yes| NONCE{"Nonce fresh and unused?"}
    NONCE -->|"no — §6.5"| EBAD["badNonce (400)"]
    NONCE -->|yes| H["Handler"]

Only once all of this succeeds is the request handed to the axum handler. Hoisting these checks into the extractor makes them structural: a new signed route cannot forget them, and no handler repeats a four-line preamble.

Three extractors build on that core: AcmeRequest<T> (decode and deserialize the payload), AcmePostAsGet (require an empty payload, else malformed), and AcmeOptionalPayload<T> — the last exists for the authorization resource, where one URL serves both a POST-as-GET read and a §7.5.2 deactivation.

Two security properties worth not breaking

jwk and kid are mutually exclusive (RFC 8555 §6.2) — the branch in the middle of the diagram. Both present, or neither, is malformed. An embedded jwk is verified and re-encoded as DER SPKI; a kid is resolved to its account and verified against the account’s stored SPKI.

The verification algorithm never rests on alg alone. On the kid path, the stored SPKI’s own AlgorithmIdentifier OID is checked against the client’s claimed alg before verification — that is the badSignatureAlgorithm exit. EC coordinates must additionally be exactly 32 octets (RFC 7518 §6.2.1.2) — a short or long one parses as a different point, which would register one key as two accounts.

Database & persistence

The server uses sqlx over SQLite or PostgreSQL, chosen by database.url’s scheme (ADR 0014). The SQL is written once, in crates/store/src/sql.rs, the only module that names either driver.

The connection pool is private to crates/store/src/. Everything else reaches the database through a table module, Database::transaction() (a Tx whose conn() is the connection the table methods take) or Database::pool_stats() (the metrics gauge), so SQL and its dialect stay in one module tree. Database::raw_pool() exists only for test fixtures, and tests/layering.rs fails the build when production code calls it.

Migrations

Migrations are embedded with sqlx::migrate!(), frozen once committed, and applied only by acme-proxy migrate/init and a process running the worker role; every other entry point refuses a database whose schema is behind. The rules and the reasons are in ADR 0003, and the how-to in Contributing. The schema is also the only frozen surface before 1.0.0 (ADR 0001).

Schema details

The tables, their constraints and the reasoning behind each — profile isolation, the CHECKed state machines, the audit trail’s deliberate lack of foreign keys, and the three different ways a secret is stored — have their own page: Database Schema.

Revocation is the one piece worth naming here, because it constrains the request path rather than the schema: it writes a reason and a timestamp and deliberately does not touch the order status, since RFC 8555 defines no “revoked” order status.

Two front ends, one operation layer

src/cli/ and crates/admin/src/webadmin/ are two front ends; crates/admin/src/admin/ is the operation layer both dispatch to and neither owns.

src/cli/            crates/admin/src/webadmin/
   (clap)              (axum)
      \                 /
       \               /
        crates/admin/src/admin/ops.rs      — delete_account, revoke_order, load_order_detail…
        crates/admin/src/admin/users.rs    — create_user, authenticate, set_password…
        crates/admin/src/admin/changes.rs  — an operator change: write, audit row, sessions, notification
        crates/admin/src/admin/render.rs   — render_*_json, the one JSON shape of each object
        crates/admin/src/admin/password.rs — the KDF, shared by both

A handler in crates/admin/src/webadmin/handlers/ is a few lines over an admin::ops call and a render_*_json, the same way a src/cli/ command body is a few lines over the same call and either the same render_*_json or a render_*_line of its own, in src/cli/render.rs. That is what keeps the password policy, the duplicate check and the rehash-on-login identical between them.

Two consequences worth knowing:

  • The destructive operations come in pairs. delete_account(id, db) simply deletes; confirm_delete_account(id, yes, reader, db) asks first. The split exists because assume_yes: bool + reader: &mut impl BufRead are a terminal’s concerns — a caller with no terminal was passing true and an empty reader, asserting a confirmation that never happened. The CLI calls the wrapper; the web calls the bare form.
  • crates/admin/src/webadmin/ is not crates/admin/src/admin/web/. That would invert the dependency, putting an HTTP server inside the operation layer.

The admin listener is assembled by webadmin::build_admin_app, which takes &[Arc<Profile>] as a slice — build_app consumes the Vec, so the admin side must be built first, and the signature is where that ordering is stated rather than a borrow error to rediscover. Its state is AdminState, not AppState: the latter holds exactly one Profile, and this listener is cross-profile by nature (revoking an order needs that order’s own profile’s revocation route, which may name a different CA from any default).

Order lifecycle

sequenceDiagram
    participant Client
    participant Axum Router
    participant Filters
    participant Order Manager
    participant Job Queue
    participant Challenge Validator
    participant Signer Backend

    Client->>Axum Router: POST /newOrder
    Axum Router->>Filters: Validate Client IP & Identifiers
    Filters-->>Axum Router: Allow/Deny
    Axum Router->>Order Manager: Create Order + Authorizations + Challenges
    Order Manager-->>Axum Router: Order Object (status: pending)
    Axum Router-->>Client: 201 Created

    Client->>Axum Router: POST /chall/{id} (trigger)
    Axum Router->>Job Queue: claim + enqueue challenge_validate
    Axum Router-->>Client: 200 OK + challenge object (processing)
    Job Queue->>Challenge Validator: Validate domain control
    Challenge Validator-->>Job Queue: Pass/Fail
    Job Queue->>Order Manager: Commit challenge + authz + order
    Note over Order Manager: Order -> "ready" once every<br/>authorization is valid
    Client->>Axum Router: POST /chall/{id} (poll)
    Axum Router-->>Client: 200 OK + challenge object (either way)
    Note over Client,Challenge Validator: This exchange in detail:<br/>Challenge Validation

    Client->>Axum Router: POST /finalize (with CSR)
    Axum Router->>Filters: Re-validate identifiers from CSR
    Filters-->>Axum Router: Allow/Deny
    Axum Router->>Job Queue: claim + enqueue signer_issue
    Axum Router-->>Client: 200 OK + order (processing)
    Job Queue->>Signer Backend: Request Signature (worker only)
    Signer Backend-->>Job Queue: Signed Certificate
    Job Queue->>Order Manager: Record certificate (order -> valid)
    Client->>Axum Router: POST-as-GET order (poll)
    Axum Router-->>Client: 200 OK (Certificate URL)

Four properties hold this together. Three are transactional:

  • Order creation inserts the order, its authorizations and their challenges in one transaction. A half-written order would be finalizable for names that were never authorized.
  • A validation outcome commits the challenge, the authorization and the order together, and the “is every authorization valid?” read happens inside that transaction. From the pool, two concurrent validations of one order could each read before the other’s write landed, and neither would promote the order to ready. Every one of those writes is also guarded on the state it leaves, since the verdict is computed by a job long after the request that claimed the challenge read those rows.
  • finalize claims the order (ready → processing) and queues its signer_issue job in one transaction, so no crash can leave an order processing with nothing coming to settle it. The job runs in the worker role — the only one that builds a signing backend — which is what keeps the CA key out of the process parsing client requests.

The fourth is about the answer rather than the write: challenge validation returns 200 plus the challenge object whether it passed or failed (§7.5.1). A 4xx would surface as a transport failure to certbot’s acme library rather than as a failed challenge.

Pluggable signing keys

The SignerBackend trait is the seam for how a certificate is obtained — locally, from an upstream ACME server, or from a script. Inside the local_ca backend there is a second, narrower seam for where the private key lives, and it is worth knowing that it is rcgen’s own trait, not one this project invented.

graph LR
    ISSUE["LocalCa::issue<br/>LocalCa::revoke"] --> SB["spawn_blocking<br/>— unconditionally"]
    SB --> ISS["Issuer&lt;'static, CaSigningKey&gt;"]
    ISS --> SW["Software(KeyPair)<br/>a PEM file on disk"]
    ISS --> PK["Pkcs11(...)<br/>behind --features hsm"]
    PK --> MOD["the PKCS#11 module (.so)"]
    MOD --> TOK[("token / HSM")]

The spawn_blocking sits before the branch, not inside one arm of it: see the first consequence below.

rcgen::SigningKey is public, and every signing entry point local_ca uses is generic over it:

callsignature
csr.signed_by(&issuer)signed_by(&self, issuer: &Issuer<impl SigningKey>)
CertificateRevocationListParams::signed_by(&issuer)signed_by(&self, issuer: &Issuer<'_, impl SigningKey>)
params.self_signed(&key)self_signed(&self, signing_key: &impl SigningKey)

So LocalCa holds an Issuer<'static, CaSigningKey> — a small enum in crates/signer/src/local_ca/key.rs with a Software(KeyPair) variant and, behind the hsm feature, a Pkcs11(..) one. Adding a key source (a cloud KMS, a remote signer daemon) means adding a variant that implements two rcgen trait methods: sign, der_bytes/algorithm. Nothing in issue, revoke, crl_der or the CSR sanitisation changes, because none of it ever names the key type.

Two consequences worth preserving:

  • Signing runs on the blocking pool. rcgen::SigningKey::sign is synchronous and called from deep inside signed_by, so there is nothing to await through. A key source that talks to hardware or a network would otherwise stall a runtime worker for its whole round trip, so LocalCa::issue and the CRL rebuild in revoke both go through spawn_blocking unconditionally — not gated on the variant, which would be a branch someone eventually gets wrong.
  • The key type must be Send + Sync. It lives inside an Arc<dyn SignerBackend>. Where the underlying handle is not (cryptoki’s Session is Send but not Sync), a std::sync::Mutex is the right wrapper: the signing call never awaits, so an async mutex would buy nothing.

Architecture Decision Records

An architecture decision record (ADR) holds the why behind a design: the problem it answers, the choice made, and what that choice rules out. The rest of the documentation says what the server does. The module docs (//!) say what each piece of code guarantees. The ADRs are where the argument lives, so that the argument is not restated wherever the decision is mentioned.

Read the relevant ADR before changing something it covers. If a change reverses a decision, write a new ADR that supersedes the old one and change the old one’s status. Do not rewrite its argument after the fact.

ADRDecisionStatus
0001Before 1.0.0, only the database schema is a compatibility promiseAccepted
0002One binary over a layered workspace of lockstep cratesAccepted
0003Migrations are append-only and applied only by the schema ownersAccepted
0004Row ids are UUID v7 stored as BLOBs, and their type says where they came fromAccepted
0005A Rust enum owns each vocabulary, and SQL checks only the closed onesAccepted
0006A request does no slow or privileged work; it queues itAccepted
0007One binary runs as role processes, and only the worker holds the CA keyAccepted
0008State that more than one process can see lives in the databaseAccepted. One exception stands, the web admin’s login limiter, for as long as
0009Dependencies are pure Rust on ring, add no global state, and earn their placeAccepted; metrics clause superseded by 0011, crypto clause narrowed by 0014
0010Errors derive thiserror, carry their whole message, and panic only at startupAccepted. This was re-argued more than once before it was written down, which
0011Metrics are built on prometheus-client, one registry per scrapeAccepted
0012Images are built natively per architecture, uncached, and published only past a guardAccepted
0013main is the trunk, and a patch line is a release branch cut when a fix needs oneAccepted
0014PostgreSQL is chosen by the URL’s scheme, over one set of queriesAccepted

Decisions argued elsewhere

Some decisions are explained where their subject lives, and have no record here, because a second copy would drift from the first.

In the book:

In a module’s own documentation (//!), where the decision concerns that module alone:

  • crates/jobs/src/jobs/mod.rs: why background work is a durable queue and not a tokio::spawn, and the Retry/Failed split every handler must honour.
  • crates/signer/src/local_ca/crl.rs: how the CRL number stays monotonic across processes.
  • crates/policy/src/filter/policy.rs: Kleene logic, and why a rule’s stages are an intersection.
  • crates/policy/src/ipam/mod.rs: why an inventory never denies, so an outage cannot fail open.
  • crates/core/src/script_hook.rs: the one contract every custom script runs under.
  • crates/core/src/templating.rs: .html escapes and .j2 does not, decided by the name.
  • crates/jobs/src/metrics.rs: bounded cardinality, families declared even when empty, and counters that survive a reload.
  • crates/jobs/src/notify/expiry.rs: why the expiry notice is a digest.
  • crates/server/src/reload.rs: the swap’s mechanics and what stays frozen.

Writing an ADR

Name the file NNNN-kebab-title.md, taking the next free number, and list it both in the table above and in SUMMARY.md under this page. doc/lint.py refuses an ADR that is missing from either list, or that lacks one of the five sections below.

The title is # ADR NNNN: <decision>, and the sections come in this order:

  • ## Status: Accepted, or Superseded by ADR NNNN. Name a proposal to replace it if one is on record.
  • ## Context: the problem, and the history that makes the rule necessary. History with no rule behind it does not belong here — git and CHANGELOG.md already keep it.
  • ## Decision: what was chosen, stated as rules.
  • ## Consequences: what the decision costs and what it rules out.
  • ## Enforced by: the tests, startup refusals and type-level constructs that hold the decision in place, named so that a grep finds them. Write “Review only” when nothing does.

Link to the page that owns a configuration key rather than restating its default, since the book documents each key in exactly one file.

ADR 0001: Before 1.0.0, only the database schema is a compatibility promise

Status

Accepted.

Context

A project that promises stability on every surface from its first release ends up carrying a compatibility layer for every shape it has ever had: aliases for renamed keys, dual syntaxes, lowering passes that translate an old section into a new one. Each of these is code that must be tested, documented and reasoned about for as long as the promise lasts, and each makes the next redesign harder.

acme-proxy is still finding its design. Several subsystems were replaced wholesale rather than extended: the filter chain became a policy engine (Policy), the Mattermost notifier became a generic webhook (Webhook), the acme_proxy signer became relay, and filter.netbox became [ipam]. None of these would have been worth doing if each had to keep reading the old shape.

One surface cannot be treated this way. The database holds accounts, orders and issued certificates that clients and relying parties depend on, and sqlx checksums every applied migration. An upgrade that cannot open the existing database is not an upgrade.

Decision

Before 1.0.0 the database schema is the only compatibility guarantee: crates/store/migrations/ is append-only (ADR 0003), so an upgrade is starting the new binary against the existing database.

Everything else may be renamed or removed in any release: configuration keys, profile names and the ACME URLs derived from them, the admin JSON API, log event names and the CLI. Such a change owes exactly three things:

  • An entry under the release’s ### Breaking heading in CHANGELOG.md, naming the old spelling and the new one. The changelog’s Compatibility section is the canonical statement of this policy; the README, SECURITY.md and the book each carry one sentence linking to it.
  • A startup refusal naming the removed key wherever practical, so an unmigrated configuration stops the server instead of coming up looking configured. This is a one-line error message, not a compatibility path.
  • Never an alias, a dual syntax or a legacy lowering. The old shape is deleted; the new code does not learn to read it.

The refusals themselves are removed at 1.0.0.

Consequences

  • A redesign costs one changelog entry and one diagnostic, which is what makes replacing a subsystem cheaper than accreting around it.
  • A key must still parse to be refused by name. The removed [filter] fields therefore survive in crates/core/src/config/types/filter.rs; a field that is gone would fail as an opaque serde error instead of a named one.
  • Operators must read the Breaking section before an upgrade. acme-proxy filter show builds a policy exactly as startup does, so it checks a migrated configuration before a restart.
  • Frozen means frozen: the comments inside committed migrations still say acme_proxy where the code says relay, because editing them would change their checksum (ADR 0003).

The how-to for a rename is in Contributing.

Enforced by

  • refuse_removed_keys in crates/policy/src/filter/build.rs, guarded by every_removed_key_is_refused_by_name.
  • the_old_acme_proxy_backend_name_is_refused_by_its_new_one (crates/signer/src/lib.rs) and the_removed_mattermost_backend_is_refused_by_name (crates/jobs/src/notify/mod.rs).
  • the_netbox_type_is_refused_by_name (crates/policy/src/filter/build.rs).
  • The Breaking entry and the “no alias” rule are review only.

ADR 0002: One binary over a layered workspace of lockstep crates

Status

Accepted.

Context

The server was one crate of about 106,000 lines. Every change recompiled all of it, and nothing but review stopped the web admin from issuing SQL or the notifier from reaching into a signer. Layering held by convention, and a convention is broken by the first drive-by import that compiles.

The code was not meant as a library for anyone else. It exists so the binary and the integration tests can reach it. That shapes what a split has to buy: compiler-enforced edges and faster rebuilds, not a public API.

Decision

  • One binary, nine library crates. The acme-proxy package at the repository root holds the clap command tree (src/cli/), main.rs and every suite in tests/. The libraries under crates/ are, bottom-up: core, store, net, policy, jobs, signer, protocol, admin, server. The crate map is in Architecture.
  • A crate names only crates beneath it. Cargo refuses a cycle but accepts an edge that skips across a layer, so the intended edges are pinned in a table (CRATE_DEPS in tests/layering.rs) and a test compares every member’s [dependencies] against it. Adding an edge means editing that table on purpose.
  • Lockstep, with no semver promise. Every member takes version.workspace, and every internal dependency is pinned with = in [workspace.dependencies]. An item becomes pub only because another crate needs it, not because it is an API.
  • Test scaffolding lives per crate, behind a test-util feature (#[cfg(any(test, feature = "test-util"))] pub mod testutil). Only other crates’ [dev-dependencies] turn it on, so no normal build ships it. A fixture belongs to the lowest crate its types allow.
  • Every cargo command takes --workspace. At a root that is also a package, a bare command acts on that package alone, and the library crates would drop out of lint, tests and coverage with nothing going red.

Consequences

  • A handler cannot reach the runtime that serves it, and the job queue cannot reach the filters. The compiler refuses the import.
  • A doc link cannot point up the graph, since a crate cannot link to its dependants. Such a mention is a plain code span instead.
  • A member’s unit tests run in the member’s own directory. A test that reads a repository file anchors on env!("CARGO_MANIFEST_DIR").
  • tracing targets are the crate paths (acme_proxy_store::db, …). The default filter acme_proxy=info still covers them all because EnvFilter matches targets by string prefix.
  • All ten packages are published together. A release bumps the workspace version once, and every = pin with it.

Enforced by

  • Cargo, for cycles.
  • crate_dependencies_follow_the_layers in tests/layering.rs, for edges across a layer.
  • The CI jobs, each of which passes --workspace (.github/workflows/ci.yml).

ADR 0003: Migrations are append-only and applied only by the schema owners

Status

Accepted.

Context

The schema is the one surface this project promises not to break (ADR 0001), and sqlx enforces that promise mechanically: it records a checksum for every migration it applies. Editing a committed file does not quietly diverge a deployment’s schema. It makes the next startup fail with a checksum mismatch. Before the first release, a schema change meant editing the migration in place and deleting sqlite.db. That stopped being possible once a database outside the repository had run the files.

SQLite adds its own constraints:

  • It cannot add a CHECK, UNIQUE or foreign key to an existing table.
  • It gives VARCHAR(n) TEXT affinity and enforces no length.
  • Under foreign_keys = ON, DROP TABLE performs an implicit DELETE FROM, which fires ON DELETE CASCADE into every child table.
  • It gives sqlx no migration lock.

Migrations used to run inside Database::connect. That made every subcommand an upgrade step. Once the server could run as several processes (ADR 0007), two processes starting together raced MIGRATOR::run with nothing to serialise them.

Decision

Append-only. Every file in crates/store/migrations/ is frozen, comments included — and, since ADR 0014, so is every file in crates/store/migrations-postgres/. A schema change is a new file in each set (sqlx migrate add <name>). The rules below describe SQLite’s constraints; PostgreSQL needs no rebuild for a CHECK or a width, so its file is usually the one-line ALTER TABLE the change actually is:

  • A new column is a new ALTER TABLE … ADD COLUMN file, even when it plainly belongs to an existing table.
  • A new or dropped CHECK, UNIQUE or foreign key is a table rebuild, written in a new migration.
  • A wrong declared width is also a rebuild. Where a width follows a constant in the code, a test pins the two together.
  • A rebuild re-creates the table’s indexes, since DROP TABLE takes them with it. It also names every column in its INSERT … SELECT, since a forgotten column is dropped silently. Rows in a table with children are parked in constraint-free CREATE TABLE … AS SELECT copies before anything is dropped, then put back parent-first. Otherwise the cascade empties the children.

Applied explicitly. Database::open connects and never migrates. Database::migrate has two callers:

  • acme-proxy migrate and acme-proxy init;
  • a serve process that runs the worker role (server::apply_or_require_schema).

Every other entry point calls pending_migrations and refuses by name, pointing at acme-proxy migrate. A default single-process serve includes the worker role, so it still migrates a fresh database on its own.

The pool is private to crates/store/. Everything else goes through a table module, Database::transaction() (a Tx whose conn() hands out the connection) or Database::pool_stats(). Database::raw_pool() exists for test fixtures only. On SQLite the connection pins two pragmas:

  • foreign_keys(true), because the schema’s ON DELETE CASCADE depends on it;
  • journal_mode = WAL, because every ACME response writes a nonce row, and the default rollback journal takes a database-wide exclusive lock on each write.

PostgreSQL needs neither — foreign keys are always enforced and there is no journal mode to choose — so its pool pins a connection limit instead, which is the resource several role processes there actually share.

Runtime queries. Queries use sqlx::query(...), not the compile-time query! macros, so building needs no DATABASE_URL.

Consequences

  • An upgrade is starting the new binary. There is no dump/restore procedure.
  • Several constraints were declared before anything wrote to them: admin_users.totp_* and admin_sessions.state’s 'pending_mfa'. Adding them afterwards would have cost a rebuild each.
  • Stale comments stay stale. Three migrations still call the relay backend acme_proxy. Treat grep results in crates/store/migrations/ as read-only.
  • sqlx::migrate!() embeds the migration set at compile time. Adding a file does not invalidate the build on its own; touch crates/store/src/db.rs.
  • SQL, and the dialect it is written in, lives in one crate. That is what kept a second backend a contained change when PostgreSQL arrived in ADR 0014.
  • PostgreSQL gives sqlx an advisory migration lock where SQLite gives it none, so the race above cannot happen there. The one-owner rule still holds on both: it is about which process may own the schema, which is a deployment property, not only about the race that exposed it.

Enforced by

  • sqlx’s checksums, at startup.
  • only_the_schema_owners_apply_migrations and production_code_never_reaches_the_raw_pool in tests/layering.rs.
  • Rebuild guards in crates/store/src/db.rs:
    • the_blob_migration_preserves_every_row;
    • the_audit_log_rebuild_keeps_every_row_and_relaxes_the_event_check.
  • Width pins in the same file:
    • declared_token_widths_match_random_token;
    • declared_issuer_widths_match_the_issuer_id;
    • every_id_column_is_declared_a_blob.

ADR 0004: Row ids are UUID v7 stored as BLOBs, and their type says where they came from

Status

Accepted.

Context

Row ids started as UUID v4 strings. Two costs followed:

  • A random id is written into an index at a random leaf on every insert. SQLite feels that mildly, and PostgreSQL would pay a page split and a full-page WAL write per row. That backend is issue #4.
  • The paged listings break ties on a whole-second created_at with ORDER BY created_at, id. With v4 ids, rows created in the same second come back in a different random order for each pair.

Converting ids from text to BLOBs added a third problem. SQLite never compares a bound String equal to a BLOB, so an internal lookup that kept a &str parameter would match nothing, silently, for ever. The case that exposed this was a job-retention test. It inserted a job with a hand-written id and asserted it was gone afterwards. Had Job::find_by_id taken a &str and parsed it internally, that assertion would have passed whether or not the sweep had run.

Decision

  • Every id is minted in one place (mint() in crates/store/src/id.rs), as a UUID version 7. The leading 48-bit millisecond timestamp makes ids created close together share a prefix, and makes them sort by creation.

  • Ids are stored as the 16 bytes themselves. sqlx’s uuid feature encodes a Uuid as a SQLite BLOB, and maps the same type to PostgreSQL’s native uuid.

  • Existing rows keep their v4 ids. An id is a foreign key, a kid is a credential a client stored, and an order id is part of a URL a client polls. A table can therefore hold both versions, and only the v7 ids sort by creation.

  • The parameter type records provenance. In crates/store/src/:

    • an id typed &str may be junk from outside the process, so it is parsed and an unparseable value answers “absent”;
    • an id typed Uuid was read out of a row.

    There is no third case, and no parallel _uuid variant of any function.

  • parse() is deliberately narrower than Uuid::try_parse. It accepts only the 36-character hyphenated form that Uuid::to_string() produces, so an id written in another spelling answers “not found” rather than resolving.

  • Columns that look like ids but do not point at a row stay text. The list, and how to query a BLOB id by hand, are in Database Schema.

Consequences

  • An internal lookup handed a stale string fixture is a compile error rather than a silently empty result.
  • Only a few functions keep &str, where the value genuinely arrives from outside the process (the request path and job payloads). They are listed in crates/store/src/id.rs.
  • The conversion migration (20260827120000_uuid_ids_as_blobs.sql) is the worked example of the DROP-cascade hazard in ADR 0003.
  • Some fresh values are deliberately not row ids and do not go through mint(): the x-request-id fallback, the job lease owner and a notification’s delivery_id.

Enforced by

  • The type system, for &str versus Uuid.
  • every_id_column_is_declared_a_blob and the_blob_migration_preserves_every_row (crates/store/src/db.rs).
  • The unit tests of parse in crates/store/src/id.rs.

ADR 0005: A Rust enum owns each vocabulary, and SQL checks only the closed ones

Status

Accepted.

Context

Many columns hold a word from a fixed list: an order’s status, an audit event, an operator’s role, a job’s kind. There are two places the list can be enforced: a SQL CHECK constraint, or a Rust type.

A CHECK catches a typo before it parks a row in a state nothing can reach. But SQLite cannot alter one without rebuilding the table (ADR 0003). Once a vocabulary grows, its constraint becomes a migration per new word, and a rolling upgrade breaks: an older binary refuses to write a word only a newer one knows.

String literals in Rust are worse than either. Before crates/store/src/status.rs existed, about 30 comparisons against literals were spread across handlers, storage and the relay flow. A typo compiled, and silently changed policy.

Decision

  • Every vocabulary is a Rust enum, and the enum is the authority. Values are written only through its as_str(), and read back through a parse that handles an unknown word explicitly.
  • Closed vocabularies also keep a CHECK. Every status column (accounts, orders, authorizations, challenges, jobs, upstream_orders, eab_keys, admin_users) changes only with its state machine. The enums in status.rs reproduce their columns’ existing strings byte for byte, so the frozen constraints stay valid.
  • Open vocabularies carry no CHECK. These are vocabularies that grow with features:
    • audit_log.event. 20260909120000 rebuilt the table to drop that CHECK when the administrative actions widened it. AuditEvent is the authority, and AuditEvent::outcome is an exhaustive match, so a new name must classify itself as a success or a failure.
    • admin_users.role. NULL reads as admin, so an operator created before the column existed keeps full authority. Any other unknown value folds to viewer, the least privilege.
    • jobs.kind, which is an open set by design.
  • Operator surfaces refuse an unknown value by name (Admin CLI) wherever the vocabulary is closed, since “no rows” reads exactly like “nothing is in that state”. --kind is the exception, because kinds are an open set.
  • Job::status and UpstreamOrder::status stay String in their models. An older binary must still render a row a newer one wrote. Their enums exist for the operator surface.

Consequences

  • A new audit event or role is a Rust change and a changelog line, with no migration.
  • A value written by hand into the database is handled by the parse rule, never trusted: a garbage role reads as viewer.
  • The database alone no longer guarantees that audit_log.event holds a known word. The insert path binds AuditEvent::as_str(), never a free string.

Enforced by

  • The CHECK (status IN (…)) constraints in the migrations.
  • EVENT_COUNT and ALL_AUDIT_EVENTS in crates/core/src/audit/mod.rs, a compile-time assertion that no variant is missing from the list.
  • AdminRole::from_storage in crates/store/src/admin_user.rs.
  • the_audit_log_rebuild_keeps_every_row_and_relaxes_the_event_check (crates/store/src/db.rs).

ADR 0006: A request does no slow or privileged work; it queues it

Status

Accepted.

Context

Three ACME operations used to do their real work inside the HTTP request:

  • POST /chall/{id} awaited the challenge validation. That reached out to an address the client named, over DNS, HTTP or TLS, to a host that may never answer. It held an admission permit for the whole of challenge.timeout_ms. Up to server.max_concurrent_requests clients pointing at black-holed addresses could pin every permit. It also forced a startup rule that the request timeout exceed the validation timeout.
  • finalize called the signing backend. The process parsing untrusted JWS and CSRs from the internet therefore held ca.key, a PKCS#11 login or a relay’s upstream account.
  • Revocation called the backend too, from POST /revokeCert, from the admin panel, and from order revoke on the host. Run beside a live server, the last of these could silently drop a revocation from the CRL.

RFC 8555 already has the states that queued work needs. A challenge is processing while “the server is working on it” (§7.1.6), an order is processing while “the certificate is being issued” (§7.4), and §8.2 pairs both with Retry-After. certbot, acme.sh and lego all poll.

Decision

A request claims, queues and answers. The worker does the work.

  • Validation. POST /chall/{id} claims the challenge (pending → processing, a compare-and-swap), writes a challenge_validate job, and answers 200 with the challenge reading processing plus Retry-After. See crates/protocol/src/acme/validate.rs.
  • Issuance. finalize checks the CSR and runs the filter synchronously. It then claims the order and queues signer_issue in one transaction, and answers processing for every backend. See crates/protocol/src/acme/issue.rs.
  • Revocation. A local CA’s revocation is a revocations row and the order’s stamp in one transaction; the worker then signs the CRL. A relay or custom revocation is a signer_revoke job that the request waits on, for up to server.request_timeout_ms less a second, answering 503 + Retry-After if the job is still running. See crates/protocol/src/acme/revoke.rs.
  • A verdict is terminal. A check that ran records its answer, pass or fail. A job is retried only when the attempt could not happen at all: the database is unreachable, or this process does not mount the profile. A backend’s own BadCsr makes the order invalid, not ready again, because the client was already told processing and polls for a terminal state.
  • No stranded claim. A failed enqueue releases its claim. abandon records a failure rather than leaving the client polling forever. recover re-queues a challenge left processing with no live job.
  • Each kind has one handler, covering every profile. A row names its subject, the subject names its profile, and the profile names its validators or backend. The job registry refuses a second handler for a kind anyway.

Consequences

  • challenge.timeout_ms bounds a job attempt, not an HTTP request. The only timeout rule left at startup concerns signer.custom’s read hooks, the one backend call still made inline.
  • A client sees processing on every finalize, including against a local CA. This shipped under ### Breaking.
  • The request path never names a signing backend, which is what lets the CA key leave the ACME process entirely (ADR 0007).
  • A request no longer holds a permit during outbound I/O to a host the client chose.
  • The integration harness runs a real worker. A test that triggers or finalizes must poll for the outcome rather than read it from the response.

Enforced by

  • the_request_path_never_holds_a_signer and the_cli_never_builds_a_signer in tests/layering.rs.
  • check_request_timeout, the remaining startup refusal.
  • tests/challenges.rs, tests/orders.rs and tests/revoke_cert.rs, which drive each operation through the queue.

ADR 0007: One binary runs as role processes, and only the worker holds the CA key

Status

Accepted.

Context

A certificate authority has three very different kinds of work, and they carry different risks:

  • Serving ACME means parsing untrusted JWS and CSRs from the internet.
  • Serving the web admin means holding operator sessions that can revoke certificates and mint EAB credentials.
  • Doing background work means dialling hosts the clients chose, talking to an upstream CA, sending mail, and signing with the CA key.

In one process, a bug in the first kind of work reaches the key used by the third. The obvious answer, separate binaries, would cost a second configuration surface and a second release artefact. It would also create a protocol between the binaries, where one database and one job queue already do the job.

Decision

  • One binary, one configuration, several processes. acme-proxy serve --role acme,admin,worker selects what a process does. With no --role, a process does all three, exactly as before the flag existed.

  • The three roles:

    • acme serves ACME and the root router;
    • admin serves the panel;
    • worker drains the job queue and owns the schema (ADR 0003) and the first-run material.

    Each role only enqueues work it does not do itself. Every role may serve its own /metrics.

  • Only the worker builds a signing backend. The signer is split in two (crates/signer/):

    • SignerInfo, the read side: the CA certificate, the stored CRL, a relay’s lazily discovered directory and http-01 store, and a custom script’s read hooks. Every role builds it, from public material only.
    • SignerBackend, the write side: issue and revoke. Only the worker builds it, in Assembly::build_parts.

    Profile has no backend field at all. The job handlers take their backend from GenerationParts::signers.

  • The worker stores each CA’s first CRL before it serves (store_first_crls, at startup and before publishing a reload), since the read side never signs.

  • Roles are a flag, not a configuration key, so they cannot change under SIGHUP. That is what lets a reload compare the same role set on both sides.

  • ProcessRole is not sockets::Role. The first names the three jobs a process does. The second names the three listeners it holds (acme, admin, metrics). worker holds no socket, and metrics is a socket any role may serve, so neither set fits inside the other.

Consequences

  • An acme or admin process never reads ca.key, never logs in to a token, and never registers upstream. It starts even when it cannot read the key.
  • An acme or admin process that finds no CA certificate refuses to start, naming acme-proxy init, because the first-run material belongs to the worker.
  • Without a worker in the same process, queued work waits for one elsewhere. The process logs the advisory server_role_no_worker. A process that does not own the schema and finds it behind refuses with server_schema_behind.
  • The wake-up is in-process. A worker in another process picks up a new row within jobs.poll_interval_ms, whereas the claim itself is race-free across processes.
  • Counters live per process, so each process needs its own metrics.bind_address. A role label is added when the metrics are rendered.
  • --role is parsed by clap, so an unknown name is refused before Config::load and before Database::open, which creates the database file.
  • State that several processes share cannot live in memory or in a file one of them owns (ADR 0008).

The supported topologies are in Deployment.

Enforced by

  • the_request_path_never_holds_a_signer and the_cli_never_builds_a_signer (tests/layering.rs).
  • tests/roles.rs, which starts acme and admin processes against a key_path that does not exist and drives an order to valid across three processes over one file-backed database.
  • an_unknown_role_is_refused_by_name (crates/server/src/roles.rs).
  • info_from_config_agrees_with_the_backends_own_info (crates/signer/src/lib.rs).

ADR 0008: State that more than one process can see lives in the database

Status

Accepted. One exception stands, the web admin’s login limiter, for as long as one admin process is the supported topology.

Context

Several pieces of state used to live in process memory or in files beside the binary. Each was correct only while exactly one process existed:

  • A local CA’s revocations. They lived in an in-memory ledger plus a JSON sidecar and a CRL file, and GET /crl served the in-memory DER. Each process rewrote both files at startup, and order revoke on the host wrote them beside a running server. Updates were lost, crl_number was duplicated, and the server served a stale CRL. A revocation could silently vanish from the CRL.
  • The relay’s published http-01 key authorizations, held in an in-memory token store. The relay job publishes them, and the root router serves them, which may be in another process.
  • Reload. Both of the above had to be handed from the outgoing configuration generation to the incoming one through an in-memory handover (CarriedState). Otherwise a reload would empty the CRL ledger or the token store under a live upstream fetch. Changing the profile set, [signer], [dns] or [proxy] was therefore refused on SIGHUP.
  • Notification delivery. A delivery queued by one process for a profile another process had not yet reloaded was retired as Failed for good.

Role processes (ADR 0007) turn every one of these from a latent bug into a routine one.

Decision

  • Revocation state is two tables, keyed on the issuer: the hex SHA-256 of the CA certificate’s SPKI, so two profiles over one CA are one issuer.
    • revocations holds one row per serial. A repeat does nothing, so the first revocation’s time and reason stand.
    • crls holds the current CRL for each issuer. It is replaced only through a compare-and-swap on crl_number (StoredCrl::replace_if_number), which keeps the number monotonic across processes.
    • GET /crl serves the stored row.
    • The old sidecar is imported once, in the same transaction as the first crls row, and never written again.
    • signer.local_ca.crl_path becomes an export that nothing reads back.
  • http01_tokens holds the relay’s key authorizations, keyed on the upstream’s token and reaped by an hourly sweep.
  • A backend whose configuration did not change is reused on reload; one whose configuration changed is rebuilt. The outgoing and incoming instances share the tables, so there is nothing to hand over. CarriedState is gone.
  • An unknown profile or backend is a bounded Retry, not Failed, so a delivery survives a rolling reload across processes.
  • Files stay files where that is a security property. ca.key (or its PKCS#11 token) and the upstream account key remain files. The database is readable by every role and by every backup, whereas a key file can be readable only by the worker’s uid.

Consequences

  • order revoke beside a running serve appears in that server’s GET /crl once its worker has signed, with no restart and no lost update.
  • The profile set, each profile’s signer, [dns] and [proxy] all reload. Of everything in the configuration, only database.url is still refused on SIGHUP.
  • crl_number never goes backwards (RFC 5280 §5.2.3). A client that meets a lower number than it has cached keeps the cached CRL, which means it keeps trusting what was revoked.
  • admin.login_max_attempts is still counted in memory, per process. A second admin process would get its own brute-force budget, which is why one admin process is the supported topology.

The CRL store’s concurrency rules are in crates/signer/src/local_ca/crl.rs.

Enforced by

  • concurrent_revocations_through_two_instances_are_all_kept and two_cas_over_one_database_keep_separate_crls (crates/signer/src/local_ca/mod.rs).
  • tests/crl.rs and tests/roles.rs.
  • tests/http01_responder.rs.
  • an_unknown_profile_or_backend_is_retried_within_its_budget (crates/jobs/src/notify/job.rs).

ADR 0009: Dependencies are pure Rust on ring, add no global state, and earn their place

Status

Accepted. The metrics clause is superseded by ADR 0011: latency histograms were the case where a metrics library would earn its place, and they arrived. The crypto clause is narrowed by ADR 0014: it governs the crypto this project calls, not what a driver carries for a wire protocol of its own.

Context

A certificate authority’s dependency graph is part of its attack surface, and this one is audited by cargo deny at all-features = true and inventoried in a committed SBOM. Every crate added is one more crate to audit, one more licence to allow, and one more release to track.

Two kinds of cost are easy to miss when choosing a dependency:

  • A second crypto stack or a C toolchain. aws-lc-rs or OpenSSL makes the build depend on a C compiler, and it means two implementations of the same primitives to keep patched.
  • Process-global state. A library that installs a global (a rustls CryptoProvider::install_default, a metrics recorder, a global HTTP client) makes behaviour depend on which code ran first. In a test suite that means two tests sharing counters or a provider, with the outcome depending on test order.

Decision

  • ring is the only crypto backend this project calls. rcgen, rustls, tokio-rustls, hickory-proto’s TSIG signing and lettre’s SMTP TLS are all built on it, as is the PostgreSQL driver’s TLS (tls-rustls-ring). There is no aws-lc-rs, no native-tls and no OpenSSL. subtle supplies the one constant-time comparison ring no longer offers.

    A driver may carry its own for a wire protocol of its own. sqlx-postgres brings RustCrypto (sha2, hmac, md-5, stringprep) because PostgreSQL authenticates with SCRAM-SHA-256, and no configuration of the driver avoids it. The narrowing is deliberate and bounded: that code is reachable only from the driver’s own handshake, never from anything here, and the alternative was to have no second backend at all. A dependency that wanted a second stack for work this project does — signing, hashing, comparing — is still refused.

  • No global installs. The rustls provider is passed to every config builder explicitly and never installed as the process default. The metrics registry is a value held by the assembly, never a global recorder.

  • Prefer an edge to a crate. Before a new crate is added, check whether the capability is already in the graph through something else. hyper (via axum), hickory-proto (via the resolver), x509-parser (via rcgen), percent-encoding (via url) and subtle (via rustls) were each promoted to direct dependencies rather than replaced by a fresh one.

  • Hand-roll a small, stable protocol rather than import a stack. Examples:

    • RFC 6238 TOTP on ring;
    • password hashing with ring’s PBKDF2 rather than four crates for Argon2;
    • the relay’s ACME client on hyper;
    • the Prometheus text format.

    Each is a page of code against a frozen specification, with its own tests.

  • Choose by maintenance and safety as well as features. For example, cryptoki rather than pkcs11, because it wraps sessions in safe types and tracks the OASIS specification.

  • Each dependency’s own reason stays beside it in Cargo.toml, as a one-line comment.

Consequences

  • cargo build needs only a Rust toolchain.
  • A test can build two registries or two TLS configurations without them interfering with each other.
  • A hand-rolled protocol is code this project owns. Each one carries its own RFC test vectors (HOTP and TOTP) or an end-to-end test against a real counterpart (the relay client against the e2e lab’s ACME server).
  • The validators deliberately trust no root store (RFC 8555 §8.3, RFC 8737 §3: the responder’s certificate is the proof, not an identity). The relay client is the opposite case, a client of a real CA, and uses webpki-roots.
  • When a library would genuinely earn its place, as histograms would for metrics, adopting it is a new decision that supersedes this one for that subsystem.

Enforced by

  • cargo deny check with all-features = true (deny.toml), and the SBOM drift check (the supply-chain and sbom CI jobs).
  • Otherwise review only.

ADR 0010: Errors derive thiserror, carry their whole message, and panic only at startup

Status

Accepted. This was re-argued more than once before it was written down, which is why it is written down.

Context

The workspace has a couple of dozen error types, and each needs a Display. Written by hand, that is roughly 250 lines of match and write! in which each message sits several screens from the variant it describes. That distance is what let messages drift from their variants, more than once.

anyhow carries startup errors up to the CLI, where CliError::failed(error.to_string()) prints them. to_string() renders only an anyhow::Error’s outermost message. An error wrapped with .context() would therefore print its context and silently lose the underlying cause.

The ACME error type, Problem (RFC 8555 §6.7), is the Err of nearly every request-path function. clippy’s result_large_err lint flags a large Err in every one of them.

Decision

  • Every error type derives thiserror::Error. No hand-written Display or std::error::Error impl remains, and a new one should not add any. Three shapes need care:
    • A field literally named source is taken as the #[source] and must implement Error. ScriptError::Spawn’s is a formatted String, so it is named detail instead.
    • A variant whose wording depends on an inner enum needs a function, not a format string. For example, RevokeError::Signer goes through signer_detail.
    • A variant that renders through a method of its own takes an expression: #[error("{}", self.reason())].
  • anyhow without .context(). Every anyhow error is built with anyhow!("… {error}"), so its message already contains its cause. That is what keeps CliError::failed(error.to_string()) lossless.
  • Problem stays small. Every constructor goes through a private Problem::build. The optional members (identifier, subproblems, and type-specific extras such as badSignatureAlgorithm’s algorithms) live behind an Option<Box<_>>. identifier renders only inside subproblems, because RFC 8555 §6.7.1 forbids it at the top level, and the serializer enforces that rather than trusting callers.
  • Error handling splits by phase. A startup path (database connect, migrations, reading the configuration) may panic! or unwrap to fail fast. A request path returns Result and degrades:
    • a database error becomes a 500 Problem;
    • a nonce that fails to save means the response goes out without Replay-Nonce, and the failure is logged.

Consequences

  • A variant and its message sit on adjacent lines.
  • Error types add one proc-macro crate to the audited graph. syn, quote and proc-macro2 were already there, via serde_derive and async-trait.
  • A startup error prints its whole chain on one line. Adding a .context() anywhere would silently start truncating that line.
  • A handler that needs a dynamic status or header returns Result<Response, Problem> and builds the response itself. That is how keyChange’s 409 gets its Location header.

Enforced by

  • clippy’s result_large_err, run with -D warnings in CI.
  • Problem::to_value, for where identifier may appear.
  • Otherwise review only: nothing mechanically forbids .context() or a hand-written Display.

ADR 0011: Metrics are built on prometheus-client, one registry per scrape

Status

Accepted. Supersedes the metrics clause of ADR 0009, which listed the Prometheus text format among the protocols this project writes by hand.

Context

The exporter was written by hand: a BTreeMap of counters per family and a write! per series. ADR 0009 accepted that for counters and a gauge, and named latency histograms as the point where a library would earn its place.

A histogram is where hand-writing stops being small. Each series needs its buckets, a running sum and a count, and every bucket is rendered cumulatively with a final +Inf bucket. A mistake in any of that produces a scrape the collector reads wrongly, or rejects outright.

Of the Rust clients, most keep a process-wide registry or recorder: the metrics facade installs a global recorder, and the prometheus crate has a default registry. ADR 0009 refuses global state, because tests sharing one registry would count each other’s requests. prometheus-client has no global state at all: a Registry is an ordinary value. It adds dtoa and a derive macro to the graph; itoa and parking_lot were already there. Its licence is Apache-2.0 OR MIT.

Decision

  • crates/jobs/src/metrics.rs is built on prometheus-client. The families are fields of Metrics, which server::Assembly holds across reloads, as before.
  • The registry is built per scrape. Metrics::render builds a Registry over clones of the families (a clone shares its series) with the role label as a registry label. Nothing registers into shared state, and the roles stay a builder step on Metrics.
  • Every family is declared even when empty. The library leaves an empty family out of the exposition; a wrapper keeps it in, so a dashboard can tell “has not happened yet” from a misspelled name.
  • The output is OpenMetrics, the library’s only text format. Series names are unchanged; a counter’s # TYPE line drops _total, and the body ends with # EOF.
  • Label values are escaped before they reach the library, which writes them verbatim.
  • Two histograms: request latency by profile and route, and issuance latency (finalize accepted to certificate stored) by profile. Neither carries status or reason, since a histogram multiplies every label by its buckets.

Consequences

  • Histograms and their buckets are the library’s to get right, not this project’s.
  • A new family is a field and a register call in render. It appears on an empty registry at once, which tests/grafana_dashboard.rs relies on.
  • Series order within a family is no longer sorted, so tests match series individually rather than comparing the whole output.
  • A scraper that only reads the old text format and ignores the content type would misread # EOF and the _total-less # TYPE lines. Prometheus negotiates OpenMetrics and reads both.

Enforced by

  • The unit tests in crates/jobs/src/metrics.rs, which assert the exposition as text, including an empty registry’s families and the closing # EOF.
  • tests/grafana_dashboard.rs, both directions, per family.
  • cargo deny check and the SBOM drift check, as for every dependency.

ADR 0012: Images are built natively per architecture, uncached, and published only past a guard

Status

Accepted.

Context

The repository shipped a Containerfile and documented a container deployment, but published no image, so every container user compiled the crate. A contribution (#3) added a workflow that published one on a release tag. Its approach was right: a tag-only trigger, GHCR with the workflow’s own token, and actions pinned by SHA. Four of its details were not.

  • It built the lab’s binary. The Containerfile compiled --profile e2e, which is release without fat LTO, tuned for the e2e lab’s inner loop. Nothing outside the binary tells the two profiles apart, so a published image of the wrong one would not have been noticed.
  • It emulated arm64 with QEMU. The release profile is fat LTO with one codegen unit. Under emulation that is the slowest build this project has, with no timeout.
  • It configured a build cache that could not help. type=gha is scoped to the ref, so a tag never reads another tag’s entries. The build’s real cache is two RUN --mount=type=cache mounts, which no cache exporter preserves. And COPY . . sits directly above the one expensive RUN, so layer reuse buys nothing. The cost was real, though: mode=max writes gigabytes into the repository’s shared 10 GB Actions cache, and evicts the rust-cache entries every CI job depends on.
  • Nothing checked the tag. CI runs on pushes to main, not on tags. A tag that did not match the manifest’s version, or that pointed at a commit CI had never passed, would have published.

An image is also the one prebuilt artifact this project distributes, and its users run it as their certificate authority. That calls for provenance an operator can check.

Decision

  • The Containerfile takes the cargo profile as a build argument, CARGO_PROFILE, defaulting to release. The lab passes e2e explicitly. The default is the distribution build, so a hand-run podman build . reproduces the published image instead of a near miss.
  • Each architecture builds on a native runner, ubuntu-latest and ubuntu-24.04-arm, as a matrix. Each leg pushes one single-architecture image by digest, with no tag. A final job joins the two digests into one manifest list and tags that, after checking there are exactly two.
  • No build cache, and the workflow says why, since an absent cache is the first thing a reader would add.
  • A guard job runs before any build. It refuses a tag that differs from [workspace.package].version, or from any crate’s =x.y.z pin. It refuses a tag off its release line, main for X.Y.0 and release/X.Y for a patch (ADR 0013). It also refuses a tag whose commit has no successful push run of ci.yml on that branch. It fails rather than waits: the release procedure tags only once CI is green.
  • Build provenance is attested once, on the manifest list’s digest, and pushed to the registry. BuildKit’s own per-image attestations are off: with them on, each leg pushes an index instead of an image, and joining those indexes would carry attestation manifests that nothing references.
  • A release tag publishes X.Y.Z, X.Y and latest. The floating X.Y is the newest release of its line. It never crosses a minor, so it never picks up a breaking change, and an operator following it gets patch releases unattended. There is no floating X: before 1.0 a minor release is where breaking changes land. Both floating tags move only on a tag push, so a manual republish of an older tag moves neither. latest also needs the tag to be the highest release, so a patch to an older line leaves it alone.
  • Every push to main publishes edge and sha-<commit>, once all of ci.yml has passed on it: ci.yml calls this workflow as its last job. The build is the same release build, attested the same way; only the tags differ.
  • The image carries the default feature set, the same binary cargo install acme-proxy produces. hsm needs a build of one’s own.

Consequences

  • An uncached release build takes tens of minutes per architecture, and now runs on every merge to main as well as on every release. A public repository’s runners are not billed, and timeout-minutes is set to stop a wedged builder, not as an estimate. Merges to main queue rather than cancel each other, since a cancelled publish leaves orphaned manifests.
  • edge and the sha- tags accumulate a version per merge in GHCR. Pruning them is a registry policy, and it must keep untagged versions (below).
  • A hand-run podman build . is now as slow as a release build. Contributors building the lab image by hand pass --build-arg CARGO_PROFILE=e2e, as tests/e2e/common.rs does.
  • The attestation covers the manifest list. Verifying by tag finds it, since a tag resolves to the list. Verifying the digest of one architecture’s image finds nothing. If that ever matters, the fix is a second attestation per leg, not moving this one.
  • The per-architecture images show in GHCR as untagged versions. The manifest list references them, so a cleanup policy must never prune untagged versions.
  • A package that the workflow’s token creates under an organisation starts private. The first release needs a one-time change to inherit the repository’s visibility, and the workflow’s run summary says so.
  • A failed architecture fails the release with no partial publish. The two legs are independent (fail-fast: false), so the healthy one still shows whether the fault is the architecture or the change.

Enforced by

  • guard in .github/workflows/release.yml: the version and pin check, the release-line check, the CI check, and the highest-release check that gates latest.
  • The image job in .github/workflows/ci.yml, which needs: every other job before it publishes edge.
  • The digest-count check in that workflow’s publish job.
  • tests/e2e/common.rs, whose image build names CARGO_PROFILE=e2e; nothing else selects the lab’s profile.

ADR 0013: main is the trunk, and a patch line is a release branch cut when a fix needs one

Status

Accepted.

Context

Every change landed on main, and every release was tagged there. That left no way to ship a fix alone. Once main held work towards the next minor, some of it breaking under ADR 0001, a bug in the last release could only be fixed by releasing the next minor with it. An operator who needed the fix had to migrate their configuration to get it.

The only image was a release’s. Nothing let an operator try the next release before it was cut, short of building it themselves.

Decision

  • main is the trunk. It is the default branch. Every pull request, feature or fix, targets it, and it is releasable at every commit: CI is green before a merge.
  • A minor release, X.Y.0, is a tag on main.
  • A patch line is the branch release/X.Y, cut from the X.Y.0 tag when the first fix needs to ship on it, not at release time. A minor that never needs a patch never gets a branch.
  • A fix lands upstream first. It merges on main, then is cherry-picked with git cherry-pick -x onto a topic branch and merged into release/X.Y through a pull request. A fix is never made on a release branch alone, so the next minor cannot regress it.
  • A patch release, X.Y.Z with Z above 0, is a tag on release/X.Y. Its version bump, pins, SBOM and changelog section are made on that branch. The changelog section is then cherry-picked to main, so the trunk’s changelog lists every release.
  • Images follow ADR 0012. A release tag publishes X.Y.Z, X.Y and latest. Every merge to main publishes edge. A push to a release branch publishes nothing, because its image is its next patch tag’s.
  • main and every release/* branch are protected by a repository ruleset: changes arrive by pull request with the CI checks passing, and the branch can be neither force-pushed nor deleted.

Consequences

  • A fix reaches operators without the next minor’s changes, and X.Y gives them patch releases without a configuration change.
  • Each backport is a second pull request, and a cherry-pick can conflict once main has moved away from the release. The -x trailer records where each one came from.
  • Only the newest release line is maintained as a rule. An older one gets a branch only if a fix is worth the backport, and the nightly advisory scan covers the newest release branch alone.
  • The book tracks main, so between releases it can describe behaviour no release has yet.
  • A patch tag on main, or a tag on a commit its branch does not contain, is refused before anything builds. The version in main’s Cargo.toml stays at the last minor until the next one is cut, so edge reports that version.

Enforced by

  • guard in .github/workflows/release.yml: “The tag is on its release line”, and “CI passed on the tagged commit” against that branch.
  • The push trigger in .github/workflows/ci.yml, which covers main and release/[0-9]+.[0-9]+ so that the guard has a run to read.
  • The repository rulesets on main and release/*, configured in the repository settings rather than in the tree.
  • The procedures in Contributing: review only.

ADR 0014: PostgreSQL is chosen by the URL’s scheme, over one set of queries

Status

Accepted.

Context

SQLite across processes is safe on one local disk and not across hosts. The role split (ADR 0007) lets acme, admin and worker run as separate processes, but only as separate processes on one filesystem, so the deployment page promised PostgreSQL for as long as it did not exist.

Two earlier decisions had already paid for most of it. ADR 0003 keeps the pool private to crates/store/, so SQL and the dialect it is written in live in one crate. ADR 0004 chose UUID v7 partly because a v4 primary key costs PostgreSQL a page split and a full-page WAL write per row. What was left was a real fork: two pool types, two row types, two parameter syntaxes, and around 350 bind sites.

The obvious answers were both bad. sqlx::Any cannot carry a Uuid at all — its type set is Null/Bool/SmallInt/Integer/BigInt/Real/Double/ Text/Blob — and it does not translate SQL. Writing every statement twice doubles the SQL and guarantees the two copies drift, which is the one failure mode nothing here would catch: a query used by an operator listing can be wrong for months.

What made a single set of queries possible is that almost nothing in this schema is dialect-specific to begin with. Every timestamp is an epoch-second integer, so there is no strftime, julianday, datetime() or date type anywhere. There is no CAST, no || concatenation, no LIKE, no GROUP_CONCAT, no IFNULL, no CTE and no window function. RETURNING, ON CONFLICT … DO NOTHING, DO UPDATE … excluded.*, partial unique indexes and bound LIMIT/OFFSET are already spelled the way both accept.

Decision

  • The scheme of database.url picks the backend, at Database::open. sqlite: creates the file; postgres:/postgresql: expects the database to exist, because creating one is an operator’s act and not something a server does to a cluster it was pointed at. Any other scheme is refused by name.
  • One set of queries, in crates/store/src/sql.rs. A statement is a sql::Query carrying its SQL and a Vec<Value> until a sql::Exec says which driver is on the other end. sql::Row hides which row came back and sql::Builder replaces sqlx::QueryBuilder. Nothing outside that module names either driver.
  • Statements keep ? and the seam rewrites to $1…$n. Writing $n in the source would have worked on both — sqlx’s SQLite driver parses a $N marker and binds argument N — but three things here build SQL by concatenation: the live_certificate! predicate spliced into the middle of three statements, the IN (?, ?, …) lists expanded per element, and the format!ed fragments in job::claim_next and job::settle. Every number would then be a hand-maintained constant. One rewrite at the edge cannot drift.
  • An absent value carries the type it would have had. SQLite has no typed null; PostgreSQL sends a type OID per parameter and refuses column "eab_kid" is of type uuid but expression is of type bigint. Every bind site knows the type statically, so Value::Null(NullKind) costs nothing and is declared nowhere twice.
  • Three things fork, and each asks Dialect. The identifier search (json_each/json_extract/instr against jsonb_array_elements/->>/strpos); the unique-violation matchers; and the _sqlx_migrations probe, which asked by reading the table and swallowing the error, where on PostgreSQL a failed statement aborts the surrounding transaction. Nothing else may fork without a line here.
  • strpos, never position(needle in haystack). It takes its arguments the other way round, so one push_bind sequence would bind the two dialects in different orders — a wrong answer rather than an error.
  • Two migration sets, both append-only. migrations/ for SQLite, frozen since 0.1.0; migrations-postgres/ from its own first release. The PostgreSQL set is not a transcription: the SQLite files carry three table rebuilds that exist only because SQLite cannot add a CHECK, a UNIQUE or a foreign key to an existing table, plus a text-to-blob id conversion, and no PostgreSQL deployment has that history to replay. Every declared width is transcribed literally.
  • A unique constraint PostgreSQL must name is named in the migration. SQLite reports the offending columns and gives sqlx no constraint name; PostgreSQL reports the constraint and never the columns. A matcher passes both spellings, so the index name is part of the schema rather than whatever the server happened to generate.
  • A database is one backend or the other, and acme-proxy transfer is the way across. Not a dual-write mode and not a sync: an offline copy of every row, refused unless the target is migrated and empty. It exists because the order row is a certificate’s only record — a deployment that moved to PostgreSQL by starting empty would leave every certificate it had issued impossible to revoke, which is the outcome live_certificates_refusal exists to prevent. The copy is driven by a declared column manifest rather than by reading the source’s shape, because the seam decodes into a known Rust type and “read this column as whatever it is” would mean deciding at runtime whether SQLite’s untyped BLOB is a bytea or a uuid. The manifest’s own hazard — a column added to the schema and forgotten here — is answered the way ADR 0003 answers it for a table rebuild: by introspecting the live schema and refusing a manifest that has drifted.
  • The database URL is redacted wherever it is printed. A DSN carries user:password@; the startup log line and the SIGHUP refusal both go through logfields::redact_url.

Consequences

  • One binary and one container image serve both, and a deployment moves from SQLite to PostgreSQL by changing one key and running acme-proxy transfer. What that command cannot check is that the source is stopped, so it says so in its prompt: a copy taken while a worker is issuing is a torn snapshot that looks exactly like a good one.
  • Tx no longer derefs to SqliteConnection. tx.conn() is what &mut *tx was, and JobQueue::enqueue_in takes a sql::Exec — the one SQLite-typed signature that had leaked outside crates/store/.
  • A declared VARCHAR(n) is now enforced. SQLite ignores the width, which is how nonces.value stayed VARCHAR(36) after the nonce became a 43-character token; on PostgreSQL that would have rejected every nonce the server mints. The width pins in db.rs are what keep the two honest.
  • sqlx/postgres brings RustCrypto (sha2, hmac, md-5, stringprep) for SCRAM-SHA-256. That is a second crypto stack in the graph, which ADR 0009 argues against; its clause is narrowed rather than worked around, because there is no configuration of the driver that avoids it. TLS stays on ring (tls-rustls-ring), and sqlx builds its ClientConfig with builder_with_provider rather than install_default, so the “nothing installs process-global state” rule is untouched.
  • PostgreSQL gives sqlx an advisory migration lock, which SQLite does not. The one-owner rule in ADR 0003 is therefore belt and braces there rather than load-bearing — it stays, because the rule is about which process may own the schema, not only about the race.
  • The coverage floor cannot see this backend: a dialect arm not taken is not an uncovered line. That is what the CI job and its REQUIRE guard are for.

Enforced by

  • The whole acme-proxy-store suite, on both backends. Every test there calls Database::connect_for_test(), which is PostgreSQL when TEST_POSTGRES_URL names one — the coverage that found MAX(x, 0), SQLite’s scalar two-argument max, in a path no dialect-specific test would have singled out. connect_in_memory() means SQLite, and is the opt-out for the seven tests that are about SQLite.
  • tests/postgres.rs, which runs the dialect-sensitive paths against both backends, and postgres_is_available_when_it_is_required, which fails rather than skips when ACME_PROXY_REQUIRE_POSTGRES is set.
  • transfer::tests::the_manifest_names_every_column and …every_table, against the live schema on whichever backend is running, plus a_database_survives_a_round_trip_through_the_other_backend, which seeds all fifteen tables and compares values after a copy out and back.
  • The postgres job in .github/workflows/ci.yml, which sets that variable and also runs roles and reload — several processes over one database, which is the deployment this exists for.
  • sql::tests::numbering_is_contiguous_from_one and the literal-skipping cases beside it.
  • production_code_never_reaches_the_raw_pool and only_the_schema_owners_apply_migrations (tests/layering.rs), unchanged.
  • declared_token_widths_match_random_token, declared_issuer_widths_match_the_issuer_id and every_id_column_is_declared_a_blob (crates/store/src/db.rs).
  • logfields::tests, for the redaction, and reload::tests::a_refusal_over_a_dsn_keeps_the_host_and_drops_the_password.

Database Schema

acme-proxy stores everything in one database — accounts, orders, the audit trail and the web admin’s own operators. There is no second datastore and no cache. This page describes what is in it and why, for anyone reading the database directly, writing a migration, or trying to understand what a delete cascades to.

Two backends, one schema. SQLite is the default and PostgreSQL is what a multi-node deployment needs; the scheme of database.url picks between them. Everything below describes both — the tables, the constraints, the cascades and the reasoning are the same either way. Where the two differ it is noted inline, and the differences are three: the column types (BLOB/uuid, INTEGER/bigint), the AUTOINCREMENT spelling, and the recipes for reading it by hand at the bottom of this page. The SQL the server issues is written once; crates/store/src/sql.rs is the seam and its //! says what had to fork.

A database is one backend or the other; there is no dual-write mode. Moving between them is acme-proxy transfer, which copies every row of every table and is guarded by crates/store/src/transfer.rs’s column manifest — a migration that adds a column adds it there too, or the copy would silently leave it behind.

There are two migration sets, one per dialect, and both are frozen and append-only — SQLite’s as of 0.1.0, PostgreSQL’s from its first release. A schema change is a new sqlx migrate add file in each, never an edit to a committed one. The PostgreSQL set is deliberately not a transcription of the SQLite one: those files carry table rebuilds that exist only because SQLite cannot add a CHECK, a UNIQUE or a foreign key to an existing table, and no PostgreSQL deployment has that history to replay. The schema is also the only surface frozen before 1.0.0 — the freeze says nothing about configuration keys, which may still be renamed. See Contributing for the three consequences that catch people out.

The tables at a glance

erDiagram
    accounts ||--o{ orders : "account_id"
    orders ||--o{ authorizations : "order_id"
    authorizations ||--o{ challenges : "authz_id"
    orders ||--o| upstream_orders : "order_id"

    admin_users ||--o{ admin_sessions : "user_id"
    admin_users ||--o{ admin_recovery_codes : "user_id"

    accounts {
        blob id PK
        text profile "UNIQUE(profile, pubkey)"
        blob pubkey
        text status "CHECK valid|deactivated|revoked"
        blob eab_kid "no FK - see below"
        text created_ip
        text last_seen_ip
    }
    orders {
        blob id PK
        text profile
        blob account_id FK
        text status "CHECK pending|ready|processing|valid|invalid"
        text identifiers "JSON array"
        text replaces "RFC 9773 certID"
        text certificate "PEM chain"
        text cert_serial
        integer cert_not_after "what the leaf says - see below"
        integer revoked_at
    }
    authorizations {
        blob id PK
        blob order_id FK
        text identifier "JSON, UNIQUE(order_id, identifier)"
        text status "CHECK pending|valid|invalid|deactivated|expired|revoked"
    }
    challenges {
        blob id PK
        blob authz_id FK
        text type "CHECK http-01|dns-01|tls-alpn-01, UNIQUE(authz_id, type)"
        text token
        text status "CHECK pending|processing|valid|invalid"
    }
    upstream_orders {
        blob order_id PK "also the concurrency guard"
        text upstream_order_url
        blob csr_der
        text client_ip "parked request context"
    }

    eab_keys {
        blob kid PK
        blob secret "retrievable on purpose"
        text profile "NULL = every endpoint"
        text status "CHECK active|revoked"
    }
    nonces {
        text value PK
        integer created_at
    }
    audit_log {
        integer id PK "AUTOINCREMENT"
        text event "no CHECK - the Rust enum"
        text outcome "CHECK success|failure"
        text account_id "no FK, deliberately"
        text order_id "no FK, deliberately"
        text identifiers "frozen into the row"
    }

    jobs {
        blob id PK
        text kind "no CHECK - see below"
        text dedup_key "partial UNIQUE(kind, dedup_key)"
        text payload "JSON, the subject's identity"
        text status "CHECK ready|running|done|failed|cancelled"
        integer run_at "the durable schedule"
        integer attempts "incremented at claim"
        integer deadline "give up after this"
        integer lease_until "when a dead runner's row is reclaimed"
        text lease_owner "which runner holds it"
    }

    admin_users {
        blob id PK
        text username UK
        text password_hash "one-way"
        blob totp_secret
        text status "CHECK active|disabled"
        text role "no CHECK - NULL reads as admin"
    }
    admin_sessions {
        text token_hash PK "SHA-256 of the token"
        blob user_id FK
        text state "CHECK pending_mfa|active"
        integer mfa_attempts
    }
    admin_recovery_codes {
        blob id PK
        blob user_id FK
        text code_hash
        integer used_at "stamped, not deleted"
    }

    revocations {
        text issuer PK "SHA-256 of the CA's SPKI"
        text serial PK
        integer revoked_at "the first one stands"
        integer not_after "NULL is never pruned"
    }
    crls {
        text issuer PK
        integer crl_number "moves only by compare-and-swap"
        blob der "what GET /crl serves"
        integer next_update
    }
    http01_tokens {
        text token PK "the upstream's own token"
        text key_authorization
        integer expires_at "a backstop, swept hourly"
    }

The diagram has three clusters, and the two things worth noticing are the edges that are not drawn:

  • The ACME graph — accounts → orders → authorizations → challenges, with upstream_orders hanging off an order and eab_keys and nonces standing alone.
  • audit_log, jobs and a local CA’s revocations, crls, plus the relay’s http01_tokens, all connected to nothing. That is policy in each case, not an omission. The audit trail and the revocation ledger must outlive what they describe, the job queue is generic (both below), and the last three are state every role process shares (ADR 0008).
  • The admin island — admin_users and its two children — which never joins to accounts. An admin_users row is an operator of this server; an accounts row is a client key that asks it for certificates. They are different populations and the schema says so.

Profiles are a database boundary

accounts.profile and orders.profile are NOT NULL, and accounts is keyed UNIQUE(profile, pubkey). One client key presenting itself at two endpoints is two independent accounts with separate orders and separate authorizations — see Profiles & Routing.

eab_keys.profile is the one nullable member of the set, and NULL means “valid at every endpoint” rather than “unknown”.

Request-path lookups always take the profile. The admin CLI deliberately uses unscoped lookups (find_any_by_id, find_any_by_kid), because an operator holding an id wants the row, not a reminder about which endpoint it belongs to.

Every foreign key is indexed and cascades

SQLite indexes primary keys and UNIQUE constraints and nothing else, so before 20260727120000_indexes_and_constraints.sql every order read and every challenge trigger was a full table scan. That migration rebuilt the four ACME tables to add both halves at once:

  • ON DELETE CASCADE on every foreign key, so an account or an order can genuinely be deleted. On SQLite this depends on the foreign_keys pragma, which crates/store/src/db.rs pins on for every connection; PostgreSQL always enforces them.
  • An index on every foreign key: idx_orders_account_id, idx_authorizations_order, idx_challenges_authz.

Two later indexes serve one query each: idx_orders_cert_serial on (profile, cert_serial), which POST /revokeCert uses on every request, and idx_orders_created_at / idx_orders_status_created_at, which the web admin’s newest-first cross-account listing needs and the ACME path never did.

idx_orders_replaces_claim is different — it is a partial unique index on (profile, replaces) where replaces IS NOT NULL AND status != 'invalid'. It is not a lookup index at all; it is RFC 9773 §5’s “already replaced?” rule enforced in SQL, which is what makes 409 alreadyReplaced race-free and what lets an order that fails release its claim. See Renewal Information.

CHECK constraints hold the state machines

Every status column carries a CHECK (status IN (…)) — accounts, orders, authorizations, challenges, jobs, upstream_orders, eab_keys and admin_users — as do challenges.type, admin_sessions.state, and audit_log’s outcome and actor_kind.

They are there because a typo in a status would otherwise park a row in a state nothing can read back, and the row would look fine. With the constraint it is a failed write at the moment of the mistake.

Open vocabularies carry none: audit_log.event (its CHECK was dropped by 20260909120000), admin_users.role and jobs.kind are validated by a Rust enum instead, because each grows with features and a CHECK would cost a table rebuild per new word. See ADR 0005.

This is also why a new CHECK is expensive: SQLite cannot add one to an existing table, so it needs a full table rebuild in a new migration. Several constraints were therefore declared before anything wrote them — admin_users.totp_secret/totp_pending_secret/totp_last_step and admin_sessions.state’s 'pending_mfa' value are the worked example, added in 20260808120000 and only used once the second factor shipped.

The audit trail has no foreign keys, deliberately

audit_log names an account_id and an order_id with no constraint behind either. An audit row has to survive the account or order it describes being deleted — a CASCADE there would destroy the evidence along with its subject, which is the one thing an audit trail may not do. The identifiers are frozen into the row for the same reason, rather than being read back through a join that may no longer resolve.

Two more consequences of that decision:

  • id is INTEGER PRIMARY KEY AUTOINCREMENT, not a plain rowid. An operator types this id, and AUTOINCREMENT is what stops SQLite handing out the rowid of a purged row a second time.
  • outcome is denormalized from event and written from the single definition in AuditEvent::outcome, so “show me everything that was refused” is an index lookup rather than event LIKE '%_failed' written out in three front ends.

Rows are only ever INSERTed. There is no setter and no UPDATE against this table anywhere in the crate; the only statement that removes anything is the retention sweep. See Audit Trail.

revocations follows the same rule for the same reason. A local CA’s revocation must outlive an order an operator deletes, or the serial would drop off the CRL, so it has no foreign key to orders either.

accounts.eab_kid is a similar deliberate non-key: it records which credential was used at registration, but an EAB credential is revocable and the account outlives it, so there is no constraint tying the two together.

The job queue is generic, and its schema says so

jobs is the second table with no foreign key, for a different reason than audit_log’s. A queue is generic: payload names whatever kind means — a local order today, a certificate serial or nothing at all tomorrow — so a typed foreign key would either be wrong for every other kind or force one nullable column per kind. A job whose subject was deleted is retired by its handler (“the order no longer exists”), which is a terminal outcome recorded in last_error, not an orphan nothing sweeps.

Three more shapes worth knowing before touching it:

  • kind carries no CHECK, unlike every other enum-ish column here. A kind is registered in code by whichever subsystem owns it, and SQLite cannot alter a CHECK without a table rebuild, so every future kind would cost one. The runner claims only the kinds its registry holds, so an unrecognised one is left alone rather than mis-run — which is also what lets an older binary meet a row a newer one wrote.
  • The identity index is partial: UNIQUE(kind, dedup_key) WHERE status IN ('ready', 'running'). Only a live job holds an identity. A plain UNIQUE would let one finished job block its own key for ever, which is fatal for a periodic kind whose key is a constant and wrong for an order retried after a failure.
  • attempts increments when the row is claimed, not when it completes, so a job that reliably kills the process still exhausts its budget instead of crash-looping. The same reasoning is why the reclaim sweep leaves the counter alone.

status = 'cancelled' is written by jobs cancel and its panel and API twins. It was declared before anything wrote it — the admin_sessions.state = 'pending_mfa' treatment, where a CHECK was written before anything filled it precisely so no rebuild would be needed later.

Two expiry columns on orders, and neither is the order’s

orders.not_after is the validity the client asked for in newOrder (RFC 8555 §7.4) — usually NULL, and clamped by the signer when set. orders.cert_not_after is what the issued leaf actually says, stamped by Order::finalize from the same DER that cert_serial and cert_pubkey come from. The order object’s own expires is a third thing again, and is not a column here.

cert_not_after has three meaningful states:

  • an epoch second;
  • NULL — issued before the column existed; the expiry sweep backfills it;
  • a negative sentinel — the sweep looked and the chain would not parse. Writing NULL back would have it re-parsed on every pass for ever.

It is optional where cert_serial is not: a chain whose serial cannot be read cannot be revoked, so it is a failed issuance, while an unreadable validity is only housekeeping. Its index is partial on certificate IS NOT NULL AND revoked_at IS NULL, which is the expiry digest’s own predicate.

An order holding a live certificate is never deleted

The order row is a certificate’s only record: revokeCert and order revoke find it by serial, the expiry digest lists it, and renewal information (RFC 9773) is derived from it. Deleting it would make a certificate that is still trusted impossible to revoke.

So account delete, order delete and eab delete --delete-accounts are refused, on every surface, while any order they would remove holds a live certificate — issued, not revoked, and not yet expired (cert_not_after NULL, negative or in the future). The check runs before the confirmation prompt and again inside the DELETE itself (live_certificate! in crates/store/src/order.rs), so a certificate issued in between is not lost. There is no override flag. The daily order_sweep follows the same rule: a valid order is never swept, whatever its age.

Secrets are stored three different ways, on purpose

The three storage shapes in this schema are not an inconsistency — each one follows from what the server has to do with the value later.

ColumnShapeWhy
eab_keys.secretRaw bytes, retrievableHMAC verification needs the same secret back on every request. A lost one is replaced, never recovered — eab create prints it once.
admin_users.password_hashOne-way KDF (PBKDF2-HMAC-SHA256), unreadableA password is only ever compared. No code path can read it out.
admin_sessions.token_hashhex(SHA-256(token)), no KDFA 256-bit CSPRNG token has no dictionary to slow down. The hash exists solely so a database read yields nothing replayable.

admin_recovery_codes.code_hash follows the password shape — a recovery code is only ever compared. It is a table rather than a JSON column so that consuming one is UPDATE … WHERE id = ? AND used_at IS NULL with rows_affected deciding a race, the same primitive nonces uses. used_at is stamped rather than deleted, so “7 of 10 remaining” is a count and a spent code leaves a trail.

admin_users.totp_secret and totp_pending_secret are plaintext BLOBs on purpose: verification recomputes the HMAC, so the server needs the same bytes back every attempt. That is eab_keys.secret’s situation, not a password’s, and any wrapping key would live in the same directory as the database.

Columns nothing ever compares against

accounts.created_ip/created_ptr/last_seen_ip/last_seen_ptr, orders.created_ip/created_ptr, admin_sessions.created_ip/user_agent, and audit_log’s client_ip/client_ptr/user_agent are forensics only. No code path compares a live request against any of them.

That is a decision, not an oversight. Pinning an identity to an address breaks CGNAT and mobile clients; pinning it to a User-Agent breaks on the next browser update. They answer “who asked for this, and from where” after the fact, and nothing else.

One column is compared, and only to decide whether to send a message: admin_users.known_login_ips, the operator’s last five distinct sign-in addresses, decides whether a sign-in is reported as coming from a new address. It never allows or denies anything.

None of them reaches an ACME object either — the wire format is RFC 8555’s and stays that way. They surface through the admin CLI and the web admin only.

Ids are UUID v7, stored as bytes

Every row this server creates is keyed by a UUID version 7 (RFC 9562 §5.7), minted in one place, acme_proxy_store::id::mint. Ids created close together share a prefix, and they sort by creation; why that matters, and the rule that an id’s Rust type says where it came from, are in ADR 0004.

On SQLite the column holds the sixteen bytes, not the thirty-six characters of the rendering; on PostgreSQL it is a native uuid, and sqlx maps the same Rust type to both. Nothing on the wire changes either way — an id is still rendered by Uuid::to_string, so account URLs, kids, order URLs and every admin API member are the same lower-case hyphenated form they always were.

What changes is an ad-hoc query, and only on SQLite, where an id column prints as a blob and wants hex():

# Readable ids.
sqlite3 sqlite.db "SELECT lower(hex(id)), profile, status FROM accounts;"

# Looking one up by the id from a URL or a log line.
sqlite3 sqlite.db "SELECT status FROM orders
                    WHERE id = unhex(replace('6ba7b810-9dad-41d1-80b4-00c04fd430c8','-',''));"

PostgreSQL needs neither, since it prints and parses the hyphenated form:

psql -c "SELECT id, profile, status FROM accounts;"
psql -c "SELECT status FROM orders WHERE id = '6ba7b810-9dad-41d1-80b4-00c04fd430c8';"

A few columns look like ids and are not, so they stay text: orders.replaces is an RFC 9773 certID, audit_log.actor_id may be an account id or an admin username, audit_log.account_id and order_id name a row that may already be gone, and request_id is whatever the caller sent.

Rows created before this changed were converted in place and keep their v4 ids, so a table holds both versions and only the v7s sort by creation.

Reading it directly

On SQLite, the file is sqlite.db by default and is opened in WAL mode, so there are normally sqlite.db-wal and sqlite.db-shm beside it. Copying only sqlite.db gives you a database missing every recent write; back up all three, or use sqlite3 sqlite.db ".backup backup.db", which is consistent by construction. On PostgreSQL, back it up the way you back up any other database — pg_dump is consistent by construction and there are no sidecar files to miss.

Ids are stored as bytes on SQLite, so they need hex() on the way out and unhex() on the way in; on PostgreSQL they are a native uuid and need neither. See Ids are UUID v7, stored as bytes.

The recipes below are SQLite’s. On PostgreSQL the ids need no wrapping and the epoch-second columns read with to_timestamp(created_at) in place of datetime(created_at,'unixepoch').

# What has this account been issued?
sqlite3 sqlite.db "SELECT lower(hex(id)), status, identifiers,
                          datetime(created_at,'unixepoch')
                     FROM orders
                    WHERE account_id = unhex(replace('…','-',''))
                    ORDER BY created_at DESC;"

# Everything refused in the last day.
sqlite3 sqlite.db "SELECT datetime(created_at,'unixepoch'), event, profile, client_ip, reason
                     FROM audit_log WHERE outcome = 'failure'
                      AND created_at > strftime('%s','now','-1 day');"

# Which migrations have run.
sqlite3 sqlite.db "SELECT version, description, success FROM _sqlx_migrations;"

Read-only inspection of a running server is safe under WAL, and on PostgreSQL by its own MVCC. Writing to the database behind the server’s back is not — the CHECK constraints will catch a bad status, but nothing will re-sync the in-memory state a handler is holding. Use the Admin CLI instead.

Custom Plugins Examples

This section provides complete examples of custom plugin scripts that can be integrated into acme-proxy. These scripts must be marked as executable (chmod +x).

Custom signer script

A custom signer that passes the CSR to a fictional internal API to obtain a certificate.

#!/bin/bash
# /etc/acme-proxy/signer/internal-pki.sh
set -e

# We only handle the "issue" hook in this example
if [ "$ACME_SIGNER_HOOK" = "issue" ]; then
    # Read the JSON payload from stdin
    PAYLOAD=$(cat)

    # Extract the base64 encoded CSR
    CSR_B64=$(echo "$PAYLOAD" | jq -r '.csr_der_base64')

    # Call internal PKI API
    # The API is expected to return a JSON with a 'certificate_pem' field.
    RESPONSE=$(curl -s -X POST https://pki.internal.company.com/api/sign \
        -H "Content-Type: application/json" \
        -d "{\"csr\": \"$CSR_B64\", \"order_id\": \"$ACME_SIGNER_ORDER_ID\"}")

    # Extract PEM from response
    PEM=$(echo "$RESPONSE" | jq -r '.certificate_pem')

    if [ -n "$PEM" ] && [ "$PEM" != "null" ]; then
        # Output the PEM chain to stdout (leaf first, then issuers)
        echo "$PEM"
        exit 0
    else
        # Exit 1 (any non-zero other than 3) = internal failure -> the client
        # gets a 500 and the order is marked invalid.
        #
        # Exit 3 is RESERVED for "this CSR is bad" -> the client gets a 400
        # badCSR and the order stays "ready" so it can retry with a corrected
        # CSR. Only use 3 when the PKI rejected the CSR itself, never for an
        # API outage like this one.
        echo "internal PKI API did not return a certificate" >&2
        exit 1
    fi
fi

# Hooks this example does not implement. `revoke` must succeed or the proxy
# leaves the order un-revoked, so returning non-zero here would be wrong for a
# real deployment;
# implement it, or set supports_crl/supports_renewal_info = false (the default)
# so those hooks are never invoked at all.
exit 1

Custom filter script

A custom filter that checks the client IP against a threat intelligence feed before allowing the connection.

#!/bin/bash
# /etc/acme-proxy/filters/threat-intel.sh

if [ "$ACME_FILTER_HOOK" = "connection" ]; then
    # Skip checking local IPs
    if [[ "$ACME_FILTER_CLIENT_IP" == 10.* ]] || [[ "$ACME_FILTER_CLIENT_IP" == 192.168.* ]]; then
        exit 0
    fi

    # Query threat intel API
    STATUS=$(curl -s -o /dev/null -w "%{http_code}" "https://threat.internal/api/check?ip=$ACME_FILTER_CLIENT_IP")

    if [ "$STATUS" = "200" ]; then
        # IP is clean
        exit 0
    else
        # IP is flagged, deny the request
        echo "Client IP $ACME_FILTER_CLIENT_IP is flagged in Threat Intel"
        exit 1
    fi
fi

exit 0

Custom IPAM script

A custom IPAM backend that reads the permitted names for an address out of a CSV the estate already maintains, one address,name,name,... row per machine.

#!/bin/bash
# /etc/acme-proxy/ipam/lookup.sh
set -u

INVENTORY="/etc/acme-proxy/ipam/inventory.csv"

# The address arrives twice — in the environment and in the JSON on stdin.
# This script uses the environment, so it never has to read stdin at all; a
# script that exits without reading it is fine and is not an error.
ROW=$(grep -m1 "^${ACME_IPAM_CLIENT_IP}," "$INVENTORY")

if [ -z "$ROW" ]; then
    # 3 is RESERVED: "this inventory holds no record of that address". The
    # `ipam` check words its own refusal for it, distinct from the one below.
    exit 3
fi

# Exit 0 with the permitted names, one per line. They are lowercased and
# stripped of a trailing dot for you, so print whatever form the file holds.
# Printing nothing here would mean "recorded, and entitled to nothing" — a
# different answer from exit 3, and also a refusal.
echo "$ROW" | cut -d, -f2- | tr ',' '\n'
exit 0

Every other non-zero exit — a missing inventory file, a grep that could not run, a timeout — is reported as a retryable 500, never as a denial. That is deliberate: an inventory this server cannot reach has decided nothing, so issuance stops rather than failing open. Do not use a non-zero exit to refuse a client; refuse by not printing the name.

acme-proxy filter explain really runs the policy, so it executes this script too.

Custom notification script

A custom notification script that sends a Slack message when a certificate is revoked.

#!/bin/bash
# /etc/acme-proxy/notify/slack.sh

WEBHOOK_URL="https://hooks.slack.com/services/AAAAAAAAA/BBBBBBBBB/XXXXXXXXXXXXXXXXXXXXXXXX"

# The JSON payload is always on stdin; read it before anything else. For
# certificate_revoked it carries the reason code, which has no env var.
PAYLOAD=$(cat)

if [ "$ACME_NOTIFY_HOOK" = "certificate_revoked" ]; then
    # ACME_NOTIFY_IDENTIFIERS is only populated for certificate_issued, so it
    # would render empty here — take the order id from the environment and the
    # reason from the payload instead.
    REASON=$(echo "$PAYLOAD" | jq -r '.reason // "unspecified"')
    MESSAGE="🚨 Certificate Revoked! Serial: \`$ACME_NOTIFY_CERT_SERIAL\` | Order: \`$ACME_NOTIFY_ORDER_ID\` | Reason: \`$REASON\`"

    curl -s -X POST -H 'Content-type: application/json' \
        --data "{\"text\": \"$MESSAGE\"}" \
        "$WEBHOOK_URL"
fi

exit 0

A notification script’s exit code is only logged — it can never fail the ACME request that triggered it.

Testing & Coverage

acme-proxy relies on a multi-layered testing strategy combining lightning-fast unit/integration tests with real-world End-to-End (E2E) scenarios.

Prerequisites

  • cargo-nextest: The project requires cargo nextest to execute the integration suite. nextest runs each test in its own isolated process. This is load-bearing because tests involving the custom scripts exec generated bash files. Under standard cargo test (which runs in threads), file descriptor sharing causes intermittent ETXTBSY failures.
  • llvm-cov: For coverage reporting.
  • Podman / Docker: Required for running the E2E suite.

Install the required Rust tools:

cargo install cargo-nextest cargo-llvm-cov
rustup component add llvm-tools-preview

Running the unit & integration suite

To run the complete in-memory test suite:

cargo nextest run --workspace

These tests use an in-memory SQLite database and an in-memory local CA, and nothing reaches a real network. A few suites write to a temporary directory or bind a loopback socket, each because the thing under test needs one: roles and reload (real processes, ports and a config.toml), filters and custom_signer (scripts, and the IPAM mocks), and revoke_cert (a CA on disk that two processes share).

PostgreSQL. Set TEST_POSTGRES_URL to a server’s URL and every crates/store/ test that calls Database::connect_for_test() runs against it instead of SQLite; tests/postgres.rs runs the dialect-sensitive paths against both, and skips without it. CI’s postgres job sets ACME_PROXY_REQUIRE_POSTGRES=1 as well, which turns that skip into a failure.

A test that calls Config::load() holds ENV_LOCK (acme_proxy_core::config::ENV_LOCK, or testutil::EnvGuard, which holds it for you). ACME_PROXY_* and ACME_PROXY_CONFIG are process state: a test setting one while another loads makes the second read the first’s variables. There is one lock for every crate on purpose, since a per-module lock would serialise a module against itself and nothing else.

A test that triggers a challenge or finalizes an order must poll for the result. Validation and issuance run in the job queue, and every test app runs a real worker, so the response only says processing. await_order and await_challenge in tests/common/ are the helpers.

The hsm feature (PKCS#11)

crates/signer/src/local_ca/pkcs11.rs is behind the hsm feature, so the command above neither compiles nor lints it — --all-targets does not enable features. Run it explicitly:

cargo nextest run --workspace --features acme-proxy-signer/hsm
cargo clippy --workspace --all-targets --features acme-proxy-signer/hsm -- -D warnings

The PKCS#11 tests create a SoftHSM2 token in a temporary directory, generate a P-256 key inside it, self-sign a CA certificate through the token, and then drive the real LocalCa end to end — issuing a leaf that must verify against that CA, and a CRL that must too. The key is generated through cryptoki itself, so softhsm2 is the only prerequisite; opensc/pkcs11-tool is not needed.

# Debian/Ubuntu
sudo apt install softhsm2
# Arch
sudo pacman -S softhsm

When no SoftHSM2 module is found the PKCS#11 tests skip with a message rather than failing, so --features hsm stays green without it. CI has a dedicated hsm job — separate from test so the coverage floor, which a feature-gated file sits outside of entirely, does not fight the feature.

cargo nextest matters more than usual here: SOFTHSM2_CONF is process-global and read at C_Initialize, and the PKCS#11 context is cached per module for the life of the process. Process-per-test isolation is what keeps those from leaking between tests.

Code coverage

CI enforces a hard floor of 97% of lines, over every package in the workspace (main.rs is excluded — it is pure socket and exit wiring). The shortest way to see the same number locally:

cargo llvm-cov nextest --workspace --summary-only

CI splits that in two, because it wants several views of one test run: the run itself with --no-report, then lcov.info, an HTML tree and the summary that gates, each generated from the profiles left on disk.

--workspace has to reach the report, and the report subcommand cannot take it. cargo llvm-cov report rejects the flag, and with no package selection it measures the package cargo picks — at a root that is also a package, the root package alone. Reporting from saved profiles at workspace scope is cargo llvm-cov --no-run --workspace, which is what CI uses:

cargo llvm-cov --no-run --workspace --summary-only \
  --ignore-filename-regex 'src/main\.rs' --fail-under-lines 97

Not -p once per member either: a crate built twice under different features contributes two coverage maps that way, and its lines are counted twice.

Gotcha: a handler annotated with #[instrument] reports far lower coverage than it actually has. The attribute moves the body into a generated async block, so the signature lines show zero hits and the body lines carry no region at all — handlers/authz.rs sits around 40% while tests/challenges.rs drives nearly every branch in it. Check cargo llvm-cov report --text for the file before writing tests against the percentage. (Installing a tracing subscriber in tests does not fix this; measured, it moves the total by 0.03 points.)

Which is why crates/admin/src/webadmin/ carries no #[instrument] at all. It is a rule for that module, not a preference: the access middleware already opens the request span, so the attribute would buy nothing and cost the module’s reported coverage.

The password KDF is slow on purpose

admin::password runs PBKDF2-HMAC-SHA256 at 600 000 iterations — roughly 85 ms in a release build. Unoptimised, ring takes ~1.1 s for the same hash, and the admin suites pay it at least twice per test (the harness creates an operator and signs in). With twenty of them in parallel that was most of the suite’s CPU time and a 40-second critical path. So the workspace Cargo.toml builds ring at opt-level = 3 in the dev and test profiles ([profile.dev.package.ring]), which brings a debug build to ~90 ms per hash. Only ring is raised: the loop is compiled entirely inside it, and optimising the workspace crates instead measurably changes nothing.

The override lives in Cargo.toml because nothing else reaches the build nextest runs: cargo --config … nextest run and CARGO_PROFILE_DEV_PACKAGE_RING_OPT_LEVEL are both silently ignored there.

The admin::password unit tests still mostly go through a private hash_with_iterations at a cheap setting — the same code path, the same salt generation and encoding, without 600 000 rounds dozens of times over. Two deliberately pay the real cost: the encoding must reflect the real constants, and the dummy hash must cost what a real row costs, or an unknown username would answer faster and enumerate the operator table.

If you add a test that signs in, expect it to cost one real hash.

Testing the web admin

tests/admin_api.rs drives the real build_admin_app through tower::ServiceExt::oneshot, the same way tests/orders.rs drives the ACME side. The harness helpers live in tests/common/mod.rs:

Helper
admin_config()a Config with [admin] enabled
test_admin_app(config)the admin router + its database
test_admin_app_with_signer(config)also returns the signer, for tests that must issue before revoking
test_admin_app_logged_in(config)creates one operator, signs in, returns an AdminSessionHandle
admin_request(app, method, path, session, body)one request, optionally authenticated
admin_login, session_cookie_token, json_body

test_admin_app and test_app_full share one_profile, so the two cannot drift into mounting subtly different endpoints.

The CSRF table is the regression suite. mutating_endpoints() in tests/admin_api.rs lists every unsafe method and path, and two tests assert each of them refuses a missing, wrong, and foreign token. AuthenticatedWrite already makes the check structural — a mutating handler cannot reach a session without it — but the residual risk is a new handler taking Authenticated by mistake, and that table is what catches it. A new endpoint under /api that is not in that list is a review catch.

E2E testing (real clients)

The E2E suite spins up complete environments using testcontainers-rs to run real ACME clients (certbot, acme.sh, lego) against the proxy.

The E2E suite is #[ignore]d by default to keep the main test cycle fast. You must have Podman or Docker running.

Run the E2E suite with:

cargo nextest run -E 'binary(e2e)' --run-ignored all
# or, with plain cargo:
cargo test --test e2e -- --ignored

Do not run cargo nextest run e2e. nextest’s bare positional filter matches against test names, not binary ids, and none of this suite’s test names contain the substring “e2e” — so that command silently matches nothing and reports 0 tests run rather than failing. The -E 'binary(e2e)' expression is what selects the binary.

Rootless Podman is auto-detected: the harness points DOCKER_HOST at the user’s podman socket if unset, and fails with a clear message naming systemctl --user start podman.socket rather than starting it itself.

The tests/e2e/common.rs harness automatically builds the necessary container images from the Containerfiles in the repository, provisions a dedicated podman network, and asserts on the container logs. It tests complex scenarios like Key Rollover (via lego), NetBox filter mocks, and full TLS-ALPN-01 responses.

Contributing to acme-proxy

Thank you for your interest in contributing to acme-proxy! Whether you’re fixing a bug, adding a new feature, or improving documentation, your help is welcome.

Development environment

To start developing, ensure you have the following installed:

  • Rust (latest stable version)
  • sqlite3 (for database inspection, though sqlx handles migrations)
  • mdBook (if you want to build this documentation locally)

Initial setup

Clone the repository and build the project:

git clone https://github.com/acme-proxy/acme-proxy.git
cd acme-proxy
cargo build

Testing

The suite is what holds RFC 8555 compliance in place, and CI enforces a coverage floor, so a change that adds a branch generally has to add a test for it.

Before submitting a pull request, run the full suite with nextest:

cargo nextest run --workspace

--workspace is not optional. The repository root is both the acme-proxy package and the root of a workspace of library crates under crates/, and a bare cargo command at such a root acts on the root package alone — the unit tests of every library crate would simply not run.

Use cargo nextest run --workspace, not cargo test. This is a requirement, not a preference: several tests execute a script file they have just written, and under cargo test — which runs tests as threads of a single process — another thread’s Command::spawn can fork while the file’s write descriptor is still open, failing with ETXTBSY roughly one run in three. nextest’s process-per-test isolation removes the race entirely. See Testing & Coverage.

What CI will check

Your pull request has to pass all of these:

cargo fmt --all --check
cargo clippy --workspace --all-targets -- -D warnings
cargo llvm-cov nextest --workspace --summary-only --fail-under-lines 97
cargo test --workspace --doc   # llvm-cov skips doc-tests
cargo deny check               # supply-chain audit, against deny.toml
RUSTDOCFLAGS="-D warnings -A rustdoc::private_intra_doc_links" \
  cargo doc --workspace --no-deps --all-features   # every intra-doc link
mdbook build doc/ && python3 doc/lint.py           # this book

cargo test --doc compiles the doc examples but not the intra-doc links, of which the workspace has a great many; cargo doc -D warnings is what catches a link a rename broke. Private intra-doc links are allowed on purpose: the library exists for the binary and the tests, and a public item explaining itself by naming the private thing it delegates to is the good outcome.

doc/lint.py holds the book to its own conventions: 80-column prose, no numbered headings, every fence tagged, every relative link and anchor resolving, every ADR listed, and no configuration key documented in two files — two copies of a default drift silently.

Four more jobs check what the ones above cannot:

  • msrv reads rust-version out of Cargo.toml and runs cargo check --locked --workspace --all-targets --all-features on exactly that toolchain, so the minimum stated there is one CI has verified.
  • hsm runs clippy and the suite with --features acme-proxy-signer/hsm against SoftHSM2. --all-targets enables no features, so without this job the PKCS#11 code would be neither linted nor tested; it is not folded into the coverage job, whose floor a feature-gated file sits outside of.
  • postgres runs the whole acme-proxy-store suite, tests/postgres.rs, roles and reload against a real PostgreSQL server, with ACME_PROXY_REQUIRE_POSTGRES=1 so a skipped test is a failure. It is separate from the coverage job for hsm’s reason.
  • e2e runs nightly, not on a push: a subset of the container lab in tests/e2e/, with real certbot, acme.sh and lego clients.

The sbom job additionally regenerates sbom.cdx.json and fails if it differs from the commit — see Changing dependencies.

Note the coverage floor is enforced, so new code generally needs new tests. cargo test --doc is the only thing that compiles the startup example in src/lib.rs.

Writing tests

  • Unit Tests: Keep them close to the code (in the same file, in a mod tests).
  • Integration Tests: Located in the tests/ directory. These tests spin up a full in-memory axum router and SQLite database to test the entire ACME flow.

See the Testing & Coverage page for more details.

Code style

  • Format your code using cargo fmt.
  • Ensure all lints pass by running cargo clippy --workspace --all-targets -- -D warnings.
  • Document public APIs using rustdoc comments (///).
  • Comments, doc comments and error-message strings are written in English, as are identifiers and log messages.
  • Every tracing call carries event = "<subsystem>_<object>_<outcome>" as its first field, as a string literal rather than a computed value, so the name stays greppable. Several are asserted by the end-to-end suite — grep before renaming one.
  • The crate is edition 2024; see rust-version in Cargo.toml for the minimum toolchain.

Changing the database schema

Both migration directories are append-only: crates/store/migrations/ (SQLite, since 0.1.0) and crates/store/migrations-postgres/ (PostgreSQL). Add a migration to each; never edit a committed one:

sqlx migrate add --source crates/store/migrations add_widget_table
sqlx migrate add --source crates/store/migrations-postgres add_widget_table

sqlx tracks each migration by a checksum, so editing a file that has already run turns every existing deployment into a startup failure. One build-system trap while you work: sqlx::migrate!() embeds the set at compile time and adding or removing a file under either directory does not on its own invalidate the build, so a test can be run against the previous set — touch crates/store/src/db.rs after changing the directory. This reverses the rule that held before the first release, when the server had never been deployed and a schema change meant editing the migration and running rm -f sqlite.db*.

Three consequences:

  • A new column is a new file, even when it plainly belongs to an existing table. ALTER TABLE ADD COLUMN is cheap; putting it in the original CREATE TABLE is what breaks. Name it in crates/store/src/transfer.rs’s manifest too, or acme-proxy transfer drops it; the_manifest_names_every_column refuses a manifest that has drifted.
  • A new CHECK, UNIQUE or foreign key needs a table rebuild in the SQLite set, because SQLite cannot add one to an existing table; PostgreSQL’s ALTER TABLE … ADD CONSTRAINT needs none. Write the rebuild in the new migration, and remember the two things a rebuild loses silently: an INSERT … SELECT drops any column you forget to name, and DROP TABLE takes the table’s indexes with it — including ones declared in an earlier migration, which will not run again to put them back.
  • A wrong declared width is a rebuild too. SQLite gives VARCHAR(n) TEXT affinity and enforces no length, so a width that no longer matches its data costs nothing at runtime and is wrong everywhere else — in what .schema tells an operator, and in any port to a dialect that does check. 20260826120000_declared_widths_for_random_tokens.sql is the worked example. Where the width follows a constant in src/, pin the two together with a test; that file’s VARCHAR(43) is TOKEN_BYTES and nothing else, so a change to the constant has to reach the schema.

Adding a configuration key

A key is a field on one of the section structs under crates/core/src/config/types/, with a #[serde(default)] that makes the whole section optional. Beyond the field itself, a new key owes:

  • Documentation in exactly one book page, as a ### Reference entry naming its environment variable, plus an entry in config.toml.example (which a test deserializes, so it cannot rot into invalid TOML). doc/lint.py refuses a key documented in two pages.
  • A decision about scope. A section listed in PROFILE_SECTIONS is per-profile and inherited key by key (see Profiles); one describing the process — [jobs], [audit], [metrics], [proxy], [admin] — is not.
  • A decision about reload. A reload rebuilds everything from the new configuration, so a key reloads unless something snapshots it at startup. Only database.url is refused on SIGHUP (FROZEN in crates/server/src/reload.rs); a new key joins it only with a reason.

A list-valued key has one more obligation, and one thing to know:

  • #[serde(deserialize_with = "string_list")] on the field. An environment variable can only carry a string, and this is what splits a,b into a list, at any depth: inside a profile or inside a named table ([filter.check.<name>], [notify.webhook.<name>], …) alike. It also reads a variable set to the empty string (a ${VAR:-} shell default) as []. Without it the key loads from a file and fails from the environment; every_list_field_reads_a_comma_separated_string refuses the omission.
  • A value containing a literal comma, such as a regex with {2,3}, can only be set from the file, since the comma is the separator.

The environment source pins prefix_separator("_"). Without it, config reuses the nested separator __ after the prefix and silently ignores every ACME_PROXY_* variable.

Changing a configuration key

The schema is the only frozen surface. Before 1.0.0, renaming or removing a configuration key is a normal change rather than one to design around — that is what keeps the code free of a compatibility layer for every shape a section has ever had. What such a change owes:

  • An entry in the changelog under the release’s ### Breaking heading, naming the old spelling and the new one. See Compatibility.
  • A startup error naming the replacement, where practical, so an unmigrated configuration stops the server instead of coming up looking configured and doing nothing. crates/policy/src/filter/build.rs’s refuse_removed_keys and the signer.backend = "acme_proxy" arm in crates/signer/src/lib.rs are the worked examples. A key must still parse to be refused by name, which is why the removed [filter] fields survive in crates/core/src/config/types/filter.rs; a field that is gone fails as an opaque serde error instead.
  • No alias, no dual syntax, no legacy lowering. Delete the old shape. The refusals themselves are one-line diagnostics and go away at 1.0.0.

Changing dependencies

sbom.cdx.json at the repository root is a committed CycloneDX 1.5 inventory of the dependency closure that ships in the binary — the artifact ASVS 5.0 V15.1.2 asks for, alongside the cargo deny gate. It is scoped --all-features --target all, so the hsm/cryptoki path and every platform-gated crate are covered; dev-dependencies are excluded, since they cannot reach a released build.

Regenerate it after any change to Cargo.toml or Cargo.lock, and when cutting a release (it records the crate version). The sbom CI job runs the same recipe and fails on any difference:

export SOURCE_DATE_EPOCH=0
cargo metadata --locked --format-version 1 >/dev/null
cargo cyclonedx --all-features --target all --spec-version 1.5 \
  --format json --override-filename sbom.cdx -q
jq --arg from "path+file://$PWD" --arg to "path+file:///acme-proxy" \
  'walk(if type == "string" then ((if startswith($from) then $to + .[($from | length):] else . end) | gsub("path\\+file:///acme-proxy#acme-proxy@"; "path+file:///acme-proxy#")) else . end) | del(.metadata.timestamp)' \
  sbom.cdx.json > sbom.cdx.json.tmp
mv sbom.cdx.json.tmp sbom.cdx.json
rm -f crates/*/sbom.cdx.json

The tool writes one document per workspace member. The committed one is the binary’s, whose closure already names every library crate, so the per-member copies are deleted rather than committed.

cargo install cargo-cyclonedx@0.5.9 --locked provides the generator; keep the version in step with the pin in .github/workflows/ci.yml, since it is written into the document. SOURCE_DATE_EPOCH makes the output reproducible (it also suppresses the otherwise-random serialNumber); the jq pass drops the wall-clock timestamp and rewrites the bom-ref values the tool derives from the checkout path — both the absolute directory it embeds and the name@ segment it drops when that directory’s basename happens to equal the crate name, so the file is identical whether it was regenerated in a worktree named acme-proxy or anything else.

Cutting a release

main is the trunk: every pull request targets it, and a minor release is a tag on it. A patch release comes from a release/X.Y branch, cut from the X.Y.0 tag when the first fix needs to ship. ADR 0013 argues the model.

Every crate of the workspace is published to crates.io together, at the binary’s version: the library crates are internal, with no semver promise of their own, and exist on crates.io only so cargo install acme-proxy can build.

A minor release

  1. On main, bump version in [workspace.package] of the root Cargo.toml and every =x.y.z pin on an acme-proxy-* crate in [workspace.dependencies], together. The exact pins are what keep the crates in step.

  2. Regenerate sbom.cdx.json (above); it records the version.

  3. Check the whole set packages and builds from its own archives, then publish it, in dependency order:

    cargo publish --workspace --dry-run
    cargo publish --workspace
    

    Publishing a workspace in one command needs cargo 1.90 or later, below the minimum supported Rust version.

  4. Once that commit is on main and its CI run is green, push the bare version as a tag:

    git tag -a 0.6.0 -m 0.6.0 && git push origin 0.6.0
    

    The tag triggers .github/workflows/release.yml, which builds the image on an amd64 and an arm64 runner and publishes ghcr.io/acme-proxy/acme-proxy as 0.6.0, 0.6 and latest, with a provenance attestation. ADR 0012 explains its shape. Its guard job stops the release, with nothing published, in four cases:

    • The tag is not the workspace version, or a crate pin is stale. The tag is most likely mistyped: delete it and push the right one. If step 1 was incomplete, fix the manifest first.
    • The tag is off its line. X.Y.0 must be on main, and X.Y.Z on release/X.Y. Move the tag.
    • There is no CI run on that branch for the commit. The tag is on a commit that was never pushed to the branch. Move the tag.
    • CI is pending or failed. Wait for it, or fix it, then use “Re-run all jobs” on the release run for the same tag.

    To rehearse the workflow without publishing, run it from a branch under “Run workflow”. It builds both architectures and pushes nothing.

    The first time the package is published, it is private, even though the repository is public. In the package’s settings, make it inherit access from the repository, then check that podman pull works with no credentials.

A patch release

  1. If release/X.Y does not exist yet, cut it from the minor’s tag and push it. The branch ruleset lets a new branch be created without a pull request:

    git switch -c release/0.6 0.6.0 && git push origin release/0.6
    
  2. Backport every fix it ships, as below.

  3. On a topic branch off release/X.Y, bump the version and the pins, regenerate the SBOM, and give the changelog its ## [X.Y.Z] section, holding the backported entries. Merge it into release/X.Y by pull request.

  4. Publish the crates from that commit, as in step 3 above.

  5. Once its CI run on release/X.Y is green, tag it: git tag -a 0.6.1 -m 0.6.1 on the branch’s head, then push the tag. The image is published as 0.6.1 and 0.6, and as latest only when no higher release exists.

  6. Cherry-pick the changelog section onto main, so the trunk’s changelog lists every release, and take the fixed entries out of [Unreleased] there.

Backporting a fix

A fix is merged on main first, and reaches a release branch as a cherry-pick, never the other way around:

git switch -c backport/0.6/fix-name origin/release/0.6
git cherry-pick -x <commit on main>

Open the pull request against release/0.6. The -x trailer names the commit on main the fix came from. When the cherry-pick conflicts, resolve it on the topic branch and say in the pull request what differs from the original.

Trying the next release

Every merge to main publishes ghcr.io/acme-proxy/acme-proxy:edge, once the whole of ci.yml has passed on it. The image job at the end of ci.yml calls release.yml for that.

Submitting a pull request

  1. Fork the repository and create your branch from main, fixes included. A fix is backported to a release branch after it merges (see above).
  2. Write clear, descriptive commit messages.
  3. If you’ve added code that should be tested, add tests.
  4. If you’ve changed APIs, update the documentation in this mdBook.
  5. Open a PR, describing the problem you’re solving and how you fixed it.

Architecture guidelines

If you are proposing a large feature (like a new Signer or Filter), please review the Architecture & Design documentation first. It’s often best to open an Issue to discuss the design before writing extensive code.

Open work

Planned and deferred work lives in the issue tracker, one issue per item, labelled by subsystem (server, store, webadmin, signer, ipam, notify). Several of them record why something was investigated and not built, which is worth reading before proposing it again.