Introduction
acme-proxy is an ACME server: the thing certbot, acme.sh, lego, Traefik and
Caddy talk to when they ask for a certificate. It serves the whole RFC 8555 flow
— account, order, authorization, challenge, finalize, certificate — plus
revocation, account key rollover (§7.3.5) and Renewal Information (RFC 9773).
What it deliberately does not implement is listed on
Protocol Support.
What it does behind that interface is the point. Clients prove control of
their names to acme-proxy, under whatever policy you configure, and it decides
how the certificate is actually produced — signing locally, relaying to a public
CA, or handing the request to a script.
The problem
Internal servers and IoT devices need TLS certificates, and cannot easily get them:
- Public CAs (like Let’s Encrypt) require DNS or HTTP validation, which internal servers hidden behind firewalls cannot easily satisfy.
- Handing out the company’s global DNS API credentials to every internal server to perform DNS-01 challenges is a massive security risk.
- Legacy internal PKI systems often do not speak the ACME protocol, forcing operators to write custom bash scripts to rotate certificates.
- When using external or commercial CAs, organizations are sometimes restricted to a single account or a limited number of External Account Binding (EAB) credentials validated for specific domains. Distributing these scarce upstream credentials directly to hundreds of internal servers is both impractical and risky.
The solution
acme-proxy solves this by standing between your internal clients and the
actual Certificate Authority. By registering a single account with the upstream
CA, it acts as an account multiplexer—allowing you to issue unlimited local
EAB credentials to your internal teams without exhausting your upstream limits.
graph LR
subgraph clients["Your network"]
A["Internal server<br/>certbot"]
B["IoT device<br/>acme.sh"]
C["Traefik / Caddy"]
end
P{{"acme-proxy"}}
subgraph backends["One of these actually signs"]
LOCAL["Embedded CA<br/>key on disk or in an HSM"]
UP["Upstream ACME CA<br/>Let's Encrypt, ZeroSSL, commercial"]
SCRIPT["Your script<br/>legacy PKI, internal API"]
end
A --> P
B --> P
C --> P
P --> LOCAL
P --> UP
P --> SCRIPT
Clients speak ordinary ACME to acme-proxy and prove control of their names to
it. Which of the three backends produces the certificate is a configuration
choice they never see — and it can differ per endpoint, so one process can serve
a local CA at /profile/dev and a Let’s Encrypt relay at /profile/prod.
- As a Proxy: It intercepts ACME requests from internal clients, opens a corresponding order with Let’s Encrypt, and safely solves the external DNS-01 challenges on the client’s behalf using a single, centrally secured DNS TSIG key.
- As a Filter: It inspects the client’s IP and requested DNS names against strict policies — including asking your IPAM (NetBox or phpIPAM) whether that address owns those names — before allowing the request to proceed.
- As a Local CA: It can operate entirely offline, issuing from an embedded ECDSA CA whose key is a file on disk or a PKCS#11 token.
- As a Multi-Tenant Server: Through its Profile system, a single binary can
host a Local CA on
/profile/devand a strictly-filtered Let’s Encrypt relay on/profile/prod. - As a Record: Every issuance and every refusal is written to an append-only audit trail, naming the actor, the address it came from and the names it asked for — the question “who got this certificate, and from where” has an answer.
It is built on axum, tokio and sqlx, and stores everything in one SQLite
file in WAL mode. There is no second datastore, no message queue and no
scheduler to operate.
It is published on crates.io, so
cargo install acme-proxy gets you the whole thing — server and admin CLI in
one binary. See Installation.
Core Concepts & Glossary
Eight words carry most of the meaning in the rest of this book. This page defines them once, in the order you meet them, so that every other page can use them without re-explaining.
Profile
A profile is an independent ACME endpoint, and the isolation boundary
everything else sits inside. Rather than running one process per environment,
you define several profiles in one; each is served at
/profile/<name>/directory.
Accounts and orders are isolated per profile — the same client key at two
profiles is two unrelated accounts — and each profile carries its own signer,
filters, challenge validation and EAB policy. A dev profile backed by a local
CA can sit beside a prod profile relaying to Let’s Encrypt under strict NetBox
filtering, in one process, over one socket and one database.
See Profiles & Routing.
Signer
A signer is what actually produces the certificate once a client has been authorized. Which one runs is a per-profile configuration choice, and the client never sees the difference.
- Local CA — an embedded certificate authority signing directly. The issuing key is a file, or a PKCS#11 token.
- Relay — opens its own order with an upstream ACME CA and returns what that CA signs.
- Custom script — anything else: a legacy PKI, an internal API, a CA that does not speak ACME.
See Signers.
Filter
A filter is a policy applied to a request before anything is signed. Filters answer “may this client ask for this?”, which challenge validation does not: a client can genuinely control a name and still have no business holding a certificate for it from you.
[filter] is a small policy engine rather than a list of switches. A check
is one named question about a request — “is this address in the management
network?” — and a rule is a boolean expression over check names plus what a
match means. filter.rules says which rules run and in what order, and the
first match wins; a stage where a rule was applicable and none matched falls
to filter.default.
Rules act at two points — on the connection, and on the identifiers, the latter
running again at finalize against the names in the CSR.
See Filters.
EAB (External Account Binding)
External Account Binding (RFC 8555 §7.3.4) makes newAccount require a
credential you minted out of band — a key identifier and an HMAC secret.
Reaching the directory is then no longer enough to register: an operator has to
have issued that client a credential first.
It runs in the other direction too. A commercial CA that granted you one scarce EAB credential is exactly the case the relay backend exists for: one upstream credential, any number of local ones.
ARI (ACME Renewal Information)
ACME Renewal Information (RFC 9773) lets the CA tell a client when to renew, rather than leaving it to guess from the expiry date. Two things follow: a fleet spreads its renewals across a window instead of stampeding at the same moment, and a CA that needs certificates replaced early can say so and be listened to.
See Renewal Information.
Order
An order is a client’s request for a certificate. It names the identifiers wanted and progresses through the states RFC 8555 defines:
pending— created; one or more authorizations still need to be satisfied.ready— every authorization isvalid; the client may nowfinalize.processing— issuance is under way but not finished.valid— the certificate is available.invalid— terminal failure.
stateDiagram-v2
[*] --> pending: newOrder
pending --> ready: every authorization valid
ready --> pending: an authorization is deactivated (§7.5.2)
ready --> processing: finalize
processing --> valid: the worker signed
processing --> invalid: signer refused or gave up
pending --> invalid: an authorization failed, or expires passed
valid --> [*]
invalid --> [*]
acme-proxy enforces these states and transitions in the database itself. The
set of states is a CHECK constraint on the status column — see
Database Schema
— and every transition is an UPDATE guarded on the state it leaves. A
validation or a signing that finishes after the order moved on, because the
client deactivated an authorization or a sibling challenge already decided it,
therefore changes nothing above its own challenge.
ready → pending is the one backwards edge, and it exists only so §7.5.2 can
hold: deactivating an authorization on an order that already reached ready has
to demote it, or the order would be finalizable for a name no longer authorized.
Two details are easy to trip on:
- Every
finalizeanswersprocessing. Signing needs the CA key, which only theworkerrole holds, sofinalizechecks the CSR, claims the order and queues the signing, and the client polls until the order isvalid. Withlocal_cathat is a moment; withrelayit is as long as the upstream CA takes. A CSR the backend itself rejects makes the orderinvalidwith abadCSRerror, since the client is already polling by then; a CSRfinalizecan refuse on its own leaves the orderreadyfor a corrected one. - Revocation is orthogonal to this machine. RFC 8555 defines no “revoked”
order status, so a revoked order’s
statusstaysvalid. The revocation timestamp and reason are recorded separately, and both admin front ends show them and can revoke —acme-proxy order show/order revoke, and the order detail page in the panel. See Revocation & CRL.
Job
A job is one unit of work the server owes itself: a relayed issuance to finish, a notification to deliver, a table to sweep. Jobs are rows in the same SQLite file as everything else, drained by one runner per process, so they survive a restart and need no scheduler beside the server.
What is worth carrying away is how a handler reports failure. Retry says the
attempt decided nothing — a refused connection, a proxy, a 503 — and the job
goes back in the queue under a growing backoff; Failed says the other side
stated a reason and is believed at once. That split is what keeps a client’s
order processing through a five-second upstream blip rather than terminally
invalid, and it is why an order that is not progressing is a question for
acme-proxy jobs list before it is a question for anything else.
Challenge
A challenge is the concrete proof that a client controls an identifier: serving a token over HTTP, publishing a DNS TXT record, or presenting a special certificate in a TLS handshake.
Each authorization carries one challenge per enabled type, and satisfying any
one of them makes the authorization valid — the others stay pending for
ever, which is correct rather than a stuck state.
Triggering one queues the check rather than performing it, so a triggered
challenge answers processing and the client polls it.
See Challenge Validation.
Quick Start
Get acme-proxy running in a couple of minutes with an auto-generated Local CA.
Step 1 — Write a configuration file
acme-proxy serves ACME only through profiles, so at least one enabled
profile is required — the server refuses to start without one. Everything else
has a working default.
By default the server also performs real domain-control validation (HTTP-01). To
test a client against the proxy locally without routing port 80 or configuring
DNS, this quick start turns validation off with challenge.bypass.
Create config.toml in your current directory:
[challenge]
# Testing only. See "Moving to production" below.
bypass = true
[profiles.default]
[profiles.default]is not empty by accident — a profile’senabledkey defaults totrue, so naming the profile is all that is required. The profile name becomes part of the URL, and of everykida client stores.
Step 2 — Run the server
acme-proxy serve
serve is also the default, so a bare acme-proxy does the same thing.
There is no --config flag. The server reads config.toml from the current
working directory; point it elsewhere with ACME_PROXY_CONFIG:
ACME_PROXY_CONFIG=/etc/acme-proxy/config.toml acme-proxy serve
(The extension may be omitted — the format is then inferred.) A missing configuration file is not an error: the server falls back to defaults, which is why an environment-only deployment works.
Individual keys can also be overridden with ACME_PROXY_* environment
variables; see the Configuration Reference.
The server binds [::]:3000 and serves the ACME directory at
http://localhost:3000/profile/default/directory. On first run it creates
sqlite.db, ca.pem and ca.key in the working directory.
Step 3 — Request a certificate
Point any standard ACME client at the profile’s directory URL. Examples for
internal.example.com:
certbot
certbot certonly \
--server http://localhost:3000/profile/default/directory \
--standalone \
--domain internal.example.com \
--email admin@example.com \
--agree-tos \
--no-eff-email
acme.sh
acme.sh --issue \
--server http://localhost:3000/profile/default/directory \
-d internal.example.com \
--standalone
lego
lego --server http://localhost:3000/profile/default/directory \
--email admin@example.com \
--domains internal.example.com \
--http \
run
Traefik
In Traefik’s static configuration (traefik.yml):
certificatesResolvers:
myresolver:
acme:
caServer: http://localhost:3000/profile/default/directory
email: admin@example.com
httpChallenge:
entryPoint: web
Caddy
In your Caddyfile:
internal.example.com {
tls admin@example.com {
ca http://localhost:3000/profile/default/directory
}
respond "Hello, world!"
}
What just happened?
- The client fetched the directory from
acme-proxy. - It registered an account and created a new order.
acme-proxyoffered an HTTP-01 challenge.- Because
challenge.bypass = true, the proxy marked the challengevalidthe moment the client triggered it, without making any network request back to the client’s responder. - The proxy signed the CSR with its auto-generated local ECDSA CA and returned the certificate.
The certificate is signed by a CA nothing trusts yet. See Trusting the
CA for how to install ca.pem where your clients will
accept it.
Step 4 — Moving to production
The defaults are safe; this quick start deliberately relaxed them. Before exposing the server:
- Remove the bypass. Delete
challenge.bypass = truesoacme-proxyactually validates domain control. With bypass on,[filter]is the only access control there is — which is exactly why bypass is not the default. See Challenge Validation. - Enable filters. Configure
allowed_ip,identifiersoripamto restrict which clients may request which names. See Filters & Policies. - Choose where state lives. The defaults write
sqlite.dband the CA material to the current working directory. Setdatabase.urlandsigner.local_ca.cert_path/signer.local_ca.key_pathto permanent paths — note these are the CA’s files, distinct fromserver.tls.cert_path/server.tls.key_path, which belong to the HTTPS listener. - Serve over HTTPS. RFC 8555 §6.1 expects ACME over HTTPS: either put the server behind a reverse proxy, or turn on TLS termination.
Installation
acme-proxy is a Rust application, published on
crates.io. There are four ways to get
it: cargo install, a source build, the published container image, or a
container image you build yourself. Only the published image comes prebuilt; no
standalone prebuilt binaries are published.
The result is a single binary carrying both the server and the admin CLI, so a deployment never needs a second tool.
Prerequisites
- Rust toolchain: the crate is edition 2024 and declares a
rust-versioninCargo.toml(currently 1.97). That file is the source of truth;cargorefuses to build with anything older. - Cargo: the Rust package manager.
SQLite is not a prerequisite: the driver is bundled with sqlx, the database
file is created automatically, and DATABASE_URL is not needed to compile.
The PostgreSQL driver is bundled the same way, so the same binary speaks both —
what it does need is a server and a database that already exist. A
sqlite3 binary is only useful if you want to inspect the database by hand.
You can install Rust via rustup:
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
From crates.io
cargo install acme-proxy
This fetches the published crate, compiles it, and puts the binary in
~/.cargo/bin — which needs to be on your PATH. Confirm it with:
acme-proxy --version
Tab completion and a man page come out of the binary itself, so there is nothing extra to download:
acme-proxy completions zsh > ~/.zfunc/_acme-proxy
acme-proxy man | sudo tee /usr/share/man/man1/acme-proxy.1 > /dev/null
Both are generated from the command tree, so regenerate them when you upgrade. The per-shell paths are in the admin CLI chapter.
To pin a version, or to move to a specific one later, name it:
cargo install acme-proxy --version 0.6.0
Upgrading is cargo install acme-proxy --force. Before doing so across a minor
version, read the ### Breaking section of the
changelog.
Before 1.0.0 the database schema is the only compatibility guarantee, so the
data survives an upgrade but a configuration key may have been renamed.
Building from source
Prefer this if you intend to change anything, or want the test suite and the book sources alongside the binary.
-
Clone the repository:
git clone https://github.com/acme-proxy/acme-proxy.git cd acme-proxy -
Build the project in release mode for production use:
cargo build --release
The binary will be located at target/release/acme-proxy.
Optional features
The default build has no optional features. One is available:
| Feature | What it adds |
|---|---|
hsm | PKCS#11 support for the Local CA’s issuing key, so it can live in a YubiKey or an HSM instead of a file — see Hardware Keys. |
cargo install acme-proxy --features hsm # or, from a clone:
cargo build --release --features hsm
It is off by default because it pulls in cryptoki and its bindings, which a
deployment signing with an on-disk key does not need. The PKCS#11 module itself
is loaded at runtime, so enabling this adds no build-time C toolchain
requirement. Configuring signer.local_ca.key_source = "pkcs11" on a binary
built without it is a startup error naming the feature, never a silent fallback
to the file key.
The published container image is the default
build, without hsm. A deployment that needs it builds its own binary.
Container (Docker / Podman)
Every release is published to the GitHub Container Registry as a
multi-architecture image, for linux/amd64 and linux/arm64:
podman pull ghcr.io/acme-proxy/acme-proxy:0.6.0
| Tag | Points at |
|---|---|
0.6.0 | That release. It never moves. |
0.6 | The newest 0.6.x release: its fixes, never a new minor. |
latest | The highest release, whatever its minor. |
edge | The head of main, rebuilt on every merge. Not a release. |
sha-… | One commit of main, as edge was when it was built. |
Run a version tag, or 0.6 to take patch releases without a change on your
side. Avoid latest for an unattended deployment. Before 1.0.0 a minor
release may rename a configuration key, so a pull of latest can stop a server
from starting; the ### Breaking sections of the
changelog
list every such change. edge is for trying what the next release will hold,
never for a certificate authority anyone depends on.
Every image holds the release build with the default features: the binary
cargo install acme-proxy produces.
To build the image yourself instead, use the Containerfile in a clone of the
repository:
podman build -t acme-proxy .
That is the same release build as the published image, fat LTO included, so it takes tens of minutes.
The image’s working directory is /data and its entrypoint is the acme-proxy
binary, so mount a volume there for the SQLite database, the configuration and
the CA key material — all of which default to paths relative to the working
directory. The image runs as a non-root user, so the mounted directory must be
writable by it — the :U flag below is the rootless-Podman way; see
Deployment for Docker.
podman run -d \
-p 3000:3000 \
-v ./data:/data:U \
ghcr.io/acme-proxy/acme-proxy:0.6.0
Drop a config.toml into ./data (it must define at least one profile — see
the Quick Start), or configure the container entirely through
ACME_PROXY_* environment variables:
podman run -d \
-p 3000:3000 \
-v ./data:/data:U \
-e ACME_PROXY_PROFILES__DEFAULT__ENABLED=true \
-e ACME_PROXY_SERVER__BASE_URL=https://acme.example.com \
ghcr.io/acme-proxy/acme-proxy:0.6.0
Verifying the image
Each published image carries a signed build provenance attestation. It records the repository, the commit and the workflow run that built the image. Check it with the GitHub CLI before you run the image:
gh attestation verify oci://ghcr.io/acme-proxy/acme-proxy:0.6.0 \
--repo acme-proxy/acme-proxy
The attestation is on the multi-architecture index, the object a tag resolves to. The digest of one architecture’s image, on its own, has no attestation to find.
Trusting the CA
With the local_ca signer, acme-proxy mints
certificates from a CA it generated itself. Those certificates are perfectly
valid, but nothing on your network trusts the CA that signed them yet — so
browsers, curl, and every TLS library will reject them until you install the
root.
This page covers distributing that root. It does not apply when you use the
relay backend to relay to a public CA, whose
roots are already trusted everywhere.
Getting the root certificate
Each profile serves its own CA material unauthenticated at
{base_url}/profile/<name>/ca.pem, as application/x-pem-file:
curl -o internal-root.pem https://acme.internal/profile/default/ca.pem
These are exactly the bytes appended to every certificate that profile issues, so a client that fetches them here and one that reads the tail of its own chain end up trusting the same anchor.
Two things worth knowing about the route. It is not advertised in the ACME
directory — it is CA infrastructure rather than an ACME resource, so a client
will not find it on its own and you distribute the URL yourself. And it is
served inside the profile router, which means it sits behind that profile’s
filter policy: if you restrict the endpoint by address, the hosts that most
need the root — the ones that do not have it yet — may be exactly the ones
refused. Add a path check allowing /ca.pem if so.
The same file is on the server’s disk at signer.local_ca.cert_path, ca.pem
by default in the working directory, which is the way to get it when the server
is not reachable or is not running:
scp acme-host:/var/lib/acme-proxy/ca.pem ./internal-root.pem
Inspect it before distributing it:
openssl x509 -in internal-root.pem -noout -subject -issuer -dates -ext basicConstraints
A freshly generated root is self-signed (subject equals issuer) and carries
CA:TRUE, pathlen:0.
If cert_path holds a bundle — an intermediate followed by a root, as in
the multi-tier
setup — then
the last certificate in the file is the root, and it is the only one your
clients need to trust. The intermediate is shipped with every issued certificate
and does not need installing.
# Split a bundle into its constituent certificates.
csplit -z -f cert- -b '%02d.pem' ca_bundle.pem '/-----BEGIN CERTIFICATE-----/' '{*}'
Installing it
Debian / Ubuntu
The file must have a .crt extension, and must be PEM despite the name.
sudo cp internal-root.pem /usr/local/share/ca-certificates/acme-proxy-root.crt
sudo update-ca-certificates
RHEL / Fedora / CentOS
sudo cp internal-root.pem /etc/pki/ca-trust/source/anchors/acme-proxy-root.pem
sudo update-ca-trust extract
Alpine
sudo cp internal-root.pem /usr/local/share/ca-certificates/acme-proxy-root.crt
sudo update-ca-certificates
Verify
curl -v https://internal.example.com 2>&1 | grep -i 'SSL certificate verify'
# or, without a server:
openssl verify -CAfile internal-root.pem issued-cert.pem
Applications with their own trust store
Updating the system store is not enough for everything. These maintain their own:
| Runtime | How to add the root |
|---|---|
| Firefox | Its own store, always. Settings → Privacy & Security → Certificates → View Certificates → Authorities → Import. Enterprise deployments can use the Certificates policy in policies.json. |
| Chrome / Edge | Uses the system store on Windows and macOS; on Linux it reads the NSS database — certutil -d sql:$HOME/.pki/nssdb -A -t "C,," -n acme-proxy-root -i internal-root.pem. |
| Java / JVM | keytool -importcert -trustcacerts -alias acme-proxy-root -file internal-root.pem -keystore "$JAVA_HOME/lib/security/cacerts". |
| Node.js | Ignores the system store by default. Set NODE_EXTRA_CA_CERTS=/path/to/internal-root.pem. |
Python requests | Uses certifi, not the system store. Set REQUESTS_CA_BUNDLE (or SSL_CERT_FILE for ssl/urllib). |
| Go | Uses the system store on Linux; no action needed after update-ca-certificates. |
| Containers | Each image has its own store. Mount the root in and run the distribution’s update command in your Dockerfile, or bake it into a base image. |
Distributing at scale
Installing a root by hand does not survive a fleet. In practice:
- Ansible / Puppet / Chef — ship the file and run the update command as a handler. This is the common approach for Linux estates.
- Active Directory Group Policy — Computer Configuration → Windows Settings → Security Settings → Public Key Policies → Trusted Root Certification Authorities.
- MDM (Jamf, Intune, …) — deploy as a certificate payload.
- Golden images — bake the root into your base image so new hosts trust it from first boot.
Whichever you use, deploy the root before you start issuing certificates from it, or the first clients to renew will break.
Revocation
If you revoke certificates, clients need to be able to see the CRL. It is served
unauthenticated at {base_url}/profile/<name>/crl as application/pkix-crl:
curl -o internal.crl https://acme.internal/profile/default/crl
openssl crl -in internal.crl -inform DER -noout -text
Like /ca.pem, the CRL is not advertised in the ACME directory and sits behind
the profile’s filter policy. Issued certificates carry a CRL distribution point
only when you set signer.local_ca.crl_distribution_points; leave it unset and
a client will not find the CRL automatically, so distribute the URL alongside
the root if your validation policy needs it. See
Revocation & CRL.
Planning ahead
The root’s validity is finite, and replacing it later means touching every host that trusts it. Two things make that easier:
- Use an intermediate. Keep an offline root and hand
acme-proxyonly an intermediate. The root you distribute then long outlives any single signing key, and a compromised proxy costs you an intermediate rather than your whole trust anchor. See Multi-Tier PKI. - Distribute early, rotate overlapping. Trust stores accept multiple roots, so push a replacement root well before it is needed and remove the old one only after nothing is signed by it.
Deployment
While acme-proxy can be run in a container, many organizations prefer running
infrastructure components directly on standard Linux VMs using systemd.
systemd service setup
Below is an example systemd service file that runs acme-proxy securely.
-
Create a dedicated user:
sudo useradd -r -s /bin/false acme-proxy -
Prepare directories:
sudo mkdir -p /etc/acme-proxy sudo mkdir -p /var/lib/acme-proxy sudo chown acme-proxy:acme-proxy /var/lib/acme-proxy -
Create the service file: Create
/etc/systemd/system/acme-proxy.service:[Unit] Description=ACME Proxy Server After=network.target [Service] Type=simple User=acme-proxy Group=acme-proxy ExecStart=/usr/local/bin/acme-proxy serve ExecReload=/bin/kill -HUP $MAINPID WorkingDirectory=/var/lib/acme-proxy # Configuration. The extension may be omitted, in which case the format # is inferred. Environment="ACME_PROXY_CONFIG=/etc/acme-proxy/config.toml" Environment="ACME_PROXY_DATABASE__URL=sqlite:///var/lib/acme-proxy/acme.db" # Security / Sandboxing ProtectSystem=strict ReadWritePaths=/var/lib/acme-proxy ProtectHome=true PrivateTmp=true NoNewPrivileges=true Restart=on-failure RestartSec=5 [Install] WantedBy=multi-user.target
ProtectSystem=strict makes the whole filesystem read-only except
ReadWritePaths, so everything the server writes must land in
/var/lib/acme-proxy. With WorkingDirectory set there, the defaults already
do: signer.local_ca.cert_path (ca.pem), key_path (ca.key), crl_path
(ca.crl) and the lock beside it are all resolved relative to the working
directory, as are server.tls.cert_path / key_path if you enable TLS.
If you set any of them to an absolute path, add that path to ReadWritePaths
too.
acme-proxyshuts down gracefully onSIGTERM(and on Ctrl+C when run in a terminal), sosystemctl restartandsystemctl stoplet in-flight requests finish rather than cutting them off. Both listeners stop together. It also reloads its configuration onSIGHUPwithout restarting, which is what theExecReloadline above wires up — see Reloading the Configuration for what a reload may change and what it refuses. One case still deserves a quiet period: a request that waits — acustomscript’scrlorrenewal_infohook, or a relay revocation waiting on its job — can take up toserver.request_timeout_ms. If systemd’sTimeoutStopSec(90 s by default) is shorter than that, systemd sendsSIGKILLfirst and the graceful path is skipped — raise it, or lower the request timeout. Challenge validation and issuance are not among these: they run in the job queue, and a job left unfinished by a restart is reclaimed by lease expiry.
- Enable and start the service:
sudo systemctl daemon-reload sudo systemctl enable --now acme-proxy
Where each socket belongs
acme-proxy opens two listeners when the web admin is enabled, and they
belong on different sides of your boundary. The ACME listener answers
unauthenticated clients by design; the admin listener has no filter chain and no
admission control, and its only access controls are the bind address, TLS and
the session.
graph TD
subgraph internal["Internal network"]
CLIENTS["ACME clients<br/>certbot, acme.sh, Traefik"]
OPS["Operator workstation"]
end
subgraph host["The acme-proxy host"]
RP["Reverse proxy (optional)<br/>sets X-Forwarded-For"]
ACME[":3000 — ACME listener<br/>filters, admission, nonces"]
ADMIN[":3001 — admin listener<br/>loopback by default"]
DB[("sqlite.db + WAL")]
CAKEY[["ca.key — 0600, or a PKCS#11 token"]]
end
CLIENTS --> RP --> ACME
OPS -.->|"SSH tunnel or VPN,<br/>NOT an open port"| ADMIN
ACME --> DB
ADMIN --> DB
ACME --> CAKEY
ACME -->|"challenge validation:<br/>back to the client, :80 / :443 / DNS"| CLIENTS
ACME -->|"upstream ACME, DNS updates, SMTP"| OUT(["Egress"])
Two edges are the ones people get wrong:
- The dotted one. If the admin listener is reachable from anywhere but
loopback, startup refuses unless
admin.tls.enabledis on — and even then, a tunnel is the better answer. See below. - The validation edge points back at the client. With
challenge.bypass = false, the server opens connections to the machines asking for certificates. A firewall that only permits inbound traffic leaves orders sitting atpending.
Reverse proxy (optional)
acme-proxy acts as an HTTP server, typically binding to port 3000. You can
bind it directly to 80 (requires CAP_NET_BIND_SERVICE) or place it behind a
reverse proxy like Nginx or Traefik, which can provide TLS termination for the
ACME API itself.
Two things to get right when proxying:
- Set
server.base_urlto the public URL. It is what the directory advertises and what every signed request is checked against (RFC 8555 §6.4), so a mismatch rejects every client. It is never derived from the request. - If you want IP-based filters to see the real client rather than the proxy, set
filter.trusted_proxiesto the proxy’s addresses and, if it is notx-forwarded-for,filter.forwarded_header. Note these are[filter]keys, not[server]keys.
Alternatively, skip the reverse proxy and let acme-proxy terminate TLS itself
— see TLS Termination.
Exposing the web admin (or rather, not)
The Web Admin is a second listener and is off
by default. When you turn it on, it binds 127.0.0.1:3001 and stays there
unless you say otherwise.
The recommended way to reach it is an SSH tunnel, which needs no configuration change and no second certificate:
$ ssh -N -L 3001:127.0.0.1:3001 ca.example.com
Then open http://localhost:3001. admin.base_url stays at its default,
because from the browser’s point of view the panel really is on localhost.
If you must bind it to a real interface, TLS is mandatory — startup refuses
a non-loopback bind while admin.tls.enabled is false, because the session
cookie is sent Secure and a browser silently declines to store one over plain
HTTP anywhere but localhost:
[admin]
enabled = true
bind_address = "0.0.0.0:3001"
base_url = "https://admin.example.com:3001"
[admin.tls]
enabled = true
Two things this listener does not have, deliberately: admission control and
a filter chain. Access control here is the bind address, TLS, and the session.
Note also that it does not honour X-Forwarded-For — behind a reverse proxy
the sign-in rate limiter counts the proxy, which is one more reason to prefer
the tunnel.
Under systemd, nothing extra is needed: the panel shares the process, the unit and the database. Bootstrap the first operator once, before or after enabling it:
$ printf '%s' "$PASSWORD" | acme-proxy admin user create alice
Container deployments (Docker / Podman)
For containerized environments, you can run acme-proxy using Docker Compose or
Podman.
Each release is published as ghcr.io/acme-proxy/acme-proxy:<version>, for
linux/amd64 and linux/arm64, and the examples below pin one.
Installation covers why to pin,
building the image yourself from the Containerfile, and verifying its
provenance:
podman pull ghcr.io/acme-proxy/acme-proxy:0.6.0
The image’s working directory is /data and its entrypoint is the binary
itself, so /data is where the database, the CA key material and the CRL land
unless you override their paths.
The image runs as a non-root user (acme-proxy, uid/gid 1000), so the
directory you mount at /data must be writable by that uid. How you arrange
that depends on the runtime:
- Rootless Podman — add
Uto the mount flags (-v ./data:/data:U, or:Z,Uon SELinux systems). Podman then chowns the volume’s contents to the user the container runs as. Alternatives:podman unshare chown -R 1000:1000 ./databeforehand, or use a named volume (-v acme-proxy-data:/data), which Podman initializes with the right owner. - Docker / Docker Compose — create the directory owned by uid
1000before the first run:mkdir -p ./data && sudo chown 1000:1000 ./data. Or override the uid to your own (--user "$(id -u):$(id -g)", or a Composeuser:line) and own./datayourself — the server only needs to read and write that one directory.
Docker Compose
Create a docker-compose.yml file:
services:
acme-proxy:
image: ghcr.io/acme-proxy/acme-proxy:0.6.0
container_name: acme-proxy
restart: unless-stopped
ports:
- "3000:3000"
volumes:
- ./data:/data
environment:
- ACME_PROXY_PROFILES__DEFAULT__ENABLED=true
- ACME_PROXY_DATABASE__URL=sqlite:///data/acme.db
- RUST_LOG=acme_proxy=info
Create the data directory with the right owner before the first run:
mkdir -p ./data && sudo chown 1000:1000 ./data
Run the stack using:
docker compose up -d
ACME_PROXY_PROFILES__DEFAULT__ENABLED=true is what defines the profile when
there is no configuration file — the server serves nothing without at least one.
For anything beyond a single default profile, mount a config.toml into /data
instead.
Podman (rootless)
Under rootless Podman you can run the container directly. The U flag chowns
the mounted directory to the non-root user the container runs as; add Z as
well on SELinux-enabled systems (RHEL/Fedora) for the mount label.
podman run -d --name acme-proxy \
-p 3000:3000 \
-v ./data:/data:U \
-e ACME_PROXY_PROFILES__DEFAULT__ENABLED=true \
-e ACME_PROXY_DATABASE__URL=sqlite:///data/acme.db \
ghcr.io/acme-proxy/acme-proxy:0.6.0
Running the roles as separate processes
acme-proxy serve runs three jobs in one process: serving ACME, serving the web
admin, and draining the job queue. --role splits them across processes of the
same binary, reading the same configuration.
| Role | Does | Holds |
|---|---|---|
acme | Serves ACME to certificate clients | The ACME listener |
admin | Serves /ui and /api | The admin listener |
worker | Drains the job queue; signs and revokes; owns the schema and the first-run material | The CA key or token, a relay’s upstream account; no listener |
All-in-one is still the default and nothing about it changes: acme-proxy serve with no --role behaves exactly as it always did. The split is worth
doing when you want privilege separation — the process parsing untrusted JWS and
CSRs from the internet is then not the one holding operator sessions, and
neither is the one making outbound connections to client-chosen hosts. Each can
run under its own uid and its own systemd sandbox.
Only the worker holds signing material. finalize queues the signing and
answers processing, a revocation is a database row or a queued job, and the
CA’s certificate, its CRL and renewal information are served from ca.pem and
the database. So the acme and admin processes never read ca.key, never log
in to a PKCS#11 token and never use a relay’s upstream account. Make ca.key
(0600) readable by the worker’s uid alone; the others need ca.pem and the
database. A custom signer’s script must still be present where acme runs if
it serves the CRL or renewal information, since those hooks answer a request.
Three things to get right:
- Initialise once, first.
acme-proxy initmigrates the database and generates the CA key, the upstream account and any self-signed TLS certificate. Run it as the uid that should own those files. A process that does not runworkerrefuses to start against a schema that is behind, namingacme-proxy migrate, and without the CA certificate, namingacme-proxy init— so starting the others before the schema or the CA exists fails loudly rather than racing, and never generates a second CA. - Give each process its own
metrics.bind_address. The counters are per-process memory, so three processes are three scrape targets; sharing one address means the second one to start fails to bind. SetACME_PROXY_METRICS__BIND_ADDRESSper unit, point each at its own file withACME_PROXY_CONFIG, or turnmetrics.enabledoff where you do not want it. Every series carries arolelabel naming the roles that process runs, so one scrape config can tell them apart. - Run at least one worker. A process without it logs
server_role_no_workerat startup; a deployment without one issues nothing, because challenge validation, issuance, CRL signing, relay and custom revocations, notifications and the periodic sweeps are all queued work. A worker in another process picks a row up within the job poll interval, so clients see a second or so moreprocessingthan all-in-one; a revocation for a relay or custom profile still running when its request’s deadline nears answers503withRetry-After, and asking again follows it.
A worked topology, one systemd unit per role:
# acme-proxy-worker.service
ExecStart=/usr/local/bin/acme-proxy serve --role worker
Environment=ACME_PROXY_METRICS__BIND_ADDRESS=127.0.0.1:3002
# acme-proxy-acme.service
ExecStart=/usr/local/bin/acme-proxy serve --role acme
Environment=ACME_PROXY_METRICS__BIND_ADDRESS=127.0.0.1:3012
# acme-proxy-admin.service
ExecStart=/usr/local/bin/acme-proxy serve --role admin
Environment=ACME_PROXY_METRICS__BIND_ADDRESS=127.0.0.1:3022
One host, one filesystem — on SQLite. SQLite across processes is fine on a
local disk in WAL mode, and busy_timeout is already set — it is not safe on
NFS or across nodes. Multi-node needs PostgreSQL: point
database.url at a server instead of
a file, run acme-proxy migrate once, and the three roles can then live on
different hosts. Nothing else about the deployment changes.
An existing SQLite deployment moves across with its accounts, orders and audit
trail intact — stop the server, create and migrate the target, then
acme-proxy transfer --to <url>.
Starting the new deployment empty instead would leave every certificate it has
already issued impossible to revoke, so this is not an optional step.
Run one admin process either way; its login rate limiter is in memory, so two would each get their own budget.
Each process reloads independently on SIGHUP, so a configuration change means
reloading all three.
Upgrading
Replace the binary, migrate, and restart. The schema is append-only as of 0.1.0 — a new release only ever adds migrations, never rewrites the ones your database has already applied.
Migrations no longer run as a side effect of opening the database. acme-proxy serve running the worker role (which the default does) still applies them at
startup, so a single-process deployment can simply restart; anything else — an
admin command, or a split deployment’s acme/admin process — checks the
schema and refuses by name until acme-proxy migrate has run.
systemctl stop acme-proxy
install -m 0755 acme-proxy /usr/local/bin/acme-proxy
acme-proxy migrate # explicit; the default `serve` would also do it
systemctl start acme-proxy
journalctl -u acme-proxy -n 50
A container is upgraded the same way, with the image tag in place of the binary:
pull the new tag, migrate with it against the same /data, then recreate the
container on it. With Docker Compose, after changing the image: line:
docker compose pull
docker compose stop acme-proxy
docker compose run --rm acme-proxy migrate
docker compose up -d
docker compose logs -n 50 acme-proxy
migrate after the service name replaces the image’s default serve command
for that one container. A single-container deployment would also migrate as it
starts; running it as a step of its own stops the upgrade on the error, instead
of leaving a server in a restart loop. A split deployment must migrate before
any of its acme or admin containers start on the new tag. The advice below
applies unchanged, the database backup first of all.
Worth knowing before you do it:
- Take a copy of the database first. SQLite in WAL mode means three files;
copy them together with the server stopped, or use
sqlite3 acme.db ".backup backup.db"on a running one. Migrations are not reversible, so a downgrade means restoring this copy. db_migration_failedat startup means the process is not serving. The most likely cause is running an older binary against a database a newer one has already migrated.- Files outside the database are untouched. The CA key and certificate, the
exported
ca.crl, the upstream account key and its.kidsidecar all persist across an upgrade — back them up on the same schedule as the database, since the CA key is the one thing that cannot be regenerated without redistributing trust. See Trusting the CA. - Read the changelog’s
Breakingsection first. Before 1.0.0 the schema is the only compatibility guarantee: configuration keys, profile names, the admin JSON API, log event names and the CLI may all have moved, and every such change is listed there. See Compatibility. A renamed key is normally refused by name at startup — the server stops with an error naming the replacement rather than coming up looking configured — soacme-proxy filter showand a--helpare cheap pre-restart checks.
TLS Termination
ACME strictly expects traffic over HTTPS (RFC 8555 §6.1).
While you can place acme-proxy behind a reverse proxy (like Nginx, HAProxy, or
Traefik) and let it terminate TLS, acme-proxy is fully capable of terminating
TLS itself via the highly secure rustls crate.
Architectural security (Slowloris protection)
Terminating TLS in async frameworks requires care to avoid denial-of-service
vectors. In acme-proxy, handshakes run off the accept path. When a TCP
connection arrives, a background task spawns the TLS handshake into a bounded
channel. The handshake is strictly bounded by handshake_timeout_ms. If this
were done inline, a single stalled client (e.g., a Slowloris attack) could block
every other connection for the length of the timeout.
Client IP preservation
Under the hood, acme-proxy wraps the TLS listener in a TapIo struct. This is
load-bearing. Without this wrapper, the Axum HTTP layer would lose visibility
into the underlying TCP socket’s peer address once TLS is wrapped around it. By
using TapIo, acme-proxy ensures that the allowed_ip and reverse_dns
filters can correctly identify the true client IP, failing closed if it cannot
be determined.
Configuration
To serve HTTPS directly, configure the [server.tls] section.
Note: Setting
server.tls.enabled = truereplaces the cleartext HTTP listener entirely. You do not get both HTTP and HTTPS simultaneously.
[server]
base_url = "https://acme.internal:3000"
bind_address = "[::]:3000"
[server.tls]
enabled = true
# The PEM encoded certificate chain (leaf first) and private key
cert_path = "server.pem"
key_path = "server.key"
# Budget for one TLS handshake
handshake_timeout_ms = 10000
If cert_path or key_path files are missing on disk at startup, acme-proxy
will automatically generate a self-signed certificate for the host specified
in server.base_url and write it to disk.
The web admin listener
[admin.tls] is the same mechanism on a second socket: the same
load-or-generate provisioning, the same TapIo wrapper preserving the peer
address, the same 0600 on a generated key. Only the defaults differ.
[admin.tls]
enabled = true
cert_path = "admin.pem" # not server.pem
key_path = "admin.key"
The paths are separate on purpose: the two listeners answer to different names
(admin.base_url versus server.base_url, and a generated certificate takes
its name from whichever applies), and sharing one certificate between them
should be a decision, not an accident. The log lines carry a listener field
("acme" or "admin") so certificate churn on one is distinguishable from the
other.
Unlike the ACME listener, TLS here is not optional once the panel leaves loopback — startup refuses that combination outright. See Web Admin.
Profiles
acme-proxy is a multi-tenant ACME server. It serves ACME entirely through
Profiles.
A single process, running on a single port with a single database, can host
multiple isolated ACME endpoints. Each profile is mounted under the
/profile/<name>/directory namespace.
Why use profiles?
- Serve an internal self-signed CA at
/profile/local/directoryfor dev environments. - Serve a strict Let’s Encrypt relay at
/profile/prod/directoryfor production services. - Apply different network filters (e.g., strict IP allowlists for production, bypass for dev) without needing to run multiple binary instances.
One process, several endpoints
Everything below the router is per profile — its own signer, filters, challenge validators and EAB policy. Everything above it is shared: one socket, one database, one process.
graph TD
REQ["Incoming request"] --> ROOT["Root router<br/>/health, /, tracing, hardening headers"]
ROOT -->|"/profile/dev/*"| PDEV["Profile: dev"]
ROOT -->|"/profile/prod/*"| PPROD["Profile: prod"]
ROOT -->|"/profile/staging/*"| PSTG["Profile: staging"]
PDEV --> SDEV["signer: local_ca<br/>filters: none"]
PPROD --> SPROD["signer: relay<br/>filters: ipam"]
PSTG --> SSTG["signer: local_ca<br/>filters: none"]
SDEV --> B1["Arc<dyn SignerBackend> #1"]
SSTG --> B1
SPROD --> B2["Arc<dyn SignerBackend> #2"]
B1 --> DB[("one SQLite file<br/>rows tagged by profile")]
B2 --> DB
Note the fan-in: dev and staging have identical [signer] sections, so
they share one backend instance rather than constructing two. Two profiles
sharing ca.key while differing elsewhere is a startup error: one CA key under
two configurations would issue under two policies from one identity — two
different crl_distribution_points for one CRL, say — and that is refused
rather than left to half-work.
Hard database isolation
Profiles act as a strict isolation boundary in the SQLite database.
The accounts table uses a constraint: UNIQUE(profile, pubkey). This means if
a client registers a key at /profile/default, and then uses the exact same
cryptographic key to connect to /profile/le, the database creates two
independent ACME accounts. This ensures endpoints cannot cross-pollinate
authorizations, orders, or nonces.
Inheritance and configuration
Profiles inherit from the base (global) configuration keys. A profile only needs to override the specific keys that differ.
Eight sections can be overridden, and no others: signer, filter,
ipam, challenge, eab, order, notify and meta. Everything else — the
listen socket, the database, logging, audit, the admin listener — is
process-wide.
graph LR
ENV["ACME_PROXY_* env"] --> BASE
FILE["config.toml"] --> BASE["Base configuration"]
BASE --> MERGE{{"merge, per key"}}
OVR["[profiles.prod]<br/>only the keys that differ"] --> MERGE
MERGE --> EFF["Effective configuration<br/>for profile 'prod'"]
Inheritance is per key, not per section. A profile that sets only
challenge.bypass keeps the global challenge.enabled rather than reverting
it to the compiled default. Arrays, however, replace wholesale — they never
append. Precedence is: profile key, then global key, then compiled default.
That split is why [filter] is shaped the way it is. filter.rules is an
array, so a profile naming its own rules replaces the sequence outright — which
is right, because order is the policy. [filter.check.<name>] and
[filter.rule.<name>] are tables, so they merge per key, and a profile can
dry-run one rule without restating anything:
[profiles.staging.filter.rule.inventory-owned]
mode = "warn"
A profile inherits every globally defined check and cannot remove one, which
costs nothing: a check no selected rule names is never built. A global
[filter] section can therefore carry a library of checks and each profile pick
the subset its own filter.rules uses.
[signer]
backend = "local_ca"
[filter]
rules = ["corp-only"] # The base selection; each profile may replace it
# The library. Both checks and both rules are declared once, globally.
[filter.check.corp-net]
type = "allowed_ip"
allow = ["10.0.0.0/8"]
[filter.check.corp-names]
type = "identifiers"
allow = ["*.corp.example.com"]
[filter.rule.corp-only]
when = "corp-net"
then = "allow"
[filter.rule.named-and-corp]
when = "corp-net and corp-names"
then = "allow"
# Profile 1: Uses the global local_ca, and the inherited address-only rule.
# `corp-names` is named by no rule it selects, so it is never built.
[profiles.default]
enabled = true
# Profile 2: Relays to Let's Encrypt, and replaces the selection with the
# stricter rule — which is what pulls `corp-names` into existence here.
[profiles.le]
enabled = true
signer.backend = "relay"
signer.relay.directory_url = "https://acme-v02.api.letsencrypt.org/directory"
filter.rules = ["named-and-corp"]
acme-proxy filter show --profile <name> prints the built policy for one
profile, which is the quickest way to confirm a profile selected what you
intended. A check that no selected rule names is reported as
filter_check_unused — an advisory, not an error.
At runtime, acme-proxy deduplicates the signer backends in memory so that two
profiles sharing the exact same signer configuration (e.g., two profiles using
the same local_ca) don’t duplicate background polling threads or memory
overhead.
Signers
The signer is what actually produces a certificate, once the client has proved control of its names and the filters have allowed the request. Everything before this point is the same whichever backend you choose; everything after it is the backend’s business.
There are three, and they answer three different questions.
| Backend | Use it when | The certificate is signed by |
|---|---|---|
| Local CA | The certificates only need to be trusted by machines you control. | This server, from a CA key on disk or in a PKCS#11 token. |
| Relay | You need publicly trusted certificates, but your clients cannot reach a public CA — or you have one scarce upstream credential to share. | A real upstream ACME CA, which this server becomes a client of. |
| Custom Script | The authority already exists and does not speak ACME. | Whatever your script talks to: a legacy PKI, an internal API, an offline process. |
Choosing one
Start from what has to trust the certificate:
- Only your own machines?
local_ca. You distribute the CA certificate once (see Trusting the CA) and the whole thing works offline, including revocation via the CRL. - Browsers, partners, anything you do not control? You need a public CA, so
relay. Your internal clients keep proving control to this server — over HTTP, DNS or TLS, whichever suits them — while the upstream challenge is solved once, centrally, with a credential no client ever holds. - An existing corporate PKI that issues by ticket, script or API?
custom. It is the escape hatch, and it is deliberately a shell contract rather than a plugin API, so anything that can be scripted can be a signer.
Nothing stops you from running more than one. [signer] is a per-profile
section, so a local_ca at /profile/dev can sit beside a relay at
/profile/prod in the same process.
What every backend has to provide
All three implement the same trait, and the shape of it is worth knowing because
it is what the rest of the server can rely on. It comes in two halves, split by
who holds the key: the backend — issue, revoke — is built only by the
process running the worker role, and the read side — the CRL, the trust
anchor, renewal information — is built by every process from public material,
so the one parsing client requests never holds signing material:
issue— the only required capability. It receives the order’s identifiers and the client’s CSR and returns a chain, or refuses withbadCSR.revoke— must be idempotent. Revoking twice is not an error, because the server cannot always know whether a previous attempt reached the authority.crl_derandrenewal_info— optional, and default to “nothing to say here”.local_capublishes a CRL; the relay passes the upstream’s renewal window through,explanationURLand all.
issue never runs inside a client’s request: finalize queues it, answers the
order processing, and the worker calls the backend. A backend may itself
answer processing rather than a certificate, meaning “this finishes
later”. Only the relay does — signing upstream takes as long as the upstream
takes — and the order simply stays processing until it has.
Configuration
[signer]
backend = "local_ca" # or "relay", or "custom"
Each backend then reads its own table — [signer.local_ca],
[signer.relay], [signer.custom] — documented on its own page.
Reference
backend (String) — Default: "local_ca" | Env: ACME_PROXY_SIGNER__BACKEND
Which backend issues certificates: local_ca, relay or custom. Any
other value is a startup error.
Backends are shared by configuration, not per profile
Two profiles whose [signer] sections are identical share one backend
instance rather than constructing two. Revocations and the CRL live in the
database, keyed by the CA’s key, so two instances over one CA would still agree
on what is revoked; what they could not agree on is everything else in the
section.
Two profiles sharing ca.key while differing anywhere else in [signer] is
therefore a startup error, not a race to discover later. See
Profiles & Routing.
Local CA Signer
The local_ca backend uses an internal ECDSA key to act as a fully functional
Certificate Authority. It is capable of generating its own self-signed root, or
it can act as a subordinate (Intermediate) CA if provided with an existing key
and certificate.
Reach for it for development environments, CI pipelines, and isolated internal networks — anywhere the certificates only need to be trusted by machines you control.
Security constraints
When processing a Certificate Signing Request (CSR) from a client, the
local_ca is deeply distrustful of the requested extensions:
- Overwriting Extensions: The Local CA overwrites every extension the CSR
asked for before signing — a fresh random serial, fixed key usages, and its
own validity window. This matters more than it sounds: the CSR parser
otherwise copies a requested
basicConstraints/keyUsagestraight into the signed leaf. - Basic Constraints: The issued leaf is never a CA. Without the reset
above, a client authorized for one name could submit a CSR carrying
CA:TRUEkeyCertSignand receive a working intermediate CA, which it could then use to mint arbitrary trusted certificates. (In implementation terms the leaf is built withIsCa::NoCarather thanExplicitNoCa— an explicitCA:FALSEbroke certbot’s chain parser — but the security property is the same.)
- Subject Alternative Names (SANs): The DNS SANs in the CSR must be
exactly the set of identifiers the order authorized — no more, no fewer —
or issuance fails with
badCSR. - Non-DNS SANs are rejected, not stripped: a CSR carrying an IP, email or
URI SAN is refused outright with
badCSR. Do not expectlocal_cato quietly drop them. - The subject is emptied: the issued leaf carries no distinguished name at all, so a Common Name in the CSR cannot leak into it. (A CN that looks like a domain the order does not cover is separately rejected earlier, at finalize — see Troubleshooting.)
Certificate validity
Leaves are valid for leaf_validity_days, starting one hour in the past to
absorb clock skew between the CA and its clients.
A client may narrow that window using the order’s notBefore/notAfter (RFC
8555 §7.4), but never widen it: a requested start is honoured only if it is
later than the default start, and a requested end only if it is earlier than
the default end. A request whose clamped window would be empty or inverted is
discarded whole, the policy default is used, and a
local_ca_requested_validity_discarded warning is logged.
Configuration
[signer]
backend = "local_ca"
[signer.local_ca]
cert_path = "ca.pem"
key_path = "ca.key"
key_type = "ecdsa-p256"
crl_path = "ca.crl"
leaf_validity_days = 90
Reference
cert_path (String) — Default: "ca.pem" | Env: ACME_PROXY_SIGNER__LOCAL_CA__CERT_PATH
The path where the CA certificate is stored. If this file does not exist, a new self-signed Root CA is generated on startup and saved here.
key_path (String) — Default: "ca.key" | Env: ACME_PROXY_SIGNER__LOCAL_CA__KEY_PATH
The path where the CA private key is stored. A generated key is created with
mode 0600 at creation time, not chmod’ed afterwards, so there is no window in
which another local user could read it.
key_type (String) — Default: "ecdsa-p256" | Env: ACME_PROXY_SIGNER__LOCAL_CA__KEY_TYPE
Algorithm for a generated CA key. "ecdsa-p256" is currently the only
accepted value; anything else is a startup error. (An existing key supplied on
disk is used as-is, whatever its type.)
crl_path (String) — Default: "ca.crl" | Env: ACME_PROXY_SIGNER__LOCAL_CA__CRL_PATH
Where the current Certificate Revocation List (RFC 5280) is exported as
PEM, for publishing from a static web server. The revocations and the CRL
GET /crl serves live in the database; this file is rewritten whenever a new
CRL is stored and is never read back. Writers lock ca.json.lock beside it,
which stays empty. The path also locates the JSON ledger a CA kept before the
database did (ca.crl → ca.json), imported once; see
Revocation.
crl_distribution_points (Array<String>) — Default: [] | Env: ACME_PROXY_SIGNER__LOCAL_CA__CRL_DISTRIBUTION_POINTS
Where a relying party can fetch that CRL. Each URL is written into every
issued leaf as cRLDistributionPoints (RFC 5280 §4.2.1.13); empty — the default
— emits no extension at all, which is why a certificate from this CA says
nothing about revocation until you set this.
Nothing derives it, deliberately. The URL is frozen into every certificate
signed while it is set, so a value read from server.base_url would silently
break certificates already issued the day that changed. And this server’s own
copy is served at {base_url}/profile/<name>/crl, inside the profile router
and therefore behind that profile’s filter policy — an address-based rule would
refuse it to exactly the relying parties the extension exists for. Name a URL
you know is publicly reachable; a webroot or CDN copy of crl_path is the usual
answer.
Several entries mean one CRL reachable in several places, not several
different CRLs. http:// is idiomatic and gets no warning: fetching a signed
CRL over TLS means validating that connection’s certificate first, which is the
loop this extension exists to break. Credentials in the URL, a non-http(s)
scheme, and any value the URL parser would normalize (a missing trailing /, a
leading space from an environment-variable list) are each a startup error naming
the key and the value.
Note that two profiles sharing one CA share this list too: they share one
[signer] section, one set of revocations and one CRL, so there is one place
that CRL is published. Giving them different URLs while they share ca.key is
refused at startup — see Profiles & Routing.
ca_issuer_urls (Array<String>) — Default: [] | Env: ACME_PROXY_SIGNER__LOCAL_CA__CA_ISSUER_URLS
Where a relying party can fetch this CA’s own certificate, written into
every issued leaf as authorityInfoAccess with the caIssuers access method
(RFC 5280 §4.2.2.1). Empty — the default — emits no extension. Same reasoning,
same validation and the same startup errors as crl_distribution_points above.
The sibling access method, id-ad-ocsp, is never written: this server runs no
OCSP responder, and a pointer at one that does not exist is worse than none.
leaf_validity_days (Integer) — Default: 90 | Env: ACME_PROXY_SIGNER__LOCAL_CA__LEAF_VALIDITY_DAYS
The validity period (in days) for issued leaf certificates. Distinct from
order.validity_seconds, which bounds the ACME order object, not the
certificate.
key_source (String) — Default: "file" | Env: ACME_PROXY_SIGNER__LOCAL_CA__KEY_SOURCE
Where the issuing private key lives. "file" is everything described on this
page: a PEM key at key_path, loaded or generated. "pkcs11" puts the key in a
hardware token instead, and reads the [signer.local_ca.pkcs11] table — see
Hardware Keys. Any other value is a startup error,
and so is "pkcs11" on a binary built without --features hsm: there is
deliberately no silent fallback, since falling back would hand an operator who
asked for hardware a software key with no indication of it.
[signer.local_ca.subject]
The X.509 Subject (Distinguished Name) of the CA certificate acme-proxy
generates for itself. It is read only on the startup that generates the CA —
supplying your own cert_path means supplying your own DN, and editing this
table afterwards changes nothing, because the certificate already exists. To
change it, retire the CA and generate a new one.
Every key is optional and String-valued. An unset or empty value is omitted
from the Subject rather than written blank; the empty string is treated as
absent because the environment source cannot distinguish the two. The one
exception is common_name, which falls back to "acme-proxy local CA" so the
CA always carries one. Leaving the whole table out therefore yields a
CommonName-only Subject — which is exactly what this backend produced before the
table existed.
This is the CA’s own DN, not the leaf’s. Issued certificates deliberately carry
an empty Subject and identify themselves through subjectAltName, which is what
every modern client reads.
common_name — Default: "acme-proxy local CA" | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__COMMON_NAME
organization — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__ORGANIZATION
organizational_unit — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__ORGANIZATIONAL_UNIT
country — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__COUNTRY
A two-letter ISO 3166-1 country code, e.g. "US".
state — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__STATE
locality — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__SUBJECT__LOCALITY
[signer.local_ca.subject]
common_name = "Example Corp Internal CA"
organization = "Example Corp"
country = "US"
Hardware-backed keys
The CA key described above is a file, and a file can be copied. If the CA is one your fleet actually trusts, the issuing key can instead live in a PKCS#11 token — a YubiKey, an enterprise HSM, or SoftHSM2 for development — where it is created once and can never be read back out.
Everything on this page still applies; only where the signature comes from
changes. See Hardware Keys (PKCS#11). Note that it requires a
build with --features hsm, which is not the default.
Multi-tier PKI (using an intermediate CA)
For production internal deployments, you should avoid using an auto-generated
Root CA directly on the server. Instead, you can use a Multi-Tier Hierarchy:
create an offline Root CA, use it to sign an Intermediate CA, and hand the
Intermediate CA to acme-proxy.
Here is how you can do this in practice using OpenSSL.
Step 1 — Create the offline root CA
Generate a private key and a self-signed Root certificate (keep this key highly secure and offline):
# Generate the Root private key
openssl ecparam -genkey -name prime256v1 -out root_ca.key
# Create the self-signed Root certificate (valid for 10 years)
openssl req -x509 -new -nodes -key root_ca.key -sha256 -days 3650 \
-out root_ca.pem \
-subj "/CN=My Company Offline Root CA"
Step 2 — Create the intermediate CA for acme-proxy
Generate the private key and CSR for the Intermediate CA:
# Generate the Intermediate private key
openssl ecparam -genkey -name prime256v1 -out acme_intermediate.key
# Create the CSR
openssl req -new -key acme_intermediate.key -out acme_intermediate.csr \
-subj "/CN=My Company ACME Intermediate CA"
Create an OpenSSL extension file (v3_ext.cnf) to ensure the Intermediate CA is
allowed to sign other certificates:
[ v3_intermediate_ca ]
subjectKeyIdentifier = hash
authorityKeyIdentifier = keyid:always,issuer
basicConstraints = critical, CA:true, pathlen:0
keyUsage = critical, digitalSignature, cRLSign, keyCertSign
(Note: pathlen:0 ensures this intermediate can issue leaf certificates, but
cannot issue further intermediate CAs).
Now, sign the Intermediate CSR using the Offline Root CA:
openssl x509 -req -in acme_intermediate.csr -CA root_ca.pem -CAkey root_ca.key \
-CAcreateserial -out acme_intermediate.pem -days 1825 -sha256 \
-extfile v3_ext.cnf -extensions v3_intermediate_ca
Step 3 — Configure acme-proxy
Finally, point acme-proxy to your newly minted Intermediate CA. The
cert_path must contain the Intermediate certificate followed by the Root
certificate (the bundle), so clients can verify the full chain.
cat acme_intermediate.pem root_ca.pem > ca_bundle.pem
Update your config.toml:
[signer.local_ca]
cert_path = "ca_bundle.pem"
key_path = "acme_intermediate.key"
Now acme-proxy issues certificates signed by the Intermediate CA, mirroring a
standard enterprise PKI hierarchy.
The order inside the bundle is load-bearing. The signing issuer is parsed from the first PEM block in
cert_path; the remaining blocks are only appended to the chain served to clients. Concatenating root-first instead would not fail loudly — it would sign with the root’s identity using the intermediate’s key, producing certificates nothing can verify. Alwayscat intermediate.pem root.pem, never the reverse.
Two further caveats:
- Nothing checks that
key_pathactually corresponds to the certificate incert_path. A mismatched pair produces unverifiable certificates rather than a startup error. (This caveat is specific tokey_source = "file"; the PKCS#11 path does verify the pair at startup.) - The whole bundle is emitted with every issued certificate, root included. Most
clients tolerate this, but if you would rather not ship the root, put only the
intermediate in
cert_pathand distribute the root out of band — see Trusting the CA.
Hardware Keys (PKCS#11)
By default the Local CA’s issuing key is a PEM file on disk, protected by
nothing but its 0600 permissions. key_source = "pkcs11" moves that key into
a hardware token — a YubiKey, a network HSM, or SoftHSM2 for development — where
it is created once and can never be read back out. acme-proxy sends the token
the bytes to be signed and receives a signature; the private key never enters
this process’s memory.
PKCS#11 rather than a vendor-specific PIV library, so a YubiKey today and an
enterprise HSM tomorrow are the same configuration with a different
module_path.
What this protects, and what it does not
Everything else about the Local CA is unchanged: the same CSR
sanitisation, the same leaf_validity_days
clamping, the same CRL and revocations. Only where the signature comes
from moves.
It protects the CA issuing key — the one that, if stolen, lets an attacker
mint certificates your fleet trusts. It does not protect the ACME account
keys, the TLS server key (server.tls.key_path), or the database; those stay on
disk.
Requirements
-
A build with the
hsmfeature. It is off by default, so the stock binary does not have it andkey_source = "pkcs11"on one is a startup error naming the feature:cargo build --release --features hsm -
A PKCS#11 module (
.so), loaded at runtime — nothing is linked at build time. -
An existing CA certificate, for the reason below.
Two rules that differ from the software path
Both are startup errors, so you will meet them immediately rather than in production.
The CA is never generated
With key_source = "file", a missing cert_path/key_path means “generate a
CA and write it here”. With key_source = "pkcs11" there is no such thing: the
private key is created inside the token by its own tooling, and this server
cannot produce one that a token would then hold. So cert_path must already
exist, and key_path is neither read nor written.
Both walkthroughs below cover creating that certificate.
The key and the certificate are cross-checked
At startup, the token key’s SubjectPublicKeyInfo is compared against the one
in cert_path. A mismatch — almost always a wrong key_label — stops the
server.
This is stricter than the file path, where (as Local
CA warns) nothing checks
that key_path corresponds to cert_path, and a mismatched pair simply
produces certificates that verify nowhere. Here a typo is caught before the
first certificate is issued rather than discovered by a client days later.
Reference
Reaching this table at all takes signer.local_ca.key_source = "pkcs11", which
is documented with the rest of [signer.local_ca] in Local
CA. Everything below is [signer.local_ca.pkcs11], read
only when that key is set.
pkcs11.module_path (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__MODULE_PATH
The PKCS#11 module to load. Required. See each walkthrough for the usual paths.
pkcs11.token_label (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__TOKEN_LABEL
Which token to use, by label. Preferred over slot_id: slot numbers are
assigned dynamically and change across reboots and re-plugs on most drivers
(SoftHSM2 will hand you something like 276468771).
pkcs11.slot_id (Integer) — Default: unset | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__SLOT_ID
Which slot to use, for tokens with no usable label. Consulted only when
token_label is empty.
pkcs11.key_label (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__KEY_LABEL
The private key’s CKA_LABEL. Required. On a YubiKey the labels are fixed by
the driver, so this is something you look up rather than choose — see below.
pkcs11.key_id (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__KEY_ID
The key’s CKA_ID as hex (01, or 01:ff), to disambiguate a token holding
several keys under one label. Optional; two keys sharing a label and no key_id
to separate them is a startup error rather than a coin flip.
pkcs11.pin_file (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__PIN_FILE
A file holding the user PIN. Trailing whitespace is trimmed, so a PIN written
with echo works. The file is checked for permissions and warns if it is
world-readable, exactly as ca.key does.
pkcs11.pin (String) — Default: "" | Env: ACME_PROXY_SIGNER__LOCAL_CA__PKCS11__PIN
SENSITIVE. The PIN directly. Prefer pin_file, or set this through the
environment variable; a PIN in config.toml is a long-lived secret in a file
that tends to get copied around. pin_file wins when both are set, and having
neither is a startup error.
A PIN is not a password. Tokens block after a small number of wrong attempts — three on a YubiKey PIV applet, after which you need the PUK. This is why
acme-proxyretries a failed signature at most once, and why the trailing newline in your PIN file is worth getting right.
Walkthrough A — SoftHSM2
SoftHSM2 is a software token: no hardware needed, and the same setup the project uses in CI. Use it to try the feature before committing to hardware.
# Debian/Ubuntu
sudo apt install softhsm2
# Arch
sudo pacman -S softhsm
Step 1 — Create a token
softhsm2-util --init-token --free --label acme-ca --so-pin 3737 --pin 1234
--free takes the first uninitialised slot. Note that the token is reassigned
to a new slot number afterwards — which is exactly why token_label is the
selector to use, not slot_id.
Step 2 — Create the CA key and certificate
The key must exist inside the token, and cert_path must hold a certificate for
it. For SoftHSM2 the simplest route is to generate both locally, import the key,
and destroy the local copy:
# The CA key and its self-signed certificate
openssl genpkey -algorithm EC -pkeyopt ec_paramgen_curve:P-256 -out ca.key
openssl req -x509 -new -key ca.key -sha256 -days 3650 -out ca.pem \
-subj "/CN=Example Corp Issuing CA/O=Example Corp" \
-addext "basicConstraints=critical,CA:true,pathlen:0" \
-addext "keyUsage=critical,keyCertSign,cRLSign"
# Move the key into the token, then remove it from disk
softhsm2-util --import ca.key --token acme-ca --label ca-key --id 01 --pin 1234
shred -u ca.key
For a real HSM, generate the key in the token instead so it never exists outside it —
pkcs11-tool --module <module> --token-label acme-ca --login --keypairgen --key-type EC:prime256v1 --label ca-key --id 01(from theopenscpackage), then certify that public key with your offline root. The import above is a development convenience, and the reason it is acceptable here is that a SoftHSM2 token is a directory of files anyway.
pathlen:0 matches what the Local CA generates for itself: it may issue leaves
but no further CAs.
Step 3 — Configure
[signer]
backend = "local_ca"
[signer.local_ca]
cert_path = "ca.pem"
crl_path = "ca.crl"
key_source = "pkcs11"
[signer.local_ca.pkcs11]
module_path = "/usr/lib/softhsm/libsofthsm2.so"
token_label = "acme-ca"
key_label = "ca-key"
pin_file = "/etc/acme-proxy/hsm.pin"
printf '1234' > /etc/acme-proxy/hsm.pin
chmod 600 /etc/acme-proxy/hsm.pin
If SoftHSM2’s token store is not in its default location, SOFTHSM2_CONF must
be set in the server’s environment — it is read by the module, not by
acme-proxy.
Step 4 — Confirm it is really using the token
RUST_LOG=info acme-proxy serve
INFO acme_proxy::signer::local_ca::pkcs11: the local CA's issuing key is on a PKCS#11 token
event="local_ca_pkcs11_opened" module=/usr/lib/softhsm/libsofthsm2.so
slot=276468771 key_label=ca-key algorithm=PKCS_ECDSA_P256_SHA256
mechanism=CKM_ECDSA_SHA256
INFO acme_proxy::signer::local_ca: event="local_ca_pkcs11_loaded" cert_path="ca.pem" key_label=ca-key
local_ca_pkcs11_opened is the line that proves it: it names the module, the
slot the token actually landed in, the curve read off the key, and the mechanism
chosen. If you see local_ca_loaded or local_ca_generated instead, the
configuration is still on the file path.
Then issue something and check it chains:
openssl verify -CAfile ca.pem /path/to/issued/cert.pem
# cert.pem: OK
Walkthrough B — YubiKey (libykcs11)
A YubiKey 5 exposes its PIV applet through libykcs11, shipped with
yubico-piv-tool.
# Debian/Ubuntu → /usr/lib/x86_64-linux-gnu/libykcs11.so
sudo apt install yubico-piv-tool
# Arch → /usr/lib/libykcs11.so
sudo pacman -S yubico-piv-tool
Both paths are in circulation; check which one you have before configuring
module_path.
Step 1 — Generate the key and certificate on the device
Use slot 9c (Digital Signature). Its PIV policy requires the PIN for every private-key operation, which is the right posture for a CA key and the reason to prefer it over 9a.
# Generate the key inside the YubiKey — it never leaves
yubico-piv-tool -s 9c -a generate -A ECCP256 -o ca_pub.pem
# Self-sign a certificate for it, on the device
yubico-piv-tool -s 9c -a verify-pin -a selfsign-certificate \
-S '/CN=Example Corp Issuing CA/O=Example Corp/' \
--valid-days 3650 -i ca_pub.pem -o ca.pem
# Store the certificate in the slot as well (optional, but conventional)
yubico-piv-tool -s 9c -a import-certificate -i ca.pem
Copy ca.pem to wherever cert_path points.
Touch policy must be
neverfor the CA slot. If the slot is provisioned to require a touch, every issuance blocks until somebody physically touches the key. That is correct for an offline root and catastrophic for an ACME server expected to issue unattended.
Step 2 — Find the key label
You do not choose the label on a YubiKey — libykcs11 assigns fixed ones per
PIV slot. Read it off the device:
pkcs11-tool --module /usr/lib/libykcs11.so --list-objects --login
Slot 9c reports as Private key for Digital Signature; 9a as Private key for PIV Authentication. Use that string verbatim.
Step 3 — Configure
[signer.local_ca]
cert_path = "ca.pem"
crl_path = "ca.crl"
key_source = "pkcs11"
[signer.local_ca.pkcs11]
module_path = "/usr/lib/libykcs11.so"
token_label = "YubiKey PIV #12345678"
key_label = "Private key for Digital Signature"
pin_file = "/etc/acme-proxy/hsm.pin"
The PIN is the PIV PIN (factory default 123456), not the PIV management
key and not the FIDO PIN.
Step 4 — Expect CKM_ECDSA
libykcs11 does not offer CKM_ECDSA_SHA256, so acme-proxy computes the
SHA-256 digest itself and asks the token to sign that. The startup line reads:
mechanism=CKM_ECDSA+SHA256
This is normal and not a downgrade — the same signature, with the hashing done on this side of the USB cable.
Performance
A YubiKey signature takes roughly 50–300 ms, and signings are serialised by a mutex. That is comfortable for hundreds of certificates a day and is not a throughput solution; the signing call runs on the blocking thread pool, so it does not stall the rest of the server while it waits. For higher volumes use a networked HSM, or the Custom Script signer against a KMS.
Operations
Backup and disaster recovery
The key cannot be backed up. That is the point of the feature, and it makes recovery something to plan before you need it. Two workable approaches:
- Two tokens, one offline root. Keep an offline root CA, use it to certify
an intermediate held on each of two tokens, and hand
acme-proxyone of them. A lost token is replaced by provisioning a new intermediate; clients trust the root and never notice. See Multi-Tier PKI. - Accept re-enrolment. For a small internal fleet, losing the CA and distributing a new one is survivable — just make sure it is a decision rather than a discovery.
Revocations live in the database like every other CA’s, so backing up the
database backs them up; crl_path is only an export of the current CRL.
When the token disappears
If the session drops — the YubiKey is unplugged, a network HSM times out —
acme-proxy reopens the session, logs back in and retries the signature
once. The relevant log lines are local_ca_pkcs11_session_lost followed by
either a successful issuance or local_ca_pkcs11_reconnect_failed.
If that fails, finalize requests return serverInternal (500) and clients
retry, which is the right behaviour: the order stays valid and issuance resumes
once the token is back. GET /crl keeps working throughout — the current CRL is
read from the database and serving it signs nothing.
Sharing one token between profiles
Several profiles may use the same module, and even the
same key. acme-proxy opens one PKCS#11 context per module for the whole
process and shares it, so this works without special configuration. Two profiles
naming the same token key with otherwise different signer settings is refused
at startup, for the same reason two profiles sharing ca.key are: one key under
two configurations would issue under two policies from one identity.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
key_source = "pkcs11" … built without | The binary has no PKCS#11 support. Rebuild with --features hsm. |
CKR_PIN_INCORRECT at startup | Usually a stray character in pin_file. Trailing newlines are trimmed, but leading or embedded whitespace is not. Check with xxd. Do not retry blindly — see the PIN warning above. |
CKR_PIN_LOCKED | Too many wrong attempts. A YubiKey PIV PIN is unblocked with the PUK (yubico-piv-tool -a unblock-pin). |
is not the key certified by … | The SPKI cross-check failed: key_label/key_id resolve to a different key than cert_path describes. List the token’s objects and compare. |
no PKCS#11 token labelled … | The label is wrong, or the token is not plugged in. The message lists the labels actually present. |
N PKCS#11 private keys are labelled … | Set key_id to pick one. |
unsupported PKCS#11 curve | Only P-256 and P-384 are supported. The message prints the CKA_EC_PARAMS it found. |
supports neither … nor CKM_ECDSA | The token cannot do ECDSA signing at all, or will not report its mechanisms. |
| Certificates that verify nowhere | Should not happen — the SPKI cross-check catches the usual cause at startup. If it does, capture the local_ca_pkcs11_opened line and the failing certificate and open an issue. |
See also Maintenance & Troubleshooting.
Relay
The relay backend keeps this server answering ACME to its own clients while a
real upstream CA does the signing. It captures internal ACME requests and
fulfils them through an upstream external CA (Let’s Encrypt, ZeroSSL, or a
commercial CA).
How it works
There are two ACME conversations, and keeping them apart is the whole trick
to reading this page. acme-proxy is the server in one and the client in
the other. Each has its own account, its own account key, its own order and its
own challenges, and nothing crosses between them.
sequenceDiagram
autonumber
participant C as Internal client<br/>(certbot, acme.sh)
participant P as acme-proxy
participant U as Upstream CA<br/>(Let's Encrypt)
Note over C,P: Conversation 1 — acme-proxy is the SERVER.<br/>Client's own account key, client's own order.
C->>P: newOrder (example.internal)
P-->>C: 201, order pending
C->>P: trigger challenge
P->>C: validate control (http-01 / dns-01 / tls-alpn-01)
P-->>C: 200, challenge valid — order ready
C->>P: finalize (CSR)
Note over P,U: Conversation 2 — acme-proxy is the CLIENT.<br/>Its OWN upstream account key, a SECOND order.
P-->>C: 200, order processing
P->>U: newOrder (same identifiers)
U-->>P: 201, upstream order pending
alt challenge_strategy = dns01
P->>P: publish TXT via RFC 2136 + TSIG
else challenge_strategy = http01
P->>P: publish token on its own /.well-known route
else challenge_strategy = bypass
P->>U: trigger — the upstream already trusts this account
end
U->>P: validate (against the proxy's thumbprint)
P->>U: finalize (the client's CSR, relayed unchanged)
U-->>P: certificate
Note over C,P: Back in conversation 1.
C->>P: poll order
P-->>C: 200, order valid + certificate URL
Two consequences fall straight out of the diagram:
- The key authorization at the upstream uses the proxy’s own thumbprint, never the client’s. They are different accounts on different servers, so the client could not answer the upstream’s challenge even in principle.
finalizestaysprocessingfor longer. Every backend’s finalize answersprocessingand is signed by the worker, but here conversation 2 takes as long as the upstream takes, so the client polls for minutes rather than for the momentlocal_caneeds — see Core Concepts.
Challenge strategies
The strategy names the one challenge type the proxy answers upstream. A CA offers several — Let’s Encrypt currently poses four — and every type the strategy does not name is ignored, including ones this server has no implementation for at all. Nothing falls back: if the upstream authorization does not offer the type the strategy names, the relay says so and stops, rather than trying a type it could not finish.
dns01 (RFC 2136 TSIG)
This is the strategy that can prove a wildcard. The proxy intercepts internal
HTTP-01 or DNS-01 challenges, but to satisfy the external CA, the proxy solves
the external DNS-01 challenge itself. It does this using an rfc2136 provider
powered by hickory-proto. It securely authenticates with the DNS server using
TSIG (Transaction Signature) to publish the TXT record. Note: The TXT record
uses the thumbprint of the proxy’s upstream account key, not the internal
client’s key.
DNS alias mode
By default the record is written at _acme-challenge.<domain>, so the update
key needs write access to every zone a profile issues for. Alias mode moves the
record to one name in a zone set aside for it, the way acme.sh’s
--challenge-alias does. Delegate each domain once with a CNAME, then point
zone and challenge_alias at the alias zone:
_acme-challenge.www.example.com. CNAME _acme-challenge.acme-alias.net.
_acme-challenge.api.example.org. CNAME _acme-challenge.acme-alias.net.
[signer.relay.dns01]
challenge_alias = "acme-alias.net."
[signer.relay.dns01.rfc2136]
zone = "acme-alias.net."
The CA follows the CNAME itself; nothing changes on its side. Every domain of the profile shares the one record name, which is safe because values are added and removed one by one. The alias is configured rather than found by following the CNAME, because this server’s resolver need not see what the CA sees, and a record published at the wrong name invalidates the upstream authorization for good.
Whoever can write the alias name can pass dns-01 for every domain pointing at
it. Domains that must not share that power belong in separate profiles, each
with its own alias and key.
http01
The proxy answers the upstream’s http-01 challenge by serving the key
authorization itself, from a route on its own root router at
/.well-known/acme-challenge/<token>.
Like dns01, the value served is derived from this proxy’s upstream account
thumbprint, not the internal client’s — the two are different accounts on
different servers, so the client cannot answer it even in principle. Unlike
dns01, the body is the key authorization verbatim rather than its SHA-256
digest (RFC 8555 §8.3 versus §8.4).
This is the opposite direction of the inbound http-01
challenge, which is this server validating its own
clients. The two share the well-known path and nothing else.
Two constraints are worth knowing before choosing it:
- It needs a forwarder.
acme-proxydoes not open a second listener and does not bind port 80. The upstream CA fetcheshttp://<identifier>:80/.well-known/acme-challenge/<token>, so something must forward or redirect that path toacme-proxy— see Deploying the http-01 responder below. - It cannot prove a wildcard. Nothing answers HTTP on the name
*.example.com. An upstream authorization for a wildcard is refused with an error namingdns01, which is the strategy that can.
There is no [signer.relay.http01] table: setting challenge_strategy = "http01" is the whole configuration.
bypass
Used when the upstream CA implicitly trusts the proxy’s account (e.g., a commercial CA with pre-validated domains). The proxy simply tells the upstream “I have validated this”, bypassing external challenges entirely.
This is the one strategy that does not name a type: it triggers whatever the authorization offers. It prefers a challenge it could have satisfied had it needed to, since an upstream that validates out of band still decides against the challenge it was pointed at.
Deploying the http-01 responder
Only relevant with challenge_strategy = "http01".
The upstream CA fetches on port 80 of each name being issued, so put the
existing web server for that name in front of acme-proxy and hand it that one
path. Either shape works — RFC 8555 §8.3 explicitly permits following redirects,
and every real CA does, so the target need not share the name:
# nginx, on the name being certified. Proxy it:
location /.well-known/acme-challenge/ {
proxy_pass http://acme-proxy:3000;
}
# ...or just redirect it, which needs no upstream block:
location /.well-known/acme-challenge/ {
return 301 http://acme-proxy:3000$request_uri;
}
# Caddy
handle /.well-known/acme-challenge/* {
reverse_proxy acme-proxy:3000
}
# Traefik, as a dynamic-configuration router
http:
routers:
acme-challenge:
rule: "PathPrefix(`/.well-known/acme-challenge/`)"
service: acme-proxy
priority: 100
services:
acme-proxy:
loadBalancer:
servers:
- url: "http://acme-proxy:3000"
The route is mounted on the root router, beside GET /health — it is not
under a profile’s /profile/<name> prefix, carries no filter chain, mints no
nonce, and answers a plain 404 rather than an ACME problem document for an
unknown token. It exists only while some profile’s signer uses this strategy;
with any other backend the path is not routed at all. A
http_01_responder_mounted line at startup confirms it is live.
Configuration
[signer]
backend = "relay"
[signer.relay]
directory_url = "https://acme-staging-v02.api.letsencrypt.org/directory"
account_key_path = "upstream_account.key"
contact = ["mailto:admin@example.com"]
challenge_strategy = "bypass"
poll_interval_ms = 2000
poll_timeout_secs = 300
Several relaying profiles
[signer] is a per-profile section, so one server can relay to several
upstreams at once — a Let’s Encrypt endpoint beside a commercial CA, or one
internal CA per environment. Each profile gets its own [signer.relay], and
they must not share an account_key_path: two backends over one upstream
account key would overwrite each other’s registration, and startup refuses it
by name.
[profiles.public.signer]
backend = "relay"
relay.directory_url = "https://acme-v02.api.letsencrypt.org/directory"
relay.account_key_path = "public_account.key"
[profiles.partner.signer]
backend = "relay"
relay.directory_url = "https://acme.commercial-ca.example/directory"
relay.account_key_path = "partner_account.key"
Two profiles whose [signer] sections are byte-for-byte identical share one
backend and one upstream account; anything that differs makes them independent.
See Profiles for what else a profile separates.
Reference
directory_url (String) — Default: "" | Env: ACME_PROXY_SIGNER__RELAY__DIRECTORY_URL
The upstream ACME server’s directory URL.
Starting the server registers an account. The first
acme-proxy servewith this backend configured contactsdirectory_urland performsnewAccountthere, writing the assigned account URL to the.kidsidecar. There is no confirmation step and no dry-run — merely booting a configuration that names a production CA creates a real account at it, and account creation is itself rate limited (Let’s Encrypt allows 10 per IP address per 3 hours). Pointdirectory_urlat a staging endpoint (https://acme-staging-v02.api.letsencrypt.org/directory) while you are still working out a configuration, and switch to production only once it is settled. Subsequent starts reuse the.kidsidecar and do not contact the upstream.
account_key_path (String) — Default: "upstream_account.key" | Env: ACME_PROXY_SIGNER__RELAY__ACCOUNT_KEY_PATH
Path to this proxy’s own account key at the upstream CA. If the file is absent,
an ECDSA P-256 key is generated on startup. The assigned kid is stored beside
it with a .kid extension.
contact (Array) — Default: [] | Env: ACME_PROXY_SIGNER__RELAY__CONTACT
Optional contacts sent with newAccount to the upstream CA.
challenge_strategy (String) — Default: "bypass" | Env: ACME_PROXY_SIGNER__RELAY__CHALLENGE_STRATEGY
How the proxy satisfies the upstream’s domain-control checks: bypass (the
upstream validates nothing), dns01 (publish the TXT record the upstream asks
for) or http01 (serve the challenge file from this server’s own root router,
which requires a reverse proxy in front of it and cannot prove a wildcard). Any
other value is a startup error.
poll_interval_ms (Integer) — Default: 2000 | Env: ACME_PROXY_SIGNER__RELAY__POLL_INTERVAL_MS
How often to poll an upstream order/authorization while it resolves.
poll_timeout_secs (Integer) — Default: 300 | Env: ACME_PROXY_SIGNER__RELAY__POLL_TIMEOUT_SECS
Total budget (in seconds) for one upstream issuance before the local order is marked invalid.
[signer.relay.dns01]
Only consulted when challenge_strategy = "dns01".
provider (String) — Default: "rfc2136" | Env: ACME_PROXY_SIGNER__RELAY__DNS01__PROVIDER
DNS provider used to publish the upstream TXT record. rfc2136 is currently the
only implementation.
challenge_alias (String) — Default: "" | Env: ACME_PROXY_SIGNER__RELAY__DNS01__CHALLENGE_ALIAS
A domain, e.g. acme-alias.net., under which every challenge record is
published as _acme-challenge.<alias> — see DNS alias mode.
Empty publishes at each domain’s own _acme-challenge name. The alias must lie
inside rfc2136.zone; one outside it, a wildcard, or a value that already
starts with _acme-challenge. is a startup error.
[signer.relay.dns01.rfc2136]
All default to "" and are required once the dns01 strategy is selected.
server — Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__SERVER
host:port of the nameserver accepting the dynamic update, e.g. 10.0.0.53:53.
This is the update target, distinct from dns.resolver, which governs
lookups.
zone — Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__ZONE
The zone the update is sent for, fully qualified with a trailing dot, e.g.
internal.company.com..
tsig_key_name — Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_KEY_NAME
Name of the TSIG key the update is signed with.
tsig_key_secret — Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_KEY_SECRET
The TSIG shared secret, in standard base64 — note this differs from EAB
secrets, which are base64url. A value that is not valid base64 is a startup
error, not a runtime one. This key is legitimately long-lived, so unlike a
one-shot EAB credential it belongs in configuration; still prefer the
environment variable over a file on disk.
tsig_algorithm (String) — Default: "" (read as hmac-sha256) | Env: ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_ALGORITHM
The TSIG algorithm: hmac-sha256, hmac-sha384 or hmac-sha512. Must match
the key as your nameserver defines it.
Records are added and removed by value, so other TXT values at the same name —
another order’s for that name, or ones this server did not write — are left
alone. The challenge name must lie inside zone; one outside it is refused
before anything is sent.
Updates are sent over UDP and retried over TCP when the response is truncated —
a TSIG-signed update readily exceeds 512 bytes, so the TCP path is a normal
occurrence rather than an edge case.
Only a UDP answer from server itself is accepted.
A successful answer counts only when it is TSIG-signed by the configured key,
for this update, within the server’s time window (RFC 8945); anything else fails
the update. A refusal is reported as the server sent it, with its TSIG error
explained: BADKEY means the server does not know tsig_key_name under
tsig_algorithm, BADSIG that tsig_key_secret does not match, and BADTIME
that the two clocks disagree.
[signer.relay.dns01.propagation]
What the dns01 strategy waits for between publishing the TXT record and
asking the upstream to validate it. The upstream looks once: a record it cannot
see yet makes the authorization invalid for good, which fails the client’s
order and, at a public CA, counts against its failed-validation limit.
mode (String) — Default: "none" | Env: ACME_PROXY_SIGNER__RELAY__DNS01__PROPAGATION__MODE
none triggers the challenge right after the update succeeds, which is right
when the update server is itself what the CA asks. delay sleeps delay_secs
first, for a provider that accepts an update before serving it (a DNS API behind
an RFC 2136 bridge) or secondaries that lag their primary. Any other value is a
startup error.
delay_secs (u64) — Default: 30 | Env: ACME_PROXY_SIGNER__RELAY__DNS01__PROPAGATION__DELAY_SECS
Seconds to wait under delay; ignored under none. Zero is refused (use
none), and so is a value not less than poll_timeout_secs, which bounds the
whole relay attempt. The delay runs once for each name in the order, so a
multi-name order needs poll_timeout_secs above the delay times the number of
names plus the time the upstream takes.
The wait is a fixed delay rather than a poll of a public resolver on purpose. A CA resolves through the zone’s authoritative nameservers itself, while a recursive resolver answers from whichever one it reached and caches a “no such record” from a lagging one for the zone’s negative TTL — and never sees an internal or split-horizon zone at all.
[signer.relay.eab]
An upstream External Account Binding credential supplied in configuration
rather than through acme-proxy upstream register. Both keys are empty by
default, which means “no configuration-file credential”. Read only by
acme-proxy serve, and only on a startup that finds no .kid sidecar beside
account_key_path — see EAB considerations below for
which of the two mechanisms to prefer.
kid (String) — Default: "" | Env: ACME_PROXY_SIGNER__RELAY__EAB__KID
The key id the upstream’s operator issued alongside the secret.
hmac_key (String) — Default: "" | Env: ACME_PROXY_SIGNER__RELAY__EAB__HMAC_KEY
Sensitive. The shared secret, in base64 — url-safe, unpadded url-safe or
standard are all accepted, the same three forms acme-proxy upstream register
takes. A value that decodes as none of them is a startup error, as is
setting either key without the other. Prefer the environment variable to a file
on disk, and clear it once registration has succeeded: while it stays non-empty
the server logs a signer_relay_eab_secret_in_config warning on every startup.
EAB considerations
An External Account Binding (EAB) credential is a one-time use token that
authorizes a single newAccount request and is useless afterwards —
registration itself only ever runs once, guarded by the .kid sidecar that ends
up next to account_key_path. There are two ways to supply it:
The Admin CLI, registering the proxy with the upstream CA out of band:
acme-proxy upstream register --profile prod --eab-kid "..." --eab-hmac-key-file /path/to/secret
The secret may also be piped on stdin. It is never accepted as a command-line
argument, because argv is visible to every user on the host via ps. Nothing
about this credential is ever written to disk — only the resulting account kid
persists.
--profile is required whenever the configuration defines more than one
profile: [signer] is a per-profile section, so registering “the upstream”
without saying which one would be registering nothing.
[signer.relay.eab] in configuration, read by acme-proxy serve
itself on the first startup with no .kid sidecar yet:
[signer.relay.eab]
kid = "..."
hmac_key = "..." # base64: url-safe, unpadded url-safe, or standard
This is the trade-off the CLI path exists to avoid: a bootstrap secret sitting
in configuration for the life of the server, in exchange for not needing a
separate imperative step — useful when config.toml is already populated by a
secrets manager or a templated deployment. Once registration succeeds, serve
logs a signer_relay_eab_secret_in_config warning on every startup for
as long as hmac_key stays non-empty, the same treatment challenge.bypass and
ipam.netbox.insecure_skip_verify get — clear it out once acme-proxy upstream show confirms a kid is stored. Setting kid without hmac_key, or
vice versa, is a startup error.
If the upstream requires EAB and neither mechanism supplies a working
credential, acme-proxy serve fails at startup naming both.
The outbound client validates the upstream’s TLS certificate against
webpki-roots — unlike challenge.http_01, which deliberately does not
validate the responder’s certificate. Here the certificate is the only thing
identifying the CA being handed your CSRs.
Custom Script Signer
The custom signer backend delegates certificate issuance, revocation and
metadata retrieval to an external script (Bash, Python, Go, …). Use it to
integrate acme-proxy with legacy PKI systems, HSMs, or internal APIs that do
not speak ACME natively.
acme-proxy still serves ACME to its own clients and still enforces domain
control, filters and EAB; only the signing step is handed off.
Configuration
[signer]
backend = "custom"
[signer.custom]
script_path = "/usr/local/bin/legacy-pki-bridge.sh"
timeout_ms = 15000
args = []
supports_crl = false
supports_renewal_info = false
Reference
script_path (String) — Default: "" | Env: ACME_PROXY_SIGNER__CUSTOM__SCRIPT_PATH
Path to the executable. An empty value is a startup error once this backend is selected.
timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_SIGNER__CUSTOM__TIMEOUT_MS
Budget for one invocation. issue and revoke run in the job queue, so there
it bounds one attempt. The crl and renewal_info hooks answer a request
inline, so while either is enabled this must stay below
server.request_timeout_ms — the server refuses to start otherwise.
args (Array) — Default: [] | Env: ACME_PROXY_SIGNER__CUSTOM__ARGS
Static arguments passed on every invocation. These are the only command-line arguments the script receives; the hook is not passed as an argument.
supports_crl (Boolean) — Default: false | Env: ACME_PROXY_SIGNER__CUSTOM__SUPPORTS_CRL
Whether the script implements the crl hook. While false, the hook is never
invoked — no process is spawned at all — and GET /crl has nothing to serve.
With it on, the answer is cached for a minute and re-served, so a burst of
requests to the unauthenticated GET /crl is one run of the script rather than
one per request. At most four read hooks (crl and renewal_info together)
run at a time, whatever the request rate.
supports_renewal_info (Boolean) — Default: false | Env: ACME_PROXY_SIGNER__CUSTOM__SUPPORTS_RENEWAL_INFO
Whether the script implements the renewal_info hook. While false, the hook
is never invoked and GET /renewalInfo/{certID} falls back to the server’s
own local estimate.
These two flags default to
falseand gate the hooks entirely. Arenewal_infohook written without settingsupports_renewal_info = truewill simply never run, with no error to explain why.
Hooks
The hook is selected by the ACME_SIGNER_HOOK environment variable, not by
a command-line argument. Every hook receives a JSON object on stdin.
| Hook | stdin | stdout | Gated by |
|---|---|---|---|
issue | {"hook":"issue","order_id":…,"identifiers":[{"type":"dns","value":"…"}],"csr_der_base64":"…"} | PEM certificate chain, leaf first | always |
revoke | {"hook":"revoke","cert_der_base64":"…","reason":<int|null>} | ignored | always |
crl | {"hook":"crl"} | raw DER of the CRL | supports_crl |
renewal_info | {"hook":"renewal_info","cert_der_base64":"…"} | see below | supports_renewal_info |
csr_der_base64 and cert_der_base64 are standard base64 of the DER
bytes — not PEM, and not ACME’s base64url.
issue
Exit codes are the contract:
0— stdout is the PEM chain (leaf first, issuers after). Trailing whitespace is trimmed and exactly one newline re-appended, since a strict parser needs a newline after the final-----END CERTIFICATE-----.3— reserved: the CSR is bad. The order becomesinvalidwith abadCSRerror, which the client reads when it polls. Do not use this exit code for backend failures.- anything else — an internal failure. The issuance is retried under the
job queue’s attempt budget, and the order is marked
invalid(terminal, but pollable) once it runs out.
The script runs in the worker role, in the signer_issue job finalize
queues: the client is answered processing and polls until the certificate is
there. It never runs inside a client’s request, so it may take as long as its
timeout_ms without holding one open.
The order’s requested
notBefore/notAfter(RFC 8555 §7.4) are not passed to the script — there is no contract for it, and inventing one would break existing scripts. Your script decides validity on its own. (Thelocal_cabackend does honour them, clamped.)
revoke
Exit 0 means revoked. Any non-zero exit is an internal failure, and
acme-proxy then leaves the order un-revoked so the operation can be retried —
the CA-side action is authoritative.
Revocation must be idempotent: acme-proxy may call this hook for a
certificate your PKI already considers revoked, and that must succeed rather
than error.
renewal_info
stdout drives RFC 9773:
- empty — no opinion; the server falls back to its own estimate.
<start> <end>— the renewal window, as epoch seconds.<start> <end> <explanationURL>— additionally supplies RFC 9773 §4.2’s optionalexplanationURL. The URL is last and optional so an existing two-token script keeps working unchanged.
Any other token count, or a non-integer timestamp, is an internal failure.
crl
stdout is the raw DER of the CRL, served by GET /crl. Empty stdout means “no
CRL”. Failures here are logged and swallowed — a broken crl hook degrades to
no CRL rather than taking the endpoint down.
Environment variables
| Variable | Set for | Value |
|---|---|---|
ACME_SIGNER_HOOK | every hook | issue, revoke, crl or renewal_info |
ACME_SIGNER_ORDER_ID | issue | The order being finalized |
ACME_SIGNER_IDENTIFIERS | issue | Comma-joined identifier values |
ACME_SIGNER_REASON | revoke | RFC 5280 reason code, empty when none given |
There is no ACME_SIGNER_PROFILE; the signer backend is never told which
profile it is serving. (Backends are shared between profiles with identical
[signer] configuration, so there would not always be one answer.)
Security & process isolation
- Environment clearing:
env_clear()is called. The script inherits a minimalPATHplus theACME_SIGNER_*variables above — nothing else. The server’s own environment may hold the NetBox token, SMTP password or RFC 2136 TSIG key, and a signing script has no business reading them. - Zombie protection: the child runs with
kill_on_drop(true)under atokio::time::timeout. A timeout alone only drops the future, so without this a hung script would outlive its deadline and leak a process per request. - Failure reporting: on a non-zero exit, the first non-empty line of stdout (falling back to stderr) is used as the error detail.
See Custom Plugins Examples for a complete script.
Challenge Validation
A challenge is how a client proves to acme-proxy that it controls the
identifiers it is asking for. Every authorization created by newOrder carries
one challenge per enabled type; satisfying any one of them makes the
authorization valid, and an order becomes ready once every authorization is.
acme-proxy implements all three challenge types RFC 8555 and RFC 8737 define:
http-01— serve a token over HTTP.dns-01— publish a TXT record.tls-alpn-01— present a special certificate in a TLS handshake.
This is validation acme-proxy performs against its own clients. It is
separate from how the relay signer backend
satisfies an upstream CA’s challenges on your behalf; the two are configured
independently and need not match.
Configuration
[challenge]
# Types offered, in this order. Empty or unknown = startup error.
enabled = ["http-01", "dns-01"]
# Skip validation entirely. Testing only.
bypass = false
# Budget for one validation attempt.
timeout_ms = 5000
[challenge] is a per-profile section: each endpoint can offer a different
set of types. See Profiles & Routing.
Reference
enabled (Array) — Default: ["http-01"] | Env: ACME_PROXY_CHALLENGE__ENABLED
Which types each new authorization offers, and in what order. Clients pick one.
Valid values are http-01, dns-01 and tls-alpn-01. An empty list, or an
unrecognised name, is a startup error — even when bypass is on, because a
bypassing server still has to advertise a challenge for the client to trigger.
bypass (Boolean) — Default: false | Env: ACME_PROXY_CHALLENGE__BYPASS
Mark a triggered challenge valid immediately, with no network check. With this
on, [filter] is the only access control there is — which is why it is not
the default.
timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_CHALLENGE__TIMEOUT_MS
Budget for one validation attempt, applied at the registry level whatever the
type. It bounds a job attempt in the runner, not a request — see
Validation runs in the job queue — and it
must stay below server.request_timeout_ms, which
server::profile::build_all refuses to start otherwise.
max_in_flight_per_account (Integer) — Default: 32 | Env:
ACME_PROXY_CHALLENGE__MAX_IN_FLIGHT_PER_ACCOUNT
How many of one account’s challenges may be validating at once. 0 is no
limit.
A validation is queued work that reaches out to an address the client named,
and an account can create as many orders as it likes. Without a cap, one busy
— or hostile — account fills the runner with outbound probes while signings,
revocations and CRL regenerations wait behind them. A trigger over the cap is
answered 429 rateLimited with a Retry-After, and the challenge is left
pending: the client re-triggers it once one of its own validations has
settled, and loses no order.
Raise it for a deployment that renews many certificates at once from one
account; the ceiling that matters is jobs.max_concurrent, which is how many
validations actually run in parallel.
Per-type keys live under [challenge.http_01] and [challenge.tls_alpn_01];
see those pages. dns-01 has no table of its own — it is governed by
dns.resolver.
Bypass is not a shortcut
With
challenge.bypass = true,[filter]is the only access control the server has. Anyone who can reach the endpoint can obtain a certificate for any name it will accept, without proving anything.
Bypass exists for two legitimate cases: local testing (as in the Quick
Start), and a deployment where an
IPAM-backed filter such as ipam is genuinely the
authority on which host may hold which name, making a network round-trip
redundant.
It defaulted to true early in this project’s life. That was reconsidered: an
empty filter.rules plus the default bind on every interface made the
combination an open CA, so the default is now false.
The two state machines
An authorization and its challenges are separate objects with separate statuses,
and the edge between them is the one worth internalising: an authorization
becomes valid as soon as any one of its challenges does. The siblings stay
pending for ever and that is correct.
stateDiagram-v2
direction LR
state "authorization" as A {
[*] --> a_pending: created with the order
a_pending --> a_valid: any one challenge valid
a_pending --> a_invalid: a challenge failed
a_pending --> a_deactivated: §7.5.2
a_valid --> a_deactivated: §7.5.2
a_pending --> a_expired: expires passed
}
state "challenge" as C {
[*] --> c_pending: created with the authorization
c_pending --> c_valid: proof accepted
c_pending --> c_invalid: proof refused — terminal
}
c_valid --> a_valid: promotes its parent
a_invalid and c_invalid are terminal: the client must create a new order,
and re-triggering the same challenge will not retry it. Deactivation is the
operator- or client-initiated exit — see Deactivation below.
Validation runs in the job queue
Triggering a challenge with POST /chall/{id} does not perform the check.
It claims the challenge, writes a challenge_validate job and answers straight
away with the challenge in the processing state, plus a Retry-After. The
job runner performs the outbound check and records the verdict.
RFC 8555 has the states for exactly this. §7.1.6: challenges “transition to the
processing state when the client responds to the challenge”. §8.2 pairs that
with a Retry-After on the challenge resource, which is what the client polls
against. certbot, acme.sh and lego all poll.
sequenceDiagram
participant C as ACME client
participant P as acme-proxy
participant W as job runner
participant T as The name being proven<br/>(port 80 / 443 / DNS)
participant D as SQLite
C->>P: POST /chall/{id}
P->>D: claim: pending → processing
P->>D: enqueue challenge_validate
P-->>C: 200 + challenge object (processing)<br/>+ Retry-After + Link: rel="up"
W->>D: claim the job
rect rgb(240, 240, 240)
Note over W,T: in the runner, under challenge.timeout_ms
W->>T: fetch token / query TXT / TLS handshake
T-->>W: answer, or timeout
end
W->>D: one transaction:<br/>challenge + authorization + order
Note over D: "is every authorization valid?"<br/>is read INSIDE this transaction
C->>P: POST /chall/{id} (retry — not a state change)
P-->>C: 200 + challenge object — valid or invalid
Three consequences:
challenge.timeout_msbounds a job attempt, not an HTTP request. It is therefore independent ofserver.request_timeout_ms, and the server no longer refuses to start when it exceeds it.- A client that points a name at an unreachable host no longer occupies one of
server.max_concurrent_requestswhile the server waits for it. - The server still needs egress to the client. For
http-01andtls-alpn-01it must be able to open a connection back to the machine requesting the certificate — a common source of “the order just sits atpending” in firewalled networks.
A validation is attempted once: a check that ran records its verdict, pass or fail. The job is retried only when the attempt could not happen at all — the database was unreachable, or the endpoint’s profile is not mounted by the process that picked the row up.
Both outcomes are 200
A validation failure returns 200 OK with the challenge object, its status
set to invalid and an error member describing what went wrong. It is not a
4xx.
This follows RFC 8555 §7.5.1, and it is load-bearing: certbot’s acme library
surfaces an HTTP error status as a transport failure, which would obscure the
actual reason the challenge failed. Read the challenge object’s status, not
the HTTP status.
Responses also carry a Link: rel="up" header pointing at the authorization,
which that same library requires.
A challenge that reaches invalid is terminal. The client must create a new
order; re-triggering the same challenge will not retry it.
Wildcards
A wildcard identifier such as *.example.com is accepted only when dns-01
is among enabled — it is the only challenge type that can prove control of a
whole subtree. Otherwise newOrder refuses with rejectedIdentifier, naming
dns-01.
For a wildcard identifier:
- The authorization is created on the base name (
example.com), with"wildcard": truein the authorization object. - It offers
dns-01alone, even if other types are enabled.
Ordering example.com and *.example.com together therefore produces two
authorizations on the same base name, and the TXT record for each goes to the
same _acme-challenge.example.com. acme-proxy matches any TXT record at
that name, so publishing both values side by side works.
Only a single leading *. is legal. *.*.example.com and foo.*.example.com
are rejected as malformed.
What happens on success
Success is committed as one transaction covering the challenge, its
authorization, and the order — including the “is every authorization now valid?”
read that promotes the order to ready. Doing that read inside the transaction
is deliberate: two concurrent validations of one order could otherwise each read
before the other’s write landed, and neither would promote the order.
Note the promotion depends on every authorization being valid, not every challenge. An authorization with three challenges needs only one of them.
Deactivation
A client can deactivate an authorization it no longer wants by POSTing
{"status": "deactivated"} to the authorization URL (§7.5.2). If the order had
already reached ready, it is demoted back to pending.
Deactivation is refused once the order is valid — at that point the
certificate exists, and revocation, not
deactivation, is what undoes it.
HTTP-01
The default challenge type. The client serves a token at a well-known path over
plain HTTP, and acme-proxy fetches it.
Not to be confused with the
http01upstream strategy. This page is aboutacme-proxyvalidating its own clients. The relay signer has achallenge_strategy = "http01"that runs the same challenge type in the opposite direction —acme-proxyserving the file, to prove itself to an upstream CA. The two share the well-known path and nothing else, and are configured independently.
How it works
- The client is given a random
tokenon the challenge object. - It computes the key authorization:
token + "." + base64url(SHA256(JWK thumbprint of its account key)). - It serves that string at
http://<identifier>/.well-known/acme-challenge/<token>. - It triggers the challenge with
POST /chall/{id}. acme-proxyresolves the identifier, fetches the URL, trims whitespace from the body, and compares it to the key authorization it computed independently.
The body is compared to the key authorization verbatim — unlike dns-01,
which compares a SHA-256 digest of it. Serving the digest here is a common
mistake when hand-rolling a responder.
Every ACME client implements a responder for this type, usually via a
--standalone mode or a webroot.
Limitations
- No wildcards.
*.example.comcannot be proven this way; usedns-01. - The name must resolve, and be reachable.
acme-proxyconnects to the client, usingdns.resolverto find it. A host behind a firewall that blocks inbound port 80 from the proxy cannot be validated.
Configuration
[challenge]
enabled = ["http-01"]
[challenge.http_01]
port = 80
https_port = 443
follow_redirects = true
max_redirects = 5
max_response_bytes = 4096
Reference
port (Integer) — Default: 80 | Env: ACME_PROXY_CHALLENGE__HTTP_01__PORT
Port the challenge is fetched from. RFC 8555 fixes this at 80 for the public Internet; it is configurable here because internal deployments frequently cannot bind low ports.
https_port (Integer) — Default: 443 | Env: ACME_PROXY_CHALLENGE__HTTP_01__HTTPS_PORT
Port used when a redirect sends the fetch to https.
follow_redirects (Boolean) — Default: true | Env: ACME_PROXY_CHALLENGE__HTTP_01__FOLLOW_REDIRECTS
Follow 3xx responses. Required by the specification, and commonly needed in practice — many hosts redirect all HTTP to HTTPS.
max_redirects (Integer) — Default: 5 | Env: ACME_PROXY_CHALLENGE__HTTP_01__MAX_REDIRECTS
Hop limit before the validation fails.
max_response_bytes (Integer) — Default: 4096 | Env: ACME_PROXY_CHALLENGE__HTTP_01__MAX_RESPONSE_BYTES
Cap on how much of the response body is read. A key authorization is under 100 bytes; this exists so a client cannot make the server read an unbounded stream.
The TLS certificate presented after a redirect to HTTPS is not validated (RFC 8555 §8.3) — at validation time the client by definition does not yet have a trusted certificate for the name.
Redirects are an SSRF surface
Following redirects means a client can steer the server’s fetch at an address of its choosing. Boulder’s usual mitigation — refusing to connect to RFC 1918 space — cannot apply here, because serving private networks is the entire point of this server.
What contains it instead:
- Only
httpandhttpsschemes are followed. - Only the two configured ports (
portandhttps_port) are connected to. - At most
max_redirectshops. - The shared
challenge.timeout_msbounds the whole attempt, redirects included. follow_redirects = falseturns the surface off entirely.- Nothing the fetch learned is echoed back to the client. The error a
client sees names its kind —
connectionorunauthorized— and the identifier, and says the server’s log has the reason. The status, the body length, the socket error and any redirect target stay inchallenge_validation_failed; a truncated body preview is logged atdebug. This is what stops the challenge from becoming a read primitive or a port scanner against your internal network. What remains is the kind itself: a client can still tell “nothing answered” from “something answered wrongly”.
If your threat model does not tolerate this, disable redirects, or use
dns-01, which makes no outbound connection to the client at all.
Troubleshooting
Look for these events in the log:
| Event | Meaning |
|---|---|
challenge_http_01_loaded | The fetch succeeded; a body was read. |
challenge_http_01_matched | The body matched. Validation passed. |
challenge_http_01_mismatch | The responder answered with the wrong content. Check that it is serving the key authorization, not just the token, and not the digest. |
challenge_http_01_redirect | A redirect was followed; the target is logged. |
challenge_validation_failed | The attempt failed. The detail says whether it was a connection error, a timeout, or a mismatch, with the status or socket error — the client is told only the kind. |
A connection failure usually means one of: the name does not resolve through
dns.resolver; nothing is listening on port; or a firewall blocks the proxy’s
egress to the client.
DNS-01
The client publishes a TXT record proving control of the domain. This is the only challenge type that can authorize a wildcard, and the only one that requires no inbound connectivity to the client at all.
How it works
- The client is given a random
token. - It computes the key authorization:
token + "." + base64url(SHA256(JWK thumbprint)). - It publishes a TXT record at
_acme-challenge.<identifier>whose value isbase64url(SHA256(keyAuthorization)). - It triggers the challenge, and
acme-proxyqueries that name throughdns.resolver.
The value is the digest, not the key authorization. This is the single most common mistake when writing a DNS responder by hand:
http-01serves the key authorization verbatim,dns-01serves the base64url-encoded SHA-256 of it. The two are not interchangeable.
acme-proxy matches any TXT record present at the name. That is required
rather than lenient: an order covering both example.com and *.example.com
produces two authorizations that publish two different values to the same
_acme-challenge.example.com, and both must be able to validate.
Configuration
There is no [challenge.dns_01] table — unlike the other two types, dns-01
has no per-type key. Enable it, and point dns.resolver wherever your
authoritative data lives:
[challenge]
enabled = ["dns-01"]
[dns]
# host:port of the nameserver every lookup goes through.
resolver = "10.0.0.53:53"
Leaving dns.resolver unset uses the system configuration (/etc/resolv.conf).
The resolver is deliberately uncached
The shared resolver performs no caching. This matters here more than anywhere else: a client typically publishes its TXT record and triggers the challenge seconds later, and a cached negative answer — an NXDOMAIN or empty NOERROR from the moment before publication — would defeat validation for the whole negative-TTL window.
(The one component that does cache is filter.reverse_dns, which builds its
own resolver. PTR lookups for an address that keeps connecting are exactly what
a cache is for.)
What this does not remove is propagation delay in your own DNS infrastructure.
If dns.resolver points at a recursive resolver rather than the authoritative
server, the record still has to reach it. Pointing straight at the authoritative
nameserver is the reliable choice for an internal deployment.
Wildcards
[challenge]
enabled = ["dns-01"] # required for *.example.com to be accepted at all
Without dns-01 enabled, newOrder refuses a wildcard identifier with
rejectedIdentifier. With it enabled:
- the authorization is created on the base name, flagged
"wildcard": true; - it offers
dns-01and nothing else, regardless of what else is enabled; - the TXT record still goes to
_acme-challenge.example.com— there is no_acme-challenge.*.example.com.
Client support
Every major client implements dns-01, but each needs a provider plugin or hook
script to write the record: certbot’s --manual with an auth hook or a DNS
plugin, acme.sh --dns dns_<provider>, lego’s --dns <provider>.
If you would rather your clients did not hold DNS credentials, that is exactly
what the relay signer backend exists for:
your clients prove control to acme-proxy however you like, and acme-proxy
holds the single RFC 2136 TSIG key that answers the upstream CA.
Troubleshooting
| Event | Meaning |
|---|---|
challenge_dns_01_loaded | TXT records were retrieved for the name. |
challenge_dns_01_matched | One of them matched. Validation passed. |
challenge_validation_failed | No record matched, or the lookup failed. |
If the lookup returns nothing, check in this order: the record exists at
_acme-challenge.<name> (not at the name itself); its value is the digest;
dns.resolver can actually see the zone; and the record has finished
propagating to whatever dns.resolver points at.
A TXT record split into multiple character-strings is reassembled by concatenation before comparison, so a long value chunked by your DNS server is handled correctly.
TLS-ALPN-01
Defined by RFC 8737. The client proves control by answering a TLS handshake on port 443 with a special self-signed certificate, rather than by serving anything over HTTP.
Its appeal is operational: validation happens entirely within the TLS layer, so a host that terminates TLS but serves no plain HTTP at all can still be validated, and nothing needs to be routed on port 80.
How it works
- The client is given a random
tokenand computes the key authorization as usual. - It generates a self-signed certificate that contains:
- a single
dNSNameSAN equal to the identifier being validated, and - a critical
id-pe-acmeIdentifierextension (OID1.3.6.1.5.5.7.1.31) whose value isSHA256(keyAuthorization).
- a single
- It arranges for that certificate to be presented when a handshake arrives
with SNI set to the identifier and ALPN protocol
acme-tls/1. acme-proxyperforms that handshake and inspects the certificate it gets back.
The proof is verified without ever completing an application-layer exchange: the certificate presented during the handshake is the answer.
What acme-proxy checks
- Exactly one
dNSNameSAN, matching the identifier. Not “at least one”. - The
id-pe-acmeIdentifierextension is present and marked critical, as RFC 8737 requires. - Its value equals the expected digest, compared in constant time.
The presented certificate is otherwise untrusted by design — it is self-signed and issued by the entity being challenged, so there is no chain to validate.
Configuration
[challenge]
enabled = ["tls-alpn-01"]
[challenge.tls_alpn_01]
port = 443
Reference
port (Integer) — Default: 443 | Env: ACME_PROXY_CHALLENGE__TLS_ALPN_01__PORT
Port the validation handshake connects to. RFC 8737 fixes this at 443 for the public Internet; it is configurable here for internal deployments that cannot bind it.
Client support
Support is thinner than for the other two types:
- lego implements a responder (
--tls). This is what the project’s end-to-end suite uses. - certbot and acme.sh do not implement a
tls-alpn-01responder. Listing this type is harmless for them — they simply pick another enabled type — but it cannot be their only option. - Servers that manage their own certificates, such as Caddy and Traefik, generally do support it natively.
If tls-alpn-01 is the only entry in challenge.enabled, certbot and acme.sh
clients will not be able to complete an order at all.
Limitations
- No wildcards. Use
dns-01. - Requires inbound connectivity from
acme-proxyto the client onport, the same constraint ashttp-01. - The responder must serve the ACME certificate only for the
acme-tls/1ALPN protocol, and its normal certificate otherwise. Serving it unconditionally would break ordinary traffic to that host for the duration.
Troubleshooting
| Event | Meaning |
|---|---|
challenge_tls_alpn_01_loaded | The handshake completed and a certificate was obtained. |
challenge_tls_alpn_01_matched | The certificate carried the right digest. Validation passed. |
challenge_validation_failed | Handshake failed, or the certificate did not satisfy the checks above. |
A handshake that fails outright usually means the responder did not negotiate
acme-tls/1 — many TLS servers simply fall through to their normal certificate
when they do not recognise the ALPN protocol, and the certificate then has no
id-pe-acmeIdentifier extension to find.
Filters
Filters provide access control for acme-proxy. They restrict which clients can
reach the server and which identifiers (e.g. DNS names) those clients can
request certificates for.
[filter] is a small policy engine with two halves:
- a check is one named question about a request — “is this address in the
management network?”, “does the inventory say this address owns this name?”.
Each is a
[filter.check.<name>]with atypesaying which question it asks. - a rule is a boolean expression over check names plus what a match means.
Each is a
[filter.rule.<name>], andfilter.ruleslists the ones to evaluate, in order. First match wins.
Everything is a named check — custom included — and two checks of the same
type are ordinary rather than impossible.
[filter]
rules = ["mgmt-bypass", "inventory-owned"]
[filter.check.mgmt-net]
type = "allowed_ip"
allow = ["10.0.0.0/8"]
[filter.check.corp-names]
type = "identifiers"
allow = ["*.corp.example.com", "corp.example.com"]
[filter.check.inventory]
type = "ipam"
[filter.rule.mgmt-bypass]
when = "mgmt-net"
then = "allow"
[filter.rule.inventory-owned]
when = "corp-names and (inventory or mgmt-net)"
then = "allow"
message = "this address owns no such name in the inventory"
See Policy: rules and conditions for the condition language and how rules are evaluated, and Checks for the types and their keys.
The two hooks
A check can act at two points, and both default to “pass” for a check that does not implement them:
- the connection stage — runs on every request, before anything else.
Refusal is
403 access_denied. - the identifier stage — runs at
newOrderand again atfinalizeagainst the names projected out of the CSR. It runs after the account is resolved, so a check can bind names to an account as well as to an address. Refusal is403 rejectedIdentifieratnewOrder,400 badCSRatfinalize.
Checking again at finalize is what stops a client from passing a benign
newOrder and then smuggling extra names into the CSR:
graph LR
REQ["Any request"] --> CONN["connection stage<br/>rules over address/path checks"]
CONN -->|deny| D403["403 access_denied"]
CONN -->|allow| ROUTE{"which resource?"}
ROUTE -->|"newOrder"| ID1["identifier stage<br/>names from the order"]
ROUTE -->|"finalize"| ID2["identifier stage<br/>names from the CSR"]
ROUTE -->|"anything else"| OK["handler"]
ID1 -->|deny| DREJ["403 rejectedIdentifier"]
ID2 -->|deny| DCSR["400 badCSR"]
ID1 --> OK
ID2 --> OK
Both stages must allow. They are evaluated independently, each over the
subset of rules that can run there, so a connection-stage allow does not skip
the identifier stage. A stage no rule applies to allows without consulting
filter.default — otherwise a policy made entirely of identifiers checks
would refuse every request before a name had been mentioned.
Available check types
| Type | Stage(s) | Purpose |
|---|---|---|
allowed_ip | both | CIDR allow/deny on the client address |
path | connection | Glob allow/deny on the request path |
reverse_dns | connection | PTR lookup with optional forward confirmation |
identifiers | identifiers | Glob or regex allow/deny on requested names |
eab | identifiers | Which EAB credential the account registered under |
ipam | identifiers | Ask an IPAM whether the client owns the names |
custom | both | Shell out to an operator-supplied script |
allowed_ip reads nothing but the client address, which both stages carry, and
that is what makes mgmt-net or inventory writable at all: the address half is
still answerable at the point where the inventory is consulted.
Seeing what a policy does
acme-proxy filter show prints the resolved policy, with every condition
re-parenthesized so you can see what the parser understood. acme-proxy filter explain evaluates it against a hypothetical request and reports each check’s
verdict, which rule matched, and the HTTP answer. See
the CLI reference.
Reference
rules (Array) — Default: [] | Env: ACME_PROXY_FILTER__RULES
Which [filter.rule.<name>] entries to evaluate, and in what order. Empty means
no filtering at all — anyone who can reach the server can obtain a certificate
if they satisfy the challenges (or, with challenge.bypass on, with no proof at
all). The server logs a filter_disabled warning at startup saying so.
default (String) — Default: "deny" | Env: ACME_PROXY_FILTER__DEFAULT
What happens at a stage where a rule was applicable and none of them matched:
"allow" or "deny". Never consulted at a stage no rule applies to.
trusted_proxies (Array) — Default: [] | Env: ACME_PROXY_FILTER__TRUSTED_PROXIES
CIDRs (or bare addresses) whose forwarded-for header is believed. A request from
any other peer is attributed to its own address, and its forwarded header is
ignored. Note this is a [filter] key — placing it under [server]
silently does nothing.
forwarded_header (String) — Default: "x-forwarded-for" | Env: ACME_PROXY_FILTER__FORWARDED_HEADER
The header consulted for the forwarded client address.
Fail-closed semantics
Every check that depends on the client address fails closed when it cannot determine one — including in deny-only (blocklist) mode. An address the server cannot see is not “absent from the deny list, therefore fine”.
This makes one deployment detail load-bearing: the server must be served with
connection info attached, which it is by default. Behind a reverse proxy, set
trusted_proxies — otherwise every request is correctly, but unhelpfully,
attributed to the proxy.
A check that cannot reach its authority — a DNS timeout, an unreachable
inventory — is a third answer, neither pass nor fail, and becomes a retryable
500 rather than a refusal the client would believe permanent. What that means
for a policy built out of or is worth reading in
Policy.
Keys the policy engine replaced
Each of these is refused by name at startup, so a configuration written against the older shape stops the server rather than coming up looking configured and filtering nothing. The refusal is a diagnostic and nothing more: none of these keys is still read, and the errors themselves go away at 1.0.0. Before then a configuration key may be renamed in any release, with every such change listed in the changelog.
| Removed | Replacement |
|---|---|
filter.enabled | Declare each filter as a [filter.check.<name>] with a type, write a [filter.rule.<name>] naming them, list it in filter.rules. |
filter.exempt_paths | A path check plus a rule — which can also combine the path with an address, and can glob. |
filter.custom_enabled | custom is an ordinary check type; filter.rules already says which run and in what order. |
[filter.allowed_ip], [filter.reverse_dns], [filter.identifiers], [filter.custom.<name>] | The type’s keys move onto its [filter.check.<name>] entry. |
[filter.netbox] | [ipam.netbox], read by a type = "ipam" check — see IPAM. |
Policy: rules and conditions
A rule is a boolean expression over check names plus what a match
means. filter.rules lists the rules to evaluate, in order, and the first
match decides.
[filter]
rules = ["mgmt-bypass", "inventory-owned"]
default = "deny"
[filter.rule.mgmt-bypass]
when = "mgmt-net"
then = "allow"
[filter.rule.inventory-owned]
when = "corp-names and (inventory or mgmt-net)"
then = "allow"
message = "this address owns no such name in the inventory"
mode = "enforce"
The condition language
expr := term ( "or" term )*
term := factor ( "and" factor )*
factor := "not" factor | "(" expr ")" | name
name := [a-z0-9-]+
not binds tightest, then and, then or; and and or are
left-associative, so a or b and c means a or (b and c). The three keywords
are matched case-insensitively and cannot be used as check names. A parse error
names the column it gave up at:
filter.rule.r.when: expected a check name, `not` or `(` at column 9 in "net and )"
acme-proxy filter show re-prints every condition with the grouping made
explicit, which is the quickest way to confirm the parser read what you meant.
Where a rule runs
A rule is evaluated at the intersection of the stages its checks can decide at — never the union. Evaluating a rule at a stage where one of its checks cannot run would silently treat that check as passing and change the boolean answer, so the intersection is the only composition that cannot lie.
The consequence worth knowing: a rule combining a connection-only check with an identifiers-only one has no stage at all. That is a startup error naming both sides, not a rule that quietly never fires.
filter.rule.strict combines `has-ptr` (connection only) and `corp-names`
(identifiers only), so there is no point in a request where both can be
evaluated. Give `has-ptr` stages = ["identifiers"] if it can decide there, or
split the rule in two.
Both stages must allow, and each evaluates its own applicable subset
independently. A stage no rule applies to allows without consulting
filter.default.
When a check cannot decide
A check has three possible answers, not two: it passed, it failed, or it could not decide. The third covers a DNS timeout, an unreachable inventory, a script that would not spawn — cases where the server learned nothing, which is not the same as learning “no”.
Conditions combine them with three-valued logic, where an unknown propagates only if it could change the answer:
and | or | |
|---|---|---|
| pass, pass | pass | pass |
| pass, fail | fail | pass |
| fail, fail | fail | fail |
| fail, unknown | fail | unknown |
| pass, unknown | unknown | pass |
| unknown, unknown | unknown | unknown |
The two bold rows are the point. mgmt-net or inventory keeps working through
an inventory outage, because a disjunction whose other side already passed does
not care what the unknown would have been. inventory on its own still becomes
a retryable 500 — the or buys resilience for the addresses it names and
nothing more.
The same principle applies at the rule level. A rule whose condition came back
unknown is not skipped: it is remembered, and once the policy reaches an answer
it is asked whether that would have mattered. If the unknown rule’s effect
differs from the effect actually reached, the whole stage is a 500. If it
agrees, the answer stands. This is what stops rule order from deciding
whether an outage is survivable.
Warn mode
mode = "warn" makes a matching rule log filter_rule_warned and not
decide — evaluation continues to the next rule. It is how a tightened policy is
rolled out: deploy it in warn mode, watch for the event, and switch to
enforce once no legitimate client trips it.
[filter.rule.inventory-owned]
mode = "warn"
A policy of nothing but warn rules therefore falls through to filter.default.
Because rules are a map rather than an array of tables, a profile can dry-run one rule and inherit the rest:
[profiles.staging.filter.rule.inventory-owned]
mode = "warn"
What the client is told
In order: the matching rule’s message if it has one; otherwise the first check
that actually refused, in evaluation order; otherwise a generic sentence. A
message is the way to say “ask the network team” instead of exposing which
check bit.
A check that could not decide never reaches the client — the response is a
plain 500, and the specifics stay in the logs.
Reference
filter.rule.<name>.when (String) — Required | Env: ACME_PROXY_FILTER__RULE__<NAME>__WHEN
The condition, in the language above. Empty is a startup error: a rule with no
condition is what filter.default is for.
filter.rule.<name>.then (String) — Required | Env: ACME_PROXY_FILTER__RULE__<NAME>__THEN
"allow" or "deny". No default — a rule that does not say what a match means
is one whose author has not finished writing it.
filter.rule.<name>.message (String) — Default: "" | Env: ACME_PROXY_FILTER__RULE__<NAME>__MESSAGE
Shown to the client verbatim in place of whichever check failed.
filter.rule.<name>.mode (String) — Default: "enforce" | Env: ACME_PROXY_FILTER__RULE__<NAME>__MODE
"enforce" or "warn". A warn rule matches, logs and does not decide.
Startup refusals
Every one of these stops the server rather than producing a policy that does not mean what it says.
| Configuration | Refusal |
|---|---|
filter.rules names a rule with no [filter.rule.<name>] | names the missing entry |
when names a check with no [filter.check.<name>] | names the missing check |
A rule is defined but filter.rules is empty | says to list the rules to evaluate |
when will not parse | names the column, and quotes the expression |
then missing, or not allow/deny | names the value |
mode not enforce/warn | names the value |
filter.default not allow/deny | names the value |
| A rule’s checks share no stage | names both sides and suggests stages |
A check is named and, or or not | says they are the language’s own words |
Checks
A [filter.check.<name>] is one named question about a request. It takes a
type saying which question, plus that type’s own keys:
[filter.check.mgmt-net]
type = "allowed_ip"
allow = ["10.0.0.0/8"]
[filter.check.corp-names]
type = "identifiers"
allow = ["*.corp.example.com", "corp.example.com"]
Two checks of the same type are ordinary — two identifier lists with different
rules, an address list per network, several script hooks — which is the main
thing the older filter.enabled shape could not express.
Naming
Each name must match ^[a-z0-9-]+$: lowercase letters, digits and -, the same
restriction profile names have. The reason is that a name is also an environment
variable segment (ACME_PROXY_FILTER__CHECK__<NAME>__…), and the config crate
lowercases those, so MgmtNet in a file and MGMTNET in the environment would
silently become two entries instead of one overriding the other.
and, or and not are the condition language’s own words and
cannot name a check.
Only what a rule names is built
A check defined here but mentioned by no selected rule is never constructed.
It opens no connection, spawns no client and validates nothing; it is reported
once at startup as filter_check_unused.
That is deliberate, and it is what makes profile inheritance usable: a global
[filter] section can carry a library of checks, every profile inherits all of
them, and each profile’s filter.rules picks the subset it actually wants
without paying for — or failing startup on — the rest.
Keys by type
type and stages are universal. Everything else belongs to one type, and
setting a key that belongs to a different type is a startup error naming
both — the keys are one flat namespace, so without that check script_path on
an allowed_ip check would simply be read by nothing.
| Type | Keys | Page |
|---|---|---|
allowed_ip | allow, deny | allowed_ip |
path | allow, deny | path |
reverse_dns | allow, deny, allow_regex, deny_regex, require_forward_confirm, timeout_ms | reverse_dns |
identifiers | allow, deny, allow_regex, deny_regex, allowed_types, allow_wildcards | identifiers |
eab | allow, deny, allow_regex, deny_regex, kids, require_active | eab |
ipam | (none — configured by the [ipam] section) | ipam |
custom | script_path, timeout_ms, pass_stdin, args | custom |
Defaults, where a type has one: require_forward_confirm = true,
timeout_ms = 2000 for reverse_dns and 5000 for custom,
pass_stdin = true, allow_wildcards = false, require_active = false,
allowed_types = ["dns", "cn"].
Every list defaults to empty, and empty always means “this type’s natural
default”, never “none”. That is not a style choice: an unset list environment
variable arrives as an empty list rather than as absent, so stages = [] has to
mean “infer” and allowed_types = [] has to mean ["dns", "cn"].
Matching: globs first, regexes on request
The name-matching checks take globs in allow/deny, where * matches one
label:
*.example.commatchesa.example.com; it does not matcha.b.example.com, and it does not matchexample.com. List the bare name too, exactly as you would in a certificate.- Everything else is literal. No
?, no character classes, no**. - Matching is case-insensitive.
allow_regex/deny_regex take regexes instead, automatically anchored as
^(?:…)$, and are unioned with the globs — so a policy can be mostly globs
with one regex where a glob will not do. Anchoring is not optional: the regex
crate searches rather than matches, so an unanchored example\.com would also
accept example.com.evil.net, which is precisely the bypass an allowlist exists
to prevent. Write .*\.example\.com for a suffix.
On allowed_ip the two lists are CIDRs or bare addresses instead; on path
they are path globs where * stops at /.
Allow and deny
One rule, shared by every check that has the pair:
denyis checked first and wins. Plain membership, not longest-prefix-match: a/32inallowdoes not beat a/8indeny.- An empty
allowimposes no constraint, so a deny-only configuration is a working blocklist rather than a list that refuses everything.
Which gives three usable shapes: allow-only (a strict allowlist), deny-only (a blocklist, everything else served), or both (an allowlist with holes punched in it).
Stages
There are two hook points — the connection, and the identifiers — and each type answers at the ones it can:
| Type | Default stages | Capable of |
|---|---|---|
allowed_ip, custom | connection + identifiers | the same |
path | connection | connection |
reverse_dns | connection | connection + identifiers |
identifiers, ipam, eab | identifiers | identifiers |
reverse_dns is the one whose default is narrower than its capability: it could
answer at the identifier stage from the same address, but a PTR plus
forward-confirmation exchange at newOrder and again at finalize triples
the lookups for an answer that has not changed. Opt in when you need it in an
identifier-stage rule:
[filter.check.has-ptr]
type = "reverse_dns"
stages = ["identifiers"]
Naming a stage the type cannot serve is a startup error saying why — an ipam
check at the connection stage would query the inventory on every newNonce.
That allowed_ip answers at both stages is load-bearing rather than
incidental: it is what makes mgmt-net or inventory a rule that can be
evaluated at all, since the address half must still be answerable at the point
where the names are known.
Reference
filter.check.<name>.type (String) — Required | Env: ACME_PROXY_FILTER__CHECK__<NAME>__TYPE
Which check type this instance is: allowed_ip, path, reverse_dns,
identifiers, eab, ipam or custom. An unknown value is refused by name,
as is the old netbox, which became ipam with its settings in
[ipam.netbox].
filter.check.<name>.stages (Array) — Default: [] | Env: ACME_PROXY_FILTER__CHECK__<NAME>__STAGES
Override where this instance decides: "connection", "identifiers", or both.
Empty infers from the type, per the table above.
Allowed IP Filter
The allowed_ip filter provides network-level access control based on the
client’s IP address. It can operate as an allowlist, a blocklist, or both.
How it works
This filter implements standard allow/deny semantics using CIDR matching:
- Deny wins: The
denylist is checked first. If a client matches any CIDR in the deny list, the request is immediately rejected. - Allow list: If an
allowlist is provided, the client must match at least one CIDR. If the list is empty, no allow constraint is imposed (functioning purely as a blocklist).
Client IP resolution & proxies
acme-proxy resolves the client IP securely. If the server is behind a reverse
proxy, the proxy must be trusted. Trusted proxies are declared under
[filter], not [server]:
[filter]
trusted_proxies = ["10.0.0.0/8"]
forwarded_header = "x-forwarded-for" # the default
Get the section right. Unknown keys are ignored rather than rejected, so
trusted_proxieswritten under[server]is silently dropped — no warning, no startup error. The forwarded header is then never believed, and every request is attributed to the reverse proxy’s own address instead of the client’s.
When a connection originates from a trusted proxy, acme-proxy walks the
forwarded-for header right-to-left, skipping trusted hops until it finds the
true client IP. Requests arriving from an address that is not in
trusted_proxies are attributed to their peer address, and any forwarded header
they carry is ignored — which is what makes a spoofed X-Forwarded-For header
useless. IP addresses are canonicalized internally (IpAddr::to_canonical()),
meaning IPv4-mapped IPv6 addresses (e.g., ::ffff:192.168.1.1) are properly
treated as IPv4.
Important (Fail Closed): The filter subsystem operates with strict fail-closed semantics. If the client IP cannot be determined (e.g., misconfigured reverse proxy or missing
TapIowrapper),ConnectionContext::require_client_ip()fails. The filter will deny access rather than assuming the client is safe.
Configuration
[filter]
rules = ["internal-only"]
[filter.check.internal-nets]
type = "allowed_ip"
# Deny external bad actors
deny = ["203.0.113.9", "198.51.100.0/24"]
# Allow internal networks
allow = ["192.168.1.0/24", "10.0.0.0/8", "fd00::/8"]
[filter.rule.internal-only]
when = "internal-nets"
then = "allow"
allow and deny take CIDRs or bare addresses (a bare address becoming a host
route), IPv4 or IPv6. Both empty while a rule names the check is a startup
error: an empty allow imposes no constraint, so the check would accept
everything, and an operator who configured it did not mean to turn on something
inert.
The shared allow/deny semantics, and the keys themselves, are documented under Checks.
This check answers at both stages, since it reads nothing but the client
address — which is what lets a rule say internal-nets or inventory and still
have the address half answerable once the requested names are known.
Path Check
type = "path" matches the request path, so a rule can be about what is being
asked for as well as who is asking.
[filter.check.public-paths]
type = "path"
allow = ["/crl"]
[filter.rule.public]
when = "public-paths"
then = "allow"
Connection stage only. By the identifier stage the path is always /newOrder or
/finalize/{id}, so a rule combining this with a name check would be asking a
question with a constant answer.
The /crl and /ca.pem trap
Both are served by the profile router, which means they sit behind the
filter policy exactly like /newOrder does. Turn on an address-based check
without accounting for them and two things break quietly:
- every relying party outside your allowlist loses revocation checking, and relying parties are precisely not the ACME clients you allowlisted;
- a host that has not installed the root yet cannot fetch it — which is the one moment it needs to, and the refusal looks like the CA being down.
So any address-based policy wants a companion rule:
[filter]
rules = ["public", "mgmt-only"]
[filter.check.public-paths]
type = "path"
allow = ["/crl", "/ca.pem"]
[filter.check.mgmt-net]
type = "allowed_ip"
allow = ["10.0.0.0/8"]
[filter.rule.public]
when = "public-paths"
then = "allow"
[filter.rule.mgmt-only]
when = "mgmt-net"
then = "allow"
Because public comes first and first match wins, both are served to anyone
while everything else still requires the management network.
Server-level routes — GET /health, GET /, and the http-01 responder — are
served by the root router, which no profile’s policy ever sees. They are
already unfiltered and need no entry here.
Paths are profile-stripped
Matching is against the path with the /profile/<name> prefix removed, so
/directory — not /profile/default/directory — is the value to list, and one
check covers every endpoint the process serves.
Globs
* matches one or more characters other than /, i.e. one path segment:
| Glob | Matches | Does not match |
|---|---|---|
/crl | /crl | /crl/extra |
/renewalInfo/* | /renewalInfo/abc123 | /renewalInfo/a/b, /renewalInfo/ |
/* | /directory, /newOrder | /renewalInfo/abc |
Everything else in a glob is literal, so a path containing regex syntax cannot
smuggle a pattern in. allow/deny follow the shared rule described in
Checks: deny wins, an empty allow imposes no
constraint.
A check with both lists empty is a startup error — it would match every request, which is a check that does nothing.
Combining with other checks
The reason this is a check rather than the flat filter.exempt_paths list it
replaces is that a path is rarely the whole answer:
# Only the operator network may revoke.
[filter.check.revocation]
type = "path"
allow = ["/revokeCert"]
[filter.rule.no-remote-revocation]
when = "revocation and not mgmt-net"
then = "deny"
message = "revocation is restricted to the management network"
The old list could only say “skip the connection stage entirely for this exact
string”, and could not express /renewalInfo/* at all.
Reverse DNS Filter
The reverse_dns filter provides advanced access control by resolving the
client’s IP address back to its associated hostnames via DNS PTR records.
This allows you to write policies based on the physical infrastructure naming convention rather than static IP addresses, which is incredibly useful in environments with dynamic IP allocation.
Resolution logic
When a client connects, the filter performs the following:
- It queries the DNS for PTR records associated with the client’s IP.
- Forward Confirmation (Optional): If
require_forward_confirm = true(the default), it takes the resulting hostnames and queries their A/AAAA records to ensure they point back to the original client IP. This prevents malicious actors from setting up a fake PTR record on an IP they control to spoof an internal hostname. - The resulting, validated hostnames are then checked against the
allowanddenyregex lists.
Crucial Detail:
acme-proxyapplies thedenylist across every PTR candidate returned by the DNS query, not just the one that would otherwise be accepted. If a client’s IP resolves togood.internalANDevil.hacker.net, and your deny list blocksevil, the connection is denied, even ifgoodmatches the allow list.
Timeout budget
DNS queries happen on the hot path — this is a connection-level filter, so it
runs on every non-exempt request, newNonce included. The filter operates
with a strict timeout_ms budget across all DNS queries so slow nameservers
cannot tie up the ACME server.
To keep that affordable, reverse_dns builds its own, caching resolver —
deliberately unlike every other DNS consumer in the server, which shares one
uncached resolver. A PTR lookup for an address that keeps connecting is exactly
what a cache is for, whereas the shared resolver must stay uncached so a
dns-01 TXT record published moments before a challenge is triggered is not
defeated by a cached negative answer. Both honour dns.resolver.
Configuration
[filter]
rules = ["known-hosts"]
[filter.check.has-ptr]
type = "reverse_dns"
# Require the PTR record to correctly forward-resolve back to the IP
require_forward_confirm = true
# Allow any host in the specific internal domain
allow = ["*.corp.example.com"]
# Deny the guest network infrastructure
deny = ["*.guest.example.com"]
timeout_ms = 2000
[filter.rule.known-hosts]
when = "has-ptr"
then = "allow"
allow/deny take globs over the resolved hostname; allow_regex/deny_regex
take anchored regexes and are unioned with them. The keys and their defaults are
documented under Checks.
Connection stage by default. It is capable of the identifier stage from the
same address, but a PTR plus forward-confirmation exchange at newOrder and
again at finalize triples the lookups for an answer that has not changed —
so opt in with stages = ["identifiers"] when a rule needs it there.
Identifiers Filter
The identifiers filter controls which domains, IPs, or URIs a client is
allowed to request a certificate for. This is crucial for preventing a
compromised client from requesting a certificate for a sensitive internal
domain.
Type flattening and cn handling
A Certificate Signing Request (CSR) can contain identifiers in multiple places
(Subject Alternative Names (SANs) and the legacy Subject Common Name (CN)).
acme-proxy flattens all of these into a single list of typed identifiers
(dns, ip, email, uri, other, cn).
- Deny applies everywhere: A
denyrule applies to every single type. If you deny*.evil.com, a client cannot sneak it into the Subject Common Name to bypass the filter. - Allow skips
cn:allowrules explicitly skip thecntype (SUBJECT_ONLY_TYPES). A CN is legacy metadata and often contains human labels (e.g.,"rcgen self signed cert"). It is not a true identifier the certificate is for, so it is exempt from strict allow-listing.
What reaches the filter
An order’s dns identifiers are checked for shape before any rule runs: a
value that is not a DNS name is malformed, and one the URL parser would read
as an IPv4 or IPv6 address — 10.0.0.5, but also 2130706433 or 0x7f.1,
which name 127.0.0.1 — is rejectedIdentifier. A rule written for names
therefore never has to anticipate an address spelled as one.
Regex anchoring
All matching is performed via Regular Expressions (Regex).
Security Notice:
acme-proxyautomatically anchors all regexes as^(?:pattern)$and makes them case-insensitive. You do not need to manually anchor your regexes. This prevents substring bypasses (e.g., an unanchoredexample\.comwould accidentally matchexample.com.evil.net).
Configuration
[filter]
rules = ["corp-names-only"]
[filter.check.corp-names]
type = "identifiers"
allowed_types = ["dns", "cn"]
allow = ["*.corp.example.com", "corp.example.com"]
deny = ["secret.corp.example.com"]
allow_wildcards = false
[filter.rule.corp-names-only]
when = "corp-names"
then = "allow"
allow/deny take globs, where * is one label — so *.corp.example.com does
not cover corp.example.com and both are listed, exactly as they would be in a
certificate. allow_regex/deny_regex take anchored regexes and are unioned
with the globs, for what a glob cannot express. The keys and their defaults are
documented under Checks.
Two instances of this type with different lists are ordinary, which is the usual way to say “these names from this network, those names from that one”:
[filter.rule.tenant-a]
when = "tenant-a-net and tenant-a-names"
then = "allow"
[filter.rule.tenant-b]
when = "tenant-b-net and tenant-b-names"
then = "allow"
EAB Check
type = "eab" matches on the External Account Binding credential the
requesting account registered under. It is the multi-tenant lever: mint one
credential per tenant, bind each to its own name space, and no tenant can
request another’s names.
[filter]
rules = ["tenant-a", "tenant-b"]
[filter.check.is-tenant-a]
type = "eab"
allow = ["tenant-a"]
[filter.check.tenant-a-names]
type = "identifiers"
allow = ["*.tenant-a.example.com"]
[filter.rule.tenant-a]
when = "is-tenant-a and tenant-a-names"
then = "allow"
Identifier stage only — at the connection stage no account has been authenticated, so there is no credential to ask about.
Why the label and not the account
The obvious handle for “which client is this” would be the account id, and it is the wrong one. An account id is a UUID v7 generated when the account is created, so a policy naming one can only be written after the fact, and you would be editing configuration in response to a client registering.
An EAB credential is the other way round: acme-proxy eab create --label tenant-a mints it before any account exists, and you choose the label. It
also survives what the account does next — a key rollover keeps the same
account, and a client re-registering under the same credential lands in the
same tenant. Credentials are deliberately reusable, so one label
naturally covers a whole team.
Blocking a single misbehaving account is a different job and already has a
lever: acme-proxy account deactivate <id>.
No credential means refused
An account registered without EAB has no credential, and this check refuses it. That is the only defensible reading of a question about which credential authorised the account — “none” cannot satisfy a tenant rule.
It follows that an eab check under a profile whose eab.enabled is false
could never do anything but refuse, so that combination is a startup error
rather than a policy.
Matching
allow/deny are globs over the label, with the usual semantics described
in Checks — deny first and winning, an empty
allow imposing no constraint. allow_regex/deny_regex take anchored
regexes and union with them.
kids pins credentials by their kid instead, for an operator who would
rather not rely on labels being unique. It is a second allow source rather
than a separate gate: either the kid being listed or the label matching is
enough.
[filter.check.is-tenant-a]
type = "eab"
allow = ["tenant-a"]
kids = ["4f1c…"] # this exact credential, whatever its label says
A credential minted with no label cannot match a label allowlist. That is the
same rule seen from the other side, and it is why eab create is worth always
giving a --label.
Making revocation retroactive
Revoking an EAB credential stops new registrations. Accounts already created
under it keep issuing, for ever — which is what the credential’s role as a
registration-time authorisation implies, and what
accounts.eab_kid’s own migration meant by calling it an audit trail.
require_active = true changes that for this check: the credential must still
be active, so acme-proxy eab revoke <kid> reaches existing accounts too.
[filter.check.live-credential]
type = "eab"
require_active = true
A deleted credential goes further than a revoked one, with or without
require_active: an account whose kid names a row that no longer exists
resolves to no credential at all, and every eab check refuses it, exactly as
it refuses an account registered without EAB. See
Deleting a credential.
Off by default, because turning it on retroactively changes what eab revoke
means for a deployment. It is the lever to reach for when a tenant’s credential
leaks — and it is a usable policy on its own, with no labels at all: “any
tenant, but not one whose credential we have withdrawn”.
Cost
Resolving the credential is two indexed reads, and they happen only when the
policy contains an eab check. A deployment without one pays nothing.
Reference
filter.check.<name>.kids (Array) — Default: [] | Env: ACME_PROXY_FILTER__CHECK__<NAME>__KIDS
Credential kids matched exactly, beside the label globs in allow. A second
allow source, not a separate gate.
filter.check.<name>.require_active (Boolean) — Default: false | Env: ACME_PROXY_FILTER__CHECK__<NAME>__REQUIRE_ACTIVE
Also require the credential to still be active, so revoking it refuses accounts already registered under it.
A check that sets none of allow, deny, kids or require_active only asks
whether the account used EAB at all, which the [eab] section already
guarantees. That is a startup error rather than a check that always passes.
Custom Script Filter
The custom filter lets operators write arbitrary scripts (Bash, Python, …) to
decide whether a connection or a CSR should be permitted. Use it for policy that
cannot be expressed with the built-in filters.
Configuration
[filter]
rules = ["scripted"]
[filter.check.check-network]
type = "custom"
script_path = "/etc/acme-proxy/filters/check-network.sh"
timeout_ms = 5000
pass_stdin = true
args = []
[filter.rule.scripted]
when = "check-network"
then = "allow"
custom is an ordinary check type: there is no separate selection list, because
filter.rules already says which checks run and in what order. Several
[filter.check.<name>] entries may point at the same script — each is told
which one invoked it through ACME_FILTER_CHECK_NAME, so one script can serve
them all and branch on it.
The keys and their defaults are documented under
Checks. An empty script_path on a check some rule
names is a startup error.
Hooks
The script is invoked at both filter hooks, distinguished by ACME_FILTER_HOOK:
ACME_FILTER_HOOK | When | Refusal becomes |
|---|---|---|
connection | Every non-exempt request | 403 access_denied |
identifiers | At newOrder and at finalize | 403 rejectedIdentifier (newOrder) / 400 badCSR (finalize) |
If your script only cares about one hook, branch on this variable and exit 0
otherwise — the script is called for both.
Data passing
Environment variables
connection hook:
| Variable | Value |
|---|---|
ACME_FILTER_HOOK | connection |
ACME_FILTER_CLIENT_IP | The resolved client address, canonicalized (an IPv4-mapped IPv6 address is flattened to IPv4). Empty when unknown. |
ACME_FILTER_METHOD | HTTP method |
ACME_FILTER_PATH | Request path |
identifiers hook:
| Variable | Value |
|---|---|
ACME_FILTER_HOOK | identifiers |
ACME_FILTER_CLIENT_IP | As above |
ACME_FILTER_ACCOUNT_ID | The authenticated ACME account |
ACME_FILTER_STAGE | newOrder or CSR |
ACME_FILTER_IDENTIFIERS | Comma-joined identifier values. Guaranteed free of commas and control characters — see below. |
ACME_FILTER_IDENTIFIERS is safe to split on , and safe to read line by
line. A request whose identifiers could not survive that join is refused before
the script runs, with badCSR, so a script never has to defend against it.
That guarantee needs stating because it is not free. At the newOrder stage
every identifier is a DNS name the server has already validated. At the CSR
stage the list also carries the certificate request’s subject CommonName,
which is arbitrary text — routinely a human label such as Example Corp Issuing CA rather than a host name, and therefore not validated as one. A
CommonName holding a comma or a newline would otherwise reach the script as
extra entries.
The typed JSON on stdin has no such ambiguity and is the better source when a
script cares which identifier is which: each entry is its own object with a
type, so a cn is distinguishable from a dns there and not here.
JSON on stdin
When pass_stdin is true (the default), a JSON object — not a bare array
— is written to the script’s standard input.
connection:
{"hook":"connection","client_ip":"203.0.113.5","method":"POST","path":"/newOrder"}
identifiers:
{
"hook": "identifiers",
"client_ip": "203.0.113.5",
"account_id": "…",
"stage": "newOrder",
"identifiers": [{"type": "dns", "value": "a.example.com"}]
}
client_ip is null rather than a string when the address is unknown.
At the CSR stage the identifiers list is the flattened projection of the
whole CSR — SANs and the subject Common Name — so entries of type ip,
email, uri, other and cn appear alongside dns. That is deliberate: a
deny rule cannot be dodged by moving a name from a SAN into the CN.
Return codes
0— permitted.- Any non-zero exit — denied. This includes exit code 255 and death by signal; the check is simply “did it exit successfully”.
- Timeout, or failure to spawn — treated as an internal error (
500), so the client retries rather than seeing a permanent refusal.
On denial, the reason sent to the ACME client is the first non-empty line of stdout, falling back to the first non-empty line of stderr, and finally to a generic “script exited with status …” message. Keep it to one line, and remember it is client-visible — do not leak internal detail into it.
Execution model and security
- Environment clearing (
env_clear): the child runs with a scrubbed environment, inheriting only a minimalPATHand theACME_FILTER_*variables above. The server’s own environment may hold secrets — the RFC 2136 TSIG key, the NetBox token, the SMTP password — and a filter script has no business reading them. - Zombie prevention (
kill_on_drop): execution is wrapped in a Tokio timeout withkill_on_drop(true). A timeout alone only drops the future, so without this a hung script would outlive its deadline and leak a process per request. - Cost: the
connectionhook runs on every non-exempt request, includingnewNonce. A script doing network I/O there will dominate your latency; prefer theidentifiershook when the policy only concerns names.
IPAM
An IPAM backend answers one question: which names does this address own?
acme-proxy asks it through the ipam filter, which
turns the answer into a decision — may this client have a certificate for the
names it is asking for? The two halves are deliberately separate. The
inventory reports what it holds; the filter alone decides what to do about
that, and it decides the same way whichever product answered.
Renamed at this release. This used to be a filter called
netbox, with its settings under[filter.netbox]. Both moved. See Migrating fromfilter.netboxbelow — the old spelling is refused by name at startup, with an error naming all three moves, rather than silently ignored.
Backends
| Backend | Product | Page |
|---|---|---|
netbox | NetBox | NetBox |
phpipam | phpIPAM | phpIPAM |
custom | an operator script | Custom Script |
One backend per profile. [ipam] is a per-profile section, so two endpoints
served by the same process may consult different inventories.
Sources
Each backend that reads an inventory itself takes a sources list naming the
places a permitted name may come from. It is validated at startup against that
backend’s own vocabulary. custom has no such list at all — the script reads
whatever it reads, so there is nothing to declare and no key to set.
| Source | What it reads | netbox | phpipam | custom |
|---|---|---|---|---|
dns_name | the address object’s own name | ✓ | ✓ | — |
custom_field | the custom field on the address | ✓ | ✓ | — |
device | the same field on the assigned device or VM | ✓ | ✓ | — |
vip | role-tagged service addresses on the same device | ✓ | — | — |
fhrp | addresses of an FHRP group the client’s interface is in | ✓ | — | — |
Two things the list does not say, and both matter:
- Order is meaningless. The result is a union of sets, unlike
filter.rules(evaluation order) orchallenge.enabled(offer order). deviceis a fallback, not a union. It is read only when the address object itself carried no value for the custom field. A value set on the address is the more specific statement, and an operator narrowing one address of a machine must not have the machine-wide list quietly widen it again.vipandfhrpare unions.
Empty, or an unknown entry, is a startup error. An inventory trusted for
nothing can never permit a name, so it is a filter that refuses everything —
and a typo that silently narrows an allowlist is worse than a refusal to boot.
A source that exists but belongs to another backend is refused by name too:
naming fhrp under [ipam.phpipam] says a check will run that phpIPAM cannot
run, and answering that with silence would leave an operator believing it does.
The default for both backends is ["dns_name", "custom_field", "device"] —
exactly what the old netbox filter always did. The two service-address
sources are off because both widen what a client may certify.
Matching is exact
Case-insensitive and ignoring a trailing dot, but otherwise literal. There is
no suffix rule and no wildcard expansion: an entry example.com does not
permit a.example.com, and a request for *.example.com requires that exact
string in the inventory. Same reasoning as the anchored patterns in
Identifiers — a rule that quietly covers more than
it says is the bypass an allowlist exists to prevent.
An ip identifier is permitted when it is the connecting address (a machine
may always certify the address it is talking from) or when it is listed like
any other name. A common name (cn) is skipped, as in identifiers. Any other
type is refused: an inventory has nothing to say about an email address or a
URI, and a filter whose job is to confirm entitlement must refuse what it
cannot confirm.
Denied versus Internal
The most consequential property of the whole subsystem.
“The inventory does not associate this name with this address” is a decision
about the client and denies the request (403 rejectedIdentifier). So is
“the inventory holds no record of this address at all”, worded differently so
an operator can tell the two apart.
“It answered 500”, “the token was refused”, “the lookup timed out” are not decisions about anybody — the server failed to reach one. Those become a 500 the client can retry.
This is enforced by the types, not by care at each call site: the error an
Ipam backend can return has no “denied” variant to reach for. An inventory
outage therefore stops issuance rather than permitting everything, and never
looks like a permanent refusal.
What it costs per request
The lookup runs inline in newOrder and again at finalize, so it is part of
those requests’ worst case. timeout_ms is one budget covering the whole
lookup however many requests the backend makes to answer it — not one per
request — and it must stay below server.request_timeout_ms.
Migrating from filter.netbox
Three changes, and the server refuses the old spelling of each by name: a check
still declared type = "netbox" names the first, a [filter.netbox] section
the other two:
- A check declared with
type = "netbox"becomestype = "ipam", plusipam.backend = "netbox". [filter.netbox]becomes[ipam.netbox]. Most keys are unchanged.- Three keys moved or changed shape:
| Was | Now |
|---|---|
filter.netbox.timeout_ms | ipam.timeout_ms |
filter.netbox.use_dns_name = true | "dns_name" in ipam.netbox.sources |
filter.netbox.device_fallback = true | "device" in ipam.netbox.sources |
Since the default sources is ["dns_name", "custom_field", "device"], a
deployment that left both booleans at their defaults needs no sources line at
all. Environment variables move from ACME_PROXY_FILTER__NETBOX__* to
ACME_PROXY_IPAM__NETBOX__*.
Startup fails on the old filter name rather than aliasing it. The section
moved too, so a silent alias would leave [filter.netbox] read by nothing
while the server came up looking configured — the same reasoning that made
signer.backend = "acme_proxy" a named refusal when it became relay. Both are
error messages rather than compatibility paths: nothing reads the old spelling,
and the refusals go away at 1.0.0. Renames like this are expected before then
and are listed in the
changelog.
Configuration
[ipam]
backend = "netbox"
timeout_ms = 5000
[ipam.netbox]
url = "https://netbox.internal.example.com"
token = "..."
sources = ["dns_name", "custom_field", "device"]
[filter]
enabled = ["ipam"]
Reference
backend (String) — Default: "" | Env: ACME_PROXY_IPAM__BACKEND
Which inventory to consult: netbox, phpipam, custom, or empty for none.
Anything else is a startup error rather than a silent fallback. Enabling the
ipam filter while this is empty is also a startup error.
timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_IPAM__TIMEOUT_MS
Budget for one whole lookup, however many requests the backend makes to answer
it — or, for custom, however long its script takes. Applied once around all
of them, so a wedged inventory cannot pin a request open. Exceeding it is
reported as a server error, not a denial.
NetBox
Reads what NetBox associates with the client’s address. Supports every source, including the two that resolve a shared service address.
What NetBox is asked
One lookup always happens:
GET <url>/api/ipam/ip-addresses/?address=<client ip>
Up to four more are made, each gated by a source:
| Query | Source |
|---|---|
dcim/devices/{id}/ or virtualization/virtual-machines/{id}/ | device |
ipam/ip-addresses/?device_id=N&role=… | vip |
ipam/fhrp-group-assignments/?interface_type=…&interface_id=… | fhrp |
ipam/ip-addresses/?fhrpgroup_id=… | fhrp |
A read-only API token is enough.
Authenticating
NetBox has two generations of API token, and this backend sends whichever it is given: the scheme is derived from the token itself, so there is nothing to configure.
| Token | Sent as | Where it comes from |
|---|---|---|
nbt_<key>.<secret> | Authorization: Bearer nbt_<key>.<secret> | v2, the default since NetBox 4.5 |
| anything else | Authorization: Token <token> | the legacy v1 token, not accepted from NetBox 4.7 |
The nbt_ prefix is NetBox’s own marker for a v2 token, and the whole string —
key, dot and secret — is displayed once when the token is created, so paste it
verbatim. A value starting nbt_ that carries no . is the key half on its
own: that is refused at startup, because NetBox would otherwise answer every
lookup with a 403, which looks exactly like a token that has been revoked.
Declaring names
Two places, and either or both can be trusted:
dns_nameon the IP address object — the ordinary case, one name per address.- A custom field (
custom_field, by defaultacme_domains) for the extra names that address may request. Configure it in NetBox as a multi-select or a text field onipam.ipaddressand — for thedevicesource — ondcim.deviceandvirtualization.virtualmachine. A single string is accepted as well as a list.
With the device source, an address that carries no value of its own falls
back to the field on its device or virtual machine, so names can be declared
once per machine rather than once per address. It is a fallback and not a
union: see Sources.
Shared and service addresses
A VRRP, CARP or keepalived pair answers on an address that belongs to the pair, not to either member — but the client connects from its own member address, so without one of the sources below it is refused a certificate for the service name. Both are unions: the member’s own names and the service address’s names are true at the same time.
Which one an estate needs depends on how it models redundancy in NetBox.
vip — a role on an address of the same device
The classic modelling: the service address is created on one of the members’ interfaces and tagged with a role.
GET <url>/api/ipam/ip-addresses/?device_id=3&role=vip&role=vrrp
vip_roles says which roles count. The role is re-checked on the answer as
well as sent as a filter — a filter parameter this server got wrong must never
degrade into “every address on the device”, which would widen an allowlist
without saying so.
fhrp — membership of an FHRP group
NetBox’s own model for first-hop redundancy: the service address is assigned to
an ipam.FHRPGroup, and each member’s interface is recorded as belonging to
that group.
client address ─▶ its interface ─▶ fhrp-group-assignments?interface_id=7
└─▶ group ids ─▶ their addresses
The direction of that chain is the membership proof. A group is only ever reached through an assignment naming the client’s own interface. Nothing is ever looked up by group name, by the service address, or by the identifier the client asked for — so there is no query that could reach a group the client is not recorded in, and no way to turn the check into a lookup of “who owns this name?” by choosing a request carefully. An interface in no group contributes nothing and costs one request.
No role filter applies here: an address assigned to an FHRP group is the
group’s service address by construction, and applying vip_roles would drop
legitimately untagged VIPs.
A client connecting from the service address needs neither source — that address object comes back from the first query with its own names attached.
TLS
Unlike the challenge validators, where the certificate is deliberately not checked because the proof is what matters, NetBox’s certificate is the only thing identifying the service whose answers decide who may have a name certified. The public roots apply, plus any operator-supplied CA. Switching that off is explicit, logged on every start, and meant to be temporary.
Configuration
[ipam]
backend = "netbox"
[ipam.netbox]
url = "https://netbox.internal.example.com"
token = "your_netbox_read_only_token"
custom_field = "acme_domains"
sources = ["dns_name", "custom_field", "device"]
Turning on service addresses:
[ipam.netbox]
sources = ["dns_name", "custom_field", "device", "vip", "fhrp"]
vip_roles = ["vrrp", "carp"]
Reference
url (String) — Default: "" | Env: ACME_PROXY_IPAM__NETBOX__URL
Base URL of the NetBox instance. Any path is kept, so an instance served under
a subpath works. Required when ipam.backend is netbox.
token (String) — Default: "" | Env: ACME_PROXY_IPAM__NETBOX__TOKEN
NetBox API token, of either generation — see Authenticating. A secret: prefer the environment variable.
custom_field (String) — Default: "acme_domains" | Env: ACME_PROXY_IPAM__NETBOX__CUSTOM_FIELD
Custom field holding the permitted names, on the address object and on its
device or virtual machine. Only read when sources names custom_field or
device.
sources (Array) — Default: ["dns_name", "custom_field", "device"] | Env: ACME_PROXY_IPAM__NETBOX__SOURCES
Where a permitted name may come from. All five sources are available here. See Sources.
vip_roles (Array) — Default: ["vip", "vrrp", "hsrp", "glbp", "carp", "anycast"] | Env: ACME_PROXY_IPAM__NETBOX__VIP_ROLES
Which NetBox address roles mark a service address. Read only when sources
names vip, so this is which roles rather than whether to look at all.
ca_cert_path (String) — Default: "" | Env: ACME_PROXY_IPAM__NETBOX__CA_CERT_PATH
Extra CA certificates (PEM) to trust on top of the public roots, for a NetBox
behind an internal PKI. Ignored when insecure_skip_verify is on.
insecure_skip_verify (Boolean) — Default: false | Env: ACME_PROXY_IPAM__NETBOX__INSECURE_SKIP_VERIFY
Skip verification of NetBox’s TLS certificate entirely. Meant as a temporary
way out of an expired NetBox certificate. Startup logs an
ipam_netbox_tls_verification_disabled warning for as long as it is set.
phpIPAM
Reads what phpIPAM associates with the client’s address.
Supports dns_name, custom_field and device; phpIPAM records no address
roles and no redundancy groups, so vip and fhrp are
refused by name at startup.
What phpIPAM is asked
One lookup always happens:
GET <url>/api/<app_id>/addresses/search/<client ip>/
One more is made when sources names device:
GET <url>/api/<app_id>/devices/<deviceId>/
Setting up the API application
In phpIPAM, under Administration → API, create an application:
- App id — becomes
app_id, and appears in every API path. - App permissions — Read is enough.
- App security — SSL with App code. The app code becomes
tokenand is sent as a baretokenheader (phpIPAM’s own scheme, notAuthorization).
The alternative SSL with User token scheme — user credentials exchanged for a six-hour session token — is not implemented: it needs a refresh loop and somewhere to keep the token, for no gain over a credential that can be rotated in the environment.
Declaring names
hostnameon the address — the direct analogue of NetBox’sdns_name, read whensourcesnamesdns_name.- A custom column (
custom_field, by defaultcustom_acme_domains). phpIPAM prefixes custom columns withcustom_, so the default carries that prefix; write whatever your column is actually called. Add it under Administration → Custom fields for IP addresses and — for thedevicesource — for Devices.
A phpIPAM custom field is a plain text column, so several names are written comma-separated:
www.example.com, api.example.com, mail.example.com
The split is on commas and is not configurable — a comma is not legal in a DNS name, so there is no estate it could need to differ for.
With the device source, an address whose column is empty falls back to the
same column on the device named by its deviceId. A fallback, not a union:
see Sources.
An unknown address is a 404
The one place phpIPAM’s wire behaviour differs in a way worth knowing about.
NetBox answers an unknown address with 200 and an empty result list; phpIPAM
answers 404 with its own envelope:
{"code": 404, "success": false, "message": "No addresses found"}
That is read as “no such address” — a denial naming the address — rather than as a transport failure, which would turn every request from an unrecorded machine into a retryable 500. Every other non-2xx status is still a failure, so a broken or misconfigured phpIPAM stops issuance rather than reading as “this address owns no names”. See Denied versus Internal.
Configuration
[ipam]
backend = "phpipam"
[ipam.phpipam]
url = "https://ipam.internal.example.com"
app_id = "acme"
token = "your_app_code"
custom_field = "custom_acme_domains"
sources = ["dns_name", "custom_field", "device"]
Reference
url (String) — Default: "" | Env: ACME_PROXY_IPAM__PHPIPAM__URL
Base URL of the phpIPAM instance. Any path is kept, so an instance served under
a subpath works. Required when ipam.backend is phpipam.
app_id (String) — Default: "acme" | Env: ACME_PROXY_IPAM__PHPIPAM__APP_ID
The API application’s identifier, which appears in every API path. One path
segment: letters, digits, - and _. Anything else is a startup error, since
a stray slash would silently retarget the API rather than fail.
token (String) — Default: "" | Env: ACME_PROXY_IPAM__PHPIPAM__TOKEN
The application’s app code, sent as a token header. A secret: prefer the
environment variable.
custom_field (String) — Default: "custom_acme_domains" | Env: ACME_PROXY_IPAM__PHPIPAM__CUSTOM_FIELD
Column holding the permitted names, on the address and on its device. Only read
when sources names custom_field or device.
sources (Array) — Default: ["dns_name", "custom_field", "device"] | Env: ACME_PROXY_IPAM__PHPIPAM__SOURCES
Where a permitted name may come from. Only these three are available here;
naming vip or fhrp is a startup error.
ca_cert_path (String) — Default: "" | Env: ACME_PROXY_IPAM__PHPIPAM__CA_CERT_PATH
Extra CA certificates (PEM) to trust on top of the public roots. Ignored when
insecure_skip_verify is on.
insecure_skip_verify (Boolean) — Default: false | Env: ACME_PROXY_IPAM__PHPIPAM__INSECURE_SKIP_VERIFY
Skip verification of phpIPAM’s TLS certificate entirely. Startup logs an
ipam_phpipam_tls_verification_disabled warning for as long as it is set.
Custom Script
Runs an operator-supplied script to answer the one question the whole subsystem asks: which names does this address own?
Use it when the inventory of record is something this server carries no client
for — a CMDB, a hosts file, an LDAP tree, a spreadsheet exported nightly, a
vendor API behind a Python wrapper. It is the same escape hatch the
custom filter and the
custom signer are for their subsystems, and it runs
under the same hardening.
It is also the backend with the least in it. There is no URL, no credential,
no TLS setting and no sources list: the script is the
inventory, and where it looks is its own business.
The contract
The script is told the client address twice — once in the environment, once in a JSON object on stdin — so neither a four-line shell script nor a Python one has to reach for the channel it finds awkward.
| Variable | Value |
|---|---|
ACME_IPAM_HOOK | Always names_for. There is one hook. |
ACME_IPAM_CLIENT_IP | The resolved client address, canonicalized (an IPv4-mapped IPv6 address is flattened to IPv4). |
The same on stdin:
{"hook": "names_for", "client_ip": "203.0.113.5"}
A script that exits without reading stdin is fine and is not an error.
The answer
stdout plus an exit code.
| Exit | stdout | Means |
|---|---|---|
0 | one name per line | The inventory holds this address, and these are its names |
0 | empty | Held, and entitled to nothing |
3 | ignored | No record of this address at all |
| anything else | the reason | The script failed — a retryable 500, never a denial |
Names are compared exactly, but the script need
not tidy them: each line is lowercased and stripped of a trailing dot, and a
blank line is ignored. So WWW.Example.COM. and www.example.com are the
same answer, and printing whatever form the inventory happens to hold is
correct.
One name per line rather than a separated list because a newline is the shell
idiom, and plain text rather than JSON because there is nothing here a
structure would carry that a list of lines does not — a contract needing jq
for what echo already does would be paid for by every script ever written
against it. The custom signer hands back its certificate chain the same way.
Why 3 is reserved
“Held, and entitled to nothing” and “no record of this address” are different
answers — the ipam check words a different refusal for each, so an operator
reading a 403 can tell them apart — and once stdout means the names, an
exit status is the only channel left to say the second one in. So it gets a
code of its own, exactly as the custom signer’s badCSR does.
Every other non-zero exit stays a failure, and that direction is the one that
matters. A script that breaks, is missing, or runs past ipam.timeout_ms
produces a server error the client can retry, never a refusal — see
Denied versus Internal. It is the same
guarantee an unreachable NetBox gets, and it is enforced by the types rather
than by care here: the error an IPAM backend can return has no denied variant
to reach for, so a broken script cannot fail open or look permanent.
An example
#!/bin/sh
case "$ACME_IPAM_CLIENT_IP" in
203.0.113.5)
echo www.example.com
echo api.example.com
exit 0
;;
203.0.113.6)
exit 0 # held, entitled to nothing
;;
esac
exit 3 # no record of this address
A fuller one, reading a real source, is in Custom Plugins Examples.
No sources
sources exists because NetBox and phpIPAM each read
several places and an operator has to say which are trusted. A script reads
whatever it reads, so there is nothing here to list and no key to set. The
vocabulary is closed and validated per backend, so a sources line under
[ipam.custom] is not a source this backend “does not support” — it is a
setting that does not exist.
What it costs
One forked process per lookup, inside newOrder and again at finalize.
ipam.timeout_ms is the budget, and it is also what kills the child at the
deadline; there is deliberately no second timeout here.
Note that acme-proxy filter explain really runs the policy, so it executes
this script — as it does the custom filter’s. That is the point of the
command, and it is reported in its side-effects list.
Security and process isolation
- Environment clearing:
env_clear()is called. The script inherits a minimalPATHplus the twoACME_IPAM_*variables above — nothing else. The server’s own environment may hold the NetBox token, the SMTP password or the RFC 2136 TSIG key, and an inventory script has no business reading them. - Zombie protection: the child runs with
kill_on_drop(true)under atokio::time::timeout. A timeout alone only drops the future, so without this a hung script would outlive its deadline and leak a process per request. - Failure reporting: on a non-zero exit other than
3, the first non-empty line of stdout (falling back to stderr) becomes the error detail. It is logged, not sent to the client: an internal error tells the client nothing about why.
Configuration
[ipam]
backend = "custom"
timeout_ms = 5000
[ipam.custom]
script_path = "/etc/acme-proxy/ipam/lookup.sh"
args = []
Reference
script_path (String) — Default: "" | Env: ACME_PROXY_IPAM__CUSTOM__SCRIPT_PATH
Path to the executable answering the lookup. Empty while ipam.backend is
custom is a startup error, the same as
signer.custom.script_path.
args (Array) — Default: [] | Env: ACME_PROXY_IPAM__CUSTOM__ARGS
Fixed arguments passed before the script is told anything about the request. One script can serve several deployments by branching on them.
There is deliberately no timeout_ms here.
ipam.timeout_ms is the budget the whole lookup runs
under, and a second one would contradict the rule the other backends follow.
Configuration Reference
This page is the structured reference for every core configuration parameter.
For deep dives into specific subsystems (Signers, Filters, Notifications, EAB), see their dedicated chapters — linked at the bottom, and the place where those sections’ keys are documented.
How configuration is resolved
Sources are layered, lowest precedence first:
- Built-in defaults — everything documented below has one, so an empty
configuration is valid apart from
[profiles]. - A configuration file —
config.tomlin the working directory, or the path inACME_PROXY_CONFIG(the extension may be omitted, in which case the format is inferred). A missing file is not an error. ACME_PROXY_*environment variables —__separates nested keys, a single_separates the prefix.server.tls.enabledisACME_PROXY_SERVER__TLS__ENABLED.
config.toml.example in the repository is the annotated companion to this page:
it lists every key with its default and its environment variable name in
context.
[profiles]is mandatory. Everything else can be left at its default, but the server serves ACME only through profiles and refuses to start without at least one enabled. See[profiles]below.
Every section, and where it is documented
Six sections are large enough to have a chapter of their own, and their keys
are documented there rather than restated here. This page stays the complete
map: every top-level section acme-proxy reads appears below, whether or
not its text lives here. A delegated section’s own sub-tables —
[signer.local_ca.subject], [ipam.netbox], [filter.check.<name>] and the
rest — are listed in the chapter that owns them.
Overridable marks the sections a [profiles.<name>] block may override.
Everything else is process-wide — one setting for the whole server, however many
endpoints it mounts.
| Section | Controls | Overridable | Documented |
|---|---|---|---|
[database] | The SQLite file or PostgreSQL server | no | below |
[server] | Listen socket, public URL, admission control | no | below |
[server.tls] | HTTPS on the ACME listener | no | below |
[admin] | The web admin listener and its sessions | no | below |
[admin.filter] | Who may reach the admin listener | no | below |
[admin.notify] | Operator security notifications | no | below |
[admin.tls] | HTTPS on the admin listener | no | below |
[nonce] | Replay-nonce freshness | no | below |
[audit] | Reverse lookups and retention for the trail | no | below |
[jobs] | The durable background-work queue | no | below |
[metrics] | The Prometheus listener | no | below |
[dns] | The resolver every outbound lookup uses | no | below |
[proxy] | The forward proxy outbound clients dial through | no | below |
[logging] | Filter, format, target | no | below |
[order] | The ACME order object’s lifetime | yes | below |
[meta] | Directory meta members | yes | below |
[profiles.<name>] | An ACME endpoint | — | below |
[signer] | How a certificate is obtained | yes | Signers |
[filter] | Who may ask, and for what | yes | Filters |
[ipam] | The inventory an ipam filter check consults | yes | IPAM |
[challenge] | How control of a name is proven | yes | Challenge Validation |
[notify] | Outbound notifications | yes | Notifications |
[eab] | External Account Binding | yes | EAB |
The criterion for the last six is having a chapter, not being overridable —
[order] and [meta] are overridable and documented here, because neither is
large enough to be worth a page. That is the whole rule; there is nothing
subtler going on.
config.toml.example in the repository carries the same list as a comment
header, with every key in context.
[database]
url (String) — Default: "sqlite://sqlite.db" | Env: ACME_PROXY_DATABASE__URL
Database connection URL, for accounts, orders, certificates and everything else this server keeps. The scheme picks the backend, and any other scheme is refused by name at startup:
sqlite://<path>— a file, created on first use. The default, and what a single-host deployment wants.sqlite:///var/lib/acme-proxy/acme.dbis an absolute path (three slashes).postgres://orpostgresql://— a server, which must already exist: creating a database is an operator’s act, not something a server does to a cluster it was pointed at. Runacme-proxy migrateonce against it.
PostgreSQL is what a multi-node deployment needs. SQLite across processes is fine on one local disk, and not safe on NFS or across hosts — see Deployment. Nothing else changes with the backend: the same binary, the same configuration and the same ACME behaviour.
A PostgreSQL URL usually carries user:password@. The password is never
logged: the startup line and the SIGHUP refusal both print it as
postgres://acme:***@host/db. It is still a credential in a configuration
file, so give that file the permissions it deserves.
This key cannot be changed by a reload — see Reload.
[server]
bind_address (String) — Default: "[::]:3000" | Env: ACME_PROXY_SERVER__BIND_ADDRESS
Network address the server binds and listens to. SIGHUP moves it without
restarting, and an address that cannot be bound refuses the reload rather than
taking the running socket down — see Configuration
Reload.
base_url (String) — Default: "http://localhost:3000" | Env: ACME_PROXY_SERVER__BASE_URL
Public base URL advertised in the ACME directory, with no trailing slash. It is
never derived from the request, and every signed request’s url field is
checked against it (RFC 8555 §6.4) — so behind a reverse proxy, or with TLS
enabled, this must be set to the public URL or every client is rejected.
max_concurrent_requests (Integer) — Default: 100 | Env: ACME_PROXY_SERVER__MAX_CONCURRENT_REQUESTS
How many ACME requests may be in flight at once before the server sheds load.
admission_wait_ms (Integer) — Default: 50 | Env: ACME_PROXY_SERVER__ADMISSION_WAIT_MS
How long a request may wait for a slot before it is refused. Past the limit a
request waits this long and then gets 503 + Retry-After — it is shed, not
queued.
request_timeout_ms (Integer) — Default: 60000 | Env: ACME_PROXY_SERVER__REQUEST_TIMEOUT_MS
Whole-request deadline. It must exceed signer.custom.timeout_ms when that
backend is installed with its crl or renewal_info hook enabled, since those
hooks run inline inside a request; the server refuses to start otherwise. It is
deliberately independent of challenge.timeout_ms and of the script’s issue
and revoke hooks, which run in the job queue — see Challenge
Validation. A relay or custom revocation that a request
queues is waited on for this long, less a second.
max_body_bytes (Integer) — Default: 131072 | Env: ACME_PROXY_SERVER__MAX_BODY_BYTES
Largest request body accepted (128 KiB).
These four keys govern the ACME routes only.
GET /healthis mounted outside all of them.
trusted_proxiesandforwarded_headerare not[server]keys — they live under[filter]. See Filters.
[server.tls]
Full treatment in TLS Termination.
enabled (Boolean) — Default: false | Env: ACME_PROXY_SERVER__TLS__ENABLED
Serve HTTPS on bind_address instead of cleartext — one listener, not two.
base_url is not rewritten for you; set it to https://… yourself or every
signed request fails the §6.4 URL check above.
cert_path (String) — Default: "server.pem" | Env: ACME_PROXY_SERVER__TLS__CERT_PATH
PEM certificate chain, leaf first. A self-signed certificate is generated and
written when either this or key_path is missing.
key_path (String) — Default: "server.key" | Env: ACME_PROXY_SERVER__TLS__KEY_PATH
Path to the private key.
handshake_timeout_ms (Integer) — Default: 10000 | Env: ACME_PROXY_SERVER__TLS__HANDSHAKE_TIMEOUT_MS
Budget for one TLS handshake. Handshakes run concurrently, off the accept path, so this never delays another client.
[admin]
The web admin interface — a second listener, on its own socket, serving no
ACME. Process-wide, so there is no [profiles.<name>].admin: an operator
manages every endpoint this process serves. Full treatment in
Web Admin.
enabled (Boolean) — Default: false | Env: ACME_PROXY_ADMIN__ENABLED
Off by default: a certificate authority should not grow a management surface
because somebody upgraded it. Bootstrap it with acme-proxy admin user create;
there is no sign-up page.
bind_address (String) — Default: "127.0.0.1:3001" | Env: ACME_PROXY_ADMIN__BIND_ADDRESS
Loopback on purpose. This listener has no admission control, filters nothing
until [admin.filter] names a rule, and — until [admin.tls]
is on — has no transport security.
Startup refuses a non-loopback bind while admin.tls.enabled is false.
The session cookie is always sent Secure, and a browser silently declines to
store one over plain HTTP on anything but localhost; the symptom would be
“sign-in works, then I am immediately signed out”, with nothing in any log to
explain it. Either enable [admin.tls], or keep the loopback bind and reach it
through an SSH tunnel:
$ ssh -N -L 3001:127.0.0.1:3001 ca.example.com
It is also an error for this to equal server.bind_address.
base_url (String) — Default: "http://localhost:3001" | Env: ACME_PROXY_ADMIN__BASE_URL
The origin the panel is reached at. Load-bearing three times over: the CSRF
origin check compares against it, a generated self-signed certificate takes its
host, and the pages build absolute URLs from it — exactly as server.base_url
does for the ACME listener. Through a tunnel this stays localhost. The
resolved origin is logged at startup (admin_origin_resolved) so a mismatch is
visible before the first refused request.
session_ttl_seconds (Integer) — Default: 43200 (12 h) | Env: ACME_PROXY_ADMIN__SESSION_TTL_SECONDS
Absolute session lifetime. Never extended by activity: past it, the operator signs in again.
session_idle_timeout_seconds (Integer) — Default: 3600 (1 h) | Env: ACME_PROXY_ADMIN__SESSION_IDLE_TIMEOUT_SECONDS
Idle lifetime, advanced on use — at most once a minute, so a polling page is not a stream of database writes. Whichever deadline comes first wins.
login_max_attempts (Integer) / login_window_seconds (Integer) — Defaults: 5 / 300 | Env: ACME_PROXY_ADMIN__LOGIN_MAX_ATTEMPTS, ACME_PROXY_ADMIN__LOGIN_WINDOW_SECONDS
Failed sign-ins allowed from one address per window, then 429 with a
Retry-After. The password hash is deliberately expensive (PBKDF2-HMAC-SHA256
at 600 000 iterations), so this is an availability control as much as a
credential one: over the limit, the hash is not computed at all.
Keyed on the client address. The forwarded-for header is believed only from a
peer listed in admin.filter.trusted_proxies — never from
filter.trusted_proxies, which governs the ACME listener. With no trusted
proxy listed, the limiter behind a reverse proxy counts the proxy.
require_mfa (Boolean) — Default: false | Env: ACME_PROXY_ADMIN__REQUIRE_MFA
Require a second factor (TOTP) of every operator. What this changes is the operator who has none: with it on, their next sign-in lands on the enrolment page and their session stays half-authenticated until they finish. An operator who already has one is challenged whether this is set or not.
It deliberately does not refuse a password-only sign-in outright: enrolling needs a session and a session would then need a factor, so that would brick the panel including the way in to fix it.
Turning it on does not retroactively end sessions that predate it — acme-proxy admin session revoke --all is the lever that does. While it is on and some
operator has no factor, every start logs admin_mfa_enrolment_pending. See
Operators and sessions.
max_body_bytes (Integer) — Default: 65536 | Env: ACME_PROXY_ADMIN__MAX_BODY_BYTES
Largest admin request body. An admin body is a small JSON object or a form, never a certificate.
page_size_max (Integer) — Default: 200 | Env: ACME_PROXY_ADMIN__PAGE_SIZE_MAX
Ceiling on ?limit= for the list endpoints; the default page size is 50. A
larger request is clamped, not refused.
template_dir (String) — Default: "" | Env: ACME_PROXY_ADMIN__TEMPLATE_DIR
Override individual page templates on disk, mirroring notify.template_dir.
Empty means the compiled-in defaults. The override is per file: a directory
holding only layout.html restyles the chrome of every page and leaves the
other fifty-five at their defaults. Every template is compiled at startup, so a
broken override refuses to start rather than serving a 500 later. Applies to
the /ui pages only; the JSON API has nothing to template. See
Customizing the Panel.
[admin.filter]
Who may reach the admin listener at all — for the deployment that cannot put a
host firewall in front of the port, such as a container. The same policy engine
and the same keys as the ACME listener’s
[filter] (rules, default,
trusted_proxies, forwarded_header, [admin.filter.check.<name>],
[admin.filter.rule.<name>]), under
ACME_PROXY_ADMIN__FILTER__…, with four differences:
- It is its own section. Nothing is inherited from the global
[filter], which is the ACME profiles’ base: an edit made for one listener must not open or shut the other. - Only the connection stage exists here, so only
allowed_ip,path,reverse_dnsandcustomchecks are accepted. Anidentifiers,eaboripamcheck, or a rule whose checks were moved to the identifier stage withstages, is refused by name at startup and on reload. - Empty
rulesfilters nothing, and logs no warning. trusted_proxiesalso keys the login limiter (seeadmin.login_max_attempts): it is the one list the admin listener believes a forwarded-for header from, and it applies whether or not a rule is written.
Every request is evaluated, /health included. A refusal is 403 — the admin
API’s JSON error (access_denied) under /api, the HTML error page elsewhere —
and is logged as filter_request_blocked with listener = "admin". See
Web Admin — Restricting who can reach it.
[admin.notify]
A whole [notify] section, process-wide, built only while [admin] is
enabled. It delivers the admin_sign_in / admin_credential_changed operator
security events (ASVS V6.3.5 / V6.3.7). Every key is the one documented for the
per-profile [notify] — under
ACME_PROXY_ADMIN__NOTIFY__… — with one difference: [admin.notify.email].to
may be empty, since each event carries the affected operator’s own contact
address as its recipient. See
Web Admin — Security notifications.
[admin.tls]
HTTPS on admin.bind_address instead of cleartext — the same
one-listener-not-two shape as [server.tls], and the same load-or-generate
provisioning.
enabled (Boolean) — Default: false | Env: ACME_PROXY_ADMIN__TLS__ENABLED
Anything but http://localhost needs this on, or the browser will not store the
session cookie at all.
cert_path / key_path (String) — Defaults: "admin.pem" / "admin.key" | Env: ACME_PROXY_ADMIN__TLS__CERT_PATH, ACME_PROXY_ADMIN__TLS__KEY_PATH
PEM chain (leaf first) and its private key. When either is missing, a
self-signed certificate for the host of admin.base_url is generated and
written at startup; the generated key is created 0600. Separate paths from
[server.tls] on purpose — the two listeners answer to different names and
should not share a certificate by accident.
handshake_timeout_ms (Integer) — Default: 10000 | Env: ACME_PROXY_ADMIN__TLS__HANDSHAKE_TIMEOUT_MS
As [server.tls].
[challenge]
Which challenge types each new authorization offers, whether they are validated
at all, and the per-type keys under [challenge.http_01] and
[challenge.tls_alpn_01].
Documented in full in Challenge Validation — this section is a per-profile subsystem with its own chapter, so its keys live there rather than being restated here.
[order]
validity_seconds (Integer) — Default: 604800 (7 days) | Env: ACME_PROXY_ORDER__VALIDITY_SECONDS
Lifetime of the ACME order object before it expires. This is housekeeping
for the order resource, not the issued certificate’s validity — that is
signer.local_ca.leaf_validity_days, or whatever the delegating backend
decides.
max_identifiers (Integer) — Default: 100 | Env: ACME_PROXY_ORDER__MAX_IDENTIFIERS
Most identifiers a single newOrder may name. A request naming more is refused
with urn:ietf:params:acme:error:malformed.
The only bound before this was server.max_body_bytes (128 KiB), which at
roughly thirty bytes per identifier admits some four thousand names in one
request. Every one of them becomes an authorization plus a challenge per offered
challenge type, all inserted in a single transaction — and SQLite has one
writer, so that transaction stalls every other write in the process while it
runs. 100 is what Let’s Encrypt allows and is far above what a real client
asks for; the point is that there is a ceiling.
Refused as malformed rather than rateLimited deliberately: the order is
malformed for this server whenever it is sent, so rateLimited — which tells a
client to come back later — would be a lie.
retention_days (Integer) — Default: 30 | Env: ACME_PROXY_ORDER__RETENTION_DAYS
Days an order is kept after it expires, before the daily order_sweep
deletes it. The authorizations and challenges beneath it go with it, through the
schema’s ON DELETE CASCADE. 0 keeps everything for ever.
A valid order is never swept, whatever its age. Its row is how
revokeCert and the CRL resolve a certificate by serial, and what RFC 9773
renewal information is derived from — deleting one would make an issued
certificate unrevokable and unrenewable, which is a far worse outcome than a
large table. Only orders that ended some other way are eligible: invalid, or
an abandoned pending/ready/processing order past its own expires, at
which point no client can act on them either.
Nothing pruned orders before this key existed, so the table and its children
grew for the life of a deployment — one order plus one authorization per
identifier plus one challenge per offered type, on a default configuration where
newAccount is open to anyone.
Per-profile like the rest of [order]. The sweep is a single job handler
(JobRegistry refuses two handlers for one kind) that applies each mounted
profile’s own value to that profile’s rows, emitting one order_reaper_swept
line per profile.
[nonce]
ttl_seconds (Integer) — Default: 300 | Env: ACME_PROXY_NONCE__TTL_SECONDS
Freshness window for JWS anti-replay nonces. Expired nonces are swept on an interval for the life of the process.
[audit]
Traceability and the CA’s audit trail. Process-wide, not per-profile — the
trail describes the CA, not one of its endpoints, so this section may not appear
under [profiles.<name>].
There is deliberately no enabled key. The address columns on accounts and
orders, and the audit_log table, are always written: recording who asked the
CA to sign something is what a CA does, not a feature to switch on. The only
thing here that can be turned off is the reverse lookup.
reverse_dns (Boolean) — Default: true | Env: ACME_PROXY_AUDIT__REVERSE_DNS
Resolve a PTR record for the client’s address and freeze it into the row beside
the address. Turn it off on an estate with no usable reverse zone: every lookup
would fail, every *_ptr column would end up NULL anyway, and all that would
be left is the round trip. Lookups go through dns.resolver like every other
DNS query this server makes.
reverse_dns_timeout_ms (Integer) — Default: 2000 | Env: ACME_PROXY_AUDIT__REVERSE_DNS_TIMEOUT_MS
Budget for one PTR lookup. Deliberately small: this runs inside a request that
has already done its real work, so a slow nameserver costs a NULL in one
column rather than latency on issuance. Every failure is a NULL, never a
refused request.
retention_days (Integer) — Default: 0 | Env: ACME_PROXY_AUDIT__RETENTION_DAYS
Delete audit_log rows older than this many days. 0 keeps everything for
ever, which is the right default for a trail whose value is that it is complete.
Any non-zero value schedules a daily audit_sweep job beside the nonce one,
running the same DELETE as acme-proxy audit cleanup --older-than <days>.
See Audit Trail.
[jobs]
The durable background-work queue: the jobs table plus the one runner that
drains it. Process-wide, not per-profile — there is one queue and one
runner for the process, so this section may not appear under
[profiles.<name>].
There is deliberately no enabled key. The queue is how the server finishes
work it has already promised a client: an order answered processing is owed a
certificate. Switching it off would not disable a feature, it would strand the
orders. What is tunable is how hard and how long the server tries.
Most of what the server does after answering a request runs here, so this section’s reach is wider than the name suggests:
- Client-visible work — challenge validation (
challenge_validate), issuance (signer_issue, andsigner_relay_issueunder therelaybackend), revocation through a relay or script (signer_revoke), and a local CA’s CRL signing (local_ca_crl_regenerate). - Deliveries — every notification
(
notify_deliver) and the expiry digest (notify_expiry_digest). - Periodic sweeps — expired nonces, orders past
order.retention_days,audit.retention_days, expired admin sessions, stalehttp-01tokens, the daily CRL refresh (local_ca_crl_sweep), and this queue’s ownretention_days.
A runner that is not running is a server that validates, issues, sweeps and
notifies nothing — job_runner_started is the line that says it is. With role
processes, only a worker process runs one.
poll_interval_ms (Integer) — Default: 1000 | Env: ACME_PROXY_JOBS__POLL_INTERVAL_MS
How often the runner looks for work nobody woke it for. This bounds only
scheduled work — a backoff coming due, a periodic sweep firing. Queueing a job
wakes the runner directly, so a relay queued by finalize starts immediately
whatever this says, which is what keeps a polling ACME client from waiting on a
tick.
max_concurrent (Integer) — Default: 8 | Env: ACME_PROXY_JOBS__MAX_CONCURRENT
How many jobs may run at once. A restart after an upstream outage that left a few thousand orders in flight would otherwise become a few thousand concurrent pollers against one CA, which is how a recoverable backlog turns into a rate-limit ban.
max_attempts (Integer) — Default: 5 | Env: ACME_PROXY_JOBS__MAX_ATTEMPTS
How many attempts a job gets before it is retired permanently. Counted when the job is claimed, so one that kills the process still exhausts its budget rather than crash-looping. With the 30-second base below and doubling, five attempts span roughly seven and a half minutes.
Unlike every other key here, this one is frozen onto each row as it is queued rather than read afresh each attempt, so raising it applies to work queued from then on and not to a backlog already waiting.
retry_base_seconds (Integer) — Default: 30 | Env: ACME_PROXY_JOBS__RETRY_BASE_SECONDS
The first retry delay; each subsequent one doubles. Longer than any single upstream round trip, so a retry is not simply the same failure again, and short enough that a blip clears inside one client poll cycle.
retry_max_seconds (Integer) — Default: 3600 | Env: ACME_PROXY_JOBS__RETRY_MAX_SECONDS
Where the doubling stops, and a real ceiling — the jitter applied to each delay only ever subtracts, so no retry is scheduled past this. An hour sits well under the default order lifetime, which keeps a job’s own deadline the binding constraint rather than this.
lease_seconds (Integer) — Default: 300 | Env: ACME_PROXY_JOBS__LEASE_SECONDS
The default budget for one attempt, and therefore how long a claim is held
before another runner may take the row. A handler needing a different one says
so itself: the relay signer asks for its own signer.relay.poll_timeout_secs.
retention_days (Integer) — Default: 7 | Env: ACME_PROXY_JOBS__RETENTION_DAYS
Delete settled job rows older than this many days. Unlike
audit.retention_days this defaults to a non-zero value: a
finished job is a receipt, not evidence, and the trail that has to be complete
is audit_log’s. 0 keeps everything for ever and stops the sweep being
scheduled at all.
[metrics]
The Prometheus exposition, served on a listener of its own — a third socket beside the ACME and admin ones, not a route on either. Process-wide, not per-profile: there is one counter set for the process, and the endpoint is a dimension of it rather than something each endpoint configures for itself.
The separate port is what settles the access question. A scrape carries no session and needs none, because reaching the port at all is the permission — so the control is your firewall, not a credential this server checks. Putting it on the ACME listener would have meant an unauthenticated route on a public socket; putting it on the admin listener would have meant an auth exemption on a listener whose rule is that every route but sign-in needs a session, plus coupling metrics to the panel being enabled.
There is deliberately no path key — /metrics is what every scrape
configuration already assumes, and this listener serves nothing else. There is
no [metrics.tls] either, unlike [server.tls] and
[admin.tls]: those carry a client’s signed requests and an
operator’s session cookie, while a scrape carries no credential and the
exposition holds no secret.
What is exposed, and how to point Prometheus at it, is in Monitoring.
enabled (Boolean) — Default: false | Env: ACME_PROXY_METRICS__ENABLED
Bind the metrics listener and serve GET /metrics. Off by default, the same
posture [admin] takes: a certificate authority should not open a new
socket because somebody upgraded it. Off means no socket at all, rather than one
answering 404.
bind_address (String) — Default: 127.0.0.1:3002 | Env: ACME_PROXY_METRICS__BIND_ADDRESS
The socket the exposition is served on — beside the other two, server on
3000 and admin on 3001.
Unlike admin.bind_address, a non-loopback value is neither refused
nor warned about. The admin listener refuses one without TLS because its
cookie is always Secure, which a browser silently declines to store over plain
HTTP, so the symptom would be an unexplained sign-out loop. Nothing here has a
cookie, and a port reachable from a Prometheus host on another machine is
exactly the intended deployment.
A value equal to server.bind_address, or to admin.bind_address while the
panel is enabled, is refused — at startup and on a reload alike — since only
one of the three could then bind.
Both keys reload: SIGHUP moves this listener, or switches it on and off,
without restarting. The new socket is bound before anything is published, so an
address that cannot be bound refuses the reload and leaves the running one
serving. See Configuration Reload.
[dns]
resolver (String) — Default: unset (system configuration, i.e. /etc/resolv.conf) | Env: ACME_PROXY_DNS__RESOLVER
host:port of the nameserver every DNS lookup this server makes goes
through: the dns-01 TXT query, the connect target that http-01/tls-alpn-01
resolve before reaching out, and filter.reverse_dns’s PTR and forward lookups.
The shared resolver is deliberately uncached, so a TXT record published
moments before a challenge is triggered is not defeated by a cached negative
answer. filter.reverse_dns is the one exception and keeps its own cached
resolver.
[proxy]
The forward proxy every outbound HTTP client dials through: the upstream CA the
relay signer talks to, the IPAM inventory, the notification webhooks, and the
http-01 and tls-alpn-01 challenge validators — the last through a CONNECT
tunnel, since it is TLS rather than HTTP.
Not everything outbound. SMTP (notify.email) and the RFC 2136 updates
signer.relay.dns01 makes are not HTTP and keep dialling directly; an estate
whose egress is proxy-only needs a separate route for those two.
Every key is empty by default, which means no proxy at all. There is no
enabled key — the presence of a URL is the switch.
http_url (String) — Default: "" | Env: ACME_PROXY_PROXY__HTTP_URL
Proxy for http:// targets, e.g. http://proxy.example.com:3128. Falls back to
$http_proxy when empty. A cleartext target is forwarded rather than
tunnelled: the request line carries the whole URL, which is what RFC 9112
§3.2.2’s absolute-form is for.
https_url (String) — Default: "" | Env: ACME_PROXY_PROXY__HTTPS_URL
Proxy for https:// targets, reached by CONNECT. Falls back to
$https_proxy, then $HTTPS_PROXY.
Normally the same http://proxy.example.com:3128 as http_url: this key names
the proxy used for https targets, not a proxy spoken to over https. An
https:// value is a startup error rather than a second TLS layer with no trust
anchor configured for it, and so is a socks5:// one.
The two keys are independent on purpose. An estate that proxies only its TLS egress is ordinary, and “it worked for http and silently did nothing for https” is the failure a single key would produce.
no_proxy (Array<String>) — Default: [] | Env: ACME_PROXY_PROXY__NO_PROXY
Targets that bypass the proxy. An entry is * (everything), a domain — which
also matches everything under it — a .domain (the same thing), an address, or
a CIDR block. Matching is case-insensitive and ignores a trailing root dot.
A network entry is compared only against a target that is already an address literal: a hostname is never resolved to test one, which would mean a DNS lookup on every outbound request and a race with the connect that follows.
An entry carrying a port is a startup error. Matching is on the host, and an entry that silently ignored half of itself is worse than one that is refused.
Loopback and localhost bypass unconditionally, before this list is consulted.
That rule exists because of the environment fallback: an operator’s inherited
shell http_proxy must not route this server’s own loopback traffic through a
corporate proxy, and the failure that would cause carries no signal at all.
Environment fallback
Each key falls back to its conventional variable when left empty:
| key | then | then |
|---|---|---|
http_url | $http_proxy | — |
https_url | $https_proxy | $HTTPS_PROXY |
no_proxy | $no_proxy | $NO_PROXY |
An empty string counts as unset in both sources, so a ${VAR:-} shell default
does not become a proxy at the empty URL.
Uppercase HTTP_PROXY is deliberately not read. Under CGI a client-supplied
Proxy: request header lands in the environment under exactly that name
(httpoxy, CVE-2016-5385 and its siblings). This server is never a CGI process,
so the vector does not reach it — but Go’s net/http dropped the variable for
this reason, matching it costs nothing, and honouring a variable purely because
everything else does is the kind of decision that is only ever wrong.
HTTPS_PROXY has no such history and is honoured.
Credentials
Userinfo in the URL becomes a Proxy-Authorization: Basic header:
http://user:password@proxy.example.com:3128. Percent-encode anything unusual —
DOMAIN%5Cuser is decoded before the header is built, since a backslash or an
@ in a proxy username is entirely ordinary and encoding it verbatim sends the
wrong credential.
On a tunnelled connection the credential is spent on the CONNECT and is
not repeated inside the tunnel, where the origin server would read it. The
configured URL is never logged with its password: every log line and every error
message renders it with the password replaced.
When the proxy is down
A configured but unreachable proxy is an error, every time. There is no fallback to a direct connection: dialling around a controlled egress path at exactly the moment the control fails is the opposite of what the setting is for. Errors name the proxy rather than the origin, so a refused connection points at the host that actually refused it.
[meta]
The optional meta members of the directory (§7.1.1). All are empty by default
and omitted from the directory when empty — never sent as an empty value.
terms_of_service (String) — Default: "" | Env: ACME_PROXY_META__TERMS_OF_SERVICE
URL of your terms of service. This one has teeth: setting it turns on
§7.3.3, so newAccount then refuses any request without termsOfServiceAgreed: true (403 userActionRequired + a Link: rel="terms-of-service" header), and
account objects begin reflecting termsOfServiceAgreed.
website (String) — Default: "" | Env: ACME_PROXY_META__WEBSITE
Informational URL about the ACME server. Advertised only.
caa_identities (Array) — Default: [] | Env: ACME_PROXY_META__CAA_IDENTITIES
Hostnames this CA recognizes in CAA records. Advertised only — this server performs no CAA checking.
[logging]
All six keys are validated at startup: an unknown value is a refusal to start with a message naming the key, never a silent fallback — a CA running at a log level or to a destination its operator did not ask for is the worse failure. The same validation runs on a reload, where a bad value refuses the whole thing rather than half-swapping the log stream.
All six also reload on SIGHUP, so raising the level
or switching to JSON mid-incident does not cost a restart.
What the resulting records actually contain, and what to alert on, is Monitoring & Observability.
filter (String) — Default: "acme_proxy=info" | Env: ACME_PROXY_LOGGING__FILTER
EnvFilter directive, and the last of three layers to be consulted. The
precedence, highest first:
--log-level, typed on the command line. A flag was asked for here and now, where the other two are ambient — the same reasoning that has--color alwaysoutrankNO_COLOR. It setsacme_proxyalone, so it never turns on a dependency’s logging by accident.RUST_LOG, when set to a non-empty value. It replaces the whole filter, so a bareRUST_LOG=debugalso turns on debug logging for every dependency.- this key.
That precedence is the same on a reload as at startup, which means editing this
key while either of the two outranks it changes nothing; the server logs
server_logging_filter_overridden, whose source field names which one, rather
than letting the edit pass for applied.
This section describes the server’s log stream. An admin command emits
nothing at all unless asked — see
the CLI’s --log-level.
json_format (Boolean) — Default: false | Env: ACME_PROXY_LOGGING__JSON_FORMAT
Emit JSON instead of the human-readable format. Set this in production if you ship logs to ELK, Loki, Datadog or similar: the structured fields become first-class keys rather than text to be re-parsed.
target (String) — Default: "stdout" | Env: ACME_PROXY_LOGGING__TARGET
Where records are written: stdout or stderr. Any other value is a startup
error.
ansi (Boolean) — Default: true | Env: ACME_PROXY_LOGGING__ANSI
ANSI colour in the human-readable format. Turn it off when the log is piped to a
file or a collector that does not strip escape sequences. Ignored when
json_format is on.
span_events (String) — Default: "none" | Env: ACME_PROXY_LOGGING__SPAN_EVENTS
Span lifecycle records: none, close or full. Any other value is a startup
error. close emits one record as each span ends, carrying the time spent busy
and idle inside it — the closest thing to per-operation timing available without
a metrics endpoint, and cheap enough to leave on. full adds
new/enter/exit and is a debugging tool.
flatten_event (Boolean) — Default: false | Env: ACME_PROXY_LOGGING__FLATTEN_EVENT
JSON only: lift a record’s own fields (event, and everything beside it) to the
top level instead of nesting them under fields. What most pipelines want; off
by default because a field can then collide with one of the format’s own keys.
[profiles.<name>]
An ACME endpoint is a profile. The server serves ACME only through them, and at least one enabled profile is required — startup fails otherwise, with a copy-pasteable minimal configuration in the error.
# The whole minimum. `enabled` defaults to true, so naming it is enough.
[profiles.default]
enabled (Boolean) — Default: true | Env: ACME_PROXY_PROFILES__<NAME>__ENABLED
Parks a profile without deleting its configuration. It also doubles as the one
key an environment-only profile needs:
ACME_PROXY_PROFILES__DEFAULT__ENABLED=true defines a working profile with no
configuration file at all.
Load-bearing rules:
- The mount path is derived from the name, never configured.
[profiles.le]answers at{base_url}/profile/le/directory. Names must match^[a-z0-9-]+$. The name is public API — it appears in everykidand order URL a client stores — so renaming a profile invalidates every client’s saved account. - Inheritance is per key, not per section. The global
[signer],[filter],[challenge],[eab],[order],[notify]and[meta]sections are the base each profile overlays. A profile that sets onlychallenge.bypasskeeps the globalchallenge.enabledrather than reverting it to the compiled default. Precedence: profile key > global key > compiled default. - Arrays replace wholesale, never append. A profile’s
filter.rulesfully replaces the global one — order is the policy, so it is stated once per profile rather than accumulated from two places. - Profiles are a database boundary, not just a URL prefix. Accounts and
orders carry a profile column, and accounts are keyed
UNIQUE(profile, pubkey)— one client key used at two endpoints is two independent ACME accounts. - Signer backends are shared by configuration. Two profiles with identical
[signer]sections share one backend instance. Two profiles sharing alocal_cakey path while differing elsewhere is a startup error.
[signer]
backend = "local_ca"
[filter]
rules = ["corp-only"]
[filter.check.corp-net]
type = "allowed_ip"
allow = ["10.0.0.0/8"]
[filter.rule.corp-only]
when = "corp-net"
then = "allow"
# Inherits everything above, but dry-runs the one rule: `[filter.rule.<name>]`
# is a table, so overriding `mode` keeps `when` and `then` from the global one.
[profiles.dev]
filter.rule.corp-only.mode = "warn"
# Overrides two keys; keeps the whole filter policy.
[profiles.prod]
signer.backend = "relay"
signer.relay.directory_url = "https://acme-v02.api.letsencrypt.org/directory"
See Profiles & Routing.
Environment variable gotchas
These bite in production and produce no error, so they are worth knowing before you configure anything through the environment.
Array-valued keys are parsed from a comma-separated string. That means a
value containing a literal comma cannot be expressed. A regex such as
^host\d{2,3}\.example\.com$ is therefore file-only — through the
environment, {2,3} splits into two list entries. In a file, write the array
form (deny = ["^host\d{2,3}\.example\.com$"]), which is taken exactly as
written; a bare string in a file is split on commas too, since by the time the
value is read there is nothing left to say where it came from.
Items are trimmed, and an empty item is a startup error. a, b is
["a", "b"], and a,,b or a trailing comma is refused by name rather than
silently dropped or kept as an entry that matches nothing.
An array set to the empty string is the empty list. Shell defaults like
ACME_PROXY_FILTER__RULES="${RULES:-}" set the variable to "", which the
configuration layer cannot distinguish from a deliberate value — so it is read
as “no values”, which is what clears a list the file set. Do not expect ""
to mean “fall back to the file”.
A numeric-looking value loses its leading zeros. The environment source
parses 007 as a number before the list is built, so it arrives as 7. A
value whose spelling matters — an argument to a custom signer, say — belongs
in the file’s array form.
Unknown keys are ignored, not rejected. A misspelled key, or a key written
under the wrong section, is silently dropped. The most common instance of this
is trusted_proxies written under [server] when it belongs to [filter] —
see
Allowed IP.
Sections documented elsewhere
The six sections with a chapter of their own, expanded — see the map above for the rest.
[signer]— Signers: local_ca and its PKCS#11 keys, relay, custom[filter]— Filters: allowed_ip, path, reverse_dns, identifiers, eab, custom[ipam]— IPAM: the inventory theipamcheck consults — NetBox, phpIPAM, a custom script[challenge]— Challenge Validation: http-01, dns-01, tls-alpn-01[notify]— Notifications: email, webhook, custom[eab]— External Account Binding
Configuration Scenarios
This page outlines complete, practical examples of configuring acme-proxy for
different real-world use cases. Each block is a whole config.toml.
All three scenarios below use the
relaybackend. Starting the server registers an account atdirectory_url— see Relay. Use a staging endpoint while you are still working the configuration out.
Let’s Encrypt relay with DNS validation
This scenario configures acme-proxy to act as an internal relay. It intercepts
ACME clients locally, but ultimately relays the issuance requests to Let’s
Encrypt.
The proxy takes the burden of solving Let’s Encrypt’s DNS-01 challenges on behalf of internal users by utilizing an RFC 2136 dynamic DNS provider. Internal clients never see a DNS credential; the single TSIG key lives here.
[server]
base_url = "https://acme.internal.company.com"
bind_address = "[::]:3000"
[signer]
backend = "relay"
[signer.relay]
directory_url = "https://acme-v02.api.letsencrypt.org/directory"
account_key_path = "le_upstream.key"
contact = ["mailto:admin@company.com"]
challenge_strategy = "dns01"
poll_interval_ms = 2000
poll_timeout_secs = 300
[signer.relay.dns01]
provider = "rfc2136"
[signer.relay.dns01.rfc2136]
server = "10.0.0.53:53"
zone = "internal.company.com."
tsig_key_name = "acme-update-key"
# MUST be standard base64 (not base64url). Prefer the environment variable
# ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_KEY_SECRET to a file.
# A non-base64 value here is a startup error, not a runtime one.
tsig_key_secret = "c2VjcmV0LXJlcGxhY2UtbWU="
tsig_algorithm = "hmac-sha256"
[challenge]
# Require local clients to prove control to the proxy via HTTP-01
enabled = ["http-01"]
bypass = false
[profiles.default]
enabled = true
The key above writes only inside internal.company.com.. For names spread over
several zones, DNS alias mode points every
domain’s challenge at one alias zone instead of needing a profile per zone.
Public CA relay with HTTP validation
The same relay as above, for an operator who has no RFC 2136 write access to the
zone but does control the reverse proxy already fronting the names being
issued. Instead of publishing a TXT record, acme-proxy serves the upstream’s
challenge file itself.
[server]
base_url = "https://acme.internal.company.com"
bind_address = "[::]:3000"
[signer]
backend = "relay"
[signer.relay]
directory_url = "https://acme-v02.api.letsencrypt.org/directory"
account_key_path = "le_upstream.key"
contact = ["mailto:admin@company.com"]
# No [signer.relay.http01] table exists — this line is the whole
# configuration. The responder is a route on this server's own root router;
# acme-proxy does NOT open a second listener or bind port 80.
challenge_strategy = "http01"
poll_interval_ms = 2000
poll_timeout_secs = 300
[challenge]
# Local clients still prove control to the proxy independently.
enabled = ["http-01"]
bypass = false
[profiles.default]
enabled = true
This only works if the upstream CA’s fetch reaches acme-proxy. Add one
location to whatever already answers on port 80 for each name being issued:
location /.well-known/acme-challenge/ {
proxy_pass http://acme-proxy:3000;
# A `return 301 http://acme-proxy:3000$request_uri;` works equally well —
# RFC 8555 §8.3 permits following redirects.
}
Note this strategy cannot issue wildcards (nothing answers HTTP on the name
*.example.com); those need scenario 1’s dns01. See
Relay for Caddy and
Traefik equivalents.
Commercial ACME CA with EAB and NetBox filtering
This scenario uses a commercial CA backend. It enforces External Account Binding
(EAB) on the internal proxy so only authorized users can register. It
additionally uses NetBox, through the ipam filter, to verify if a client’s
IP is actually permitted to request a certificate for a specific DNS name.
[server]
base_url = "https://ca.internal.company.com"
[signer]
backend = "relay"
[signer.relay]
directory_url = "https://commercial-ca.example.com/acme/directory"
account_key_path = "commercial_upstream.key"
# The commercial CA already trusts this server's upstream account implicitly,
# so the upstream challenge is bypassed.
challenge_strategy = "bypass"
[eab]
# Require all internal clients to register with an EAB credential generated by the admin
enabled = true
[filter]
# Ask the inventory about every internal request
enabled = ["ipam"]
[ipam]
backend = "netbox"
timeout_ms = 5000
[ipam.netbox]
url = "https://netbox.internal.company.com"
# Store this in ACME_PROXY_IPAM__NETBOX__TOKEN ideally
token = "your_netbox_read_only_token"
custom_field = "acme_domains"
sources = ["dns_name", "custom_field", "device"]
[profiles.default]
enabled = true
Protocol Support
acme-proxy implements RFC 8555 in full, plus the extensions an ACME client in
2026 expects to find. This page is the conformance summary: what a client can
call, which RFC section governs it, and what is deliberately not implemented.
Everything here is served per profile, under /profile/<name>/. There is no
ACME at the bare root — see Profiles & Routing.
Resources
Every row is reachable at {base_url}/profile/<name><path>. The Advertised
column says whether the directory object names it: a client is expected to find
advertised resources by reading the directory rather than by constructing paths.
| Path | Method | RFC | Advertised | Notes |
|---|---|---|---|---|
/directory | GET, POST | §7.1.1 | — | It is the entry point; §6.3 requires POST-as-GET to work too. |
/newNonce | HEAD, GET, POST | §7.2 | yes | All three forms. |
/newAccount | POST | §7.3 | yes | Find-or-create by public key: 201 for a new account, 200 for an existing one, Location either way. |
/acct/{id} | POST | §7.3.2, §7.3.6 | — | Contact update and deactivation. kid-authenticated. There is deliberately no unauthenticated GET. |
/acct/{id}/orders | POST | §7.1.2.1 | — | The account’s order list, filtered as §7.1.2.1 requires. |
/keyChange | POST | §7.3.5 | yes | Key Rollover. |
/newOrder | POST | §7.4 | yes | Accepts notBefore/notAfter and RFC 9773’s replaces. |
/order/{id} | POST | §7.1.3 | — | POST-as-GET. |
/order/{id}/finalize | POST | §7.4 | — | Takes the CSR; hands it to the configured signer. |
/authz/{id} | POST | §7.5, §7.5.2 | — | One URL serves both the read and the deactivation, told apart by whether a payload arrived. |
/chall/{id} | POST | §7.5.1 | — | Triggers validation. Both outcomes are 200. |
/certificate/{id} | POST | §7.4.2 | — | POST-as-GET; PEM chain. |
/revokeCert | POST | §7.6 | yes | Revocation & CRL. |
/renewalInfo/{certID} | GET | RFC 9773 §4.1 | yes | Unauthenticated. Advertised without the id — §4.1 has the client append it. |
/crl | GET | RFC 5280 | no | Routed but not advertised: a CRL is CA infrastructure, not an ACME resource. |
/ca.pem | GET | — | no | The profile’s trust anchor, so installing it is one curl. 404 unless the backend has one of its own. Not advertised, for the same reason as /crl. See Trusting the CA. |
Two paths sit outside every profile, on the root router:
| Path | Purpose |
|---|---|
/health | Liveness. Outside the filter chain, the admission limiter and the nonce middleware — see Monitoring. |
/.well-known/acme-challenge/{token} | Mounted only when a signer backend has an http-01 token store to publish, i.e. signer.relay.challenge_strategy = "http01". See Relay. |
Two further paths are not on this listener at all. GET /metrics has a
socket of its own (off by default, [metrics]), so firewalling that port is
what controls who can read it — see
Monitoring. The web admin’s /ui and
/api likewise have their own listener; see
Web Admin.
Protocol behaviour worth knowing
Every signed request is checked the same way. The media type
(application/jose+json), any crit header, the JWS url against the route
actually reached, and the nonce are all verified before a handler runs, so no
resource can forget one. jwk and kid are mutually exclusive per §6.2, and
the verification algorithm never rests on the client’s alg alone. See
Architecture.
Refusals are problem documents. Every rejection is
application/problem+json with an RFC 8555 URN type. A multi-identifier order
rejected on several names comes back as one compound problem with a
subproblems entry per identifier (§6.7.1); a single rejection stays its own
type, unwrapped.
A POST always carries a fresh Replay-Nonce, errors included (§6.5). GET /directory, GET /crl, GET /ca.pem and GET /renewalInfo do not mint one —
nothing asks them to, and each nonce is a committed database write.
Extensions
| Extension | RFC | Default | Page |
|---|---|---|---|
| External Account Binding | §7.3.4 | off | EAB |
| Account key rollover | §7.3.5 | always on | Key Rollover |
| Renewal Information (ARI) | RFC 9773 | always on | ARI |
| Terms of service | §7.3.3 | off | Set meta.terms_of_service and newAccount starts enforcing it. |
| Wildcard identifiers | §7.1.3 | requires dns-01 | Challenge Validation |
Directory metadata
The directory’s meta object carries only what is configured. An unset member
is omitted, never sent empty — "website": "" says less than saying
nothing.
externalAccountRequiredappears astruewhen EAB is on for that profile.termsOfService,websiteandcaaIdentitiescome from[meta].
meta.terms_of_service is the one with teeth: setting it turns on §7.3.3, so
newAccount then refuses a request without termsOfServiceAgreed: true (403 userActionRequired plus a Link: rel="terms-of-service" header), and the
account object starts reflecting the field.
Not implemented
Stated explicitly, because each is something a reader may reasonably expect:
- CAA checking.
meta.caa_identitiesis advertised to clients; this server does no CAA lookup of its own. Where the relay backend is in use, the upstream CA performs its own. - OCSP. Revocation is published as a CRL at
GET /crl. There is no OCSP responder, and noauthorityInfoAccessOCSP pointer is ever written into an issued certificate. The local CA does write thecaIssuershalf of that extension, and acRLDistributionPointspointer, once an operator names the URLs — see Local CA. - Identifier types other than
dns.newOrderaccepts DNS names, including wildcards;ipidentifiers (RFC 8738) andpermanent-identifierare not supported. - Pre-authorization (§7.4.1). The directory does not advertise
newAuthz, and it is not routed. Authorizations exist only as part of an order. POSTto/renewalInfo— RFC 9773 §4.3’s optional client-side renewal signal. TheGEThalf is implemented.
External Account Binding (EAB)
acme-proxy supports enforcing External Account Binding (RFC 8555 §7.3.4).
When enabled, any client attempting to register a new account must provide an
EAB credential that was minted out-of-band by the server operator.
This is a highly effective security mechanism for internal CA endpoints: it restricts account creation to authorized entities without relying purely on IP filtering.
Solving commercial CA EAB limits
Beyond security, EAB support in acme-proxy solves a significant operational
hurdle when using Commercial CAs (like ZeroSSL, Sectigo, or GlobalSign).
When working with external or commercial CAs, organizations are sometimes restricted to a single account or a limited number of EAB credentials validated for specific domains. It is impractical to distribute these scarce upstream credentials to hundreds of individual internal servers.
By placing acme-proxy in front of the commercial CA:
- The proxy consumes a single upstream EAB credential to register its own master account with the Commercial CA.
- The proxy then issues its own unlimited local EAB credentials to your internal servers.
- This effectively multiplexes the upstream account, allowing thousands of internal clients to securely acquire certificates without exhausting your upstream quotas or spreading sensitive upstream secrets across your infrastructure.
Configuration
[eab]
enabled = true
Reference
enabled (Boolean) — Default: false | Env: ACME_PROXY_EAB__ENABLED
When enabled, the newAccount endpoint will refuse any request that doesn’t
carry a valid, unused EAB payload. Standard onlyReturnExisting lookups are
exempt because they only query existing accounts and never create new ones.
Minting credentials via CLI
You manage EAB credentials with the Admin CLI, because the secret is sensitive and is shown exactly once.
To create a new credential for a client:
acme-proxy eab create --label "DevOps Team"
# Bind it to a single endpoint in a multi-profile deployment:
acme-proxy eab create --label "DevOps Team" --profile prod
The CLI prints a kid (Key Identifier) and an HMAC secret. The secret is
printed only once. It is stored but never displayed again, so a lost secret is
replaced, not recovered.
--profile matters in a multi-tenant deployment: omitted, the credential is
accepted at every profile, which is what an unscoped credential means. Bind it
unless you intend that.
A client then uses the credential at registration:
certbot register \
--server https://acme.internal/profile/default/directory \
--eab-kid "the-kid" \
--eab-hmac-key "the-secret"
Credentials are reusable, not single-use
A credential is not consumed by the account it creates. The same
kidcan bind any number of accounts and keeps working until you explicitly revoke it. There is nousedstate — onlyactiveandrevoked.
This is worth being deliberate about. Handing one credential to a team means any number of hosts can register with it, and a leaked credential stays valid until someone notices. If you want one-client-one-credential, mint one per client and revoke it once that client has registered.
(This differs from the upstream credential consumed by acme-proxy upstream register — or, alternatively,
signer.relay.eab in configuration. Commercial CAs typically issue
single-use EAB credentials, which is precisely why that one is consumed once by
registration and then no longer needed at all, unlike the credentials described
on this page.)
Revocation
Revoke a credential to prevent any further use:
acme-proxy eab revoke <kid>
This takes effect immediately, with no restart — credentials are read from the
live database on every newAccount. Revocation is idempotent.
Revoking does not disturb accounts that already registered with the
credential; it only stops new ones from being bound. To shut out an existing
account, deactivate it: acme-proxy account deactivate <id>.
A revoked credential keeps its row, so the accounts it registered still resolve
to it — its label keeps matching eab filter checks, and
require_active can refuse them.
Deleting a credential
Deleting removes the row. What happens to the accounts registered with it is your choice:
acme-proxy account list --eab-kid <kid> # who registered with it
acme-proxy eab delete <kid> # keep the accounts
acme-proxy eab delete <kid> --deactivate-accounts # retire them
acme-proxy eab delete <kid> --delete-accounts # remove them
- Keep (the default) leaves the accounts as they are. They still name the
deleted
kidbut resolve to no credential, so everyeabfilter check refuses them from then on. --deactivate-accountsmoves them todeactivated: they can request nothing more, and their orders are kept, so every certificate they hold can still be revoked and still appears in the expiry digest until it expires. This is the way to retire a tenant.--delete-accountsdeletes them with their orders, authorizations and challenges. It is refused while any of those orders holds a live certificate (issued, not revoked, not expired): the order is the only record of the certificate, and without it the certificate could never be revoked. Revoke those certificates or wait for them to expire, or deactivate instead. A refusal changes nothing, not even the credential.
No choice revokes a certificate. The panel’s credential page offers the same
three, and the JSON API takes them as
DELETE /api/eab/{kid}?accounts=keep|deactivate|delete. Each writes an
eab_deleted audit row, plus one account_deactivated or account_deleted
row per account it changed.
Inspecting credentials
acme-proxy eab list --json
acme-proxy eab show <kid>
Neither ever renders the secret. eab list is newest first and paged
(Admin CLI → Paging); it reads the same query
GET /api/eab and the panel’s /ui/eab do, so the three cannot come to
describe the credential set differently.
Key Rollover
acme-proxy fully supports Account Key Rollover (RFC 8555 §7.3.5) via the POST /keyChange endpoint.
This allows a client to proactively rotate the cryptographic key-pair associated with their ACME account without needing to register a new account or abandon their existing authorizations and orders.
Cryptographic verification
The key rollover process is highly secure and requires the client to prove possession of both the old and the new key simultaneously to prevent hijacking.
- Inner JWS: The client constructs a JSON Web Signature (JWS) whose payload names the account and the new key, signed by the new key itself.
- Outer JWS: That inner JWS becomes the payload of an outer JWS, signed by
the old key currently on the account and carrying its
kid. acme-proxyunwraps this nested structure and verifies both signatures (ES256 or RS256, viaring). The outer signature proves the request comes from the current account holder; the inner one proves possession of the new key, so a key the requester does not control cannot be installed.- Collision check: the proxy checks that the new key does not already
belong to another account. If it does, the request is rejected with
409 Conflictplus aLocationheader naming the account that already holds the key. - Update: the account’s stored public key is replaced in a single statement.
Because accounts are keyed UNIQUE(profile, pubkey), the collision check is
scoped per profile — the same key legitimately belonging to a different account
at a different profile is not a conflict.
Client support
Many modern clients support key rollover natively. For instance, using lego:
lego --server https://acme.internal/profile/default/directory \
--email "admin@example.com" \
accounts keyrollover
(Note: Ensure you are using a recent version of the client, as older clients
may lack support for the keyChange endpoint).
Renewal Information (ARI)
acme-proxy implements the ACME Renewal Information extension (RFC 9773). ARI
allows a CA to signal to ACME clients when they should ideally renew their
certificates. This prevents the “stampede” effect where thousands of
certificates expire simultaneously, and allows the CA to orchestrate graceful
mass-revocations by shortening the renewal window.
GET /renewalInfo/{certID}
This is an unauthenticated endpoint clients use to ask “when should I renew this certificate?”. The response provides a time window (Start and End).
The window is resolved in this order:
- Revoked certificates win. If the certificate is known to be revoked locally (looked up by its serial), the answer is a window entirely in the past, prompting a compliant client to renew immediately. This is checked before the signer backend is consulted, so a locally revoked certificate is never talked out of renewing by an upstream CA that has not noticed yet.
- The backend’s own opinion, if it has one. The
relaybackend relays the upstream CA’s answer verbatim, including RFC 9773 §4.2’s optionalexplanationURL— so if Let’s Encrypt signals early renewal, that reaches your internal clients unchanged. Thecustombackend can do the same through itsrenewal_infohook, but only whensigner.custom.supports_renewal_info = true; otherwise the hook is never invoked. - A local estimate, when the backend has no opinion — always the case for
local_ca. The certificate’snotBeforeandnotAfterare read by parsing the certificate itself, and the suggested window runs from ⅔ to ¾ of the validity period. For a 90-day certificate that is roughly day 60 to day 67, leaving a comfortable margin before expiry rather than running up to it.
Spreading clients across that window is what prevents the “stampede” effect where a whole fleet renews at the same moment.
The replaces field
During a newOrder request, an ARI-aware client can include a replaces field
containing the certID of the certificate it intends to replace.
acme-proxy validates it against all three of RFC 9773 §5’s correspondence
rules. The certID is decoded into an Authority Key Identifier and a serial
number, the predecessor order is looked up, and then:
- The AKI must match the predecessor’s issuer — the serial alone is not enough, since serials are only unique per issuer. (A certificate stored before this check existed may have no recorded AKI; that case falls back to matching on serial alone, for backwards compatibility.)
- The predecessor must belong to the requesting account. You cannot claim to replace someone else’s certificate.
- The new order must share at least one identifier with the predecessor. A renewal that covers none of the same names is not a replacement.
Failing any of them is malformed; an unknown certID likewise.
Concurrency guard: a client cannot have two orders replacing the same
predecessor. A second newOrder naming a replaces value already claimed gets
409 alreadyReplaced. An order that later becomes invalid releases its claim,
so a failed attempt does not permanently block a retry.
The accepted replaces value is stored and reflected back on the 201 response
and on every subsequent poll of the order.
Reloading the configuration
acme-proxy reloads its configuration on SIGHUP, without restarting:
sudo systemctl reload acme-proxy
# or, directly:
kill -HUP "$(pidof acme-proxy)"
Nothing is dropped. Connections stay open, in-flight ACME orders keep their state, and queued background work keeps its place. What changes is what the server does with the next request — and, if you moved a listener, on which port it answers it.
A reload is a rebuild and a swap, not a patch. Both routers, every profile’s filter chain and challenge registry, the notification backends, the job registry and both TLS acceptors are constructed fresh from the file on disk — and only once every one of them has succeeded is any of it published.
All or nothing
A reload either applies completely or changes nothing at all.
Three things stop one, and each says so in the log:
| Log event | What happened |
|---|---|
server_config_reload_refused | A key that cannot change while the process runs did change. |
server_config_reload_failed | The file did not load, or what it asks for could not be built. |
server_config_reloaded | It applied. |
In the first two cases the server carries on with exactly the configuration it already had. That matters more than it sounds: a reload that applied the half it understood would leave a running server that no file on disk describes, and “what is this thing actually running?” would stop having an answer.
Every success carries a generation — 1 for the configuration the process
started with, and one higher for each reload that landed. It is the quickest
way to answer “did my SIGHUP take?”:
journalctl -u acme-proxy | grep server_config_reloaded
What a reload cannot change
One key. A refusal names it, what the server is running, and what the file now says:
`database.url` cannot be changed while the server is running
(running with `sqlite://sqlite.db`, the file now says `sqlite://other.db`):
restart to apply it
| Key | Why a restart is needed |
|---|---|
database.url | The connection pool is open, and the accounts and orders this CA has issued against it do not follow it elsewhere. A different database is a different CA. |
Everything else reloads. That was not always true, and the four keys that came
off this table most recently are worth knowing about because they are the ones
an operator is most likely to remember as frozen: the set of enabled profiles,
each profile’s [signer] section, dns.resolver and [proxy]. See
Adding and removing endpoints below.
What a reload does change
Everything else, including the things operators reach for most:
- Access policy —
[filter]rules and checks, and the[ipam]inventory they consult. Tightening a rule takes effect on the next request. - The set of endpoints, and each one’s
[signer]— see below. - Egress —
dns.resolverand[proxy]. Every outbound client is rebuilt, including the signer backends, which used to be the reason these two were frozen. - Notification backends — a new webhook, a changed template directory, a
different
eventslist. - Challenge settings, order and account policy,
[meta],[eab]. - TLS certificates —
server.tls.cert_pathandkey_path(and the admin listener’s own pair) are re-read, so a renewed certificate is served to the next connection while established ones are undisturbed. - The listeners themselves — see below.
- The panel’s templates —
admin.template_diris recompiled. A template that does not parse fails the reload, so a mistake never reaches a browser. - Retention —
audit.retention_daysandjobs.retention_days. A sweep already scheduled keeps its current time and picks the new cutoff up on its next run. - Background work — every
[jobs]key, so slowing a retry storm, widening a lease or raising concurrency mid-incident costs nothing. See below. - Logging — every
[logging]key, so raising the level or switching to JSON mid-incident costs nothing. The swap is the first thing a reload publishes, so theserver_config_reloadedline that confirms it is already under the new settings. One caveat:RUST_LOGstill outrankslogging.filter, exactly as it does at startup, so with it set an edited filter changes nothing — the server says so withserver_logging_filter_overridden.
For what each key means, see the configuration reference.
Adding and removing endpoints
Add a [profiles.<name>] section and signal, and that endpoint starts serving:
server_config_reloaded generation=2 profiles=["le", "staging"]
profile_mounted profile=staging directory=https://ca.example/profile/staging/directory
Remove it and signal again, and it stops. Endpoints that were already running are not disturbed either way — no connection is dropped and no in-flight order loses its state, which is the whole reason this is worth doing without a restart.
Three things to know.
profile_mounted fires only for an endpoint that was not there before. It
is a lifecycle event, delivered to whichever [notify] backends are configured,
so re-firing it for every endpoint on every SIGHUP would make the notification
surface noisiest in exactly the config-managed deployments that would least want
it.
Unmounting keeps the data. The accounts and orders belonging to that
endpoint stay in the database and come back exactly as they were if you mount it
again — the profile name is what they are keyed on. Setting enabled = false is
the same thing as removing the section.
Drain a relay endpoint before removing it. Unmounting the last profile a
relaying backend serves takes its job handler with it, so any issuance still
waiting on the upstream has nothing left to finish it. Those rows sit in the
queue until the endpoint is mounted again, and the orders behind them expire.
Nothing else is affected, and no other backend has background work to lose.
Editing a signer
A profile’s [signer] section reloads, including the one an endpoint is
actively issuing with. The obvious worry — that a rebuilt local CA would forget
what it had revoked — does not arise: revocations and the CRL live in the
database, which the running CA and its replacement both read, so a revocation
that lands during the reload is not lost either, and GET /crl answers
identically across the swap. A relay serving http-01 keeps its published key
authorizations in the database the same way, so an upstream CA fetching one
mid-reload still gets it.
A backend whose section did not move is not rebuilt at all. That matters most
with key_source = "pkcs11": an ordinary reload does not log in to the token
again.
Two edges:
- Changing where the CA material lives is a new CA. Point
crl_pathat a different file and the endpoint starts from that file’s revocation history, not the old one’s. That is the intended reading of the key, but it is worth saying, because the certificates already issued do not move with it. [dns]and[proxy]rebuild every backend. They are not[signer]keys, but an outbound client caches them, so a change to either has to reach the signers to mean anything. Nothing is lost — the same handover applies.
Moving a listener
All seven keys that decide where a socket is, or whether there is one, reload:
server.bind_address, server.tls.enabled, admin.enabled,
admin.bind_address, admin.tls.enabled, metrics.enabled and
metrics.bind_address. A reload that moved one says which:
server_config_reloaded generation=2 listeners_rebound=["acme"]
Three things are worth knowing before you use it.
A bad address refuses the reload; it does not take the socket down. Every
new socket is bound before anything is published, so a port already in use, a
name that does not resolve or a privileged port you no longer have the
capability for is a server_config_reload_failed — with the listener that is
already running still answering on the address it always had. Fix the file and
signal again.
Established connections are not disturbed. A rebind replaces the socket new connections arrive on; a request already in flight finishes, and a keep-alive connection opened before the move stays usable until its client closes it. The old socket stops accepting immediately, so nothing new arrives there.
Turning TLS on or off does not move the socket at all. The mode is decided
per connection, like the certificate: the next client to connect speaks the new
protocol on the same port, and listeners_rebound stays empty. Remember to move
server.base_url with it, or every signed request fails RFC 8555 §6.4’s URL
check — the server warns tls_base_url_mismatch when it can see the two
disagree.
One caveat on the panel. Switching admin.enabled off releases the socket and
empties its router, so nothing answers on it — but it does not sign anybody out:
sessions live in the database and are waiting when you switch it back on. Use
acme-proxy admin session revoke --all if that is what you meant. Switching it
off and on again does clear the login-attempt lockout, since the limiter goes
with the panel.
Retuning the job runner
All seven [jobs] keys reload. These are the knobs you reach for while
something is going wrong — an upstream CA rate-limiting you, a backlog draining
too slowly — so a restart to apply them would have dropped exactly the in-flight
orders you were trying to save.
The runner does not restart; it picks the new values up on its next pass and says so:
job_runner_retuned poll_interval_ms=250 lease_seconds=120 max_concurrent=16
That line is the confirmation worth grepping for. server_config_reloaded means
a generation was published; this means the runner is actually running under it.
It is only emitted when the pacing really moved, so reloads that touch other
sections stay quiet.
Each key lands at its own grain, and the differences are all in the same direction — nothing already in flight is disturbed:
poll_interval_mstakes effect immediately, without waiting out the old interval first.max_concurrentwidens at once when raised. Lowered, it takes back the slots that are free and reaches the new figure as running jobs finish; no job is cancelled to get there sooner.lease_seconds,retry_base_secondsandretry_max_secondsapply to the next job claimed. One already running keeps the budget and the backoff it started under.max_attemptsis frozen onto each job when it is queued, so a change applies to work queued from then on. Raising it is not a way to rescue a backlog that is about to give up — those rows keep the budget they were queued with.retention_daysrebuilds the sweep, including registering it when it goes from0to a real value.
Systemd
Add ExecReload to the unit so systemctl reload works:
[Service]
ExecReload=/bin/kill -HUP $MAINPID
See Deployment for the rest of the unit.
Two things to expect
Neither is a problem, but both look odd if you are not expecting them.
Startup warnings repeat. A reload re-emits the advisories that describe the
configuration it just applied — challenge_validation_bypassed,
filter_disabled, tls_disabled and the rest. That is deliberate: they are
written to stay visible for as long as the condition holds.
profile_mounted does not repeat. It is a lifecycle notification meaning
“this endpoint came up”, delivered to whichever [notify] backends are
configured. Firing it on every reload would make the notification surface
noisiest in exactly the config-managed deployments that would least want it.
Reloads a restart still handles better
Two edges, both brief and both identical to what a restart does:
- A request that was already in flight when the reload landed finishes under the old configuration. That includes the notification it may queue.
- For the moment it takes in-flight requests to drain, the old and new
admission limits are both in force, so concurrency can briefly reach twice
server.max_concurrent_requests.
If a change is important enough that neither is acceptable, restart.
Admin CLI
acme-proxy embeds an administrative command-line interface in the same binary,
so a full deployment never needs a separate tool to manage its state: accounts,
orders, the audit trail, the job queue, nonces, EAB credentials, the upstream
account, the web admin’s operators and sessions, and the database itself
(migrate, init, transfer). profile and filter read the configuration
back as the server would build it.
Invoking it
serve is the default subcommand, so a bare acme-proxy starts the server.
You reach the admin CLI by naming a subcommand:
acme-proxy account list
Every command reads the same configuration as the server — config.toml in the
working directory, or ACME_PROXY_CONFIG, plus ACME_PROXY_* environment
overrides — so it must be run where the configuration points at the same
database. There is no --config flag.
Commands operate directly on the database, SQLite or PostgreSQL. Running them against a live server is safe (SQLite runs in WAL mode; PostgreSQL is built for concurrent writers), but they act immediately and are not transactional across the server’s own in-flight requests.
The schema is applied explicitly
Opening the database does not migrate it. acme-proxy migrate applies any
migrations that have not run yet, and a serve running the
worker role does the same at startup — so a default single-process acme-proxy serve against a fresh database still just works.
Everything else checks the schema and refuses by name:
database error: the schema is 3 migration(s) behind; run `acme-proxy migrate` first
This used to be a side effect of opening the database, which made every
subcommand an upgrade step — acme-proxy audit list from a newer binary
silently rewrote the schema — and let two processes starting together race the
migration runner, SQLite offering sqlx no lock to serialise them.
acme-proxy migrate is idempotent and safe to run repeatedly; it prints how
many migrations it applied, or says the schema is already up to
date.
acme-proxy init migrates and then generates whatever first-run material
the configuration calls for — the local CA key and certificate, an upstream
account for a relay profile, a self-signed TLS certificate. It is the one
command that creates key material, so a split deployment runs it once, as the
uid that should own those files, before starting anything.
Moving between backends
acme-proxy transfer --to <url> copies every row of the configured
database into another one. The scheme of each URL picks its backend, so this is
how a SQLite deployment becomes a PostgreSQL one — and the reverse is the same
command with the two swapped.
$ acme-proxy transfer --to postgres://acme@db.internal/acme
Copy 14203 row(s) from sqlite://acme.db to postgres://acme:***@db.internal/acme?
The source server must be stopped, or the copy is a torn snapshot.
Continue? [y/N] y
accounts 412
orders 9881
audit_log 3910
…
Copied 14203 row(s) into 15 table(s).
Stop the server first. Nothing can check it: a worker that issues a certificate while the copy is running writes rows the copy has already walked past, and the result looks exactly like a good one. That is the only part of this an operator has to get right unaided.
The target must already exist, be migrated and be empty. Create the
database, run acme-proxy migrate against it (this command will not — applying
a schema belongs to migrate, init and the worker role, and nothing else),
then transfer. A target that already holds rows is refused by name, listing
them: a transfer is a copy, not a merge, and there is no flag that makes it
one.
Row ids, certificate serials and audit ids all survive, because all three are
things something outside the database still refers to — a kid a client
stored, a serial the CRL carries, an id an operator typed. A certificate
issued before the move is revocable after it.
What does not travel is the schema’s own history: each backend keeps its own
migration set and checksums. And --json answers
{"tables": [{"table", "rows"}], "total"} — a report of a copy, not a listing,
so it has its own shape rather than the paged envelope.
Roles
serve takes --role, a comma-separated list of acme, admin and
worker. With no --role it runs all three in one process, which is the
default and what every deployment before the flag existed did.
| Role | Does |
|---|---|
acme | Serves ACME to certificate clients: the ACME listener and the root router. Enqueues work, runs none. |
admin | Serves the web admin, /ui and /api. Enqueues work, runs none. |
worker | Drains the job queue, and owns the schema and the first-run material. |
Splitting them puts the process that parses untrusted JWS and CSRs, the process that holds operator sessions, and the process that reaches out to client-chosen hosts in three different places, each able to run under its own uid. See Deployment.
An unknown role is refused with usage before anything is read:
error: invalid value 'wroker' for '--role <ROLES>': unknown role `wroker`
(expected one of: acme, admin, worker)
A process running no worker logs the advisory server_role_no_worker at
startup: nothing there drains the queue, so challenge validation, notifications
and the periodic sweeps all wait for a process that does.
Global flags
-y, --yes — skip the interactive “Are you sure?” prompt on destructive
commands. It is a global flag, so it may be given anywhere on the line. account delete, order delete, eab delete, jobs cancel, audit cleanup, nonce cleanup, transfer, admin user delete and admin user totp reset prompt;
nothing else is gated by it.
--json — where supported, emit JSON instead of the human-readable line
format. Single-item commands print one JSON object. Every list command prints
the same envelope the admin JSON API returns, so a script does not learn one
shape for the shell and another for the API:
{ "items": [ … ], "total": 137, "limit": 50, "offset": 0 }
total is what the same filters match unpaged, which is the difference
between having read the table and having read a page of it. See
Paging. (It is not newline-delimited JSON.)
--color <auto|always|never> — when to colour the human-readable output.
Also global. The default is auto: colour when the stream is a terminal and
NO_COLOR is unset or empty, which means a piped or redirected run is plain
without your having to say so.
alwayscolours regardless of the stream and regardless ofNO_COLOR— it was typed on this command line, so it outranks both. That is what makesacme-proxy audit list --color always | less -Rwork.nevernever colours, whatever the terminal is.- An unrecognised value is refused rather than treated as
auto.
Colour is decided separately for stdout (the output) and stderr (error
messages), since the two are redirected independently. It is semantic, never
decorative: statuses (valid, pending, revoked…), audit events that name
a refusal, filter explain’s per-check verdicts, and the standing warnings such
as eab create’s “shown only this once”. Labels, timestamps and identifiers are
never coloured.
--json output never carries colour, at any setting — it is the same bytes
a script parses today. So is every human-readable line under --color never.
Note this is not the same switch as logging.ansi, which colours the server’s
log stream and is a configuration key rather than a flag; the CLI’s colour is
not configurable, on purpose, since the right answer depends on the terminal in
front of you rather than on the deployment.
--log-level <off|error|warn|info|debug|trace> — emit log records for this
run. Also global.
An admin command prints only its own output by default: no log records at
all, on either stream. That is what makes acme-proxy account list --json | jq
work — before this flag existed, a db_migration_completed record landed on
stdout ahead of the JSON on every invocation, because [logging] was installed
for every subcommand and its target defaults to stdout.
- Records go to stderr, whatever
logging.targetsays. stdout is the answer; a diagnostic does not belong in it. - The level covers
acme-proxyalone, so--log-level debugdoes not also turn onsqlxandhyper. SetRUST_LOGfor a directive that reaches further — a non-emptyRUST_LOGturns records on by itself, with no flag. --log-level offis silence stated explicitly, which is what a script wants when the environment it runs in may carry aRUST_LOG.
It is worth reaching for on filter show, which builds the policy exactly as
startup does and so is the cheapest pre-restart check: the refusals are
printed either way, but the advisories are log records, so
filter show --log-level warn is how you see filter_disabled and
filter_check_unused before a restart rather than after it.
On serve the flag outranks both RUST_LOG and logging.filter — a flag was
typed where the other two are ambient — and [logging]’s other five keys still
decide the format, the target and the rest. It survives a SIGHUP, so an
operator who started the server at --log-level debug keeps it across a reload;
an edit to logging.filter is then a no-op the server warns about
(server_logging_filter_overridden).
--version — print the build’s own version and exit. It is the first thing
a bug report asks for, and the answer a checkout cannot give on a host where the
binary was copied in. --help is its counterpart and works at every level:
acme-proxy audit --help lists that group’s subcommands.
Exit codes
An admin command exits with one of these. A script can branch on the code
without parsing stderr; the deciding question between 1 and 3 is whether
re-running the identical command is worth it.
| Code | Meaning | Examples |
|---|---|---|
0 | Success — the command did what was asked. | |
1 | The host could not carry out the request. Worth retrying, or fixing the host and retrying. | A database that will not open, a signer or CA error, an unreadable --password-file, an unreachable upstream, a broken [dns]/[proxy] section, invalid configuration. |
2 | The command line itself was rejected. Emitted by the argument parser. | An unknown flag or subcommand, a missing argument. |
3 | The request cannot be satisfied as written. Re-running the identical command will not help. | No object with that id (no such order …); an object in the wrong state (… is already revoked, only ready or failed jobs can be cancelled); an unknown --status/--event/--outcome value, or --role on admin user create; contradictory flags (--hide-superseded without --expiring-in); nothing supplied on stdin where a password or an EAB key was asked for. |
serve exits 1 for any startup failure and otherwise runs until it is
signalled.
Shell completions
acme-proxy completions <shell> prints a completion script on stdout, for
bash, elvish, fish, powershell or zsh. It is generated from the same
command tree clap parses, so it covers every subcommand and flag, four levels
deep — acme-proxy admin user totp completes to status, reset and
recovery-codes. Flag values complete only where the flag has a fixed set
clap knows about, which today is --color, --log-level and completions’
own <shell>; --status and --outcome take a
string the command refuses by name, so a shell has nothing to offer for them.
The command reads neither the configuration nor the database, so it works anywhere, including in a shell startup file and before a deployment exists.
# bash — system-wide, or ~/.local/share/bash-completion/completions/acme-proxy
acme-proxy completions bash | sudo tee /etc/bash_completion.d/acme-proxy
# zsh — any directory on $fpath; the file must be named _acme-proxy
acme-proxy completions zsh > ~/.zfunc/_acme-proxy
# fish
acme-proxy completions fish > ~/.config/fish/completions/acme-proxy.fish
Regenerate after upgrading: before 1.0.0 the CLI is not frozen, so a script kept from an older binary can go on offering a subcommand that no longer exists. The policy is stated in the changelog.
One limitation is the generator’s rather than this CLI’s: the fish script
stops completing at three levels, so acme-proxy admin user totp offers
nothing past totp. The other four shells complete the whole tree.
Man page
acme-proxy man prints the roff source of acme-proxy.1 on stdout. Like
completions, it reads nothing and is generated from the command tree:
acme-proxy man | sudo tee /usr/share/man/man1/acme-proxy.1 > /dev/null
man acme-proxy
It is one page for the top-level command — the options, every subcommand with
its one-line purpose, the environment variables, the configuration file, and a
pointer back to this book. The per-flag detail of each subcommand lives in the
tables below rather than in the page, which is why SEE ALSO names the book.
Read it without installing anything with acme-proxy man | man -l -.
Account management
| Command | Flags |
|---|---|
account list | --profile <name>, --eab-kid <kid>, --limit <n>, --offset <n>, --json |
account show <id> | --json |
account update-contact <id> | --contact <uri> (repeatable) |
account deactivate <id> | — |
account delete <id> | (prompts) |
account list --profilerestricts the listing to one ACME endpoint. Without it, accounts from every profile are listed — the admin CLI is deliberately unscoped by default, unlike the request path, which always scopes by profile. The listing is newest first and paged; see Paging below.account list --eab-kidlists the accounts one EAB credential registered — what to look at beforeeab delete.account listshows, per account, the address its key was last seen from and that address’s reverse name (ip (ptr), the address alone when no name resolved,-when neither was recorded).account showprints one field per line and adds where the account was registered from. Nothing in the server ever compares against these — pinning an identity to an address breaks CGNAT and mobile — and none of them reaches an ACME object.account deactivateprevents the account from making any further requests. It is the operator-side equivalent of a client deactivating itself.account deletecascades: every order, authorization and challenge belonging to the account is destroyed with it. The prompt names what will go.account deleteandorder deleteare refused while a live certificate would go with them — one issued, not revoked, and not known to have expired. An order row is the only record of its certificate: without it,revokeCertandorder revokecannot find the certificate, and it drops out of the expiry digest and renewal information. Revoke it first (order revoke), or wait for it to expire. A certificate whose expiry was never recorded counts as live. There is no flag to override this.
Order management
| Command | Flags |
|---|---|
order list | --profile <name>, --account-id <id>, --status <status>, --identifier <name>, --identifier-contains <text>, --cert-serial <hex>, --expiring-in <days>, --hide-superseded, --limit <n>, --offset <n>, --json |
order show <id> | --json |
order chain <id> | — |
order delete <id> | (prompts) |
order revoke <id> | --reason <n>, --wait <seconds> (default 30) |
-
--statusis one ofpending,ready,processing,validandinvalid, and is refused by name otherwise. -
--identifier <name>finds the orders that name that identifier exactly (case-insensitive): the answer to “which order coversweb.corp.example.com”. It is an exact match on purpose —--identifier example.comwill not surfaceevil-example.com, which is the wrong thing to hand somebody hunting a misissuance.--identifier-contains <text>is the substring form for when only a fragment of the name is remembered; the two are mutually exclusive.A wildcard order stores the wildcard form, so
--identifiermatches*.example.comand nothost.example.com— “which order named this?” is not “which certificate covers this?”, and only the first is a question an exact match can answer.--identifier-contains example.comspans both. -
--cert-serial <hex>finds the order whose issued certificate carries that serial — the value an abuse report hands you, and the same oneaudit list --cert-serialfilters on. Case and separators do not matter: whatopenssl x509 -serialprints (upper case) and what a report quotes (often colon-separated) are both folded to the form the column holds. -
order list --expiring-in <days>asks a different question over a different query: the certificates this CA issued that reach their notAfter inside the window, soonest first, each annotated with whatever has already replaced it. It is the same listing the[notify.expiry]digest mails and the panel shows at/ui/expiring, so the three cannot come to disagree about what “expiring” or “already replaced” means. Add--hide-supersededto drop the rows that have a successor and leave only the ones to act on. Under--jsonthe envelope carries two more members, asGET /api/expiringdoes:hidden, the rows--hide-supersededdropped from this page, anddays, the window asked for. -
--status,--account-id,--identifier,--identifier-containsand--cert-serialare refused with--expiring-in, by name. The expiry listing is issued, unrevoked certificates by definition, has no account predicate and is ordered by expiry, so any of them would silently mean something other than it does elsewhere — the rule--statusandaudit list --eventalready follow. -
order showprints one field per line, omitting every field that was not recorded rather than rendering it empty — the shapeaudit showandaccount showhave. It covers the certificate’s serial and the leaf’s ownnotAfterbeside the requestednotAfterthe client asked for, and the revocation timestamp and reason, which are deliberately absent from the ACME JSON a client sees — revocation state is admin-visible only. -
order showandorder show --jsondescribe the same order, field for field, with one deliberate exception: the issued chain.--jsoncarries it ascertificatePemand the web panel offers it as a download; the text rendering does not, a command run to get one’s bearings being the wrong place for several kilobytes of PEM. The three URL members (authorizations,finalizeand the ACMEcertificateURL, which is reachable only by signed POST-as-GET) are likewise--json’s alone — the indented authorization tree is what a terminal reads instead. -
order chain <id>prints that chain and nothing else, so it pipes:$ acme-proxy order chain 0198f3b1-... > web.example.com.pemIt is the terminal’s spelling of the panel’s
GET /ui/orders/{id}/chain.pemdownload, and it keeps that route’s rule: an order that never reached issuance is an error, not an empty file, because zero bytes named.pemread as a broken certificate rather than an absent one. There is no--json— the PEM is the output. -
order revokeis the operator-side equivalent ofPOST /revokeCert, for an out-of-band compromise report a client cannot or will not act on. It never loads a CA key or contacts an upstream itself; the running server does the signing, as Revocation & CRL describes. It is not confirm-gated, because revocation only ever tightens trust.--reasonis the RFC 5280 reason code:0–6or8–10, since7is unused; any other value is refused. Omitted, the revocation carries no reason. For arelayorcustomprofile the revocation is queued for the running server, and--waitis how many seconds the command waits for it to land before returning (default 30);--wait 0queues it and returns at once. A local CA’s revocation is recorded immediately and does not wait.
Job queue
The background queue (relayed issuance, notification delivery, the periodic
sweeps) is what keeps an order processing through a transient upstream failure
instead of failing it. When an order is stuck, this is where to look.
| Command | Flags |
|---|---|
jobs list | --kind <k>, --status <s>, --limit <n>, --offset <n>, --json |
jobs show <id> | --json |
jobs cancel <id> | (prompts) |
jobs run-now <id> | — |
-
jobs listis paged like the other listings, newest first.--statusis one ofready,running,done,failed,cancelledand is refused by name — passed to SQL an unknown value answers “no rows”, which reads as “nothing is in that state”.--kindis not refused: a job kind is an open set, so a typo simply matches nothing. -
jobs showprints one field per line, and for a relay issuance job (kind = signer_relay_issue) it appends the upstream order it drives — the upstream URLs and the upstream’s own error text. The stored CSR is never shown. -
jobs cancelretires a job (status = cancelled) and is confirm-gated. Two things it will not do:- A
runningjob is refused — a runner owns it, and its lease will expire or it will settle. Wait it out, then cancel the resultingready/failedrow. - Cancelling a periodic sweep job (
nonce_sweep,audit_sweep,order_sweep, …) stops that sweep until the server restarts. The prompt says so.
Cancelling an in-flight
signer_relay_issuejob also abandons the ACME order: the local order is markedinvalid(so the client stops polling), the upstream mapping is abandoned (so a restart does not resume it), and acertificate_issue_failedaudit row is written naming you. This is the operator-side way to stop a relayed issuance that will never complete. - A
-
jobs run-nowmakes a job eligible immediately — it is picked up withinjobs.poll_interval_ms(it does not wake the runner). On areadyjob it just pullsrun_atforward; on afailedjob it grants exactly one more attempt (attemptsis set tomax_attempts - 1), not a fresh budget, because a full reset is what turns a permanently failing job into an infinite retry loop. It is not confirm-gated.
The web admin has the same surface at /ui/jobs (GET /api/jobs), with cancel
and run-now behind an operator-or-higher session.
Audit trail
| Command | Flags |
|---|---|
audit list | --profile <name>, --account-id <id>, --order-id <id>, --cert-serial <hex>, --event <e>, --outcome success|failure, --since-days <n>, --limit <n>, --offset <n>, --json |
audit show <id> | --json |
audit cleanup | --older-than <days> (prompts) |
audit listis paged like every listing; see Paging.- An unknown
--eventor--outcomeis refused by name, listing the values this build knows. Passed through to SQL it would answer “no rows”, which reads exactly like “nothing happened”. audit showprints one field per line, omitting every field that was not recorded rather than rendering it empty.audit cleanupis the only command in this binary that destroys audit history, so it is confirm-gated and its prompt names the row count.audit.retention_daysruns the same sweep daily.
The web admin can read this trail but not prune it — see Audit Trail.
Paging
Every listing in this binary is paged. account list, order list (both
of its queries), audit list, jobs list, eab list, upstream order list,
admin user list and admin session list all take --limit <n> and
--offset <n>, defaulting to 50 rows. orders and audit_log each grow a
row per issuance for the life of the deployment, so there is deliberately no
“everything” spelling and --limit 0 is not a way around it: on a year-old CA
that is a terminal full of scrollback and a table loaded into memory. A
nonsense window is corrected rather than refused — a --limit 0 becomes one
row, a negative --offset becomes zero.
eab list, admin user list and admin session list used to answer a bare
JSON array with no window, on the argument
that an operator mints those rows by hand a few at a time. That was true of how
the tables fill and said nothing about how long they have been filling — and it
made a script learn one shape for the shell and another for /api.
Every paged listing ends with a count, always and not only when the page is short:
$ acme-proxy order list --limit 2
...
2 of 1877 row(s).
“42 of 1877” is the difference between having read the table and having read a
page of it. Page with --offset; the listings are ordered newest first,
tie-broken on the row id, so a row cannot swap between pages and go unseen.
admin user list is the one exception, and is oldest first. The bootstrap
operator — created before there was a panel to sign in to — is precisely the
row whose position should not move as colleagues are added.
order list --expiring-in adds a third number when --hide-superseded drops
rows, because supersession is decided per row and cannot become part of the
query — so the total counts the window, not the rows printed under it:
$ acme-proxy order list --expiring-in 30 --hide-superseded --limit 20
...
6 of 8 row(s), 2 superseded hidden.
Under --json every one of them answers the same envelope the admin API
returns, member for member (order list --expiring-in adds hidden and days,
as its API twin does):
$ acme-proxy eab list --limit 2 --json
{"items":[…],"total":37,"limit":2,"offset":0}
The window is not clamped to admin.page_size_max. That key is a ceiling on
what an HTTP caller may ask the server for; this front end already answers to a
shell on the host.
Access policy
| Command | Flags |
|---|---|
filter show | --profile <name>, --json |
filter explain | --profile <name>, --client-ip <ip>, --identifier <name> (repeatable), --path <p> (default /newOrder), --account-id <id> (default explain), --json |
--profile may be omitted only when exactly one profile exists, the same rule
upstream show follows: [filter] is per-profile, so acting on “the policy”
without saying which one would be acting on nothing.
filter show prints the resolved policy — every check with its type and the
stages it decides at, then every rule in evaluation order with its condition
re-parenthesized. That last part is the point: an operator who wrote
a or b and c sees a or (b and c) printed back and has their answer about
precedence without reading the grammar.
Both commands build the policy rather than reading the file back, so every
startup refusal reaches you here too. filter show is therefore the cheapest
way to check a policy before restarting the server. Like every command but
completions and man, they open the configured database first, so on a fresh
host run acme-proxy migrate before them:
$ acme-proxy filter show
profile: default
default: deny (when a rule was applicable and none matched)
checks
inventory ipam identifiers only
mgmt-net allowed_ip connection and identifiers
rules (first match wins)
mgmt-bypass mgmt-net -> allow
evaluated at: connection and identifiers
inventory-owned inventory or mgmt-net -> allow
evaluated at: identifiers only
--json prints the same policy as a document, which is what a configuration
check in CI reads — and is the shape the web panel renders, so the two front
ends cannot come to describe one policy differently:
{
"profile": "le",
"active": true,
"defaultEffect": "deny",
"warning": null,
"checks": [
{ "name": "mgmt-net", "type": "allowed_ip", "stages": "connection and identifiers" }
],
"rules": [
{ "name": "inventory-owned", "when": "corp-names and (inventory or mgmt-net)",
"then": "allow", "mode": "enforce", "stages": "identifiers only" }
]
}
checks is name-sorted and rules is evaluation order. An endpoint with no
rules answers "active": false and the warning above, with defaultEffect
null rather than the configured word: filter.default is consulted only
where some rule was applicable, so with no rules it is not a fact about that
endpoint at all.
filter explain evaluates it against a hypothetical request and reports all
three stages — connection, newOrder and CSR — because every stage must
allow, and that is the thing most easily misread. For each it prints every
check’s verdict with its reason, which rule matched, and the HTTP answer that
stage would produce.
$ acme-proxy filter explain --client-ip 10.0.0.5 --identifier web.corp.example.com
Checks the evaluation never reached are listed as skipped: a short-circuited operand and a passing one look identical in the outcome, so this is the only way the output can answer “why did my inventory check not run”.
This really runs the policy.
filter explainexecutes yourcustomscripts and issues real IPAM and DNS requests, exactly as a request would, because a stubbed answer would be worse than nothing the first time it disagreed with production. It writes nothing, and it names the checks that reached outside the process at the end of its output (sideEffectsunder--json).That is also why
explainis a host-only command with no web-admin equivalent: the address and names are chosen by the caller, so behind a session it would be script execution and outbound requests driven from one stolen cookie.showis the opposite case and the panel does serve it — see Web Admin — because it reads an already-built policy and reaches nothing outside the process.
Nonce housekeeping
| Command | Flags |
|---|---|
nonce count | --json |
nonce cleanup | --ttl-seconds <n> (prompts) |
nonce cleanup deletes expired nonces. The server already sweeps them on an
interval for the life of the process, so this is mainly a debugging tool.
--ttl-seconds defaults to the configured nonce.ttl_seconds.
nonce count reports the table size and the window a nonce is fresh for — the
terminal’s spelling of GET /api/nonces, answering the same
{count, ttlSeconds} under --json:
$ acme-proxy nonce count
1284 nonce(s), ttl 300s.
The two numbers are only meaningful together. The count should sit near the request rate times the TTL; one far above that says the reaper is not running. Values are never listed, on either surface: a nonce is a bearer credential until it is consumed, so a listing would put live ones on a screen.
Profiles
| Command | Flags |
|---|---|
profile list | --json |
The ACME endpoints this configuration mounts, name-sorted, each with the two facts that decide whether it is safe to expose and then its directory URL — last, because it is the only field here with no bounded width, and any column after it would be ragged:
$ acme-proxy profile list
internal challenges=bypassed eab=off https://ca.example.com/profile/internal/directory
le challenges=validated eab=on https://ca.example.com/profile/le/directory
challenges=bypassed is painted as a warning, because it is one: an endpoint
that marks a challenge valid without checking anything is an open CA wherever
[filter] is empty. A profile parked with enabled = false is absent, which is
the honest answer to “what does this mount”.
Like filter show, this builds the answer the way startup does rather than
reading a file back, so a configuration serve would refuse is refused here
too — which makes it a pre-restart check as well as a listing. The panel’s
GET /api/profiles and /ui/profiles render the identical document from the
mounted profiles instead, so between an edit and its SIGHUP the two
legitimately disagree, and only this one can be pointed at a configuration the
server would not start on.
Upstream account management
Only relevant with signer.backend = "relay".
| Command | Flags |
|---|---|
upstream show | --profile <name>, --json |
upstream register | --profile <name>, --eab-kid <kid>, --eab-hmac-key-file <path> |
upstream order list | --profile <name>, --status <s>, --limit <n>, --offset <n>, --json |
upstream order show <local-order-id> | --json |
--profile is required for show/register whenever the configuration
defines more than one profile. [signer] is a per-profile section, so acting
on “the upstream” without saying which one would be acting on nothing. It may be
omitted only when exactly one profile exists. upstream order list takes
--profile as an ordinary filter and is cross-profile without it.
upstream register performs this proxy’s own newAccount at the upstream CA
and stores the resulting account URL beside account_key_path with a .kid
extension. It is the only time the account is registered: serve reads the
stored kid and never needs the EAB credential again.
Security note: the EAB HMAC secret is read from
--eab-hmac-key-file, or prompted on stdin. It is deliberately not accepted as a command-line argument, because argv is visible to every user on the host viaps. Omit--eab-kidentirely when the upstream requires no External Account Binding.
upstream order list|show reads the upstream_orders table — one row per local
order this proxy relayed to its upstream. --status is processing, valid
or invalid, refused by name. Each row carries the upstream order/finalize/
certificate URLs, the upstream CA’s own error text for a failed relay, and
the finalize request’s request_id; show takes the local order id and
cross-links to the relay job. It is read-only — to stop an in-flight relay, use
jobs cancel on the signer_relay_issue job (see Job queue),
which abandons the local order too. The panel twins are
GET /api/upstream-orders and /ui/upstream-orders.
External Account Binding (EAB)
| Command | Flags |
|---|---|
eab create | --label <text>, --profile <name>, --json |
eab list | --limit <n>, --offset <n>, --json |
eab show <kid> | --json |
eab revoke <kid> | — |
eab delete <kid> | --deactivate-accounts or --delete-accounts (prompts) |
eab createprints the generated HMAC secret once. It is stored but never shown again, so a lost secret is replaced, not recovered.--profilebinds the credential to one endpoint. Omitted, the credential is accepted at every profile — which is what an unscoped credential means, and is usually not what you want in a multi-tenant deployment.eab revoketakes effect immediately, with no restart: credentials are read from the live database on everynewAccount.eab deleteremoves the credential. By default its accounts are kept, but they no longer resolve to any credential, so everyeabfilter check refuses them.--deactivate-accountsalso deactivates them and keeps their orders, so their certificates can still be revoked.--delete-accountsalso deletes them and everything under them; likeaccount delete, it is refused while any of their orders holds a live certificate, and then nothing changes. The prompt names how many accounts and orders are involved.eab listis newest first and paged; see Paging. It reads the same queryGET /api/eaband/ui/eabdo, so the three cannot come to describe the credential set differently.
See External Account Binding for the protocol side.
Web admin operators and sessions
The web admin has no sign-up page: the first operator is created here. These
commands work whether or not [admin] is enabled, and whether or not the server
is running.
| Command | Flags |
|---|---|
admin user create <username> | --password-file <path>, --role admin|operator|viewer (default admin), --contact <address> |
admin user list | --limit <n>, --offset <n>, --json |
admin user show <username> | --json |
admin user passwd <username> | --password-file <path> |
admin user role <username> <admin|operator|viewer> | revokes the operator’s sessions |
admin user contact <username> | --contact <address> (omit, or pass empty, to clear) |
admin user delete <username> | confirm-gated; -y skips |
admin user disable|enable <username> | — |
admin user totp status <username> | --json |
admin user totp reset <username> | confirm-gated; -y skips |
admin user totp recovery-codes <username> | prints them once |
admin session list | --user <u>, --limit <n>, --offset <n>, --json |
admin session revoke | --user <u> (optionally --session <id>) or --all |
$ printf '%s' "$PASSWORD" | acme-proxy admin user create alice
Created admin user alice (bac6a47e-711b-4e8e-858e-417da905dab9), role admin.
- The password never goes in argv. There is no
--passwordflag andclaprejects one: argv is visible viapsand lands in shell history. Supply it on stdin or with--password-file(which strips one trailing newline). Typing it interactively works but echoes, and the command says so. - Minimum 12 characters. Stored as PBKDF2-HMAC-SHA256 at 600 000 iterations, and
not recoverable — a lost password is replaced with
admin user passwd. admin user passwdandadmin user disableboth revoke every session that user holds. A password changed because it may have leaked, that left the leaked session alive, would be a change in name only.--role/admin user roleset the operator’s privilege tier, which scopes what their web sessions may do —viewerreads only,operatoradds every CA action,adminadds managing other operators. It does not restrict the CLI: the host is the trusted plane. An unknown value is refused by name, andadmin user rolerevokes the operator’s sessions so a demotion takes effect at once. A row that predates the feature reads asadmin. See Web Admin — Roles.- Usernames are stored lowercased, so
Aliceandalicecannot become two logins that read as one in a log line. - There is deliberately no
admin user totp enrol. Enrolling happens in the panel, which shows the setup key once behindCache-Control: no-store; there is no way to do it from a terminal that does not put that key into scrollback and shell history — the same reasoning that keeps a password out ofargv. What the shell is for is the case the panel cannot serve:totp resetis how an operator who has lost their authenticator gets back in. It asks first, because it removes a security control rather than tightening one, and it takes the recovery codes and every live session with it. admin user showis the detail beside the listing, and adds the two things a row cannot carry: whether enrolment was started and never confirmed, and how many recovery codes are left. That first one matters because “enrolment pending” and “no factor” behave identically at the login prompt — an operator who believes they enrolled has no other way to find out.admin user totp statussays the same thing about the factor alone. It also shows the operator’scontactaddress and the recent login addresses that raise a “new address” notification.--contact/admin user contactset the address a web-admin operator receives security notifications at (a sign-in from an unfamiliar address, a refused second factor, a lockout, a credential change — see Web Admin — Security notifications). The address is validated as a mailbox; a bad one is refused.admin user contactwith no--contact, or an empty one, clears it. Notifications are delivered only whenadmin.enabledand[admin.notify]are configured. A change made from this CLI — a contact address,passwd,totp reset,recovery-codes— is notified like one made in the panel: the CLI queues the message and the running server’s worker sends it.admin session listshows a fingerprint of the stored token hash, never the hash itself. That fingerprint is the<id>admin session revoke --user <u> --session <id>takes to end one session rather than all of an operator’s — the granularity the panel’s Operators and Account pages already have.--sessionneeds--user, since the fingerprint only names a row within one operator’s sessions. Both listings are paged; see Paging, which also has the reasonadmin user listis the one listing ordered oldest first.
See Web Admin — Users & Sessions for the full treatment.
Web Admin
A browser- and script-facing management interface for the server: the accounts it has registered, the orders it has issued, the EAB credentials it honours, and the nonce table.
It is a second listener, on its own socket, serving no ACME — and it is off by default. A certificate authority should not grow a management surface because somebody upgraded it.
It has two faces over the same operations: HTML pages at /ui for a
browser, and a JSON API at /api for a script. Neither is built on the
other; both are thin layers over the same crates/admin/src/admin/ operations
the CLI calls.
[admin]
enabled = true
There is no sign-up page and never will be. The first operator is created from a shell on the host — see Users & Sessions.
Why a second listener
The ACME listener is public, unauthenticated, and frequently internet-facing. Everything about its defaults follows from that: admission control that sheds load, a filter chain, a small body limit.
The admin listener has the opposite shape. It defaults to loopback, requires a
session on every route but sign-in, and deliberately carries no admission
control (the availability concern here is credential brute force, which the
login rate limiter handles and admission control would not touch). It does not
inherit the profiles’ [filter] either: it has a policy of its own,
[admin.filter].
Access control on this listener is the bind address, TLS, [admin.filter], and
the session.
Putting both on one socket would have meant one set of defaults for two very different threat models.
Exposing it
The default binds loopback only:
[admin]
bind_address = "127.0.0.1:3001"
base_url = "http://localhost:3001"
The recommended way to reach it from elsewhere is an SSH tunnel, which needs no configuration change at all:
$ ssh -N -L 3001:127.0.0.1:3001 ca.example.com
Then open http://localhost:3001 locally. admin.base_url stays as it is,
because from the browser’s point of view the panel really is on localhost.
Binding to a real interface
Startup refuses a non-loopback bind_address while admin.tls.enabled is
false:
admin.bind_address `0.0.0.0:3001` is not loopback while admin.tls.enabled is
false: the session cookie is sent `Secure`, which a browser will not store over
plain HTTP on anything but localhost, so signing in would appear to succeed and
then fail silently. Set admin.tls.enabled = true, or bind 127.0.0.1 and reach
it through an SSH tunnel
This is an error rather than a warning on purpose. The session cookie is always
sent Secure — making that conditional is how a session cookie leaks — and
browsers accept a Secure cookie on http://localhost but silently refuse it
on http://192.0.2.10:3001. The operator would see “sign-in works, then I am
immediately signed out” with nothing in any log to explain it.
So a public bind means TLS:
[admin]
enabled = true
bind_address = "0.0.0.0:3001"
base_url = "https://admin.example.com:3001"
[admin.tls]
enabled = true
As with [server.tls], a self-signed certificate for the host of
admin.base_url is generated on first start when the files are missing, and the
key is created 0600. The paths default to admin.pem/admin.key, separate
from the ACME listener’s — the two answer to different names and should not
share a certificate by accident.
Give the panel its own host name
Every response carries Strict-Transport-Security: max-age=31536000; includeSubDomains. Once a browser has seen that over HTTPS it will refuse plain
HTTP for the whole host, and for every name under it, for a year — which is
the point when the panel has a host of its own, and a trap when it does not.
The scope is the host in admin.base_url, so
base_url = "https://admin.example.com:3001" commits admin.example.com and
anything below it, and nothing else. But
base_url = "https://example.com:3001" commits every subdomain of
example.com, including services that have nothing to do with this server and
may not speak HTTPS at all. HSTS is not scoped by port, so running on :3001
does not narrow it.
Give the panel a name of its own. There is no configuration key for the header: weakening it for everyone is the wrong trade when a dedicated host name costs a DNS record.
Nothing to worry about while admin.tls.enabled is false and the bind is
loopback — a browser ignores the header entirely over plain HTTP (RFC 6797
§7.2), which is why it is emitted unconditionally rather than gated.
Restricting who can reach it
Where a host firewall can guard the port, use it. Where it cannot — a container
runtime owns the host’s netfilter tables, so an nftables rule on the host is
not yours to write — [admin.filter] puts the same policy engine the ACME
listener uses in front of every admin request, /health included:
[admin]
enabled = true
bind_address = "0.0.0.0:3001"
base_url = "https://admin.example.com:3001"
[admin.tls]
enabled = true
[admin.filter]
rules = ["mgmt"]
[admin.filter.check.mgmt-net]
type = "allowed_ip"
allow = ["10.20.0.0/24", "127.0.0.1/32"]
[admin.filter.rule.mgmt]
when = "mgmt-net"
then = "allow"
A caller outside 10.20.0.0/24 gets 403 before any handler runs: the admin
API’s access_denied error under /api, the HTML error page elsewhere. Only
checks that read the request itself make sense here — allowed_ip, path,
reverse_dns, custom — and startup refuses the others by name. The check and
rule syntax is the one Filters documents; the section’s
own rules are in the
Configuration Reference.
Two things are worth knowing in a container:
- Container networking rewrites addresses. Published ports usually keep the
client’s address, but some setups (rootless Docker’s port forwarder, a
userland proxy) make every connection appear to come from the bridge
gateway. Check the
client_ipon the admin access line before relying on an allowlist. - A reverse proxy in front (Traefik, Caddy, nginx) is every peer. List its
network in
admin.filter.trusted_proxies, and the policy, the access line and the sign-in rate limiter all see the real client from itsX-Forwarded-For. The loopback-or-TLS rule above still applies to the bind: a proxy in another container reaches the panel over the network, so keep[admin.tls]on and have the proxy speak HTTPS to it.
A health check run inside the container (Docker’s HEALTHCHECK) arrives from
loopback, so allow 127.0.0.1/32 as above. An orchestrator probing from
outside needs a type = "path" check for /health and a rule allowing it.
Authentication
Sign-in exchanges a username and password for an opaque session token:
- The token is 256 random bits, and only its SHA-256 is stored. A database
read — a backup, a
.dump, an injection — yields nothing replayable. - It travels in a
__Host-acme_admin_sessioncookie,HttpOnly; Secure; SameSite=Strict; Path=/. The__Host-prefix is browser-enforced: it requires those attributes, so an edit that drops one breaks visibly. - Sessions have both an absolute lifetime (
session_ttl_seconds, never extended by activity) and an idle timeout (session_idle_timeout_seconds). - Passwords are hashed with PBKDF2-HMAC-SHA256 at 600 000 iterations. A failed
sign-in costs the same whether the username exists or not, so the endpoint
cannot be used to enumerate operators — and every failure returns the same
invalid_credentials, whatever the real cause. The log says which.
CSRF
Every request with an unsafe method must carry an X-CSRF-Token header matching
the session’s csrfToken, which sign-in and GET /api/session both return.
SameSite=Strict is set as well, but it is not sufficient on its own here,
and the reason is specific: SameSite is scoped to the registrable domain, not
the origin — different ports of the same host are same-site. A panel on
:3001 beside anything else on :8080 of the same box is exactly what it does
not cover.
There is also an origin gate: a request whose Origin does not match
admin.base_url, or whose Sec-Fetch-Site says cross-site, is refused. That
is what covers sign-in itself, which by definition carries no session token yet.
Both headers are checked only when present, so a script or curl is unaffected.
Roles
A session also carries the operator’s role
(Users → Roles). A mutating request is refused with
403 insufficient_role unless the role permits it: a CA action (revoke, delete
an account or order, EAB, nonce sweep) needs operator or admin; the
/operators/* and /ui/operators/* colleague-management routes (including a
colleague’s role and notification address) need admin; the own-account routes
(password, notification address, own sessions, own second factor, sign-out)
need only a live session. An operator whose row predates the feature is admin.
Reads are open to every role with one exception: the /operators surface needs
admin to read as well as to act, on both front ends. A tier that gated
only the writes would be a control over what a colleague can do and not over
what they can learn — and that surface carries every operator’s role and contact
address, the addresses each recently signed in from, and every live session’s
fingerprint. So GET /api/operators, /api/operators/{username} and
/api/operators/{username}/sessions, and the two /ui/operators pages, answer
403 insufficient_role to an operator or a viewer.
SameSite=Strictalso means clicking a link into the panel from another site will not carry your session. For an admin panel that is a feature, but it surprises people.
The panel
Open admin.base_url in a browser: / redirects to /ui/, and anything under
it that needs a session bounces to /ui/login.
| Page | What is on it |
|---|---|
/ui/ | What needs attention — failed jobs, certificates expiring within seven days, refusals in the last day, and any endpoint bypassing validation — each linking to its list; then totals for accounts, orders, EAB credentials and nonces, and the mounted endpoints |
/ui/accounts | Every account, filterable by profile and EAB credential, listed with the address its key was last seen from; a detail page carries both recorded addresses, the contact editor, deactivate and delete |
/ui/orders | Every order, filterable by profile, status, account, identifier (exact or substring) and certificate serial; a detail page shows the authorizations and challenges, offers the issued chain for download, and revokes or deletes |
/ui/eab | Credentials, paged; minting (the secret is shown once); a detail page counts the accounts it registered, links to them, and revokes or deletes — keeping, deactivating or deleting those accounts |
/ui/expiring | Certificates lapsing inside a window, soonest first, each annotated with whatever has already replaced it; filterable by profile and window, with a control to hide the replaced ones. Read-only |
/ui/jobs | The background queue — relayed issuance, notification delivery, the periodic sweeps — filterable by kind and status. A detail page cross-links a relay job to its upstream order and carries Cancel and Run now (see below) |
/ui/upstream-orders | One row per local order relayed to an upstream CA: the upstream URLs, the upstream’s own error text, the local order status. Read-only — stop an in-flight relay from /ui/jobs instead |
/ui/audit | The CA’s audit trail — every issuance and every refusal, filterable, with a detail page per row. Read-only: there is no route here that prunes it |
/ui/nonces | The table size, and a manual sweep |
/ui/profiles | The endpoints this process serves, and a warning for any that bypass validation |
/ui/profiles/{name}/filter | One endpoint’s resolved access policy: every check, and every rule in evaluation order with its condition re-parenthesized. Read-only |
The account pages surface the CA’s account-side traceability columns: where
newAccount was called from and the reverse name that address had at the time
(Created from), and when and where the key last authenticated a request
(Last seen, Last seen from). Each is shown only when it was recorded — a
reverse lookup that found nothing leaves the address alone, and an estate where
it can never succeed leaves both names blank rather than every row saying
“unknown”. Nothing in the server ever compares against them; they answer “who
asked for this certificate, and from where”, not “may this request proceed”.
How it is built
[htmx], vendored into the binary — no npm, no build step, no CDN. The templates are [minijinja] and can be overridden on disk without rebuilding.
Each list and detail route serves two representations of one URL: a whole
document for a normal navigation, and the bare fragment htmx is going to swap
when the request carries HX-Request. So /ui/orders?status=valid is a real,
bookmarkable URL whether you got there by clicking a filter or by typing it.
The CSRF token reaches the browser as an hx-headers attribute on <body> and
comes back as the same X-CSRF-Token header the API uses — the pages needed no
second CSRF mechanism, which is the main reason htmx was chosen over plain
forms.
Sign-in is the exception: a plain HTML form, no JavaScript, protected by the origin gate rather than a token (there is no session to have one yet). It works with scripting disabled.
The API
Mounted at /api, unversioned. Every response is application/json with
Cache-Control: no-store.
A filter left blank is the same as leaving it out: ?profile= is every profile,
not a profile whose name is the empty string. That matters because the panel’s
own controls are <select> elements inside a submitted form, which always send
their name — so every profile and any status reach the API as blanks.
| Method | Path | |
|---|---|---|
POST | /api/session | sign in — {username, password} |
GET | /api/session | who am I, and my csrfToken |
DELETE | /api/session[?all=true] | sign out (of this browser, or all) |
GET | /api/session/mfa | what a half-authenticated cookie still owes — {step} |
POST | /api/session/mfa | finish the sign-in — {code}, a TOTP or a recovery code |
GET | /api/mfa | {totpEnabled, enrolmentPending, recoveryCodesRemaining} |
POST | /api/mfa/totp | begin an enrolment — returns the secret, once |
POST | /api/mfa/totp/confirm | {code} — returns the recovery codes, once |
DELETE | /api/mfa/totp | turn it off; 409 while admin.require_mfa is on |
POST | /api/mfa/recovery-codes | reissue — returns them once |
GET | /api/accounts?profile=&eabKid=&limit=&offset= | eabKid: the accounts one credential registered |
GET | /api/accounts/{id} | |
GET | /api/accounts/{id}/orders?limit=&offset= | |
PATCH | /api/accounts/{id} | {contact: [...]} |
POST | /api/accounts/{id}/deactivate | |
DELETE | /api/accounts/{id} | cascades to the account’s orders; 409 live_certificates while one holds a live certificate |
GET | /api/orders?profile=&accountId=&status=&identifier=&identifierContains=&certSerial=&limit=&offset= | identifier is exact, identifierContains a substring, the two mutually exclusive; certSerial is the issued leaf’s serial |
GET | /api/orders/{id} | order + authorizations + challenges, plus certificatePem once issued |
POST | /api/orders/{id}/revoke | {reason} optional |
DELETE | /api/orders/{id} | 409 live_certificates while its certificate is live |
GET | /api/eab?limit=&offset= | never shows a secret |
POST | /api/eab | {label, profile} — returns the secret, once |
GET | /api/eab/{kid} | |
POST | /api/eab/{kid}/revoke | the row survives, moved to revoked |
DELETE | /api/eab/{kid}?accounts=keep|deactivate|delete | the row goes; 409 live_certificates for delete while an account holds a live certificate |
GET | /api/expiring?profile=&days=&superseded=&limit=&offset= | read-only; superseded=hide drops the replaced rows |
GET | /api/audit?profile=&accountId=&orderId=&certSerial=&event=&outcome=&limit=&offset= | read-only |
GET | /api/audit/{id} | one row |
GET | /api/nonces | {count, ttlSeconds}, the shape nonce count --json prints |
POST | /api/nonces/cleanup | {ttlSeconds} optional |
GET | /api/profiles | the endpoints actually mounted; profile list’s document |
GET | /api/profiles/{name}/filter | one endpoint’s resolved access policy; read-only |
GET | /health | unauthenticated, no database access |
Lists return an envelope, not a bare array — total is what the same filters
match unpaged, which is what a page control needs:
{ "items": [ … ], "total": 137, "limit": 50, "offset": 0 }
limit defaults to 50 and is clamped to admin.page_size_max rather than
refused. Every listing on the CLI answers the same four members
(Admin CLI → Paging), so a script learns one shape rather than
two.
POST /api/session for an operator with a second factor answers 200 with
{"mfaRequired": true, "step": "verify"} and no user member — a
half-authenticated session must not read operator metadata. It is not a 401:
the password was right, and a script has to be able to tell those apart. Send
the code to POST /api/session/mfa, which answers the ordinary session body and
a new cookie; the pending one is dead by then.
The two …/session/mfa routes are the only mutating endpoints that do not take
X-CSRF-Token. The sign-in page they serve is a plain form with no token to
send, exactly as POST /api/session has none, and the origin check covers both.
The issued chain
An order’s detail carries certificatePem, the chain as it was issued, and the
order page renders it with a GET /ui/orders/{id}/chain.pem download
(application/pem-certificate-chain). The ACME certificate member beside it
is a URL, and one a browser cannot follow — RFC 8555 §7.4.2 serves it by
signed POST-as-GET only, so it is there for completeness rather than as a link.
certificatePem is on the detail shape only, never on a listing: a page of
fifty orders would otherwise carry fifty chains for a column no list shows. An
order that never reached issuance has neither the field nor the download, and
the route answers 404 rather than an empty file.
On a host holding the database, acme-proxy order chain <id> prints the same
bytes on stdout under the same rule — see
Admin CLI → Order management.
The card names the leaf two more ways, and both are on the listing shape as
well since each is one short string: certSerial, which is what an abuse
report quotes and what GET /api/audit?certSerial= filters on, and
certNotAfter, the certificate’s own expiry — a different date from the Not after above it, which is the §7.4 window the client requested. Both are
omitted on an order that never issued rather than sent empty.
Revocation
POST /api/orders/{id}/revoke resolves the revocation route from the
order’s own profile. Two profiles can hold two different CAs, and revoking
against the wrong one would record nothing useful and leave the real CRL
untouched. An order belonging to a profile this process no longer mounts
answers 409 profile_not_mounted rather than guessing.
The panel holds no signing backend — only the worker role does — so it revokes
the way the CLI does. A local_ca revocation is recorded at once and the worker
signs it into the CRL. A relay or custom revocation is queued for the
worker and waited on; if it is still running when the wait ends, the API
answers 202 {"status": "queued", "job": "<id>"} and the page says so. Follow
it on the Jobs page.
The job queue
/api/jobs and /ui/jobs list the background queue and expose two mutations,
both requiring an operator-or-higher session (a viewer is refused):
POST /api/jobs/{id}/cancel— retires the job (status = cancelled). Arunningjob is refused with409 job_not_cancellable; wait out its lease and cancel the resultingready/failedrow. Cancelling asigner_relay_issuejob also marks its ACME orderinvalidand abandons the upstream mapping, with acertificate_issue_failedaudit row naming the operator — the way to stop a relayed issuance that will never complete. Cancelling a periodic sweep job stops that sweep until the server restarts.POST /api/jobs/{id}/run— makes the job eligible immediately (picked up withinjobs.poll_interval_ms; the runner is not woken). On areadyjob it pullsrun_atforward; on afailedone it grants exactly one more attempt —attemptsis set tomax_attempts - 1, not reset, so repeatedly clicking it is the only way to loop a permanently failing job.
On the page, a 409 renders as a banner beside the still-present card; a 5xx
replaces the page. The list rows carry no controls — the two buttons live on
the detail card, and only when the job is ready or failed.
/api/upstream-orders and /ui/upstream-orders are read-only, like the
audit trail below but for the queue’s reason: abandoning an in-flight relay is
POST /api/jobs/{id}/cancel, not a route here. They list one row per relayed
local order — the upstream URLs, the upstream CA’s own error text, the finalize
request’s request_id — joined to the local order for its identifiers and
status, and never carry the stored CSR. Both contribute no entry to the CSRF
test table.
The audit trail is read-only here
/api/audit and /ui/audit list, filter, page and resolve a row by id, and
that is all they do. There is no route on this listener that deletes audit
history, and that is a deliberate limit rather than a missing feature: the first
thing a stolen session would do is erase what it had just done, and a trail the
watched thing can erase proves nothing. Pruning happens on the host with
acme-proxy audit cleanup, or on a schedule via audit.retention_days.
This is also why /api/audit contributes no entry to the CSRF test table — with
no mutating verb, there is nothing to protect. Every other verb on those paths
is unroutable. See Audit Trail.
The expiry list is read-only too, for a different reason
/api/expiring and /ui/expiring answer the digest’s question on demand:
what lapses soon, and has anything replaced it? They share the query, the
ordering and the supersession rule with [notify.expiry] and with order list --expiring-in, so a page and a mail never disagree about what is about to
expire.
Neither has a mutating route, and the reason is not the audit trail’s: renewal is the client’s action, driven by its own ACME flow against a key this server does not hold. There is simply nothing here for a button to do. Both therefore contribute no entry to the CSRF test table.
Two members of the answer need reading together. total counts the rows the
window matches; hidden counts the ones this page dropped as already
replaced. They are separate because supersession is computed per row rather
than in SQL, so the count beside the page cannot follow the filter down — and a
pager whose arithmetic quietly disagrees with the rows under it would be worse
than saying so. The page says it in a line above the table.
The default window is [notify.expiry] lead_days wherever the digest is on,
and 30 days where it is off: an operator who has chosen a lead time gets that
one back.
The policy is shown, never explained
/api/profiles/{name}/filter and /ui/profiles/{name}/filter answer what
acme-proxy filter show prints: the default effect, every check with its type
and the stages it decides at, and every rule in evaluation order with its
condition re-parenthesized, so an operator who wrote a or b and c reads
a or (b and c) back. It is the missing half of the warning on /ui/profiles:
an endpoint with challenge.bypass on has [filter] and nothing else between
it and its clients, and this is where that policy can be read.
filter explain has no equivalent here and is not getting one. It really
runs the policy — it executes the operator’s custom scripts and issues real
IPAM and DNS requests, against an address and a list of names the caller
chose. Behind a session that is script execution plus SSRF from one stolen
cookie, on a listener that deliberately carries no filter chain and no
admission control. show is the opposite case: it reads an already-built
policy through four accessors, runs no check, and reaches nothing outside the
process. Neither surface has a mutating verb — a policy is configuration, and
configuration is edited in config.toml and reloaded — so both contribute no
entry to the CSRF test table, for the same structural reason /api/audit does
not.
One difference from the CLI is worth knowing, because it is the useful kind.
The panel reads the live policy, the one this process is enforcing right
now; acme-proxy filter show rebuilds one from configuration, which is what
makes it the cheapest pre-restart check. [filter] reloads on SIGHUP and the
whole policy is swapped, so between an edit and its reload the two legitimately
disagree — and a configuration that would be refused is reported by the CLI
while the panel goes on serving the last good policy. Both answers are correct;
comparing them is how an operator finds out which state they are in.
An endpoint with no rules is a state, not an error: the answer carries
"active": false and says the endpoint filters nothing, rather than a 404.
That is the one policy an operator most needs to be told about.
Security notifications
Every web-admin sign-in and every second-factor change is logged, and — when
[admin.notify] is configured — the operator it happened to is also told
(ASVS V6.3.5 / V6.3.7). Two events fire:
admin_sign_in— a completed sign-in from an address not among the operator’s recent ones (admin_users.known_login_ipskeeps the last five distinct addresses; a first-ever sign-in has no baseline and is silent), a correct password followed by a refused second factor, or a per-session second-factor lockout.admin_credential_changed— the operator’s password changed, a second factor was enrolled or removed, recovery codes were regenerated, or another administrator reset this operator’s second factor (by_self = false). Also when the operator’s notification address changed (change = "contact_address") — and that one message goes to the address that was replaced, not the new one, since whoever made the change controls the new address. A first-ever address replaced nothing and tells nobody.
[admin.notify] has exactly the shape of the per-profile [notify]
section — enabled, email / webhook / custom
backends, template_dir, per-backend events — but is process-wide and built
only while [admin] is enabled (ACME_PROXY_ADMIN__NOTIFY__ENABLED). Email
delivery goes to each operator’s own address rather than to notify.email.to,
stored in admin_users.contact_email. An operator sets their own from Your
account (it asks for the current password), an admin sets a colleague’s
from the Operators page, and the host sets anybody’s with
acme-proxy admin user create --contact <address> or
admin user contact <username> --contact <address>. See
Users → Notification address.
An operator with no address on file gets no message (the event is still
logged); [admin.notify.email].to, if set, is the fallback for that case.
Changes made from the host CLI (admin user passwd, contact,
totp reset, totp recovery-codes) notify too. The CLI queues the delivery
and exits, and the running server’s job runner sends it, so a change made
while no server runs is delivered when one starts. Nothing is queued while
admin.enabled is off, since there is then no [admin.notify] to deliver
through.
Errors
Not ACME problem documents. Every Problem type in this server is a hardcoded
urn:ietf:params:acme:error:* URN, and nothing on this listener is an ACME
error:
{ "error": "not_found", "message": "no such account: acct-1" }
error is a stable snake_case code you may branch on; message is for a human
and may change. The codes:
- Request:
bad_request,invalid_status,invalid_contact,conflicting_identifier_filter,not_found,method_not_allowed. - Session and access:
session_invalid,session_expired,session_idle,invalid_credentials,csrf_failed,rate_limited,access_denied,insufficient_role,mfa_required,mfa_not_enabled. - The row’s state:
order_not_issued,already_revoked,live_certificates,last_admin,job_not_cancellable,job_not_runnable,profile_not_mounted. - Server:
signer_failed,internal.
The pages answer the same failures as HTML carrying the same code, with one
deliberate split: a refusal that is about the row’s state — 409 already_revoked, order_not_issued — comes back as a banner beside the button
you pressed, with the record still on screen, while a server problem
replaces the page. A missing session is neither: a browser gets 303 to
/ui/login, and an htmx request gets the HX-Redirect header, because a 303
is followed by fetch before htmx ever sees it and the sign-in page would be
swapped into whatever you clicked.
Driving it with curl
$ curl -sc jar -X POST http://127.0.0.1:3001/api/session \
-H 'content-type: application/json' \
-d '{"username":"alice","password":"…"}'
{"csrfToken":"…","expiresAt":"…","user":{…}}
$ curl -sb jar 'http://127.0.0.1:3001/api/orders?limit=5'
$ curl -sb jar -X POST http://127.0.0.1:3001/api/eab \
-H "x-csrf-token: $CSRF" -H 'content-type: application/json' \
-d '{"label":"team-a"}'
Security notes
- This is a second attack surface on a certificate authority. It is off by
default, binds loopback, and needs a session — but it has no admission
control, and filters nothing until
[admin.filter]names a rule. That is worth knowing rather than discovering. - Sign-in is protected by a fixed-window rate limiter
(
admin.login_max_attemptsperadmin.login_window_seconds, keyed on the client address — an IPv6 client by its /64, which one subscriber can rotate through at will). An attempt counts from the moment it starts, so a parallel burst gets no more guesses than a sequence. Over the limit, the password hash is not computed at all — 600 000 iterations is a denial-of-service lever otherwise — and the hash that does run is on the blocking pool, off the workers that serve requests. - A forwarded-for header is believed only from
admin.filter.trusted_proxies, never from the ACME listener’sfilter.trusted_proxies; honouring it from anyone else would let a caller spoof the key the rate limiter counts on. Behind a reverse proxy that is not listed there, the limiter counts the proxy. - Every response carries
Content-Security-Policy: default-src 'none'; script-src 'self'; style-src 'self'; img-src 'self' data:; connect-src 'self'; form-action 'self'; frame-ancestors 'none'; base-uri 'none'. Nounsafe-inline, nounsafe-eval— affordable because htmx is served from this origin and drives everything throughhx-*attributes rather than inline handlers. Alongside it:no-store,nosniff,X-Frame-Options: DENY,Referrer-Policy: same-originand HSTS. htmx.min.jsis a vendored third-party file, andcargo denyaudits the crate graph and cannot see it. Its version, source URL, SHA-256 and licence are recorded incrates/admin/src/webadmin/static/README.md, which is the only provenance record there is — check it when you update.- A second factor (TOTP) is available per operator, and
admin.require_mfamakes it compulsory — see Operators and sessions. It is off by default, so the loopback bind plus an SSH tunnel remains the baseline posture and not a substitute for one. - WebAuthn is not implemented.
webauthn-rs0.5 hard-depends onopenssl/openssl-sysand is MPL-2.0, neither of which this tree carries; the design does not preclude it later (another factor is another branch in the same state machine, not a change to it).
Configuration
See the Configuration Reference for every
[admin], [admin.filter] and [admin.tls] key, and Customizing the
Panel for admin.template_dir.
[htmx]: https://htmx.org [minijinja]: https://docs.rs/minijinja
Web Admin — Users & Sessions
The web admin has no sign-up page. Operators are created from a shell on the
host, with acme-proxy admin.
Bootstrapping the first operator
$ printf '%s' "$PASSWORD" | acme-proxy admin user create alice
Created admin user alice (bac6a47e-711b-4e8e-858e-417da905dab9).
That is enough to sign in — the panel itself does not need to be running, and does not need to be enabled yet.
If the panel is enabled with no operators, startup says so:
WARN event="admin_no_users" the web admin is enabled but has no operators:
create one with `acme-proxy admin user create <username>`
The password never goes in argv
There is deliberately no --password flag, and clap refuses one:
$ acme-proxy admin user create alice --password hunter2
error: unexpected argument '--password' found
argv is visible to every process on the host via ps, and shells routinely
write it to history. The same reasoning already applies to the upstream EAB
secret (acme-proxy upstream register).
Two ways in, both of which keep it out of the process table:
$ printf '%s' "$PASSWORD" | acme-proxy admin user create alice # stdin
$ acme-proxy admin user create alice --password-file /run/secrets/pw
--password-file strips a single trailing newline, so it does not matter
whether the file was written with printf '%s' or printf '%s\n'.
Typing it interactively works too, but the password will echo — there is no
rpassword here, because echo suppression needs a real TTY and that would break
the injectable-reader design the whole admin layer is testable through. The
command warns when it notices a terminal.
Password policy
Three rules, checked in this order, each of which ends the check — one refusal names one reason:
- Length. Minimum 12 characters, maximum 1024 bytes.
- It must not name this deployment — see the word list below.
- It must not be a commonly used password — see the corpus below.
There are deliberately no composition rules. “One digit, one symbol” measurably pushes people towards weaker, more guessable passwords; the two checks above refuse the passwords composition rules were reaching for, without dictating shape.
Length is counted in characters, so a 12-character passphrase in a non-Latin script is 12, not its byte count. The maximum is in bytes, and is a denial-of-service control rather than a security one: without it a sign-in could hand 600 000 iterations a multi-megabyte input.
All three run on admin user create and admin user passwd, and none of them
runs on sign-in. A password that predates a rule change must still work, and a
corpus refresh must never lock an operator out of the panel they would need to
be signed in to fix.
The context-specific word list
A password may not contain any word that names this deployment, compared
case-insensitively. Words shorter than four characters are ignored: a subject
holding CA, or a host label io, would refuse a large share of every password
anyone typed and buy nothing.
The list is derived from your own configuration, not fixed:
| Source | What is taken |
|---|---|
| Always | acme, proxy |
| The operator’s username | Split on anything that is not a letter or a digit |
server.base_url, admin.base_url | The host only, split on . and - |
[signer.local_ca.subject] — global and each profile’s own | common_name, organization, organizational_unit, state, locality, each split into words |
Each [profiles.<name>] name | The name, which appears in every kid and order URL that endpoint issues |
So a CA at https://ca.example.com with a CommonName of “Example Corp Issuing
CA”, managed by operator, refuses any password containing acme, proxy,
example, corp, issuing or operator — acmeproxy2026! among them.
Two limits are deliberate:
- A CA already on disk is not described here.
[signer.local_ca.subject]is read only when this server generates a CA, so an adoptedca.pemcarries a subject the configuration never saw. Add those words to the subject section if you want them barred. countryis not read, being two characters and so below the floor whatever it holds.
If no profile resolves — which is every admin user create run against a
configuration that has none yet — the rest of the list still applies. A missing
section is never a reason to refuse a password change.
The common-password corpus
A password may not be (whole-string, case-insensitively) one of 13 918 known
common passwords compiled into the binary. It is not a substring test:
a-long-enough-password contains password and is accepted.
The corpus is the top 700 000 entries of SecLists’
xato-net-10-million-passwords-1000000.txt, filtered to entries of at least
12 characters. That filter is the whole reason a compiled-in list is
affordable: password, qwerty and 123456 are refused by the length rule
before this check is reached, so carrying them would cost every deployment
bytes — including the ones running with admin.enabled = false — in exchange
for nothing. Filtering turns 8.5 MB into 195 KB.
Provenance, the exact rank cut, the size budget it was derived from and the
refresh command live in crates/admin/src/admin/corpus/README.md.
Passwords are stored as PBKDF2-HMAC-SHA256 at 600 000 iterations (OWASP’s current recommendation for the non-Argon2 case), in a self-describing format:
pbkdf2-sha256$600000$<salt-b64url>$<hash-b64url>
Self-describing so the cost can be raised, or the algorithm swapped, without a migration: a row written under older parameters is re-hashed in place on its owner’s next successful sign-in. A lost password is replaced, never recovered — nothing can read one back.
Second factor (TOTP)
Off by default and per-operator. Once enrolled, signing in is two requests: the password mints a half-authenticated session that reaches nothing but the page finishing the login, and a code from an authenticator app turns it into a real one.
stateDiagram-v2
[*] --> anonymous
anonymous --> anonymous: wrong password<br/>(counts against login_max_attempts)
anonymous --> locked_out: limiter tripped, by peer address
locked_out --> anonymous: login_window_seconds elapses
anonymous --> active: password ok, no factor enrolled<br/>and require_mfa is off
anonymous --> pending_mfa: password ok, factor enrolled
anonymous --> enrolling: password ok, no factor<br/>and require_mfa is on
pending_mfa --> pending_mfa: wrong code<br/>(counts against mfa_attempts)
pending_mfa --> [*]: mfa_attempts exceeded —<br/>the pending row is DELETED
pending_mfa --> active: correct TOTP code
pending_mfa --> active: unused recovery code
enrolling --> active: enrolment confirmed
active --> [*]: sign out, expiry, idle timeout,<br/>or a revoked session
Three things the diagram is making explicit:
- Promotion mints a new session.
pending_mfa → activeis an insert plus a delete, not anUPDATE— a new cookie and a new CSRF token — because the pending token crossed the wire before authentication had finished. - Guessing is bounded twice, and the two bounds are not redundant. The
limiter is keyed on the peer address;
mfa_attemptsis keyed on the session, because apending_mfacookie is deliberately valid from any address and one IPv6 /64 supplies 2⁶⁴ fresh addresses. enrollingis only reachable with no factor. A session that owes a code can never reach an enrolment route, or the factor would be bypassable by enrolling a new one over it.
Enrolling
From the panel, not from a shell. Sign in, click your username in the top-right corner, and press Set one up:
- The page shows a base32 setup key and an
otpauth://URI. Type the key into an authenticator app, or open the URI on the device holding it. - Type the code the app shows and press Confirm. Nothing is enabled until you do — an enrolment begun and abandoned leaves you exactly where you were.
- Ten recovery codes appear. Store them now; they are shown once.
There is no QR code, and the panel’s Content-Security-Policy is not why — it
already permits a data: image. A QR renderer would be a dependency for a
convenience rather than a capability, and every authenticator has “enter a setup
key manually”.
There is deliberately no acme-proxy admin user totp enrol. There is no way
to enrol from a terminal that does not put the setup key into scrollback and the
shell’s own history — the same reasoning that keeps a password out of argv.
The codes are HMAC-SHA-1, six digits, thirty seconds, which is RFC 6238’s
default. Not a lapse: Google Authenticator ignores the algorithm= parameter of
an otpauth:// URI and always computes SHA-1, so anything else would produce an
entry that yields wrong codes forever with no diagnosis from either side.
SHA-1’s collision attacks do not weaken HMAC-SHA1.
Recovery codes
Ten, each usable once, hashed the same one-way as a password — so a lost set is replaced, never recovered:
$ acme-proxy admin user totp recovery-codes alice
New recovery codes for alice — the previous set no longer works.
Store these now; they are not recoverable.
K7QF2-3BXTM
…
Type one into the code box at sign-in exactly as you would a six-digit code: the
server tells them apart by shape, so there is no mode to choose. Case and the
- do not matter.
When somebody is locked out
A lost phone with no recovery codes left is a shell command on the host — the same place the first operator was created:
$ acme-proxy admin user totp status alice
alice totp=enabled recovery-codes=3
$ acme-proxy admin user totp reset alice
Remove the second factor and every recovery code for alice, and revoke their
sessions? [y/N] y
Removed the second factor for alice. Their sessions were revoked; they can sign
in with a password alone until they enrol again.
reset asks first, unlike most admin commands, because it removes a security
control rather than tightening one. It takes the recovery codes and every live
session with it.
status distinguishes three states, and the middle one matters: pending means
an enrolment was started and never confirmed, which behaves exactly like off
at the sign-in prompt. An operator who believes they enrolled has no other way
to find out they did not.
Requiring it of everybody
[admin]
require_mfa = true
This governs the operator who has no factor: their next sign-in lands on the enrolment page and their session stays half-authenticated until they finish. An operator who already has one is challenged whether this is set or not.
It deliberately does not refuse a password-only sign-in. Enrolling needs a session and a session would then need a factor, so refusing would brick the panel — including the way in to fix it.
Two consequences worth knowing:
- It does not retroactively end sessions that predate it. The lever that does is
acme-proxy admin session revoke --all. - Turning the factor off is refused while it is set, since the operator would simply be made to enrol again on their next sign-in.
While it is on and somebody still has no factor, every start says so:
WARN event="admin_mfa_enrolment_pending" count=2
admin.require_mfa is on and some operators have no second factor
What it costs an attacker
A half-authenticated session lives five minutes and no longer — it is a
password that has been accepted and nothing more. Code attempts share the login
rate limiter (admin.login_max_attempts per admin.login_window_seconds, per
address), so five wrong codes also lock the password out from that address: one
address, one budget. A code accepted once cannot be replayed inside its own
thirty-second window.
The password asked for again when an operator replaces or removes a live
factor shares that same budget, and is checked before the hash rather than
after it. So a stolen session cookie cannot be used to grind the account
password — which matters, because a correct guess there would let the thief
enrol their own authenticator, end every other session and void the recovery
codes. Past the budget those routes answer 429 with Retry-After, and the
account card shows it as a banner. What the budget does not bound is an
attacker with many source addresses: the cookie is deliberately valid from
anywhere, so each address buys its own admin.login_max_attempts. Rotate the
password and run acme-proxy admin session revoke --all if you believe a
cookie has been taken.
Changing your own password
Also from the panel, not only from the host. Sign in, open your username in the top-right corner, and use the Password card: current password, then the new one.
The current password is asked for unconditionally — unlike the second-factor controls above, this runs whether or not you have TOTP enrolled. ASVS 5.0 V6.2.3 is why: a live cookie is not proof you still know the password, only that you did at sign-in. Guessing it is bounded the same way and shares the same budget as What it costs an attacker — five wrong attempts lock the address out. The new one still has to satisfy Password policy.
$ acme-proxy admin user passwd alice --password-file /run/secrets/new
Password changed for alice. Every session they held was revoked.
That command, run from the host, still ends every session — there was no request to preserve. The panel’s own change is different in exactly one way: the session that submitted it stays signed in, and every other session of that operator is revoked. Rotating a credential from inside a session you are already trusted on need not sign you out of the tab that did it.
Your notification address
The address your security notifications are delivered to — a sign-in from a new address, a change to your password or second factor. Set it on Your account under Notification address, or from the API:
$ curl -X POST https://admin.example.com/api/account/contact \
-H 'Cookie: __Host-acme_admin_session=…' -H 'X-CSRF-Token: …' \
-d '{"current_password": "…", "contact": "alice@example.com"}'
An empty or absent contact clears it. It asks for your current password
although an address is not a credential, and revokes no session: it is where
the alarms go, so a stolen cookie that could change it silently would switch
off the one signal that the cookie was stolen. For the same reason, a change
is reported to the address it replaced — the new one is controlled by whoever
made the change. Setting the address you already have does nothing and sends
nothing. An address that does not parse as a mailbox is refused
(400 invalid_contact).
Roles
Every operator has one of three roles, stored in admin_users.role and
governing what that operator’s web sessions may do. The CLI is unaffected —
a shell on the host runs every subcommand whatever the row says, the same way
it already ignores disabled for read commands.
| Role | Can do |
|---|---|
viewer | Read every page and API route except the Operators surface, which is admin-only to read as well as to act. Act on their own account only: change their password, revoke their own sessions, manage their own second factor, sign out. |
operator | Everything a viewer can, plus every CA action: revoke a certificate, deactivate or delete an ACME account, delete an order, mint or revoke an EAB credential, run a nonce sweep. |
admin | Everything an operator can, plus the Operators surface — reading it as well as acting on it: disable, enable, reset a colleague’s second factor, revoke one of their sessions. |
A web session that tries something above its role gets 403 insufficient_role
(an error document in the panel, a JSON error from the API); the request never
reaches the operation.
The panel does not offer what a role cannot reach: a lower tier is shown no
Operators nav entry, and the per-row Delete/Revoke/Deactivate controls are
hidden from a viewer. That is presentation, not authorization — the write
extractors decide either way — but a control that always answers 403 is a
worse page than no control.
Demoting the last admin is refused, since the Operators surface is
admin-only on both front ends and a deployment with none could not manage
operators from the panel at all. Promote somebody else first. Nothing is
unrecoverable: this host can always set a role back.
A username must match ^[a-z0-9._-]+$. It is a URL segment on the panel,
so a name holding /, ?, # or a space would produce links and routes that
never match. Existing operators are unaffected — the rule is checked when one
is created, never when one signs in.
An operator whose row predates this feature has role admin. The column is
NULL for them, which reads as admin, so an upgrade changes nobody’s
authority and the first operator — the only way into the panel — is never
locked out of it.
Set the role when you create the operator, or change it later:
$ acme-proxy admin user create noc --role viewer --password-file /run/secrets/pw
Created admin user noc (…), role viewer.
$ acme-proxy admin user role noc operator
Role of noc set to operator. Every session they held was revoked.
--role defaults to admin and an unknown value is refused by name.
admin user role revokes the operator’s sessions, the same as passwd and
disable — a demotion that left a live admin session alive would take effect
only when that cookie expired.
An admin can also change a colleague’s role from their page on the
Operators surface (From the panel), with the same
revocation. Not their own: that surface refuses to target the caller, which is
also why the last-admin refusal cannot be reached from the panel — an admin
changing somebody else always leaves at least themselves.
Managing operators
$ acme-proxy admin user list
alice active admin totp=on 2026-08-08T13:21:18Z 2026-08-08T15:07:17Z
1 of 1 row(s).
$ acme-proxy admin user list --json
$ acme-proxy admin user show alice
id bac6a47e-711b-4e8e-858e-417da905dab9
username alice
status active
role admin
totp enabled
recovery_codes 7
created 2026-08-08T13:21:18Z
updated 2026-08-08T15:07:17Z
last_login 2026-08-08T15:07:17Z
$ acme-proxy admin user passwd alice --password-file /run/secrets/new
Password changed for alice. Every session they held was revoked.
$ acme-proxy admin user disable alice
Disabled alice. Their sessions were revoked.
$ acme-proxy admin user enable alice
$ acme-proxy admin user delete alice # asks first; -y skips
admin user list is paged like every other listing
(Admin CLI → Paging) and is the one that is oldest first:
the bootstrap operator is the row whose position should not move as colleagues
are added. admin user show adds the two things a row cannot carry — whether
enrolment was started and never confirmed, and how many recovery codes are
left. The first matters because “pending” and “no factor” behave identically at
the login prompt.
Usernames are stored lowercased, so Alice and alice cannot become two logins
that read as one in a log line.
A password change revokes every session that user held. A password changed because it may have leaked, that left the leaked session alive, would be a change in name only. Disabling does the same.
From the panel
Also from the Operators page, not only from the host — for the operations
that do not mint a credential. The page and its actions need the admin role
(Roles); an operator or viewer session is refused it. Sign in,
open Operators, and pick a colleague:
bob active off 2026-08-08T13:21:18Z 2026-08-08T15:07:17Z
Their page shows the same status, role and second-factor summary admin user show does, their notification address and the addresses they recently signed
in from, plus their own live sessions (see Sessions below). It
offers three buttons — Disable, Reset second factor, and, per session,
Revoke — and two forms: Change role (which revokes every session they
hold) and Save address (reported to the address it replaces; empty clears
it). Every one of them asks for your own password again first, whether or
not you have a second factor. Disabling a colleague’s account or ending one of
their sessions is a much larger blast radius than anything on your own account
page, and a stolen cookie alone should not be sufficient authority for it —
which is true of an operator who has enrolled no factor exactly as it is of one
who has. (The second-factor routes on your own account page make the opposite
trade, and deliberately: a first enrolment protects nothing, and a password
prompt there would stand in front of the admin.require_mfa bootstrap.)
$ curl -X POST https://admin.example.com/api/operators/bob/disable \
-H 'Cookie: __Host-acme_admin_session=…' -H 'X-CSRF-Token: …' \
-d '{"password": "…"}'
$ curl -X POST https://admin.example.com/api/operators/bob/role \
-H 'Cookie: __Host-acme_admin_session=…' -H 'X-CSRF-Token: …' \
-d '{"password": "…", "role": "viewer"}'
$ curl -X POST https://admin.example.com/api/operators/bob/contact \
-H 'Cookie: __Host-acme_admin_session=…' -H 'X-CSRF-Token: …' \
-d '{"password": "…", "contact": "bob@example.com"}'
An unknown role is 400 naming the three that exist; an address that is not a
mailbox is 400 invalid_contact. In the panel both are a banner on the card.
Two things this surface deliberately does not do, both already settled
above: create and passwd stay on the host. Minting a credential is
where “no sign-up page” already draws the line: the Operators page can disable
an account, reset its factor, end a session, move it between roles or change
where its notifications go, but never set a password or bring an account into
being. And an operator can never target themself here — GET /ui/operators/{your own username} redirects straight to
Your account, which already owns every one of
those actions for yourself, so there is exactly one page an operator manages
their own account from.
Sessions
$ acme-proxy admin session list
01234567 bac6a47e-… active 2026-08-08T15:07:17Z expires=2026-08-09T03:07:17Z 192.0.2.1
1 of 1 row(s).
$ acme-proxy admin session list --user alice --json
$ acme-proxy admin session revoke --user alice
Revoked 2 session(s) for alice.
$ acme-proxy admin session revoke --user alice --session 01234567
Revoked session 01234567 for alice.
$ acme-proxy admin session revoke --all
The id shown is a fingerprint of the stored token hash, not the hash itself —
printing the hash would put every live session’s lookup key on a terminal. It is
what revoke --session <id> takes to end one session; --session needs
--user, since the fingerprint only names a row within one operator’s sessions.
Two deadlines apply, and whichever comes first wins:
| Key | Default | |
|---|---|---|
admin.session_ttl_seconds | 43200 (12 h) | absolute; never extended |
admin.session_idle_timeout_seconds | 3600 (1 h) | advanced on use |
The idle deadline is advanced at most once a minute, so a page polling every few seconds is not a stream of database writes.
An admin_session_sweep job removes expired and idle rows for the life of the
process, starting with one pass at startup. Unlike nonces, sessions outlive a
restart, so a startup-only sweep would leak every session an operator never
explicitly signed out of.
created_ip and user_agent are recorded for forensics and are never
compared against the live request: pinning a session to an address breaks every
mobile and CGNAT operator, and pinning it to a User-Agent breaks on the next
browser update.
From the panel
admin session revoke takes a whole operator (--user), one of their sessions
(--user with --session), or the whole server (--all). The panel has the
single-session form too, at two different trust levels:
Your own sessions. Sign in, open your username in the top-right corner, and scroll to Sessions: every browser currently signed in as you, the one answering this request labelled, and a Revoke beside each of the others. No password re-entry — this is the same trust level as Sign out everywhere, which sits right below it and is now a button rather than only an API route nothing linked to.
$ curl https://admin.example.com/api/account/sessions \
-H 'Cookie: __Host-acme_admin_session=…'
Revoking the session making the request behaves exactly like signing out of just this browser: the cookie is cleared and you land back on the sign-in page. Revoking another one of your own ends it immediately, wherever it is signed in.
Another operator’s sessions. Reached from their page under Operators — the “colleague’s laptop went missing” answer that used to require SSH. Listing and revoking there behaves the same way, with one difference: it asks for your own password first, the same gate described above. A session id is a fingerprint of the stored token hash either way — printing the hash would put every live session’s lookup key on a terminal — and it only ever resolves within the one operator it was listed under, so an id copied from one operator’s page can never revoke another’s session by accident or by guessing.
What the log says
A failed sign-in returns one invalid_credentials whatever went wrong, so the
endpoint cannot be used to enumerate operators. The log keeps the distinction:
WARN event="admin_login_failed" username="alice" client_ip=… reason="wrong_password"
WARN event="admin_login_failed" username="ghost" client_ip=… reason="unknown_user"
WARN event="admin_login_failed" username="bob" client_ip=… reason="account_disabled"
WARN event="admin_login_failed" username="alice" client_ip=… reason="rate_limited"
INFO event="admin_login_succeeded" username="alice" client_ip=…
The second step keeps the same shape — one refusal to the client, the reason in
the log — and note that admin_login_succeeded is emitted at promotion, not
when the password is accepted:
INFO event="admin_login_mfa_pending" username="alice" client_ip=… step="verify"
WARN event="admin_mfa_failed" username="alice" client_ip=… reason="wrong_code"
WARN event="admin_mfa_failed" username="alice" client_ip=… reason="replayed"
INFO event="admin_mfa_verified" username="alice" method="totp"
WARN event="admin_mfa_recovery_code_used" username="alice" remaining=6
INFO event="admin_mfa_enabled" username="alice" recovery_codes=10
INFO event="admin_mfa_disabled" username="alice"
reason="replayed" is the one to look at twice: it means a correct code
arrived a second time inside its own window, which is what somebody replaying an
observed code looks like.
One more worth knowing, because nobody will guess it from a 401:
WARN event="admin_password_hash_unreadable" username="alice"
stored password hash could not be decoded; run
`acme-proxy admin user passwd` to rewrite it
A corrupt admin_users row refuses the sign-in rather than erroring the
endpoint, and the account stays unusable until the password is rewritten.
Revoking a single session, and every mutation on the Operators page, each
leave their own line — surface says which front end it came through and
target_username is absent on the self-service one, since there is nothing to
distinguish it from:
INFO event="admin_session_revoked" surface="ui" username="alice"
INFO event="admin_operator_disabled" surface="api" username="alice" target_username="bob"
INFO event="admin_operator_enabled" surface="ui" username="alice" target_username="bob"
INFO event="admin_operator_totp_reset" surface="api" username="alice" target_username="bob"
INFO event="admin_operator_session_revoked" surface="ui" username="alice" target_username="bob"
Customizing the Panel
The pages at /ui are minijinja templates compiled into the binary. Any one
of them can be replaced on disk without rebuilding, the same way notification
templates work — an operator who has already
overridden a notification should not have to learn a second scheme.
[admin]
template_dir = "/etc/acme-proxy/admin-templates"
Each name is looked for in that directory first and falls back to the
compiled-in default. The override is per file, not per directory: a
directory holding only layout.html restyles the chrome of every page and
leaves the other fifty-five exactly as shipped.
Every template is compiled at startup. A broken override refuses to start,
naming the file and the parse error, rather than serving a 500 the first time
somebody opens that page.
The files
Paths are relative to template_dir, and are also how the templates refer to
each other in {% extends %} and {% include %}.
| File | What it is |
|---|---|
layout.html | The chrome every full page extends: <head>, navigation, <body> |
login.html | Sign-in. Standalone — extends nothing, and uses no JavaScript |
mfa/challenge.html | The second sign-in step. Standalone and JavaScript-free for the same reason; branches on step between proving a code and setting one up |
mfa/_setup.html | The setup key and the otpauth:// URI. Included by both the sign-in flow and the account page, so it renders no <form> of its own |
mfa/_codes.html | A fresh recovery set, the one time it exists in the clear |
mfa/enrolled.html | Where a forced enrolment lands: the codes, then a link into the panel |
account/index.html, account/_mfa.html | The operator’s own page, and the fragment every mutation on it swaps |
account/_card.html | The second-factor card itself, with no id — so _codes.html can wrap it without nesting two elements carrying one |
account/_enrol.html, account/_codes.html | The enrolment step, and the codes plus the refreshed card |
account/_password.html, account/_password_card.html | The password-change swap target, and the form inside it |
account/_contact.html | Where this operator’s own security notifications go |
account/_sessions.html | This operator’s own live sessions, the swap target of revoking one |
index.html | The overview: four counts and the endpoint list |
partials/_flash.html | The inline banner every mutation’s answer renders |
partials/_pager.html | The previous/next controls under a list |
partials/_filter_meta.html | The tail of every list’s filter form: the loading indicator and the way back to the unfiltered list |
partials/_sessions_table.html | A table of live sessions, shared by the account page and an operator’s card |
accounts/list.html, accounts/_table.html | The account list, and the table htmx swaps |
accounts/detail.html, accounts/_card.html | One account, and the card every account mutation returns |
orders/list.html, orders/_table.html | The order list |
orders/detail.html, orders/_card.html | One order with its authorizations and challenges |
expiring/list.html, expiring/_table.html | The expiry list, its window and profile filters, and the hidden-count line |
eab/list.html, eab/_table.html | The credential list and the create form |
eab/detail.html, eab/_card.html | One credential |
eab/_created.html | The one-time HMAC secret |
nonces/index.html, nonces/_panel.html | The nonce count and the sweep control |
profiles/list.html, profiles/_table.html | The mounted endpoints |
profiles/filter.html | One endpoint’s resolved access policy. No fragment: nothing on it swaps |
audit/list.html, audit/_table.html | The audit trail |
audit/detail.html, audit/_card.html | One audit row |
jobs/list.html, jobs/_table.html | The background job queue |
jobs/detail.html, jobs/_card.html | One job, and the card its cancel and run-now actions return |
upstream_orders/list.html, upstream_orders/_table.html | The relay backend’s upstream orders. Read-only |
upstream_orders/detail.html, upstream_orders/_card.html | One upstream order, cross-linked to its job |
operators/list.html, operators/_table.html | The web admin’s operators |
operators/detail.html, operators/_card.html | One operator and their live sessions, re-rendered by every mutation on the page |
A file whose name starts with _ is a fragment: htmx swaps it on its own,
so it must not contain <html> or <body>, and it must keep the id on its
root element — that id is what the page’s hx-target points at.
Two things not to break
The extension is a security control
Every page template is named .html, and that is deliberate. minijinja decides
auto-escaping from the template name, and the notify templates are named .j2
precisely so that escaping is off for them (an email body is not markup).
Renaming a page template to .j2 — or adding a new one under a name minijinja
does not recognise as HTML — turns an account contact or an EAB label into
stored XSS.
The CSRF token has to stay on <body>
layout.html carries:
<body hx-headers='{"X-CSRF-Token": "{{ csrf_token }}"}'>
That attribute is the only route by which the token reaches a mutating request. A layout that drops it loses every write at once — which is the intended failure mode; a partial loss would be far harder to notice.
The same file sets three htmx options the other templates depend on:
<meta name="htmx-config"
content='{"includeIndicatorStyles":false,"defaultSwapStyle":"outerHTML","responseHandling":[...]}'>
includeIndicatorStyles: false stops htmx injecting an inline <style> element
that the Content-Security-Policy’s
style-src 'self' would block — the rules it would have injected live in
admin.css instead. responseHandling makes htmx swap non-2xx responses,
without which a 409 conflict would fail silently instead of showing the
operator a banner.
defaultSwapStyle: "outerHTML" means a response replaces the element it
targets. Every fragment is written for that: its root element carries the same
id as the swap target (accounts/_table.html is <div id="accounts-table">,
the target of the accounts filter form). An override of a fragment must keep
that root and its id, or the next swap on the page finds no target.
Context
Every full page gets csrf_token, user, can_write, nav (the active
navigation item) and title, plus its own data; the fragment a mutation answers
with gets the first three. Gate a control on {% if can_write %} — true for an
operator or admin session — which is also false when absent, so a control
never appears by accident. Timestamps are RFC 3339 strings, and the ago
filter renders one as 3 h ago or in 5 d, or as nothing when it is not a
timestamp:
{{ order.createdAt }} ({{ order.createdAt | ago }})
``` That data is the **same JSON the API returns** —
`render_account_json`, `render_order_detail_json`, `render_eab_json` — so `GET
/api/accounts/{id}` is an accurate description of what `account` holds in
`accounts/_card.html`. Lists additionally get `page` (`{items, total}`),
`pager`, `filters` and `profiles`.
A quick way to see a context in full is to render it:
```jinja
<pre>{{ account | tojson(indent=2) }}</pre>
Starting from the shipped version
The defaults are in the source tree under
crates/admin/src/webadmin/templates/. Copy the one you want to change:
$ mkdir -p /etc/acme-proxy/admin-templates
$ cp crates/admin/src/webadmin/templates/layout.html /etc/acme-proxy/admin-templates/
Then send SIGHUP — see Reloading the Configuration. Templates
are compiled up front, so a mistake fails the reload and the panel goes on
serving the last set that worked; it never reaches a browser. The same compile
happens at startup, where a mistake stops the process instead.
Stylesheet and scripts
admin.css and htmx.min.js are served from /ui/static/ and are not
covered by template_dir — they are embedded assets, not templates. To restyle
beyond what CSS variables allow, override layout.html and point its <link>
at your own file. Note that the Content-Security-Policy is default-src 'none'
with style-src 'self': a stylesheet must be served from this origin, and an
inline <style> block or style= attribute will be blocked.
Revocation & CRL
acme-proxy implements certificate revocation per RFC 8555 §7.6, and — with the
local_ca backend — publishes the resulting Certificate Revocation List.
POST /revokeCert
Revocation is available to a client through the standard ACME endpoint,
advertised in the directory. The request payload carries the base64url DER of
the certificate and an optional reason code.
Two ways to authorize it
RFC 8555 §7.6 names three, and acme-proxy accepts two:
- The order’s account, signing with its
kidas usual. - The certificate’s own key pair, signing with an embedded
jwkand no account at all. This is the RFC’s accountless case, and it is what lets the holder of a compromised key revoke it even if the ACME account is gone.
The third — an account holding valid authorizations for every identifier
in the certificate, without being the one that ordered it — is not
supported. It would let one account revoke another’s certificate on the
strength of authorizations obtained later, and an operator who needs that has
acme-proxy order revoke and the panel, which are attributable.
Because of the second form, this endpoint resolves authorization itself rather than going through the usual account lookup, and it is deliberately not gated on the account’s status — a deactivated account can still revoke its certificates.
How the certificate is identified
The submitted DER is decoded, the order is looked up by the certificate’s serial number, and the stored leaf is then compared to the submitted bytes for an exact DER match. A serial-only lookup would not be enough on its own; the byte comparison is the safety net.
One subtlety worth stating, because getting it wrong is a vulnerability: the key checked against the account is the one stored with the order, never a key re-derived from the submitted certificate. Re-deriving it would let anyone who merely observed the certificate on the wire revoke it, since the certificate contains its own public key.
Responses
200 OK— revoked. For alocal_caprofile the order reads revoked at once and the CRL follows as soon as the worker has signed it.503 serverInternal+Retry-After—relayorcustomonly: the revocation was queued for the worker and had not completed by the request’s deadline. Nothing failed, and it carries on; asking again waits on the same queued revocation.400 alreadyRevoked— the certificate was already revoked. This is checked after authorization, so an unauthorized caller cannot use the endpoint to probe whether a certificate has been revoked.400 badRevocationReason— the reason code is out of range.401 unauthorized— the signer is neither the order’s account nor the certificate’s key.
Reason codes
Reason codes are RFC 5280 §5.3.1 values. Codes 7 and 11 are not valid CRL
reasons, and out-of-range values are meaningless; all three are refused with
400 badRevocationReason, naming the code, so a client that meant a real
reason can send it rather than have one silently dropped. Omitting reason
entirely is always accepted, and records the revocation with no reason.
The first revocation of a certificate is the one that counts. Two at once —
a client and an operator, or acme-proxy order revoke beside a running server —
are one withdrawal of trust: whichever writes first sets the reason and the
time, on the order and on the CRL alike, and the other is answered
alreadyRevoked. The same holds for the queued path: a second request while
the first is still queued joins that job instead of queueing another, so it
inherits the first request’s reason.
Who revokes: never the request
Withdrawing trust needs what only the worker role holds — the CA key, a token
login, a relay’s upstream account — so no request, from a client or from the
panel, calls a backend itself:
local_ca: the revocation is a database row and the order’s stamp, in one transaction, and the worker signs it into the CRL (local_ca_crl_regenerate). No key is needed to record it.relay/custom: the revocation is asigner_revokejob for the worker, and the request waits on it.
For those two, the backend’s own revoke is called before the order is
marked revoked. The CA-side action is authoritative, so if the backend fails,
the order is deliberately left un-revoked and the job retries. A backend’s
revoke must therefore be idempotent — it may legitimately be called again
for a certificate it has already revoked.
Revocation is orthogonal to the order state machine
RFC 8555 defines no “revoked” order status, so a revoked order’s status stays
valid. The revocation timestamp and reason are stored in separate columns and
are not exposed in the ACME JSON a client polls.
They are visible through the admin CLI:
acme-proxy order show <id> # prints revoked, reason and the serial
acme-proxy order show <id> --json # the same as revokedAt / revocationReason
Revoking as an operator
For an out-of-band compromise report that the certificate holder cannot or will not act on:
acme-proxy order revoke <order-id> --reason 1
It is not confirm-gated — unlike order delete — because revocation only
ever tightens trust; there is no destructive outcome to protect against. It
runs the same checks and writes the same audit row and certificate_revoked
notification as the ACME endpoint, and like it never loads the CA key or talks
to an upstream itself: whatever holds the signing material is the running
server’s worker.
local_ca: the command records the revocation in the database and marks the order revoked in one transaction, then queues alocal_ca_crl_regeneratejob and prints its id. The order reads revoked at once. The server’s job runner signs the new CRL, usually withinjobs.poll_interval_ms, and with no server running the CRL catches up when one starts. A CA that no server has started with yet is refused: its revocation state is imported the first time a server meets it.relay/custom: revoking means asking the upstream CA or running the operator’s script, which is the server’s job. The command queues asigner_revokejob and waits for its answer, up to--waitseconds (default 30;0returns at once). A job that has not run by then is not an error: the command exits0naming it, andacme-proxy jobs show <id>follows it. A job that failed exits1with its error.
See Admin CLI.
GET /crl
With the local_ca backend, the CRL (RFC 5280) is served unauthenticated at
{base_url}/profile/<name>/crl, with content type application/pkix-crl.
- It is routed but deliberately not advertised in the ACME directory. A CRL is CA infrastructure, not an ACME resource, so it has no directory entry.
- A valid, correctly signed empty CRL exists from the moment a worker has started with the CA, before anything has ever been revoked. Clients fetching it do not have to special-case “no revocations yet”. Every role serves the stored CRL; none signs one but the worker.
- It is signed again for every revocation — immediately where the process
holding the key records it, and otherwise by the
local_ca_crl_regeneratejob that revocation queues — and by a daily refresh; see below.
The revocations and the current signed CRL live in the database, keyed by the CA’s key, so every process over one database — the server, a reloaded configuration, a second server — serves the same CRL. Backing up the database backs up the revocations; there is no separate file to keep in step with it.
crl_path still holds the current CRL, as PEM, rewritten every time a new one
is stored. It is an export for operators who publish the CRL from a static
web server: nothing ever reads it back, and losing it loses nothing. A second
file, ca.json.lock, sits beside it and is held while the export is written, so
two processes cannot leave an older CRL in place of a newer one. It is always
empty and needs no backup.
Upgrading from a JSON ledger
Before revocations moved into the database, they were kept in a JSON
sidecar beside crl_path — the same path with the extension swapped to
.json, so ca.crl was accompanied by ca.json. The first time a CA meets a
database with no CRL for it, the sidecar is imported: every entry becomes a
revocation, and the CRL number resumes above the last one the sidecar
published. The server logs local_ca_ledger_imported when it happens, which is
at startup through the daily refresh’s first pass, or at the first revocation
or CRL fetch if that comes sooner.
The import happens once. The sidecar is never read again nor written, so edits to it after that change nothing. Keep it with an old backup if you like, or delete it.
A sidecar that cannot be imported — a hand-edited serial that is not hex, say —
is logged as local_ca_crl_initialization_failed, naming the entry. Until it
is fixed, that CA’s revocations and GET /crl fail rather than proceed without
the history the sidecar holds; the import is tried again on the next attempt.
Expired entries are dropped
A revocation entry is not kept for ever. RFC 5280 §3.3 lets one go once it has appeared on a CRL issued after the certificate expired — nothing can present the certificate any more, and every relying party has had a CRL saying so — and this is what stops the CRL growing for the life of the deployment. The prune runs daily, starting shortly after startup.
The same daily pass re-signs a CRL that has less than half of its seven-day validity left, even when nothing was revoked or pruned, so a quiet CA’s CRL never lapses. When neither applies, nothing is signed.
Two rules are worth knowing:
- An entry is dropped an hour after the certificate’s own
notAfter, not at it. A relying party whose clock is behind yours still considers the certificate valid for a moment, and that moment is exactly when it would otherwise accept one you revoked. - An entry is dropped only once the stored CRL was issued past that
notAfter, which is §3.3’s actual condition. A certificate that expires between two signings therefore stays listed for one more pass: the newest CRL a relying party can fetch predates the expiry, and dropping the entry would leave that CRL as the last word on a certificate it still lists as valid to anyone whose clock disagrees. - An entry whose expiry is unknown is never dropped. That is any entry recorded before this server started tracking expiries (see below), and an unknown expiry is not an expired one.
Each CRL carries a crlNumber that only ever increases, including across a
restart, across a prune that shortens the list, and across several processes
signing for one CA. A client that meets a lower number than it has cached keeps
its cached CRL, so the number is stored with the CRL and a new one is only ever
stored over the one it was numbered after.
Sidecars written by 0.1.0 are a bare JSON array with no expiries and no number. They are imported as-is, with the number resuming above anything that format could have published. Their entries have no known expiry, so they stay on the CRL.
Telling clients where it is
A certificate does not point at this CRL until you say where to fetch it.
signer.local_ca.crl_distribution_points is that URL: set it, and every leaf
issued from then on carries a cRLDistributionPoints extension naming it (and
ca_issuer_urls does the same for the CA’s own certificate). Both are empty by
default, so out of the box the CRL above is reachable only by somebody who
already knows this server exists.
Two things follow from a URL being signed into a certificate:
- Certificates already issued keep the URL they were signed with, for their
whole validity. Changing the key changes nothing that is already out there,
which is why the value is yours to choose rather than something derived from
server.base_url. - The URL above is not automatically a good answer.
/crlis served by the profile router, so it sits behind that profile’s filter policy; an address-based rule will refuse it to relying parties outside the allowlist. Either add apathcheck permitting/crl(see Thepathcheck) or publish a copy ofcrl_pathsomewhere unconditionally reachable and name that.
See Local CA for both keys.
Other backends
custom— the CRL comes from the script’scrlhook, and only whensigner.custom.supports_crl = true. Otherwise there is nothing to serve. See Custom Script Signer.relay— the upstream CA publishes its own CRL or OCSP responder; this server does not republish it.
Interaction with renewal information
A certificate acme-proxy knows to be revoked is reported through
ARI with a renewal window entirely in the
past, prompting a compliant client to renew immediately.
That check happens before the signer backend is consulted, so a locally revoked certificate is never talked out of renewing by an upstream CA that has not yet noticed.
Notifications
The notify subsystem alerts operators on lifecycle events within the ACME
server, and web-admin operators on security events on their own account.
Supported events
| Event | Fired when |
|---|---|
profile_mounted | A profile is initialized at startup. |
account_created | A client registers a new account. |
account_deactivated | An account is deactivated. |
certificate_issued | An order is finalized and a certificate is minted. |
certificate_revoked | A certificate is revoked, via the ACME API or the admin CLI. |
challenge_failed | A domain-control validation attempt fails. |
certificates_expiring | The periodic expiry digest, one per profile. |
admin_sign_in | A web-admin sign-in from an unfamiliar address, a refused second factor after a correct password, or a lockout. |
admin_credential_changed | A web-admin operator’s password or second factor changed. |
These nine names are the only valid values wherever a backend’s events list
is configured. An unrecognised name is a startup error, not a silently
ignored entry.
certificates_expiring is the one that is not a thing that just happened.
Every other event describes a single subject at the moment it changed; this one
is a digest sent on a schedule, listing the certificates on one profile that
expire inside a configured window. It sends nothing until
notify.expiry.lead_days is set, so leaving it in an
events list costs nothing.
admin_sign_in and admin_credential_changed are the web-admin
operator security events
(ASVS V6.3.5 / V6.3.7). They fire only on the process-wide [admin.notify]
dispatcher — never on a profile’s [notify] — so listing either in a
per-profile backend’s events costs nothing. Their email goes to the affected
operator’s own address rather than to notify.email.to.
Backends
- Email — SMTP, via
lettre. - Webhook — any HTTP endpoint, with the URL, method, headers and body all configured. This is how Slack, Mattermost, Teams, Telegram and Matrix are reached: they differ in those four values and nothing else, so each is a configuration entry rather than a backend of its own.
- Custom Script — shell out to a local script, for a channel that is not an HTTP request at all.
Email and webhook render their messages with MiniJinja templates you can override; see Customizing Templates.
Configuration
[notify]
# Which backends are active. Empty (the default) means no notifications at all.
enabled = ["email", "webhook"]
# Which [notify.webhook.<name>] entries to POST to, when "webhook" is listed
# above.
webhook_enabled = ["slack"]
# Which [notify.custom.<name>] entries to run, when "custom" is listed above.
custom_enabled = []
# Optional directory of template overrides, checked per template file before
# falling back to the compiled-in default.
template_dir = "/etc/acme-proxy/templates"
Reference
enabled (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__ENABLED
Active backends: any of email, webhook, custom. Empty means the
subsystem is off. "mattermost" was removed in favour of webhook and is
refused by name.
webhook_enabled (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__WEBHOOK_ENABLED
Which entries under [notify.webhook.<name>] to POST to, and in what order.
Listing "webhook" in enabled while leaving this empty is a startup error,
as is naming an entry that has no table.
custom_enabled (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__CUSTOM_ENABLED
Which entries under [notify.custom.<name>] to run, and in what order. The
same two startup errors apply.
template_dir (String) — Default: "" | Env: ACME_PROXY_NOTIFY__TEMPLATE_DIR
Directory searched for template overrides. Lookup is per file, so overriding
one message (say email/certificate_issued.body.j2) leaves every other message
at its compiled-in default. Empty means defaults only.
Each backend additionally takes its own events list and timeout_ms; see the
backend pages.
Expiry digest
A certificate approaching its notAfter is the one thing the events above cannot
report: nothing happens when a certificate is a fortnight from expiring. The
[notify.expiry] table adds a periodic sweep that looks, and sends one
message per profile listing what it found.
Deliberately not one message per certificate. A renewal is a new order, so the certificate it replaced still reaches its own expiry on schedule — a per-certificate reminder therefore fires for every certificate the CA has ever issued, on its way out, in exactly the deployments where the automation is working. Instead, each entry in the digest says whether something has already taken its place, and the entries where nothing has are the ones worth acting on.
That annotation is drawn from two signals, and the message says which was used:
the successor order’s own replaces field (RFC 9773 §5 — exact, but only from
clients that send one), or a later, unrevoked certificate of the same
account covering all of the same names. Both are deliberately narrow. A
certificate wrongly marked as already renewed is one an operator skips over
while it lapses; one wrongly left unmarked is a line of noise.
expiry.lead_days (Integer) — Default: 0 | Env: ACME_PROXY_NOTIFY__EXPIRY__LEAD_DAYS
How far ahead to look. 0 is off — the sweep is never scheduled at all,
the same shape audit.retention_days and jobs.retention_days use.
expiry.interval_days (Integer) — Default: 7 | Env: ACME_PROXY_NOTIFY__EXPIRY__INTERVAL_DAYS
How often the digest is sent. There is deliberately no per-certificate rate limit beside it: the digest is the rate limit. A digest with nothing to report is not sent, so the absence of a message is what “everything is renewed” looks like.
expiry.max_entries (Integer) — Default: 50 | Env: ACME_PROXY_NOTIFY__EXPIRY__MAX_ENTRIES
The most certificates one message lists. The number that matched is carried whole regardless, so a truncated digest still says how many it did not name.
The schedule is a row in the durable job queue rather than a timer, so it
survives a restart: a server restarting more often than interval_days still
sends its digest on time instead of resetting the clock each start.
Delivery semantics
Dispatch is fire-and-forget: the event is written to the durable job queue and the ACME response proceeds immediately. A notification backend can never delay or fail the request that triggered it.
Delivery itself is a notify_deliver job, one row per backend per event, so
one flaky webhook is retried without re-sending through an email backend that
already succeeded. Two consequences worth planning around:
- A row outlives the process that wrote it. A notification generated moments before a restart is delivered by whoever starts next, rather than lost. There is no drain at shutdown to configure or wait for.
- A failure is retried, unless it never could have worked. A refused SMTP
connection, a timeout, a 429 or a 5xx from a webhook goes back in the queue
under
jobs.max_attemptsand the shared backoff. A template that does not render, aurlthat does not parse and any other 4xx are refused on the first attempt — retrying would reach the same answer four more times and delay the log line saying so.
Every attempt logs notify_delivered or notify_delivery_failed. When the
attempts run out, one notify_delivery_abandoned says the notification is
genuinely lost — that is the line to alert on. Because the custom backend’s
contract is an exit code with no way to say “never retry”, every failure of
a custom script is treated as retryable.
Customizing Templates
acme-proxy uses the MiniJinja
templating engine to render notification payloads. The server embeds sensible
default templates inside the binary, but you can override any of them by
pointing the server to a custom template directory.
Enabling custom templates
In your config.toml, define a template_dir:
[notify]
enabled = ["email", "webhook"]
template_dir = "/etc/acme-proxy/templates"
The server will look in this directory before falling back to its embedded defaults. You only need to create the files you want to override.
Template file structure
Templates are grouped by backend and event name. Email requires separate files for the subject line and body.
/etc/acme-proxy/templates/
├── email/
│ ├── account_created.subject.j2
│ ├── account_created.body.j2
│ ├── certificate_issued.subject.j2
│ └── certificate_issued.body.j2
└── webhook/
├── certificate_issued.j2
└── challenge_failed.j2
(Available event names: profile_mounted, account_created,
account_deactivated, certificate_issued, certificate_revoked,
challenge_failed, certificates_expiring, admin_sign_in,
admin_credential_changed).
A webhook/<event>.j2 renders the message, not the payload: every
[notify.webhook.<name>] entry then wraps it in its own body template. So a
file here restyles the text for every webhook target at once, and an entry’s
body restructures one target’s request without touching the text. See
Webhook.
Context variables
When rendering a template, acme-proxy passes a context object containing the
event’s data. All events include profile (the name of the profile triggered)
and most include client_ip (the IP address of the ACME client that initiated
the request).
certificate_issued
Triggered when an order is finalized and the signer mints a certificate.
profile(String)order_id(String)account_id(String)cert_serial(String) - Hex-encoded serial numberidentifiers(List of Strings) - The SANs/Domains requestedclient_ip(Option<String>)
Example (webhook/certificate_issued.j2):
✅ **Certificate Issued** on profile `{{ profile }}`
**Domains:** {{ identifiers | join(", ") }}
**Serial:** `{{ cert_serial }}`
**Requested By IP:** `{{ client_ip | default("system") }}`
challenge_failed
Triggered when an HTTP-01 or DNS-01 validation attempt fails.
profile(String)order_id(String)account_id(String)authz_id(String)challenge_id(String)challenge_type(String) - e.g. “http-01”identifier(String) - The domain that failederror(String) - The detailed error from the validation attemptclient_ip(Option<String>)
account_created / account_deactivated
Triggered on account lifecycle events.
profile(String)account_id(String)contact(List of Strings) - e.g.["mailto:admin@example.com"](only on created)client_ip(Option<String>)
certificate_revoked
Triggered via the ACME API or Admin CLI.
profile(String)order_id(String)account_id(String)cert_serial(String)reason(Option<Integer>) - RFC 5280 revocation reason codeclient_ip(Option<String>)
profile_mounted
Triggered during server startup when a profile is successfully initialized.
profile(String)
certificates_expiring
The periodic expiry digest — the one event whose subject is a list, so a template here loops where every other one interpolates.
profile(String)generated_at(Integer) - Epoch seconds; the pointdays_remainingcounts from, so a message read days later is still self-describinglead_days(Integer) - The window this digest coverstotal(Integer) - How many certificates matched, which may be more thancertificatesholds:notify.expiry.max_entriesbounds the list, and the count is what lets a truncated message say how many it did not namecertificates(List) - each withorder_id,account_id,cert_serial,identifiers(List of Strings),not_after(Integer, epoch seconds),days_remaining(Integer, floored) andsuperseded_bysuperseded_by(Option) - absent when nothing has replaced this certificate; otherwiseorder_id,cert_serial,not_afterandvia, whereviais"replaces"(the client said so, RFC 9773 §5) or"identifiers"(a later certificate of the same account covers the same names)
There is no client_ip: a digest is generated by a sweep, with no request
anywhere in scope.
Example (webhook/certificates_expiring.j2):
⏰ **{{ total }} expiring** on `{{ profile }}`
{%- for cert in certificates %}
{{ cert.identifiers | join(", ") }} — {{ cert.days_remaining }}d
{{- " (already replaced)" if cert.superseded_by }}
{%- endfor %}
The superseded_by test is what makes a digest readable: an operator scans for
the entries without it. Rendering every row identically would bury the handful
that nobody has renewed among the many that are already taken care of.
admin_sign_in
A web-admin operator sign-in worth flagging.
profile(String) — always__admin__username(String) — the operatorrecipient(Option) — the operator’s own contact address; theemailbackend sends there rather than tonotify.email.tooutcome(String) —succeeded_from_new_address,second_factor_refusedorlocked_outclient_ip(Option),user_agent(Option)at(Integer) — epoch seconds
admin_credential_changed
A web-admin operator’s password or second factor changed.
profile(String) — always__admin__username,recipient— as abovechange(String) —password,second_factor_enabled,second_factor_disabledorrecovery_codes_regeneratedby_self(Boolean) —falsewhen another administrator made the changeclient_ip(Option),user_agent(Option),at(Integer)
Email Notifications
The email notification backend sends alerts via SMTP, using the lettre crate
to format and dispatch multipart messages.
Delivery semantics
Notifications in acme-proxy are designed to be entirely non-blocking.
- When an event occurs (e.g., a certificate is issued or revoked), the notification subsystem spawns a background Tokio task.
- The task attempts to connect to the SMTP server and send the email.
- Fire and Forget: If the delivery fails (e.g., the SMTP server is
unreachable),
acme-proxylogs an error at thewarnlevel but does not retry. Notification failure does not cause the ACME transaction (likefinalize) to fail.
Template customization
By default, acme-proxy sends plain, functional emails rendered from templates
compiled into the binary. You can override any of them by setting
notify.template_dir.
The engine is MiniJinja and the files are .j2. Email needs two files
per event — one for the subject line and one for the body — under an email/
subdirectory of template_dir:
/etc/acme-proxy/templates/
└── email/
├── certificate_issued.subject.j2
└── certificate_issued.body.j2
Lookup is per file, so overriding one message leaves the rest at their defaults. See Customizing Templates for the full layout and the context variables available to each event.
Configuration
[notify]
enabled = ["email"]
# Note: the path is the templates root, NOT the email/ directory inside it —
# "email/" is appended by the loader.
template_dir = "/etc/acme-proxy/templates"
[notify.email]
smtp_host = "smtp.internal.corp"
smtp_port = 587
smtp_username = "acme_bot"
smtp_password = "super_secret_password"
smtp_security = "starttls"
from = "acme-proxy@internal.corp"
to = ["pki-admins@internal.corp", "security@internal.corp"]
events = ["certificate_issued", "certificate_revoked"]
timeout_ms = 5000
Reference
notify.enabled and notify.template_dir are subsystem-wide keys documented in
Notifications. The keys below are specific to this
backend.
smtp_host (String) — Default: "" | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_HOST
SMTP server hostname.
smtp_port (Integer) — Default: 587 | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_PORT
SMTP server port.
smtp_username (String) — Default: "" | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_USERNAME
SMTP authentication username.
smtp_password (String) — Default: "" | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_PASSWORD
SMTP authentication password. Sensitive — prefer the environment variable to a file on disk, as with every other secret in the configuration.
smtp_security (String) — Default: "starttls" | Env: ACME_PROXY_NOTIFY__EMAIL__SMTP_SECURITY
TLS requirement: "starttls", "tls", or "none".
from (String) — Default: "" | Env: ACME_PROXY_NOTIFY__EMAIL__FROM
Sender email address.
to (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__EMAIL__TO
List of recipient email addresses. Required for a per-profile [notify]
backend. Under [admin.notify] it is the fallback: the web-admin security
events (admin_sign_in, admin_credential_changed) are delivered to the
affected operator’s own contact address, and to is used only when that
operator has none — so [admin.notify.email].to may be left empty.
events (Array) — Default: every event | Env: ACME_PROXY_NOTIFY__EMAIL__EVENTS
Lifecycle events this backend reacts to.
timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_NOTIFY__EMAIL__TIMEOUT_MS
Timeout budget for the SMTP exchange.
Webhook Notifications
The webhook backend makes one HTTP request per event, with the URL, the
method, the headers and the body all stated in configuration.
That is the whole design. Slack, Mattermost, Microsoft Teams, Telegram and Matrix differ in those four values and in nothing else — the transport, the timeout, the TLS stack, the proxy and the retry rules are the same for all of them — so a chat provider is an entry in a table here rather than a backend in the binary. See Provider recipes below for a copy-pasteable entry per provider.
Configuration
Entries are named, selected and ordered by webhook_enabled, exactly like
[notify.custom]. Two entries are two independent deliveries: a
retry never re-sends through one that already succeeded.
[notify]
enabled = ["webhook"]
webhook_enabled = ["slack", "oncall"]
[notify.webhook.slack]
url = "https://hooks.slack.com/services/T00000000/B00000000/XXXXXXXX"
body = '{"text": {{ message | tojson }}}'
[notify.webhook.oncall]
url = "https://chat.internal.corp/api/v1/rooms/pki/messages"
events = ["certificate_revoked", "challenge_failed"]
[notify.webhook.oncall.headers]
Authorization = "Bearer s3cret"
An entry name must match ^[a-z0-9-]+$: it is also an environment-variable
segment, which the configuration loader lowercases, so anything else could name
one entry in a file and a silently different one through the environment.
How a body is rendered
Rendering has two stages, and the split is what lets you restyle every message without restructuring any payload, or the reverse:
- The message.
webhook/<event>.j2renders human-readable text. These are embedded in the binary and overridable file by file throughnotify.template_dir— overridewebhook/certificate_issued.j2and every other message stays at its default. - The payload. The entry’s own
bodyis a template withmessage,hook(the event name) and every field of the event in scope — the same fields Customizing Templates lists.
| tojson is not decoration. A .j2 template has auto-escaping off, on
purpose, so a message holding a quote or a newline — a challenge_failed
quotes what the validator saw — would otherwise render a payload the provider
answers 400 to. That answer is permanent, so the delivery is refused on the
first attempt, for exactly the events you most wanted to hear about.
Provider recipes
Everything below is the whole entry: no other key is needed. message is the
rendered text from stage 1.
| Provider | method | body | headers |
|---|---|---|---|
| Slack | POST | {"text": {{ message | tojson }}} | — |
| Mattermost | POST | {"text": {{ message | tojson }}} | — |
| Microsoft Teams | POST | {"text": {{ message | tojson }}} | — |
| Google Chat | POST | {"text": {{ message | tojson }}} | — |
| Telegram | POST | {"chat_id": "-1001234567890", "text": {{ message | tojson }}} | — |
| Matrix | PUT | {"msgtype": "m.text", "body": {{ message | tojson }}} | Authorization: Bearer <token> |
Notes on the two that are not simply a URL:
- Telegram’s URL is
https://api.telegram.org/bot<token>/sendMessage, and the destination is thechat_idin the body rather than anything in the URL. - Matrix’s URL is the room’s send endpoint,
/_matrix/client/v3/rooms/<room>/send/m.room.message/<txn>, on your own homeserver. The transaction id is meant to change per message; a fixed one makes the homeserver deduplicate, which is a deliberate choice worth knowing you are making.
Mattermost’s channel and username overrides, which this backend replaced,
are two more members of the same object:
body = '{"channel": "pki-alerts", "username": "ACME Proxy", "text": {{ message | tojson }}}'
Delivery semantics
Nothing here is specific to this backend — see Notifications for the queue, the retries and the one log line that means a notification was genuinely lost. Two points that bite webhooks in particular:
- A 4xx is permanent. Every 4xx except
408and429is the provider stating a reason — a webhook that has been deleted, a payload it will not accept — so it is refused on the first attempt rather than retried four more times.5xx,429,408, a timeout and a refused connection are retried. A refusal’s response body is quoted in the log line, truncated: Slack’sinvalid_payloadis usually the entire diagnosis. - Everything unusable is refused at startup, not at delivery time: an empty
or unparseable
url, a scheme that is nothttp/https, amethodoutsidePOST/PUT/PATCH, a header a wire format will not carry, and abodythat does not compile as a template.
The URL’s path and every header value routinely carry the credential, so nothing this server logs or reports ever renders more than the host and the header names.
Reference
url (String) — Default: "" | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__URL
The endpoint to call. Required once the entry is listed in webhook_enabled.
method (String) — Default: "POST" | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__METHOD
POST, PUT or PATCH, in any case. Anything else is a startup error: a
webhook is a write, and a body on GET would be a request no provider answers.
headers (Table) — Default: {} | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__HEADERS__<HEADER>
Extra request headers. Applied after the defaults, so an entry may override
content-type (application/json) or user-agent (acme-proxy). Header
names arrive lowercased from the environment, which HTTP does not care about.
body (String) — Default: {"text": {{ message | tojson }}} | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__BODY
The request body, as a template. The default is the payload Slack, Mattermost,
Teams and Google Chat all accept, so those four need a url and nothing else.
events (Array) — Default: every event | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__EVENTS
Lifecycle events this entry reacts to.
timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_NOTIFY__WEBHOOK__<NAME>__TIMEOUT_MS
Budget for one delivery attempt, connect and handshake included.
Custom Script Notifications
The custom notification backend runs a local script (Bash, Python, Go, …) when
an ACME event occurs. Use it to integrate with internal ticketing systems,
custom logging infrastructure, or alerting pipelines that a plain HTTP webhook
cannot satisfy.
Configuration
custom is a named map, like [filter.check]. Two keys switch it
on: notify.enabled activates the backend, and notify.custom_enabled selects
which scripts run, and in what order.
[notify]
enabled = ["custom"]
custom_enabled = ["ticket-creator"]
[notify.custom.ticket-creator]
script_path = "/etc/acme-proxy/scripts/ticket-creator.sh"
timeout_ms = 10000
args = []
events = ["certificate_issued", "certificate_revoked"]
Three ways to get this wrong, all of which fail at startup rather than at delivery time:
- Listing
"custom"innotify.enabledwhilecustom_enabledis empty. - Naming an entry in
custom_enabledthat has no[notify.custom.<name>]table. - Using an entry name outside
^[a-z0-9-]+$—ticket_creator(underscore) is rejected;ticket-creatoris fine.
One process is spawned per enabled script per event.
Reference
script_path (String) — Default: "" | Env: ACME_PROXY_NOTIFY__CUSTOM__<NAME>__SCRIPT_PATH
Path to the executable. Required.
timeout_ms (Integer) — Default: 5000 | Env: ACME_PROXY_NOTIFY__CUSTOM__<NAME>__TIMEOUT_MS
Maximum execution time. A script still running when this expires is killed.
args (Array) — Default: [] | Env: ACME_PROXY_NOTIFY__CUSTOM__<NAME>__ARGS
Static arguments passed to the script on every invocation.
events (Array) — Default: every event | Env: ACME_PROXY_NOTIFY__CUSTOM__<NAME>__EVENTS
Which events this script reacts to. Valid names are profile_mounted,
account_created, account_deactivated, certificate_issued,
certificate_revoked, challenge_failed, certificates_expiring. An
unrecognised name is a startup error.
Execution model and security
- Environment clearing (
env_clear): the child runs with a scrubbed environment, inheriting only a minimalPATHand the injectedACME_NOTIFY_*variables. The server’s own environment may hold secrets —notify.email.smtp_password, the RFC 2136 TSIG key, the NetBox token — and a notification script has no business reading them. - Zombie prevention (
kill_on_drop): the script runs under a Tokio timeout withkill_on_drop(true). Atokio::time::timeoutonly drops the future, so without this a timed-out script would outlive its deadline and leak a process per event. - Fire and forget: delivery runs in a background task. A failing or hanging
script can never delay or fail the ACME response that triggered it; failures
are logged (
event = "notify_delivery_failed") and dropped, with no retry.
Data passing
The script receives context both ways.
Environment variables
All seven are always set. Ones that do not apply to the event are set to the empty string rather than omitted, so a script can read them unconditionally.
| Variable | Value |
|---|---|
ACME_NOTIFY_HOOK | The event name, e.g. certificate_issued. |
ACME_NOTIFY_PROFILE | The profile the event occurred in. |
ACME_NOTIFY_CLIENT_IP | The ACME client’s address; empty when no request was in scope (e.g. profile_mounted, or an asynchronous relay completion). |
ACME_NOTIFY_ACCOUNT_ID | The account, when the event has one. |
ACME_NOTIFY_ORDER_ID | The order, when the event has one. |
ACME_NOTIFY_CERT_SERIAL | The certificate serial, on certificate_issued / certificate_revoked. |
ACME_NOTIFY_IDENTIFIERS | Comma-joined identifier values. Only populated for certificate_issued. |
On certificates_expiring every variable but ACME_NOTIFY_HOOK and
ACME_NOTIFY_PROFILE is empty, and that is not an omission: a digest is
about a list of certificates spanning however many accounts, so there is no one
account, order or serial for a variable to hold. Read the list from the JSON on
stdin, which is the channel that carries structure.
On admin_sign_in / admin_credential_changed the only populated variables
are ACME_NOTIFY_HOOK, ACME_NOTIFY_PROFILE (__admin__) and
ACME_NOTIFY_CLIENT_IP (the web-admin client’s address). The username,
recipient, outcome / change, by_self and user_agent are on the stdin
JSON — an operator is not an ACME subject, so it has no account or order id.
There is no
ACME_NOTIFY_EVENT; the event name isACME_NOTIFY_HOOK.
JSON on stdin
A JSON object is always written to the script’s standard input. It is the
event’s own fields plus a "hook" key naming the event — the same data the
templating backends render from. For example, on
certificate_issued:
{
"hook": "certificate_issued",
"profile": "default",
"order_id": "…",
"account_id": "…",
"cert_serial": "…",
"identifiers": ["a.example.com", "b.example.com"],
"client_ip": "203.0.113.5"
}
The exact fields per event are listed in Customizing Templates — the template context and the stdin payload carry the same values.
The issued certificate itself is never passed to a notification script, on stdin or otherwise. Only its serial and the identifiers it covers are available. A script that needs the PEM must fetch it out of band.
Example
#!/bin/bash
# /etc/acme-proxy/scripts/ticket-creator.sh
set -euo pipefail
payload=$(cat) # the JSON described above
case "$ACME_NOTIFY_HOOK" in
certificate_issued)
echo "issued ${ACME_NOTIFY_CERT_SERIAL} for ${ACME_NOTIFY_IDENTIFIERS}" \
>> /var/log/acme-issuance.log
;;
certificate_revoked)
# ACME_NOTIFY_IDENTIFIERS is empty for this event — read the payload
# if you need more than the serial.
reason=$(echo "$payload" | jq -r '.reason // "unspecified"')
curl -sS -X POST https://tickets.internal/api/incidents \
-H 'Content-Type: application/json' \
-d "{\"serial\":\"$ACME_NOTIFY_CERT_SERIAL\",\"reason\":\"$reason\"}"
;;
esac
See Custom Plugins Examples for more.
Audit Trail
A certificate authority’s most important record is not what it holds but what it
did: who asked it to sign or withdraw a certificate, who administered it, from
where, and what it answered. acme-proxy writes that down in one append-only
table and surfaces it in three places — the CLI, the JSON API and the panel.
There is deliberately no audit.enabled. Recording who asked the CA to sign
something, and who changed its stored state, is not a feature of this server, it
is what a CA does. The only thing an operator can switch off is the reverse-DNS
lookup, because that one costs a network round trip and there are estates where
it can never succeed.
What gets recorded
Two different things, in two different places.
Traceability columns, on the rows themselves:
| Table | Columns |
|---|---|
accounts | created_ip / created_ptr — where newAccount was called from, frozen at creation; last_seen_at / last_seen_ip / last_seen_ptr — where the key last authenticated a request |
orders | created_ip / created_ptr — where the order was opened from |
The audit log, one row per CA action and per refusal, plus one per successful administrative action:
| Event | Written when |
|---|---|
certificate_issued | the signer returned a certificate |
certificate_issue_failed | finalize was refused after the order was ready — a bad CSR, a CSR/order mismatch, a filter denial, or the backend’s own failure |
certificate_revoked | a revocation succeeded, whether through ACME, the CLI or the panel |
certificate_revoke_failed | a revocation was refused — including alreadyRevoked, an unauthorized caller, and an unknown certificate |
The refusals are the point. A stream of certificate_revoke_failed rows naming
certificates that do not exist is somebody enumerating serials, which is exactly
the question a trail exists to answer.
Administrative actions
An operator (through the panel) or the host CLI changing stored state also
leaves a row. These are recorded only on success — a not-found or refused
admin operation is the operator being told the state of things, not the CA
turning a remote party away — and each is attributed to actor_kind = "admin"
(the operator’s username and resolved address) or "cli" (the host).
| Event | Written when |
|---|---|
account_deactivated / account_contact_updated / account_deleted | an ACME account was deactivated, had its contact list rewritten, or was hard-deleted (the row names the account and its profile; account_deleted carries the cascade count) |
order_deleted | an order row was hard-deleted (names the order, account and identifiers) |
eab_created / eab_revoked / eab_deleted | an External Account Binding credential was minted, revoked or deleted (the detail names the kid; the secret is never recorded). eab_deleted also says whether its accounts were kept, deactivated or deleted, and each account changed gets its own account_deactivated or account_deleted row |
operator_created / operator_deleted | a web-admin operator was added or removed |
operator_role_changed / operator_disabled / operator_enabled | an operator’s privilege tier or sign-in status changed |
operator_password_changed / operator_contact_updated | an operator’s password or notification address changed (self-service or by an admin) |
operator_totp_enrolled / operator_totp_disabled / operator_recovery_codes_regenerated | an operator’s second factor was enrolled, removed (including an admin reset), or its recovery codes reissued |
session_revoked | one session, all of an operator’s sessions, or every session on the server was revoked (the detail says which) |
job_cancelled / job_advanced | a background job was cancelled or nudged/revived (run-now). A cancelled relay issuance instead writes certificate_issue_failed — it abandons a certificate order |
nonce_cleanup_completed / audit_pruned | the nonce table or the audit log itself was swept by hand (audit_pruned records its own action, so a manual prune always leaves the one row that says it happened) |
database_transferred | every row was copied into another backend (acme-proxy transfer). Written to the source, which is the database that holds the trail leading up to the move; the copy has already read past this row, so the target’s own trail begins at the transfer it arrived in |
The vocabulary is defined by acme_proxy_core::audit::AuditEvent, not by a
database constraint, so a newer server writing a name an older one does not know
still loads on the older one — it simply shows the raw string. The web audit
surface stays read-only: a stolen session that could erase the trail would
make the trail prove nothing, so pruning is audit cleanup on the host or
audit.retention_days and nothing else.
Two boundaries worth knowing:
- Nothing is ever compared against any of this. Pinning an identity to an address breaks CGNAT and mobile clients, so these columns answer “who asked for this certificate, and from where”, never “may this request proceed”.
- None of it reaches an ACME object. The wire format is RFC 8555’s and stays that way. The trail is visible through the admin surfaces only.
What is not recorded
Refusals that never reached a CA action. A request rejected for a bad signature,
a replayed nonce or an unready order is protocol bookkeeping — nothing was
signed, nothing was withdrawn, and recording it would bury the rows that matter.
For the same reason, a revokeCert payload that cannot be parsed at all writes
nothing: a row naming no subject is noise.
Reading it from the CLI
acme-proxy audit list
acme-proxy audit show 4213
| Command | Flags |
|---|---|
audit list | --profile, --account-id, --order-id, --cert-serial, --event, --outcome, --since-days <n>, --limit <n>, --offset <n>, --json |
audit show <id> | --json |
audit cleanup | --older-than <days> (prompts) |
audit list is paged, defaulting to 50 rows, as account list and order list are. This table grows a row per issuance for the life of the deployment,
so it always prints N of M row(s) — a page must never be mistaken for the
whole trail. There is no “everything” spelling on purpose; on a year-old CA that
is a terminal full of scrollback and a table loaded into memory. Page with
--offset, and see Paging for the window every listing shares
and the --json envelope it answers with.
$ acme-proxy audit list --outcome failure --since-days 7
4213 2026-08-09T18:22:04Z certificate_issue_failed prod acme:acct-9f2c 10.4.1.19 (web7.corp.example) api.corp.example reason=badCSR
1 of 1 row(s).
An unknown --event or --outcome is refused by name, listing the values
this build knows. Passed through to SQL it would answer “no rows”, which looks
exactly like “nothing happened” — the single most misleading answer an audit
tool can give.
audit show prints one field per line, omitting every field that was not
recorded rather than showing it as empty:
$ acme-proxy audit show 4213
id 4213
created 2026-08-09T18:22:04Z
event certificate_issue_failed
outcome failure
profile prod
actor acme:acct-9f2c
account acct-9f2c
order ord-71ab
client_ip 10.4.1.19
client_ptr web7.corp.example
user_agent certbot/2.9.0
request_id 01J9F2K7Q4
reason badCSR
identifiers api.corp.example
The actor
Every row names who acted, as a kind plus an optional id:
| Kind | Meaning |
|---|---|
acme | an ACME client, identified by its account id |
admin | a web admin operator, identified by username |
cli | someone on the host — there is no request and no address, and the row says so |
system | the server itself, e.g. a relayed order settling in the background |
An administrative revocation is attributed to the operator, not the certificate’s owner. Recording the client there would say the opposite of what happened.
Retention
audit.retention_days = 0 — the default — keeps everything for ever, which is
the right default for a trail whose value is that it is complete. Setting it
non-zero spawns a daily sweep running the identical DELETE as:
acme-proxy audit cleanup --older-than 365
This is the only command in the binary that destroys audit history, so it is
confirm-gated and its prompt names the number of rows it is about to remove. Use
-y to skip the prompt in a cron job.
The web admin cannot prune the trail.
/api/auditand/ui/auditare read-only, and there is no route to list, because the first thing a stolen session would do is erase what it had done — and a trail that can be erased by the thing it is watching proves nothing. Pruning happens on the host, or on a schedule set in configuration.
Reading it from the web admin
| Surface | |
|---|---|
GET /api/audit?profile=&accountId=&orderId=&certSerial=&event=&outcome=&limit=&offset= | the paged envelope every list endpoint returns |
GET /api/audit/{id} | one row |
/ui/audit | the list, filterable, with a detail page per row |
The API omits absent fields rather than sending them as null, so a client can
test for presence directly. There is no since filter on this surface: a
browser filters by picking a page, and a date parser here would be a second
definition of “how far back” that the CLI already has.
Rows carry remote text. A
User-Agentand a reverse DNS name are both written by whoever is on the other end — the PTR for a client’s address is controlled by whoever runs that address’s reverse zone. The panel escapes them like any other untrusted value, andtests/admin_pages.rspins it.
Why it survives deletion
audit_log is the one table in the schema with no foreign keys, and that is
the design rather than an oversight. An audit row has to outlive account delete and order delete; a CASCADE would destroy the evidence along with
its subject. So account_id and order_id are plain columns naming a row that
may be gone, and the identifiers are frozen into the row rather than joined back
to an order that no longer exists.
The rest follows from the same rule: rows are only ever inserted — there is no
setter and no UPDATE anywhere in the crate — and the primary key is
AUTOINCREMENT, so SQLite cannot reuse the rowid of a purged row.
Configuration
See [audit] for the three keys. The
section is process-wide rather than per-profile: the trail describes the CA, not
one of its endpoints.
Monitoring & Observability
Running an ACME server in production requires visibility into its health,
request volume, and error rates. acme-proxy offers three surfaces for that: a
health endpoint, structured logging, and a Prometheus exporter.
Health checks
- Endpoint:
GET /health - Response:
200 OKwith the body{"healthy": true}
The handler takes no application state: it does not query the database, and it
does not consult the signer. It answers 200 whenever the process is alive and
able to accept a connection, and it has no other failure mode. Treat it as a
liveness probe, not a readiness probe — it will not tell you that the disk
holding sqlite.db filled up.
What makes it worth polling is where it is mounted. /health lives on the
root router, not inside a profile, which means it is deliberately outside:
- the admission-control layer, so a saturated server still answers the probe (inside the limit, the probe was starved exactly when it mattered, and a load balancer would go on reporting the node healthy right up to the point where the probe could no longer get a slot);
- every profile’s filter chain, so an IP allowlist never has to be widened for your load balancer;
- the
Replay-Noncemiddleware, so probing costs no database write; and - the
Link: rel="index"middleware.
GET / redirects to /health.
Metrics
GET /metrics serves the OpenMetrics text format
(application/openmetrics-text; version=1.0.0), which Prometheus reads
natively. It is off by
default and lives on a listener of its own — a third socket beside the
ACME and admin ones, configured by
[metrics]:
[metrics]
enabled = true
bind_address = "127.0.0.1:3002"
The separate port is the access control. A scrape carries no credential and none is checked: reaching the port at all is the permission, so restrict it the way you restrict any other internal service — a firewall rule, a network policy, or leaving it on loopback and scraping from the same host. The endpoint is absent from the ACME and admin sockets entirely, so exposing the ACME listener to the internet does not expose this.
$ curl -s localhost:3002/metrics
# HELP acme_proxy_requests Requests served, by endpoint, matched route and response status.
# TYPE acme_proxy_requests counter
acme_proxy_requests_total{role="acme,admin,worker",profile="default",route="/newOrder",status="201"} 42
acme_proxy_requests_total{role="acme,admin,worker",profile="none",route="/health",status="200"} 8613
# HELP acme_proxy_request_duration_seconds Time to answer a request, by endpoint and matched route.
# TYPE acme_proxy_request_duration_seconds histogram
# UNIT acme_proxy_request_duration_seconds seconds
acme_proxy_request_duration_seconds_sum{role="acme,admin,worker",profile="default",route="/newOrder"} 0.84
acme_proxy_request_duration_seconds_count{role="acme,admin,worker",profile="default",route="/newOrder"} 42
acme_proxy_request_duration_seconds_bucket{role="acme,admin,worker",le="0.005",profile="default",route="/newOrder"} 3
…
acme_proxy_request_duration_seconds_bucket{role="acme,admin,worker",le="+Inf",profile="default",route="/newOrder"} 42
# HELP acme_proxy_certificates_issued Certificates signed, by endpoint.
# TYPE acme_proxy_certificates_issued counter
acme_proxy_certificates_issued_total{role="acme,admin,worker",profile="default"} 41
# HELP acme_proxy_certificate_issue_failures Issuance attempts the CA refused, by endpoint and ACME problem type.
# TYPE acme_proxy_certificate_issue_failures counter
acme_proxy_certificate_issue_failures_total{role="acme,admin,worker",profile="default",reason="badCSR"} 1
# HELP acme_proxy_certificate_issue_duration_seconds Time from an accepted finalize to a stored certificate, by endpoint.
# TYPE acme_proxy_certificate_issue_duration_seconds histogram
# UNIT acme_proxy_certificate_issue_duration_seconds seconds
acme_proxy_certificate_issue_duration_seconds_sum{role="acme,admin,worker",profile="default"} 45.0
acme_proxy_certificate_issue_duration_seconds_count{role="acme,admin,worker",profile="default"} 41
…
# HELP acme_proxy_database_pool_connections Connections in the database pool.
# TYPE acme_proxy_database_pool_connections gauge
acme_proxy_database_pool_connections{role="acme,admin,worker",state="idle"} 4
acme_proxy_database_pool_connections{role="acme,admin,worker",state="busy"} 1
# EOF
A counter’s # TYPE line names its family without _total, as OpenMetrics
requires; its series, which is what a query names, keep the suffix.
| Metric | Type | Labels |
|---|---|---|
acme_proxy_requests_total | counter | role, profile, route, status |
acme_proxy_request_duration_seconds | histogram | role, profile, route |
acme_proxy_certificates_issued_total | counter | role, profile |
acme_proxy_certificate_issue_failures_total | counter | role, profile, reason |
acme_proxy_certificate_issue_duration_seconds | histogram | role, profile |
acme_proxy_database_pool_connections | gauge | role, state |
Six things are worth knowing about the numbers.
role names the roles that process runs, comma-separated —
acme,admin,worker for an all-in-one deployment, which is the default. The
counters are per-process memory by design, so a
split deployment
is several scrape targets reporting the same family names, and this is what
tells them apart. Each process also needs its own metrics.bind_address.
route is the matched route pattern, not the URI. A request for
/profile/le/order/9f3c… is counted under route="/order/{id}", and the
endpoint it reached becomes the profile label. Anything that matched no route
at all — a scanner, a typo — collapses into a single route="<unmatched>"
series rather than one series per path tried. Root-router requests such as
/health carry profile="none".
reason is the ACME problem type the CA refused with (badCSR,
serverInternal), the same vocabulary
acme-proxy audit list --event certificate_issue_failed prints.
Both are rendered from one record, so the metric and the trail cannot disagree
about what happened.
The two histograms time different things.
acme_proxy_request_duration_seconds runs from the request reaching the router
to its response head, with buckets from 5 ms to 10 s; it has no status label,
because a histogram multiplies every label by its bucket count.
acme_proxy_certificate_issue_duration_seconds runs from the finalize request
being accepted to the certificate being stored, whichever process signs it, with
buckets from 1 s to 1 h. It is measured from job timestamps, which are whole
seconds, and for a relayed order it includes the
upstream CA’s own validation. It is observed by the process that signs, so in a
split deployment it is on the worker’s scrape target.
The counters survive a reload but not a restart. SIGHUP rebuilds the
routers and keeps the registry, so a configuration change does not read as a
counter reset; a restart genuinely is a new process and starts from zero, which
is what rate() expects.
The pool gauge is read from the pool. state="idle" is connections the pool
holds but has not checked out and state="busy" those in use, so
busy is in-flight database work rather than a number this server maintains in
parallel.
A minimal scrape configuration:
scrape_configs:
- job_name: acme-proxy
static_configs:
- targets: ['ca.internal:3002']
A dashboard over all six families ships in the repository — see Grafana Dashboard.
Logging
acme-proxy uses the tracing crate, configured by the six keys in
[logging] — filter, JSON or
human-readable output, stdout or stderr, ANSI colour, span timing, and
whether JSON fields sit at the top level. Every one of them is validated at
startup: an unknown value is a refusal to start with a message naming the key,
never a silent fallback.
Two of them are worth setting deliberately in production:
[logging]
# Structured fields become first-class keys rather than text to be re-parsed.
json_format = true
# ... and at the top level, which is what most pipelines want.
flatten_event = true
The rest of this page is about what those records contain.
Log levels
RUST_LOG=acme_proxy=info # operational logging (recommended for production)
RUST_LOG=acme_proxy=debug # per-request detail, challenge validation steps,
# database writes, truncated response previews
# from http-01
There is no trace level to reach for: the crate emits nothing below debug.
RUST_LOG replaces the whole filter, so a bare RUST_LOG=debug also turns on
debug logging for every dependency. acme-proxy does not configure SQL
statement logging itself; to see sqlx’s own statement logs you must ask for
them explicitly, e.g. RUST_LOG=acme_proxy=info,sqlx=debug.
Everything on this page is the server’s log stream. The admin commands emit
none of it unless asked, with
--log-level or a non-empty RUST_LOG, and what they
then emit goes to stderr rather than into the output a script is parsing.
Request correlation
Every request passes through one server-wide middleware that reads an incoming
x-request-id header or generates a UUID v7 when absent, opens the request
span, and echoes the id back on the response. Every log line emitted while
handling that request is nested under that span and carries the id, so one
client’s
failing renewal can be pulled out of a busy log in a single query — and a
reverse proxy that already assigns request ids will have its value preserved
rather than replaced. A header whose bytes are not valid ASCII cannot be echoed
back, so it is replaced by a generated id and a request_id_header_invalid line
at debug says so.
The span carries method, uri, version, request_id, client_ip and —
for a request that reached an ACME endpoint rather than /health — profile.
client_ip is the peer address of the connection, replaced by the address
resolved from filter.forwarded_header where filter.trusted_proxies says the
peer is a reverse proxy. Both are per-profile settings, so a request that never
reaches a profile — /health, the http-01 responder, anything admission
control sheds — always shows the peer. A request arriving over a socket with no
peer address at all leaves the field absent rather than reporting a placeholder.
The access line
Each request closes with exactly one request_completed record carrying
status and latency_ms. It is emitted under this crate’s own target, so it is
visible at the default filter. Its level varies:
| Case | Level | outcome |
|---|---|---|
| Response status is 5xx | warn | failure |
GET/HEAD of /health or / | debug | success |
| Everything else | info | success |
This is the only event name emitted at more than one level, and the only one
whose outcome comes from the response rather than from its own name. Liveness
probes are at debug deliberately: at one probe per second per node they would
otherwise be most of the log. Set RUST_LOG=acme_proxy=debug to see them.
Structured events
Every log record carries two fields you can build alerting on, and their shape
is enforced by a test rather than by convention (tests/logging_convention.rs).
event is a stable, greppable name, shaped <subsystem>_<object>_<outcome>
— always a literal in the source, so a name in a log greps straight back to the
line that wrote it. The subsystem prefix comes from a closed list, so a family
grep finds the whole family: event = "certificate_revoke reaches every
revocation refusal, db_ reaches the storage layer, challenge_http_01_
reaches one validator. Challenge types are always spelled with separators
(http_01, dns_01, tls_alpn_01).
outcome is one of four values, and it exists so that “show me everything
that broke” is an exact match instead of a suffix glob. Failure is spelled a
dozen ways across the several hundred event names (_failed, but also
_invalid, _mismatch, _missing, _unauthorized, _rejected), so globbing
on the name silently misses most of it:
outcome | Meaning |
|---|---|
success | The operation completed. |
failure | The operation did not. This is the field to alert on. |
progress | An operation began or was asked for; its result is not yet known. Always paired with a _started or _requested name. |
advisory | Nothing failed, but the operator should keep seeing it — a configuration posture like tls_disabled or challenge_validation_bypassed. |
Every error-level record is outcome = "failure"; an advisory sits at warn.
The events worth building alerts on:
| Event | Level | Meaning |
|---|---|---|
request_completed | info / warn / debug | One per request. See “The access line” above for the level. |
server_listening, profile_mounted | info | Startup completed; one profile_mounted per enabled profile. server_listening is repeated by a reload that moved this socket, carrying the address it moved to. A reload emits profile_mounted only for an endpoint that was not previously served. |
profile_unmounted | warn | A reload stopped serving an endpoint. Its accounts and orders stay in the database; an issuance still waiting on an upstream has no handler left to finish it. |
server_fatal_error, profile_init_failed | error | The process is not serving. |
server_socket_bind_failed, admin_socket_bind_failed, metrics_socket_bind_failed | error | The ACME, admin or metrics socket could not be bound. At startup the process does not serve; on a reload nothing is applied and whatever was already listening keeps listening. |
metrics_listening | info | The metrics listener started. Its message is the reminder that the endpoint is unauthenticated by design, so the port must be firewalled. |
metrics_config_invalid | error | metrics.bind_address names the same socket as the ACME or admin listener. The process did not start. |
server_listener_stopped | info | A listener was switched off by a reload (admin.enabled or metrics.enabled). The socket is released; established connections finish. |
server_config_reloaded | info | A SIGHUP applied. Carries generation, which counts from 1 and rises by one per reload — the quickest check that one landed — and listeners_rebound, naming any socket that moved. |
server_config_reload_refused | warn | The new file changes a key that cannot change while the process runs. Nothing was applied; the error field names the key. |
server_config_reload_failed | error | The new file did not load, or what it asks for could not be built. Nothing was applied. |
server_logging_filter_overridden | warn | A reload changed logging.filter while something outranked it, so the edit had no effect. source names which: "flag" for a --log-level on the server’s own command line, "env" for RUST_LOG. Both win on a reload exactly as they do at startup; drop the one named and reload again. |
db_migration_failed | error | Startup aborted before serving. |
request_shed | warn | A request was refused with 503 + Retry-After: 5 because server.max_concurrent_requests was saturated for longer than admission_wait_ms. Sustained occurrences mean the limit is too low, or something is retrying hot. |
request_deadline_exceeded | warn | A request exceeded server.request_timeout_ms. |
request_handler_panicked | error | A route handler panicked. The client got a 500 in the listener’s normal error shape rather than a dropped connection, and the process is unaffected — but a handler panic is always a bug. Alert on it. The listener (acme / admin) and, for the panel, surface (api / ui) fields say where; the panic message is in error. |
filter_request_blocked, filter_denied | warn | The filter policy refused a request. Expected in normal operation; a spike is either an attack or a policy change that broke a legitimate client. listener = "admin" on filter_request_blocked marks an [admin.filter] refusal. |
filter_rule_warned | warn | A mode = "warn" rule matched and did not decide. This is the line a dry-run rollout is watched on: when it stops appearing for legitimate clients, the rule is safe to switch to enforce. |
challenge_validation_failed, challenge_failed, challenge_validation_timeout | warn | Domain-control validation did not pass. The most useful signal that clients are misconfigured — or that egress to them is blocked. The timeout variant means nothing answered within challenge.timeout_ms at all. |
challenge_http_01_mismatch | warn | The responder answered, but with the wrong key authorization. The body itself is never logged here; a truncated preview goes to challenge_http_01_mismatch_body at debug. |
challenge_validation_abandoned | warn | The queue gave up on a validation — the attempts ran out, or the authorization expired under it. The challenge, its authorization and its order are marked invalid so the client stops polling. Distinct from challenge_failed, which means the check ran and the client’s setup did not satisfy it; this one means the check never reached a verdict. |
challenge_validation_enqueue_failed | error | A challenge was claimed but its validation job could not be written. The claim is released, so the client may trigger again; the request answers 500. Alert on it — it means the database refused a write on the ACME path. |
server_role_no_worker | warn | This process runs no worker role, so nothing in it drains the job queue and it holds no signing backend. Expected in a split deployment; a deployment where no process runs one issues nothing at all, because challenge validation, issuance, CRL signing, notifications and the sweeps are all queued work. |
server_schema_behind | error | A process that does not own the schema found migrations unapplied and refused to serve. Run acme-proxy migrate, or start the worker role, before the others. |
nonce_replayed | warn | A JWS carried a nonce that was unknown, already consumed or expired. Routine in small numbers (a client racing itself); a flood is a client stuck in a retry loop, or a replay attempt. |
key_change_rejected | warn | POST /keyChange refused. reason = bad_signature means the inner JWS did not verify — somebody attempted a rollover they could not prove possession for. |
order_finalize_queued | info | finalize accepted a CSR, claimed the order and queued its signing. Logged by the process serving ACME; the outcome is logged by the worker that signs. |
local_ca_leaf_issued, order_finalized | info | A certificate was issued. Logged by the worker. |
order_finalize_bad_csr, order_finalize_issuance_failed | warn / error | The signer backend rejected the CSR (the order goes invalid with badCSR), or failed to sign (the job retries). |
server_reload_supervisor_gone | error | A SIGHUP arrived but nothing is left to act on it — the process is shutting down, or the reload supervisor died. The configuration on disk is not in effect; restart to apply it. |
order_finalize_authority_withdrawn | warn | The account, or one of the order’s authorizations, was deactivated after finalize queued the signing. The order goes invalid with unauthorized and nothing is signed. |
order_finalize_abandoned | warn | The queue gave up on an issuance — the attempts ran out, or the order expired under it — and marked the order invalid so the client stops polling. Carries the last reason. A steady stream means the backend is down; alert on it. |
certificate_revoked, certificate_revoke_signer_failed | info / error | Revocation succeeded, or the signer refused it — in which case the order is left un-revoked for a retry. |
local_ca_crl_republished | info | A revocation acme-proxy order revoke recorded without the CA key was signed into the CRL by the local_ca_crl_regenerate job. Carries the CA’s issuer id. |
certificate_revoke_queued | info | A revocation for a relay or custom backend was queued for the worker by a process that holds no backend — revokeCert, the panel or the CLI — which then waits on the job. Carries order_id and job_id. |
certificate_revoke_abandoned | error | A queued revocation for a relay or custom backend was given up after its attempts ran out: the certificate is still trusted. Carries order_id and the last reason. Alert on it. |
local_ca_crl_not_stored | error | GET /crl found no stored CRL for a CA. The worker stores one before it serves, so this means no worker has run with this CA yet — start one. |
local_ca_crl_pruned | info | Revocation entries whose certificates had expired were dropped from the CRL (RFC 5280 §3.3). Carries rows_removed and the CA’s issuer id. Silent when nothing expired, which is most days. |
local_ca_crl_prune_failed | error | The daily CRL refresh failed for one CA — the database, or a CRL that could not be signed. Nothing is lost and the refresh tries again tomorrow, but the refresh is also what re-signs a CRL before its nextUpdate, so two days of this in a row deserve a look. |
local_ca_ledger_imported | info | A CA met the database for the first time and imported the JSON ledger it kept beside crl_path before. Carries rows_imported and the new crl_number. Logged once per CA; the sidecar is not read again. |
local_ca_crl_initialization_failed | error | A CA could not store its first CRL — usually a sidecar that will not import, named in error. That CA’s revocations and GET /crl fail until it is fixed, rather than proceeding without the history the sidecar holds. |
local_ca_crl_export_failed | error | The CRL was stored and is served, but writing its copy to crl_path failed. Only a static publication of that file is stale; the daily refresh writes it again. |
http_01_token_store_failed | error | The relay could not publish, retract or look up an http-01 key authorization in the database. A failed publish retries the relay attempt; a failed lookup answers the upstream’s fetch 500. |
upstream_relay_succeeded, upstream_relay_failed | info / warn | Outcome of one relayed issuance under the relay signer backend. |
job_run_completed | info | One background job finished. Carries job_kind, the attempt it succeeded on, and duration_ms. |
job_run_retried | warn | A job failed in a way that may not recur and went back in the queue. Carries the reason and the next run_at. Routine in ones; a steady stream of the same job_kind means whatever it talks to is unwell. |
job_run_abandoned | error | A job was retired permanently — the handler refused it, the attempts ran out, or its deadline passed. For a signer_relay_issue job this is the moment the client’s order goes invalid, so it is the line to alert on. |
job_run_panicked | error | A job handler panicked. The job is retried straight away rather than holding its lease, but this is always a bug — alert on it. |
db_job_leases_reclaimed | warn | Rows whose runner died holding the lease were returned to the queue. Expected once after an unclean shutdown; recurring means the runner is being killed mid-job. |
job_lease_lost | warn | A job finished after its lease had already been reclaimed, so its result was discarded and another runner will repeat the work. Means an attempt is overrunning jobs.lease_seconds. |
job_deadline_passed | warn | A job was claimed after its own deadline and retired without running. For a relay that means the local order had already expired. |
job_runner_retuned | info | A configuration reload moved the runner’s pacing, and it is now running under the new values — which server_config_reloaded alone does not tell you. Carries all five: poll_interval_ms, lease_seconds, retry_base_seconds, retry_max_seconds and max_concurrent. Silent when a reload leaves [jobs] alone. |
job_runner_started, job_runner_stopped | info | The queue runner’s lifecycle. job_runner_stopped carries how many leases it released on the way out; a missing one after a restart is why work waits out a lease instead of resuming immediately. Note the six table sweeps and every notification run through this runner, so a runner that is not started is a server that is not sweeping or notifying either. |
upstream_bad_nonce_retry | debug | Normal ACME churn against the upstream; only interesting in bulk. |
notify_delivery_failed | warn | One delivery attempt did not land. Never affects the ACME response, and no longer the end of the story: it carries retryable, and a true there means the delivery went back in the queue. |
notify_delivery_target_missing | warn | A queued delivery names a profile or backend this process does not have. It is retried rather than dropped, since another process over the same database may be on a newer configuration; a target that really is gone ends in notify_delivery_abandoned once the attempts run out. |
notify_delivery_abandoned | warn | A notification was given up on — the attempts ran out, or a backend reported a failure that could never succeed. This is the line that means an operator was not told something. Alert on it; notify_delivery_failed on its own is usually just a bad minute. |
notify_delivered | info | One delivery landed. Carries the backend, the event kind and the attempt it succeeded on. |
notify_delivery_queued | info | A delivery was written to the queue, one line per backend that wanted the event. Carries a delivery_id shared by that event’s rows, which is how they are correlated. |
tls_handshake_timeout, tls_handshake_failed | debug | Only with server.tls.enabled. Deliberately below the default filter: on a public listener these are scanner background noise, and one warn per failed handshake is a flood, not a signal. |
nonce_reaper_swept | debug | The periodic nonce cleanup ran. It is a nonce_sweep job, so its absence over a long window means the job runner is unwell — check job_runner_started. |
audit_write_failed | warn | An audit row could not be written. The failure is swallowed deliberately — a certificate the CA has already signed must not become a 500 the client retries into a second issuance — so this line is the only evidence that the trail has a hole in it. Alert on it. |
audit_reverse_dns_failed, audit_reverse_dns_timeout | debug | A PTR lookup for a client address found nothing in time. Costs a NULL in one column, never a refused request. Routine where no reverse zone exists; turn audit.reverse_dns off there. |
job_retention_sweep_failed | warn | The daily sweep of finished jobs past jobs.retention_days failed. Nothing is lost; the table grows until the next one succeeds. |
notify_expiry_digest_failed | error | The [notify.expiry] digest for one profile could not be built — usually the database. Carries profile. The digest is rescheduled at its interval regardless, so one failure costs one digest, not the schedule. |
audit_reaper_swept, audit_reaper_failed | debug / warn | The daily retention sweep, only with a non-zero audit.retention_days. audit_reaper_swept carries the rows removed and the cutoff. Runs as the audit_sweep job. |
order_reaper_swept, order_reaper_failed | info / error | The daily order-retention sweep, one line per profile, carrying that profile’s rows removed and cutoff. Runs as the order_sweep job, and only for profiles with a non-zero order.retention_days. It deletes expired, non-valid orders and cascades to their authorizations and challenges; a valid order is never swept, so revocation and renewal information stay available for every certificate actually issued. |
ipam_netbox_tls_verification_disabled | warn | Emitted on every start while insecure_skip_verify is set, deliberately not once-only. |
proxy_configured | info | Emitted once at startup when [proxy] resolves to anything, and not at all otherwise. Carries source (config, environment or both), the two proxy URLs with any password redacted, and the no_proxy rule count. Worth reading on a first start: an inherited shell https_proxy is otherwise an invisible reason for every outbound call to behave differently. |
Only with [admin] enabled — see Web Admin:
| Event | Level | Meaning |
|---|---|---|
admin_listening, admin_origin_resolved | info | The web admin started. admin_origin_resolved carries the origin the CSRF check will compare against, and the resolved bind address — check these agree with how you actually reach the panel. |
admin_config_invalid | error | [admin] cannot work; the process did not start. The message names the two keys that disagree. |
admin_tls_init_failed, admin_filter_init_failed | error | [admin.tls] or [admin.filter] could not be built — an unreadable certificate, a policy that does not parse. At startup the process does not serve; on a reload nothing is applied. |
admin_no_users | warn | The panel is enabled but has no operators — a running service with no way in. Fix with acme-proxy admin user create. |
admin_login_succeeded | info | Carries username and client_ip. |
admin_login_failed | warn | Carries reason: wrong_password, unknown_user, account_disabled or rate_limited. The client is told none of this — every failure returns one invalid_credentials — so this log line is the only place the distinction exists. A run of unknown_user from one address is somebody guessing usernames; a run of rate_limited is a brute-force attempt, or an operator locked out by their own retries. |
admin_logout | info | Carries scope = "one" or "all", and surface = "api" or "ui". |
admin_mfa_enrolled | info | An operator confirmed a second factor; carries username. |
admin_mfa_recovery_codes_regenerated | info | A fresh set of recovery codes was minted, from the panel or admin user totp recovery-codes. Carries username and minted. |
admin_mfa_step_up_refused | warn | A sensitive change asked for the operator’s password again and did not get it. reason is wrong_password or rate_limited; a run of them on a signed-in session is somebody holding a session who does not know its password. |
admin_password_hash_unreadable | warn | A stored hash could not be decoded. The account is unusable until acme-proxy admin user passwd rewrites it, and nothing else will tell you. |
db_admin_user_created, db_admin_user_deleted, db_admin_user_password_changed, db_admin_user_status_changed, db_admin_user_role_changed, db_admin_user_contact_changed | info | The operator audit trail. db_admin_user_contact_changed carries cleared (whether the notification address was removed rather than set). |
db_admin_sessions_revoked, db_admin_session_deleted | info | Sessions ended, by a password change, a disable, or an explicit revoke. db_admin_sessions_revoked carries scope: user, user_except_current (what a password change does, so the operator making it is not logged out by their own action) or all. |
admin_eab_created, admin_eab_revoked, admin_eab_deleted | info | Carries the kid and the operator who did it — never the secret. admin_eab_deleted carries accounts: keep, deactivate or delete. |
db_account_delete_blocked, db_order_delete_blocked, db_eab_delete_blocked | info | An operator delete was refused because it would have removed the only record of a live certificate. Carries live_certificates. Nothing was changed. |
admin_order_revoked, admin_order_revoke_queued, admin_order_deleted, admin_account_contact_updated, admin_account_deactivated, admin_account_deleted, admin_nonces_cleaned, admin_session_revoked | info | Admin writes, each naming the operator. Each carries surface = "api" or "ui": the JSON API and the HTML panel run one shared action, so the same write logs the same line whichever surface it came through. |
admin_operator_role_changed, admin_operator_contact_updated | info | An operator’s privilege tier (carries role) or notification address (carries contact_set) was changed in the panel. username is who made the change and target_username whose it was — the same name when an operator edits their own address. Carries surface. |
admin_job_cancelled, admin_job_advanced, admin_job_revived | info | An operator acting on the background queue through /api/jobs or /ui/jobs. (The acme-proxy jobs subcommand writes the matching audit rows but emits no log line: a non-serve invocation installs no subscriber.) Each carries surface and the operator. admin_job_cancelled on a signer_relay_issue job also carries order_abandoned = true and the order_id — the ACME order was marked invalid and the upstream mapping abandoned. admin_job_revived carries attempts (set to max_attempts - 1). |
admin_revoke_signer_failed | error | The CA-side revocation failed, so the order is left un-revoked for a retry. Answered as 502 rather than 500. |
admin_db_error | error | A database failure on an admin route. The sqlx message is here and deliberately not in the response body, which says only “internal error”. |
admin_session_reaper_swept, admin_session_reaper_failed | debug / error | The periodic session sweep ran, or failed, as the admin_session_sweep job. A failure leaves expired sessions in the table, but they are refused on use regardless. |
admin_session_orphaned | warn | A session outlived its user despite the FK cascade. Should be impossible; the session is deleted and refused. |
The full set is much broader than this table — several hundred names, of which
the ones above are the curated subset worth an alert. Because the subsystem
prefix is drawn from a closed list, an ad-hoc investigation can grep a whole
family: event = "certificate_revoke for every revocation refusal,
event = "replaces_ for RFC 9773 correspondence, event = "db_ for the
storage layer.
Suggested alerts
Two sources, and they are complementary rather than alternatives. The metrics are cheap to alert on and answer “how much”; the log stream answers “which one, and why”, and covers everything the metrics do not.
With [metrics] enabled:
# Issuance is failing, whatever the reason.
rate(acme_proxy_certificate_issue_failures_total[15m]) > 0
# Nothing has been issued over a window longer than your shortest renewal
# interval -- silent breakage, and the one nothing else will tell you.
increase(acme_proxy_certificates_issued_total[6h]) == 0
# The server is shedding load: undersized, or a client is in a retry loop.
rate(acme_proxy_requests_total{status="503"}[5m]) > 0
# 5xx as a share of everything: this server failing, not clients being refused.
sum(rate(acme_proxy_requests_total{status=~"5.."}[5m]))
/ sum(rate(acme_proxy_requests_total[5m])) > 0.01
# The pool is saturated, so requests are queueing on a connection.
acme_proxy_database_pool_connections{state="idle"} == 0
# Requests are slow: the 95th percentile over a second, on any route.
histogram_quantile(0.95,
sum(rate(acme_proxy_request_duration_seconds_bucket[5m])) by (le, route)) > 1
The up metric Prometheus synthesises per target also gives a liveness alert
for free, though /health remains the better probe for a load balancer since it
is on the ACME listener itself.
From the log stream, which reaches further. The first is the general case and subsumes most of the rest; the others are worth splitting out because each has a different response:
outcome = "failure"above its usual rate → something broke. This is the one query that catches every refusal, whatever the event is called, and it is the reason the field exists.request_shedabove zero for more than a few minutes → the server is undersized, or a client is in a retry loop.challenge_validation_failedrising against a flatlocal_ca_leaf_issuedrate → clients are asking and failing; check egress and the responder ports.- Any
upstream_relay_failed→ issuance is broken for every client behind the relay, not just one. - No
local_ca_leaf_issuedat all over a window longer than your shortest renewal interval → silent breakage. - Any
audit_write_failed→ the audit trail is silently incomplete, and nothing else will tell you. See Audit Trail. server_fatal_error,db_migration_failed,profile_init_failed→ page immediately.
Grafana Dashboard
A dashboard over the six metric families the
[metrics] listener exposes ships in
the repository at dashboards/acme-proxy.json. It is a starting point rather
than a fixed artifact — import it, then change whatever your deployment needs.
Enable the metrics listener first; the dashboard has nothing to draw otherwise. See Monitoring for the endpoint and a scrape configuration.
Importing it
Download it from the repository:
curl -O https://raw.githubusercontent.com/acme-proxy/acme-proxy/main/dashboards/acme-proxy.json
Then Dashboards → New → Import in Grafana, upload the file, and pick your Prometheus data source when asked.
For a provisioned Grafana, drop it in the dashboards directory your provider config points at:
# /etc/grafana/provisioning/dashboards/acme-proxy.yaml
apiVersion: 1
providers:
- name: acme-proxy
type: file
options:
path: /var/lib/grafana/dashboards
The data source is a template variable, deliberately, rather than the
__inputs block Grafana’s “export for sharing externally” produces: that block
is never substituted under file provisioning, so a dashboard carrying one works
when a human imports it and silently breaks when a machine does.
What is on it
Fourteen panels in three rows.
Issuance — certificates signed and refused over the dashboard’s time range,
the success ratio between them, the issuance rate split by endpoint, refusals
broken down by ACME problem type, and the p50 and p95 issuance latency by
endpoint. That last panel is the one worth
knowing about: reason says why the CA refused, in the same vocabulary
acme-proxy audit list prints, because both are rendered from one
record.
Requests — request rate by route and by status, the 5xx share, requests shed by admission control, unmatched paths, and the p50 and p95 latency by route.
Database — the database pool by state, and its saturation.
The latency panels are histogram_quantile over the buckets, so a percentile
is only as fine as the buckets around it; see
Monitoring for the bucket bounds and exactly what each
histogram times. Issuance latency is drawn from the process that signs,
which in a split deployment is the worker.
Two things to know before editing it
route is safe to group by. It is the matched route pattern, so
/order/{id} is one series however many orders exist. Grouping by a raw URI
would be a series per order, per account and per challenge — which is memory in
Prometheus for as long as it retains them. The same reason every unmatched
request collapses into a single <unmatched> series rather than one per path a
scanner tries.
The pool panels must not be filtered by $profile. Everything else on the
dashboard is scoped to the profile variable, so reaching for it on the two
database panels is the natural next edit — but
acme_proxy_database_pool_connections is process-wide and carries no profile
label at all. Filtering it matches nothing, and the panel goes blank with no
error to explain why. A test in the repository (tests/grafana_dashboard.rs)
fails if the shipped dashboard ever grows that filter.
Empty panels are not always a fault
Four cases read as “no data” and are all correct:
- Refusals by reason is empty on a CA that has refused nothing. Only
refusals the CA itself made are counted — an order rejected at
newOrder, for a wildcard on an endpoint offering nodns-01for instance, is protocol bookkeeping rather than a CA action, and appears as a403on the request panels instead. - Success ratio and 5xx share show
NaNwhen nothing has happened at all, because the ratio is genuinely undefined. They read0only when something did happen and all of it failed. - Issuance latency is empty on an
acme-only oradmin-only scrape target: only theworkersigns, so only it observes the histogram. - Everything is empty for the first scrape or two after a restart. Counters do
not survive a restart, though they do survive a
SIGHUPreload.
Keeping it honest
tests/grafana_dashboard.rs runs in CI and asserts that every metric the
dashboard queries is one this build actually emits — and the converse, that no
emitted family is missing from it. A histogram counts as shown when any of
its _bucket, _sum or _count series is. The names come from rendering a
real, empty registry rather than from a list somebody maintains by hand, so a
renamed metric fails the build instead of quietly blanking a panel that nobody
looks at until an incident.
Troubleshooting
This section covers common issues when operating acme-proxy in production.
The server refuses to start
Several configuration mistakes are deliberately fatal at startup rather than silently degrading at runtime. The log line names the problem in each case.
| Message | Cause | Fix |
|---|---|---|
no enabled [profiles] | No profile is defined, or all are enabled = false. The server serves ACME only through profiles. | Add [profiles.default], or set ACME_PROXY_PROFILES__DEFAULT__ENABLED=true. |
Unknown or empty challenge.enabled | The list is empty or names a type that does not exist. This is fatal even with challenge.bypass = true — a bypassing server still has to advertise a challenge type. | Use one or more of http-01, dns-01, tls-alpn-01. |
filter.check.<name> has neither allow nor deny entries | An allowed_ip or path check with both lists empty permits everything, so it is a no-op and almost certainly a mistake. | Populate filter.check.<name>.allow or .deny, or delete the check and any rule naming it. |
request_timeout_ms too small | It must exceed signer.custom.timeout_ms when that backend is configured with its crl or renewal_info hook on — those hooks run inline inside a request, so a smaller whole-request deadline would cut them off every time. challenge.timeout_ms and the script’s issue hook are not checked against it: both run in the job queue. | Raise server.request_timeout_ms. |
Two profiles sharing ca.key but differing elsewhere | One CA key under two [signer] configurations would issue under two policies from one identity. | Make the [signer] sections identical (they then share one backend), or give each profile its own CA files. |
| An entry name with an underscore | [filter.check.*], [filter.rule.*], [notify.custom.*], [notify.webhook.*] and profile names must match ^[a-z0-9-]+$. The name is also an environment-variable segment, which the config crate lowercases. | Use hyphens: threat-intel, not threat_intel. |
A custom notify backend enabled with no entries | "custom" is listed in notify.enabled while notify.custom_enabled is empty. | Populate notify.custom_enabled, or drop "custom". |
filter.enabled is no longer a setting | A 0.1.x [filter] section. The flat all-must-pass chain was replaced wholesale by named checks and rules; filter.exempt_paths and filter.custom_enabled are refused the same way. | Declare each filter as a [filter.check.<name>] with a type, write a [filter.rule.<name>] whose when names them, and list the rules in filter.rules. See Filters. |
[filter.rule.<name>] is configured but filter.rules is empty | A rule is declared but nothing selects it, so the policy it describes would never run. Inherited rules count: a profile that leaves filter.rules empty meets this too. | List the rule in filter.rules, or delete it. |
Upstream requires EAB, no .kid sidecar | The relay backend has never registered with the upstream. | Run acme-proxy upstream register --profile <name> --eab-kid …, or set signer.relay.eab.kid/hmac_key in configuration. |
key_source = "pkcs11" … built without | signer.local_ca.key_source = "pkcs11" on a binary with no PKCS#11 support. Deliberately fatal rather than falling back to the file key, which would silently leave the CA key on disk. | Rebuild with cargo build --release --features hsm. |
| A PKCS#11 key that is not the certified one | The token key’s public key does not match cert_path — almost always a wrong key_label. | See Hardware Keys. |
admin.bind_address … is not loopback while admin.tls.enabled is false | The session cookie is always sent Secure, which a browser will not store over plain HTTP on anything but localhost. Refused rather than warned about, because the symptom is otherwise “signing in succeeds and then immediately signs you out” with nothing in any log to explain it. | Set admin.tls.enabled = true, or bind 127.0.0.1 and reach the panel through an SSH tunnel. |
metrics.bind_address and … are both | The metrics endpoint is a listener of its own and cannot share a socket with the ACME or admin one. Checked at startup and on every reload. | Give it its own port. |
notify.webhook.<entry>.body is not a valid template (likewise .url, .method, .headers) | Every webhook entry is compiled and validated at startup: the body template, the URL, a verb outside POST/PUT/PATCH, and a header name or value a wire format will not carry. Deliberately fatal — a broken entry would otherwise become a permanent delivery failure discovered on the first event that mattered. | Fix the entry. See Webhook. |
admin.template_dir … is not a directory, or a template that does not compile | The panel’s per-file overrides are compiled at startup for the same reason, so a broken one refuses to start rather than serving a 500 later. | Point it at a directory, or leave it empty for the compiled-in defaults. See Customizing the Panel. |
A [proxy] URL that is not http://, or a no_proxy entry with a port | proxy.http_url / proxy.https_url name the proxy to dial, which is spelled http://host:port even for HTTPS targets; there is no SOCKS support. A no_proxy entry matches a host or a network, never a port. | Drop the scheme to http://, and the port from the no_proxy entry. |
ipam.phpipam.sources naming vip or fhrp | phpIPAM records neither roles nor redundancy groups, so those two sources exist only for NetBox. Refused by name rather than silently returning nothing. | Use dns_name, custom_field or device. See phpIPAM. |
Failures specific to a hardware CA key — PIN, token, slot and mechanism problems — have their own table in Hardware Keys (PKCS#11).
A wildcard order is rejected
- Symptoms —
newOrderfor*.example.comreturnsrejectedIdentifiernamingdns-01. - Cause — Wildcards can only be proven with a DNS challenge, so
acme-proxyaccepts them only whendns-01is amongchallenge.enabled. - Fix — Add
dns-01tochallenge.enabled. See Challenge Validation.
A client suddenly cannot order anything
- Symptoms — Every order-side request from one client returns
unauthorized. - Cause — The account has been deactivated — by the client itself, or by
acme-proxy account deactivate. Deactivation is permanent and blocks all issuance. - Fix — The client must register a new account.
An order sits in processing and never finishes
- Symptoms — A client polls an order that reached
processingand stays there. Only therelaybackend defers like this;local_caandcustomanswerfinalizeinline. - Cause — The issuance is a background job, and it is either waiting out a backoff after a retryable upstream failure or has been retired for good. The order object itself will not say which.
- Fix — Ask the queue:
# What is the job doing, and how many attempts has it spent?
acme-proxy jobs list --kind signer_relay_issue --limit 20
# The upstream's own error text, and the URLs it was talking to.
acme-proxy jobs show <job-id>
A job still ready is waiting for its next attempt at the run_at it
prints — acme-proxy jobs run-now <id> pulls that forward. A failed one has
spent its budget: run-now grants exactly one more attempt, and
acme-proxy jobs cancel <id> abandons the ACME order so the client stops
polling and can order again. See Job queue, and
job_run_abandoned in Monitoring for the
log line that says it happened.
SQLite database locks
- Symptoms — The server logs show
database is lockederrors during high concurrency order creation. - Cause — The database may not be utilizing Write-Ahead Logging (WAL) or your filesystem does not support proper locking mechanisms (e.g., NFS).
- Fix — Ensure
acme-proxyis running on a local filesystem andjournal_mode = WALis applied. (The server automatically attempts to enable WAL on startup).
Reading the database directly
When the CLI cannot answer a question — usually “what does the row actually say?” — the database is readable while the server runs, because WAL allows a reader alongside the writer.
# Which migrations have run, and did they all succeed?
sqlite3 sqlite.db "SELECT version, description, success FROM _sqlx_migrations;"
# Everything this server refused in the last day, and why.
sqlite3 sqlite.db "SELECT datetime(created_at,'unixepoch'), event, profile, client_ip, reason
FROM audit_log WHERE outcome = 'failure'
AND created_at > strftime('%s','now','-1 day');"
# An order that will not progress: its status and its authorizations'.
sqlite3 sqlite.db "SELECT o.status, a.identifier, a.status, c.type, c.status, c.error
FROM orders o JOIN authorizations a ON a.order_id = o.id
JOIN challenges c ON c.authz_id = a.id WHERE o.id = '…';"
Read, do not write. The
CHECKconstraints will catch an impossible status, but nothing re-syncs the in-memory state a running handler is holding. Use the Admin CLI to change anything.
Backups must include the WAL. Copying sqlite.db on its own gives you a
database missing every recent write. Either take all three files (sqlite.db,
-wal, -shm) with the server stopped, or run sqlite3 sqlite.db ".backup backup.db", which is consistent by construction and safe against a running
server.
The schema itself — every table, constraint and index, and why each is shaped the way it is — is documented in Database Schema.
Upstream Let’s Encrypt rate limits
- Symptoms — Order finalization fails with HTTP 429 Too Many Requests from the upstream CA.
- Cause — When using the
relaysigner backend, all internal clients share a single external ACME account. Let’s Encrypt applies rate limits (e.g., 50 certificates per registered domain per week). - Fix — Request a rate limit increase for the root domain you are relaying, or implement careful caching mechanisms on your internal servers to prevent excessive renewals.
Network challenge blockages
- Symptoms — Order stays in
pendingstate, or challenge verification fails. - Cause — If
challenge.bypass = false, the server must reach the internal client over HTTP (port 80) or DNS to verify domain ownership. Firewalls might be blocking this internal callback. - Fix — Ensure the host running
acme-proxyhas egress network access to reach the internal servers requesting certificates.
EAB registration fails
- Symptoms — Client receives an error stating
External Account Binding is required. - Cause — The client is connecting to a profile that requires EAB, but did
not provide the
kidandhmaccredentials. - Fix — Create EAB credentials using the
acme-proxy eab createCLI command and configure the client to use them.
Order finalization fails (413 payload too large)
- Symptoms — Finalizing the order returns a 413 error.
- Cause — The encoded Certificate Signing Request (CSR) exceeds the
server.max_body_byteslimit (default 128 KiB). - Fix — Ensure your client is not generating excessively large CSRs or increase the limit in your configuration.
Order finalization fails (badCSR: identifier mismatch)
- Symptoms — Finalizing the order returns
400 badCSRcomplaining about the requested identifiers. - Cause — The DNS Subject Alternative Names in your CSR are not exactly the set of names the order authorized. The comparison is set equality, so an extra name fails just as an omitted one does.
- Fix — Configure your ACME client to put every ordered domain — and nothing else — into the CSR’s SANs.
Order finalization fails (badCSR: common name)
- Symptoms —
400 badCSR, “CSR common name is a domain the order does not cover”. - Cause — The CSR’s Subject Common Name looks like a DNS name that the order
does not authorize.
acme-proxydoes not ignore the CN: a domain-shaped CN must be covered by the order, precisely so a name cannot be smuggled past the identifier filters by moving it out of the SANs. - Fix — Either add that name to the order, or drop the CN. A CN that is not
domain-shaped — a human label like
rcgen self signed cert— is tolerated and ignored. Note that thelocal_casigner empties the subject entirely on the issued certificate, so a CN is never carried through to the leaf.
Order finalization fails (badCSR: non-DNS SAN)
- Symptoms —
400 badCSRfrom thelocal_casigner for a CSR containing an IP, email or URI SAN. - Cause —
local_caaccepts DNS SANs only, and rejects anything else outright rather than stripping it. - Fix — Remove the non-DNS SANs from the CSR.
Connection refused or 503 service unavailable
- Symptoms — High-throughput ACME clients receive HTTP 503 errors during bursts.
- Cause — The server’s load shedder activated because the number of
concurrent in-flight requests exceeded
server.max_concurrent_requests. - Fix — Increase
max_concurrent_requestsandadmission_wait_ms, or configure your client to retry with exponential backoff.
Security Model
acme-proxy is a certificate authority, or the thing standing in front of one.
Anything that can make it sign gets a certificate your infrastructure will
trust. This page states the trust boundaries and what each secret protects, and
links to the page that owns each detail.
It is a map, not a manual. Every claim here is explained in full somewhere else; if you need to act rather than orient, go to the hardening checklist.
What has to be true for a certificate to be issued
Four gates stand between a packet arriving and a certificate coming back. They are independent, and each covers something the others do not.
| Gate | Answers | Default | Owned by |
|---|---|---|---|
| Connection filters | May this address talk to this endpoint at all? | off (filter.rules = []) | Filters |
| EAB | Is this client allowed to register an account here? | off | EAB |
| Challenge validation | Does the client actually control the name? | on | Challenge Validation |
| Identifier filters | Is this client allowed these particular names? | off | Filters |
Two of the four are off by default, which makes the third load-bearing:
With
challenge.bypass = trueand an emptyfilter.rules, every client that can reach the socket can obtain a certificate for every name it asks for. That combination is why validation is on by default.
The identifier gate runs twice — at newOrder and again at finalize
against the names actually in the CSR. That is not belt-and-braces: without the
second run, a name could be smuggled past the policy by keeping it out of the
order and putting it in the CSR. See
Filters.
What each secret protects
| Secret | Compromise gets an attacker | Stored |
|---|---|---|
The CA private key (signer.local_ca.key_path) | The ability to mint any certificate your fleet trusts, silently, for as long as the CA is trusted. There is no audit row for a signature made outside this server. | On disk at 0600, created with create_new rather than chmod’ed after the fact. Can live in a PKCS#11 token instead. |
| An EAB HMAC secret | The ability to register accounts at that profile, subject to every other gate. Revocable without a restart. | Retrievable bytes — HMAC verification needs the same secret back each time. See Database Schema. |
The upstream ACME account key (signer.relay.account_key_path) | Control of your account at the upstream CA, including revoking what it issued. | On disk at 0600, beside a .kid sidecar naming the account. |
| An RFC 2136 TSIG key | The ability to write records in the zone it is scoped to. This is the credential the relay exists to not distribute. | Configuration, or ACME_PROXY_SIGNER__RELAY__DNS01__RFC2136__TSIG_KEY_SECRET. |
| A web admin password | Nothing on its own once a second factor is enrolled. Otherwise: an operator session. | One-way KDF, unreadable. |
| A web admin session cookie | An operator session until it expires — but not the ability to change the second factor, which takes the password again. | Only hex(SHA-256(token)) is stored. |
| A NetBox API token | Read access to your IPAM. | Configuration; belongs in the environment variable. |
The CA key is the one whose loss is not recoverable by rotation: every
certificate it signed stays trusted until the CA itself is distrusted
everywhere. That is the argument for hardware-backed
keys, and for keeping an offline root and giving
acme-proxy only an intermediate — see
Local CA.
How often to roll each of these, and how, is Secret Rotation.
Two listeners, two exposure surfaces
The ACME listener and the web admin listener are separate sockets with separate TLS configuration, separate authentication and separate defaults. They are not two paths on one server, and they should not sit on one interface.
- The ACME listener answers unauthenticated clients by design. It carries the filter chain, the admission limiter and the nonce middleware.
- The admin listener is off by default, binds
127.0.0.1by default, and carries no filter chain and no admission control — its availability concern is credential brute force, which the login limiter handles. Access control is the bind address, TLS, and the session.
Startup refuses a non-loopback admin.bind_address while
admin.tls.enabled is false. The session cookie is always Secure, which
browsers silently decline over plain HTTP off localhost; the symptom would be
“login succeeds, then immediately logs out” with nothing in any log.
See Deployment for where each socket belongs relative to a firewall.
Trusting a forwarded address
Every IP-based decision — filters, the login limiter, what lands in the audit trail — rests on which address the server believes the client has.
Behind a reverse proxy that address arrives in a header, and a header is written
by whoever is talking to you. filter.trusted_proxies is the allowlist of hops
whose forwarded-for header is believed; with it empty, the peer address is used
and the header is ignored. Setting the header name without setting
trusted_proxies does not make the header trusted.
The admin listener does no forwarded-header handling at all, deliberately: trusting one without an allowlist would let any caller choose its own rate-limiter key.
See Allowed IP.
The audit trail is the record, and it is not a control
Every issuance and every refusal is written to audit_log with the actor, the
address, that address’s reverse name, the identifiers, the User-Agent and the
request id. Nothing in the server ever compares a live request against any of it
— address pinning breaks CGNAT and mobile clients, and that is a deliberate
non-feature.
Two properties make it evidence rather than logging:
- It survives deletion of its subject.
audit_loghas no foreign keys, so deleting an account or an order does not take its history with it. See Database Schema. - Nothing in the web admin can erase it. The panel’s audit surface is
read-only; pruning is
audit cleanupon the host, or theaudit.retention_dayssweep. A stolen session that could erase the trail would make the trail prove nothing.
See Audit Trail.
Where this server can be made to talk to something else
Three subsystems make outbound connections on behalf of a client’s request, which makes each one a request-forgery surface worth knowing about:
http-01validation follows redirects, because RFC 8555 requires it. Boulder’s mitigation — blocking RFC 1918 targets — does not apply here, since serving private networks is the entire point. What contains it instead: onlyhttp/https, only the two configured ports, at mostmax_redirectshops, a shared timeout, an off switch, and the fetched body is never echoed into the client-visible error. See HTTP-01.- The
customhooks (signer, filter, notify) execute an operator-supplied script with request-derived data in its environment and on stdin. They run withenv_clear(), a minimalPATH, a timeout andkill_on_drop— but the script is yours, and quoting its inputs is your job. - The relay backend talks to an upstream ACME server and, with
dns01, writes DNS records. Unlikehttp-01validation, it validates the upstream’s TLS certificate againstwebpki-roots: there, the certificate is the only thing identifying the CA being handed your CSRs.
What is out of scope
- Availability of the ACME listener beyond the admission limiter. There is no per-account quota and no rate limiting by identifier.
- Confidentiality of issued certificates. They are public objects; the audit trail records who received one.
- CAA. This server performs no CAA lookup — see Protocol Support.
- Protecting the database file from a local root. The SQLite file holds EAB secrets and TOTP secrets in a form the server can read back, which means so can anyone who can read the file. File modes are the boundary.
Hardening Checklist
Run through this before an acme-proxy deployment issues a certificate anything
depends on. Every item links to the page that explains it; nothing here is
explained only here.
The defaults are already the safe end of most of these. The items that need a decision from you are marked decide.
Before it serves anything
-
challenge.bypassisfalse. It is the default. With it on,[filter]is the only thing between a client and a certificate for any name it names. → Challenge Validation - At least one gate is configured — filters, EAB, or both. Validation alone proves the client controls the name; it does not say the client is allowed to have a certificate from you. → Filters, EAB
- decide —
server.bind_addressis the interface you meant. The default[::]:3000is every interface. → Configuration Reference -
server.base_urlmatches how clients actually reach the server, including the scheme. It is checked against the JWSurlof every signed request, so a mismatch fails every request rather than degrading. → Configuration Reference - ACME is served over HTTPS — either
server.tls.enabled = trueor a reverse proxy in front. RFC 8555 §6.1 expects it. → TLS Termination -
acme-proxy filter explainagrees with what you meant, for both a client that should be served and one that should not. A policy is easier to get subtly wrong than a list. → CLI -
/crland/ca.pemare still reachable if any check is address-based. Both are served by the profile router, so an allowlist covers them too — and neither the relying parties that fetch the CRL nor the hosts that have yet to install the root are the ACME clients you allowlisted. → Path Check
An or is a hole you opened deliberately
A check that cannot reach its authority answers “unknown” rather than “no”, and
pass or unknown is pass. That is the point — it is what keeps an inventory
outage from locking every client out — but it means an or weakens the
fail-closed property to whatever its other side says.
when = "mgmt-net or inventory"
reads as “the inventory decides, unless the address is already trusted”. If
mgmt-net is wide, the inventory is decorative for everything inside it. That
may be exactly what you want; what you must not do is write it believing both
checks apply.
The rule of thumb: an or over an address check is a bypass for that
address range, so keep the range as small as the outage you are insuring
against. and has no such property — fail and unknown is fail, so a
conjunction never becomes more permissive because something broke.
The CA key
- decide — the issuing key is an intermediate, not a root. An offline root means a compromise is recoverable by re-issuing the intermediate rather than re-trusting every endpoint. → Local CA
- decide — the key lives in a PKCS#11 token if the deployment justifies it. The key then cannot be copied, only used. → Hardware Keys
-
ca.keyis0600and owned by the service user.acme-proxycreates it that way; a key restored from a backup may not be. - The CRL is reachable by everything that validates your certificates, and the database is backed up — the revocations recorded there are the authoritative record, not the CRL file. → Revocation & CRL
Behind a proxy
-
filter.trusted_proxiesnames the hops you trust, or the forwarded header is ignored entirely. Settingfilter.forwarded_headeralone does not make it trusted. → Allowed IP - The proxy does not forward
/healthif you do not want it public — it is mounted outside the filter chain on purpose. → Monitoring - With
signer.relay.challenge_strategy = "http01", port 80 of every name being issued forwards or redirects/.well-known/acme-challenge/here. Nothing in the process can do this for you. → Relay
The web admin
Skip this section entirely if admin.enabled is false, which is the default.
-
admin.bind_addressis loopback, oradmin.tls.enabledistrue— startup refuses the other combination. → Web Admin - decide — reach it over an SSH tunnel or a VPN rather than exposing the socket. It has no filter chain and no admission control. → Deployment
-
admin.base_urlis the origin operators actually type. It is load-bearing four ways — the CSRF origin check, the generated certificate’s host, and the label an authenticator app shows. - Every operator has a second factor, and
admin.require_mfa = trueso the next one does too. It does not retroactively end sessions that predate it;admin session revoke --alldoes. → Users & Sessions - Recovery codes are stored somewhere that is not the panel. Without
them, a lost authenticator needs
admin user totp reseton the host. → Users & Sessions - No
--passwordanywhere in your provisioning. There is no such flag; the password arrives on stdin or via--password-file. → Users & Sessions
Secrets
- Tokens and TSIG keys come from the environment, not the file where the
option exists —
ipam.netbox.token,signer.relay.dns01.rfc2136.tsig_key_secret. -
[signer.relay.eab]is emptied after the first registration. It is a bootstrap credential that authorizes exactly onenewAccount; the server warns on every startup for as long as it stays set. → Relay -
insecure_skip_verifyis unset. It warns on every startup by design, so it stays visible for exactly as long as it is needed. → NetBox - The database file is
0600. It holds EAB secrets and TOTP secrets in a form the server reads back, so anyone who can read the file can too. → Database Schema
Ongoing
- decide — a rotation interval for each long-lived secret. Nothing in the server expires these on a timer; the page recommends one per secret and names the events that force a rotation early. → Secret Rotation
- Backups copy the WAL.
sqlite.dbalone is missing every recent write; use.backup, or take all three files. → Database Schema - decide —
audit.retention_days.0, the default, keeps everything for ever, which is the right default for a trail whose value is that it is complete. → Audit Trail - Something watches the logs for refusals —
certificate_issue_failedandcertificate_revoke_failedrows, and a run of unknown-certificate revocation attempts, which is somebody enumerating serials. → Monitoring - Startup warnings are read, not filtered out. Three of them repeat on
every start precisely so they cannot become background noise:
challenge_validation_bypassed,ipam_netbox_tls_verification_disabled,signer_relay_eab_secret_in_config. → Monitoring
Secret Rotation
The Hardening Checklist is run once, before a deployment issues a certificate anything depends on. This page is the other half: what to do on a schedule after that. Every secret in the Security Model’s inventory is here, with a recommended interval and the events that force a rotation early.
The intervals are a starting point, not a protocol requirement — adjust them to
what the deployment is worth to an attacker. Nothing in acme-proxy expires
these on a timer: rotation is an operator action, except for the session cookie,
which ages out on its own.
The schedule
| Secret | Rotate every | Rotate now when | How |
|---|---|---|---|
| The CA issuing key | not on a timer — see below | the root or intermediate key may have been exposed; someone who could read it leaves | re-issue the intermediate from the offline root — Local CA |
| An EAB HMAC secret | 12 months, per credential | the client system is rebuilt; the person who held it leaves | acme-proxy eab revoke <kid>, then eab create for the replacement — CLI |
| The upstream ACME account key | not on a timer | the disk may have been read; the relay profile is being decommissioned | register a fresh upstream account — new account_key_path, delete the .kid sidecar, restart — CLI |
| An RFC 2136 TSIG key | 12 months, with whoever runs the zone | anyone with the zone’s write path leaves; an update you did not make appears in the nameserver log | add the new key on the nameserver, update signer.relay.dns01.rfc2136.tsig_key_secret in the environment, SIGHUP — Relay |
| A web admin password | not on a timer, by design | it was shared, phished, or typed into the wrong window; the operator leaves | acme-proxy admin user passwd <user> — it also ends every session that user holds — Users & Sessions |
| A web admin session cookie | rotates itself — admin.session_ttl_seconds (12 h absolute), and the idle timeout | a laptop is lost; a session is suspected stolen | acme-proxy admin session revoke --user <u>, or --all — Users & Sessions |
| A TOTP secret | not on a timer | the authenticator device is lost or replaced | acme-proxy admin user totp reset <user> — it takes the recovery codes and the sessions with it — Users & Sessions |
| Recovery codes | regenerate when few remain | one has been used in anger; the printed or stored copy is exposed | acme-proxy admin user totp recovery-codes <user> — Users & Sessions |
| An IPAM API token | 12 months, or whatever the IPAM’s own policy says | the token appears in a log or a ticket; an operator with IPAM access leaves | issue a new read-only token, update ipam.netbox.token (or ipam.phpipam.token) in the environment, SIGHUP — NetBox |
Only the session cookie expires on its own — an absolute lifetime and an idle timeout, whichever comes first. The password deliberately has no forced periodic change: an operator’s stored hash is re-encoded on their next login when the KDF cost rises, so there is nothing a scheduled reset would achieve. Everything else is a manual cadence because nothing revokes it for you.
The CA key is the exception
Losing the CA key is not something rotation recovers from. Every certificate it signed stays trusted until the CA itself is distrusted everywhere, and there is no audit row for a signature made outside this server. So the practice for this one key is structural rather than scheduled.
- Give
acme-proxyan intermediate, not a root, and keep the root offline. A compromise of the online key is then recoverable by re-issuing the intermediate rather than by re-trusting every endpoint in the fleet. See Local CA and its security constraints. - Re-issue the intermediate from the root before it expires. Doing it on a
calendar, well ahead of the
notAfter, keeps the recovery path exercised rather than theoretical. - Rotating the root is a multi-year event. Distribute the replacement long before it is needed, run both roots in the trust store, and remove the old one only once nothing is signed by it. See Trusting the CA.
- An actual key compromise is a distrust-and-reissue event, not a rotation. Publish the revocation, pull the root, and re-issue what mattered under the new one. See Trusting the CA.
Rotating a secret that lives in the environment
The TSIG key and the IPAM token belong in environment variables rather than in
config.toml, and so does the relay’s bootstrap EAB secret until the first
registration clears it — the Hardening Checklist says
which. Rotating one of these is the same three steps every time.
- Stage the new value on the system that backs it — a second TSIG key on the nameserver, a fresh token in NetBox or phpIPAM — so that both the old and the new one work for a moment.
- Update the environment and send
SIGHUP(orsystemctl reload). The[signer]and[ipam]sections both reload with no restart and no dropped request. See Reloading the configuration. - Confirm from the logs, or from a test issuance, that the new credential is the one in use, then retire the old value on the far side.
The EAB HMAC secrets and the TOTP secrets are held in the database in a form the server reads back, so a rotation does not remove the need to protect the file itself — file mode is the boundary. See Database Schema.
ASVS 5.0 Assessment
A self-assessment of acme-proxy against OWASP ASVS
5.0,
whose text is vendored under rfc/asvs-5.0/ in this repository.
- Assessed at: Level 2. Every L1 and L2 requirement in scope is given a status; L3 requirements are listed too, as information rather than as a bar being claimed.
- Assessed against: the tree at release 0.6.0.
- Method: source review. The evidence column names a file, not a promise — for a control requirement, the documentation on this site is context and the code is the evidence.
This is a self-assessment and not a certification. ASVS is explicit that a verification claim means an assessor performed the work; nobody outside the project has performed it here. Read this page as the maintainers’ own answer to “which recognised controls does this meet, and where does it fall short”, which is a useful thing to have written down and a different thing from an audit.
What was assessed
Three surfaces, kept apart because their controls genuinely differ:
| Surface | What it is | Where |
|---|---|---|
| ACME listener | Unauthenticated by design, authenticated per request by JWS. Carries the filter chain, admission control and the nonce middleware. | crates/server/src/, crates/protocol/src/handlers/, crates/protocol/src/extractors/, crates/protocol/src/middlewares/ |
| Web admin | The only session-based, browser-facing surface. Off by default, loopback by default. | crates/admin/src/webadmin/, crates/admin/src/admin/ |
| CLI and process | Answers to a shell on the host and holds no session. | src/cli/, src/main.rs, crates/core/src/config/ |
Most of V3, V6 and V7 apply only to the web admin. When it is disabled — which is the default — those chapters have no surface to apply to at all.
Chapter disposition
| Chapter | In scope | Note |
|---|---|---|
| V1 Encoding and Sanitization | yes | |
| V2 Validation and Business Logic | yes | |
| V3 Web Frontend Security | yes | Web admin only |
| V4 API and Web Service | yes | GraphQL and WebSocket sections are n/a |
| V5 File Handling | partly | There is no upload feature; the sections that assume one are n/a |
| V6 Authentication | yes | Web admin and the CLI credential lifecycle |
| V7 Session Management | yes | Web admin only |
| V8 Authorization | yes | |
| V9 Self-contained Tokens | yes | The self-contained token here is the ACME JWS, not a session JWT |
| V10 OAuth and OIDC | no | No OAuth, no OIDC, no external identity provider, no token endpoint. Nothing in the chapter has a subject |
| V11 Cryptography | yes | |
| V12 Secure Communication | yes | |
| V13 Configuration | yes | |
| V14 Data Protection | yes | |
| V15 Secure Coding and Architecture | yes | |
| V16 Security Logging and Error Handling | yes | |
| V17 WebRTC | no | No WebRTC, no media, no signalling |
Summary
Counts are of the requirements enumerated in the per-chapter tables below; every requirement of every in-scope chapter is present, so the totals match ASVS 5.0’s own counts. L1+L2 is the bar being assessed; the L3 column is reported for information.
| Chapter | L1+L2 met | partial | gap | n/a | L3 (met / short / n/a) |
|---|---|---|---|---|---|
| V1 Encoding and Sanitization | 17 | 1 | 0 | 9 | 2 / 0 / 1 |
| V2 Validation and Business Logic | 11 | 0 | 0 | 0 | 0 / 1 / 1 |
| V3 Web Frontend Security | 16 | 1 | 0 | 2 | 6 / 4 / 2 |
| V4 API and Web Service | 4 | 0 | 0 | 6 | 6 / 0 / 0 |
| V5 File Handling | 4 | 0 | 0 | 5 | 0 / 0 / 4 |
| V6 Authentication | 26 | 1 | 0 | 8 | 8 / 1 / 3 |
| V7 Session Management | 15 | 1 | 0 | 2 | 0 / 1 / 0 |
| V8 Authorization | 7 | 0 | 0 | 0 | 4 / 2 / 0 |
| V9 Self-contained Tokens | 7 | 0 | 0 | 0 | 0 / 0 / 0 |
| V11 Cryptography | 11 | 2 | 0 | 1 | 5 / 2 / 3 |
| V12 Secure Communication | 6 | 1 | 0 | 2 | 0 / 2 / 1 |
| V13 Configuration | 9 | 4 | 0 | 0 | 5 / 3 / 0 |
| V14 Data Protection | 9 | 0 | 0 | 0 | 2 / 1 / 1 |
| V15 Secure Coding and Architecture | 11 | 1 | 0 | 1 | 8 / 0 / 0 |
| V16 Security Logging and Error Handling | 15 | 1 | 0 | 0 | 1 / 0 / 0 |
| Total | 168 | 13 | 0 | 36 | 47 / 17 / 16 |
The short version. There is no L1 or L2 gap. The four password-policy
requirements that used to sit here — V6.2.4 at L1, and V6.1.2 / V6.2.11 /
V6.2.12 at L2 — were one missing control seen from four angles, and closed as
one: check_password_policy now refuses a password that names this deployment,
or that appears in a compiled-in corpus of common passwords. V6.2.2 and V6.2.3,
both L1, closed the same way: one self-service password-change route on the
account page, gated the way the second-factor routes already were, rather than
two separate fixes.
Everything else that falls short of met is either a partial — a control that exists but does not reach everywhere the requirement asks — or a documented deviation, where the project has knowingly chosen otherwise and argued the choice already. Both have their own sections below.
Two chapters deserve a note on their shape. V6 Authentication carries the most n/a rows because the web admin has one authentication pathway and no out-of-band, biometric or federated factors — most of the chapter has no subject here. V16 Security Logging is the only chapter with no L1 requirements at all and is met almost entirely, which is what you would hope for in a certificate authority: the audit trail is the product.
V1 Encoding and Sanitization
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 1.1.1 | Decode into canonical form once, before processing | 2 | met | The JWS protected header and payload are base64url-decoded exactly once in crates/protocol/src/extractors/acme.rs, before any check reads them; a DNS identifier passes normalize_dns_name (crates/protocol/src/acme/rules.rs) once, before storage and before the filter sees it |
| 1.1.2 | Output encoding as the final step, or by the interpreter | 2 | met | minijinja escapes at render time. The rule is per template name: .html auto-escapes, .j2 does not — see crates/core/src/templating.rs |
| 1.2.1 | Context-correct output encoding for HTTP/HTML | 1 | met | Every panel template is .html and therefore auto-escaped; auto_escaping_is_on_for_pages_and_off_for_notify pins both directions |
| 1.2.2 | Encode untrusted data in dynamically built URLs; safe protocols only | 1 | met | Panel URLs are built from server-side ids. The one place an untrusted URL is followed — an http-01 redirect — is checked against a scheme allowlist in Http01Validator::redirect_allowed (crates/net/src/challenge/http_01.rs) |
| 1.2.3 | Encode when building JavaScript or JSON | 1 | met | All JSON is produced by serde_json; no template writes into a <script> block |
| 1.2.4 | Parameterized database queries | 1 | met | Every statement in crates/store/src/ is a runtime sqlx::query with .bind(). No query is assembled with format! |
| 1.2.5 | Protection against OS command injection | 1 | met | ScriptHook::run uses Command::new(path) with an argv vector and no shell (crates/core/src/script_hook.rs); payloads go to stdin as JSON |
| 1.2.6 | LDAP injection | 2 | n/a | No LDAP client |
| 1.2.7 | XPath injection | 2 | n/a | No XPath |
| 1.2.8 | LaTeX injection | 2 | n/a | No LaTeX |
| 1.2.9 | Escape special characters in regular expressions | 2 | met | compile_anchored and the glob translation both run regex::escape over everything that is not the wildcard (crates/policy/src/filter/mod.rs) |
| 1.2.10 | CSV and formula injection | 3 | n/a | No CSV or spreadsheet export; the CLI emits text or JSON |
| 1.3.1 | Sanitize untrusted HTML from editors | 1 | n/a | No rich-text input anywhere |
| 1.3.2 | Avoid eval() and dynamic code execution | 1 | met | No dynamic code execution. The one place operator-supplied code runs is a custom script hook, which is a configured executable, not evaluated input |
| 1.3.3 | Sanitize before a dangerous context; trim over-long input | 2 | met | Contacts reject control characters, User-Agent is truncated to 256 characters before storage (crates/core/src/audit/mod.rs), identifier lists are capped by order.max_identifiers |
| 1.3.4 | Sanitize user-supplied SVG | 2 | n/a | No user-supplied images |
| 1.3.5 | Sanitize user-supplied scriptable or template content | 2 | n/a | template_dir overrides are operator-supplied files on the host, not user input |
| 1.3.6 | SSRF protection by allowlist of protocols, domains, paths, ports | 2 | partial | Scheme, port and hop count are allowlisted for http-01 redirects; destination addresses deliberately are not. See Documented deviations |
| 1.3.7 | No templates built from untrusted input | 2 | met | Template sources come only from the embedded table or template_dir; untrusted values are only ever bound as context |
| 1.3.8 | JNDI injection | 2 | n/a | No JNDI |
| 1.3.9 | Sanitize before memcache | 2 | n/a | No memcache |
| 1.3.10 | Sanitize format strings | 2 | met | Rust format strings are compile-time literals; a runtime string can never become one |
| 1.3.11 | Sanitize before mail systems (SMTP/IMAP injection) | 2 | met | contact_shape_error rejects control characters, hfields and multiple addresses before a contact can reach a notify template (crates/protocol/src/acme/rules.rs) |
| 1.3.12 | Regular expressions free from exponential backtracking | 3 | met | The regex crate has no backtracking and guarantees linear time; patterns are operator configuration, not request input |
| 1.4.1 | Memory-safe strings and copies | 2 | met | Safe Rust. The unsafe blocks in the tree are std::env::set_var in tests and one PKCS#11 Send impl (crates/signer/src/local_ca/pkcs11.rs) |
| 1.4.2 | Prevent integer overflow | 2 | met | Time and TTL arithmetic uses saturating_add/saturating_sub throughout crates/store/src/; release builds are not built with overflow checks disabled beyond the default |
| 1.4.3 | Release memory and resources; no dangling pointers | 2 | met | Ownership and Drop. Script hooks additionally set kill_on_drop so a timed-out child is reaped |
| 1.5.1 | Restrictive XML parser configuration (XXE) | 1 | n/a | No XML parser in the dependency graph |
| 1.5.2 | Safe deserialization of untrusted data | 2 | met | serde into concrete structs. No polymorphic or client-chosen types; crates/core/src/jws/mod.rs deliberately does not use deny_unknown_fields because RFC 8555 §6.2 allows extra header parameters, and every field it acts on is named |
| 1.5.3 | Consistent parsers for one data type | 3 | met | One JSON parser (serde_json) and one URL parser (url) in the tree |
V2 Validation and Business Logic
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 2.1.1 | Documented input validation rules | 1 | met | Identifier syntax and normalisation are specified in Filters; the ACME wire formats are RFC 8555’s and the deviations are listed in Protocol Support |
| 2.1.2 | Documented rules for combined data items | 2 | met | The order/CSR consistency rule — the CSR may name only what the order named — is stated in Filters and enforced at finalize |
| 2.1.3 | Documented business logic limits | 2 | met | order.max_identifiers, server.max_body_bytes, nonce TTL, order and authorization lifetimes and the admission limiter are all in Configuration Reference with their defaults |
| 2.2.1 | Validate input against expectations | 1 | met | Identifiers are normalised and type-checked, wildcards are refused when dns-01 is off, contacts are shape-checked, the CSR is parsed and compared against the order |
| 2.2.2 | Validation enforced at a trusted service layer | 1 | met | Every check is server-side. The panel’s only client-side code is htmx attribute dispatch |
| 2.2.3 | Combinations of related data items are reasonable | 2 | met | a_wildcard_identifier_is_rejected_when_dns_01_is_disabled and a_csr_requesting_ca_powers_yields_a_leaf_without_them in tests/security.rs are two of the pinned cases |
| 2.3.1 | Business logic flows only in the expected step order | 1 | met | The order state machine refuses out-of-sequence transitions: an_order_missing_an_authorization_never_becomes_ready, an_expired_order_cannot_be_finalized, a_deactivated_account_cannot_finalize_a_ready_order (tests/security.rs) |
| 2.3.2 | Business logic limits implemented as documented | 2 | met | an_order_naming_more_identifiers_than_the_limit_is_refused (tests/security.rs) |
| 2.3.3 | Transactions succeed in full or roll back | 2 | met | Multi-row writes run inside pool.begin()/commit() — order creation with its authorizations (crates/protocol/src/handlers/order.rs), challenge validation (crates/protocol/src/handlers/authz.rs), session promotion (crates/store/src/admin_session.rs) |
| 2.3.4 | Locking prevents double-booking of limited resources | 2 | met | Single-use resources are claimed by UPDATE … WHERE … AND <unused> and decided by rows_affected == 1, never by read-then-write: nonces, recovery codes (crates/store/src/admin_recovery_code.rs), the TOTP replay step (crates/store/src/admin_user.rs) and session promotion |
| 2.3.5 | Multi-user approval for high-value flows | 3 | gap | Issuance and revocation are single-actor operations. There is no second-operator approval, and no plan to add one — a CA that needs a quorum to sign is a different product |
| 2.4.1 | Anti-automation on expensive functions | 2 | met | The ACME listener carries an admission limiter with a queue budget and a request deadline (crates/protocol/src/middlewares/admission.rs); the admin login path is rate-limited per address — an IPv6 client per /64 — counted from the moment an attempt starts, and the KDF runs on the blocking pool (LoginLimiter, crates/admin/src/webadmin/session.rs) |
| 2.4.2 | Business flows require realistic human timing | 3 | n/a | Every consumer of the ACME API is a machine; timing gates would break the protocol |
V3 Web Frontend Security
Applies to the web admin. With admin.enabled = false, the default, none of
this is exposed at all.
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 3.1.1 | Documented expected browser security features | 3 | partial | Web Admin states the cookie and CSRF requirements and the Secure-over-plain-HTTP failure mode; there is no statement of what the panel does when a browser lacks a feature |
| 3.2.1 | Prevent content being rendered in the wrong context | 1 | met | default-src 'none' plus X-Content-Type-Options: nosniff on every response; the API is nested under /api with its own JSON fallback so a page path never returns an API body |
| 3.2.2 | Text rendered as text, not HTML | 1 | met | Auto-escaping templates; no innerHTML outside htmx’s own fragment swap of server-rendered HTML |
| 3.2.3 | Avoid DOM clobbering | 3 | met | The panel ships no application JavaScript — htmx is the only script, and everything is driven by hx-* attributes |
| 3.3.1 | Secure attribute and a __Secure-/__Host- prefix | 1 | met | __Host-acme_admin_session, which browsers accept only with Secure, Path=/ and no Domain (crates/admin/src/webadmin/session.rs) |
| 3.3.2 | SameSite set according to purpose | 2 | met | SameSite=Strict on both the session cookie and its clearing form |
| 3.3.3 | __Host- prefix unless shared with other hosts | 2 | met | Same as 3.3.1 |
| 3.3.4 | HttpOnly for values scripts must not read | 2 | met | HttpOnly is set; the CSRF token travels in the page and the x-csrf-token request header, never in a readable cookie |
| 3.3.5 | Cookie name and value under 4096 bytes | 3 | met | A 32-byte token, base64url-encoded, plus a fixed name |
| 3.4.1 | HSTS on all responses, ≥ 1 year, includeSubDomains for L2 | 1 | met | max-age=31536000; includeSubDomains, applied by the shared security_headers() constructor to both listeners (crates/protocol/src/router.rs). Emitted unconditionally: a browser ignores it over plain HTTP (RFC 6797 §7.2), so gating it on TLS would remove only a header that is already inert. includeSubDomains makes the host in admin.base_url load-bearing — see give the panel its own host name |
| 3.4.2 | CORS Access-Control-Allow-Origin fixed or allowlisted | 1 | met | No CORS layer exists on either listener, so no Access-Control-Allow-Origin is ever emitted |
| 3.4.3 | CSP with object-src 'none' and base-uri 'none' | 2 | met | default-src 'none'; script-src 'self'; style-src 'self'; img-src 'self' data:; connect-src 'self'; form-action 'self'; frame-ancestors 'none'; base-uri 'none' — no unsafe-inline, no unsafe-eval (crates/admin/src/webadmin/mod.rs). object-src falls back to default-src 'none' |
| 3.4.4 | X-Content-Type-Options: nosniff | 2 | met | security_headers() |
| 3.4.5 | Referrer policy | 2 | met | Referrer-Policy: same-origin on the admin listener |
| 3.4.6 | CSP frame-ancestors on every response | 2 | met | frame-ancestors 'none', with X-Frame-Options: DENY alongside for older clients |
| 3.4.7 | CSP reports a violation-reporting location | 3 | gap | No report-to/report-uri. For a single-origin panel with no inline script, the report channel would have no consumer |
| 3.4.8 | Cross-Origin-Opener-Policy on document responses | 3 | gap | Not set. frame-ancestors 'none' covers framing but not shared Window access from a popup |
| 3.5.1 | Anti-forgery tokens or non-safelisted header fields | 1 | met | A per-session CSRF token in the x-csrf-token header on every unsafe method, plus an Origin check against admin.base_url — the module doc in crates/admin/src/webadmin/session.rs explains why SameSite=Strict alone is not enough here |
| 3.5.2 | Functionality cannot be called without a preflight | 1 | n/a | The panel does not rely on CORS preflight; it uses the token in 3.5.1 |
| 3.5.3 | Sensitive functionality uses unsafe HTTP methods | 1 | met | Every mutating route is POST/DELETE; GET routes are read-only. mutating_endpoints() and mutating_page_endpoints() are the lists that make this checkable |
| 3.5.4 | Separate applications on different hostnames | 2 | partial | The ACME and admin surfaces are separate sockets with separate TLS and separate defaults, and the admin binds loopback unless TLS is on. They are usually two ports on one host, and cookies are not port-scoped — see Documented deviations |
| 3.5.5 | Validate postMessage origins | 2 | n/a | No postMessage |
| 3.5.6 | No JSONP | 3 | met | None |
| 3.5.7 | No authorized data in script resources | 3 | met | The only script served is a static, unauthenticated copy of htmx |
| 3.5.8 | Authenticated resources embeddable only when intended | 3 | met | Sec-Fetch is not inspected, but frame-ancestors 'none', SameSite=Strict and the CSRF token together mean no cross-origin embed carries the session |
| 3.6.1 | SRI for externally hosted client assets | 3 | met | Nothing is externally hosted. htmx is vendored under crates/admin/src/webadmin/static/ and served from the same origin |
| 3.7.1 | Only supported, secure client-side technologies | 2 | met | HTML, CSS and htmx. No plugins |
| 3.7.2 | Automatic redirects only to allowlisted hosts | 2 | met | The panel’s redirects are fixed relative paths (/ui/, the sign-in page); there is no next= parameter and no open-redirect surface |
| 3.7.3 | Notify before redirecting outside the application | 3 | n/a | The panel never redirects off-origin |
| 3.7.4 | HSTS preload | 3 | n/a | The panel is an internal service with an operator-chosen hostname; preloading is a decision for the operator’s domain, not this software |
| 3.7.5 | Documented behaviour on browsers lacking security features | 3 | gap | See 3.1.1 |
V4 API and Web Service
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 4.1.1 | Accurate Content-Type with charset | 1 | met | ACME responses are application/json or application/problem+json; panel pages are text/html; charset=utf-8 via axum::response::Html; the two static assets set their own type with an explicit charset (crates/admin/src/webadmin/pages/assets.rs) |
| 4.1.2 | Only user-facing endpoints redirect HTTP to HTTPS | 2 | met | Neither listener redirects. server.tls.enabled makes the socket speak TLS instead of cleartext, not alongside it (crates/net/src/tls.rs) |
| 4.1.3 | Intermediary-set header fields cannot be overridden by the user | 2 | met | A forwarded-for header is believed only from a hop in filter.trusted_proxies; with the list empty the header is ignored entirely and the peer address is used (crates/core/src/client.rs). The admin listener has its own list, admin.filter.trusted_proxies, empty by default (crates/admin/src/webadmin/filter.rs) |
| 4.1.4 | Only supported HTTP methods are usable | 3 | met | axum routes declare their methods and answer 405 otherwise; both routers carry an explicit method_not_allowed_fallback so the refusal is a proper problem document rather than an empty body |
| 4.1.5 | Per-message digital signatures for highly sensitive requests | 3 | met | Every state-changing ACME request is a JWS signed by the account key, verified against a nonce and the request URL (RFC 8555 §6.2) — this is the protocol’s own design, not an addition |
| 4.2.1 | Correct HTTP message framing (request smuggling) | 2 | met | hyper performs the framing and rejects conflicting Content-Length/Transfer-Encoding; the application never parses framing itself |
| 4.2.2 | Generated Content-Length matches the body | 3 | met | Response bodies are axum types; the length is computed, never asserted |
| 4.2.3 | No connection-specific header fields over HTTP/2 or HTTP/3 | 3 | met | hyper enforces this. No handler sets Transfer-Encoding |
| 4.2.4 | Reject CR/LF in HTTP/2 and HTTP/3 header fields | 3 | met | http::HeaderValue rejects control bytes on construction, in both directions |
| 4.2.5 | Avoid generating over-long URIs or header fields | 3 | met | Outbound URLs are built from configuration plus a bounded token or id; the http-01 validator additionally caps redirect hops |
| 4.3.1 | GraphQL query cost limiting | 2 | n/a | No GraphQL |
| 4.3.2 | GraphQL introspection disabled | 2 | n/a | No GraphQL |
| 4.4.1 | WebSocket over TLS | 1 | n/a | No WebSocket |
| 4.4.2 | WebSocket handshake Origin check | 2 | n/a | No WebSocket |
| 4.4.3 | Dedicated WebSocket session tokens | 2 | n/a | No WebSocket |
| 4.4.4 | WebSocket tokens obtained through the authenticated session | 2 | n/a | No WebSocket |
V5 File Handling
There is no file upload feature. The only client-supplied structured input is a CSR inside a signed JWS, which is assessed under V2 and V11 rather than here. The two paths that touch the filesystem on a request’s behalf are the embedded static-asset allowlist and the certificate-chain download.
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 5.1.1 | Documented permitted file types, extensions and sizes | 2 | n/a | No upload feature to document |
| 5.2.1 | Only accept files of a processable size | 1 | met | server.max_body_bytes (128 KiB) and admin.max_body_bytes (64 KiB) bound every request body, applied as DefaultBodyLimit at each router root |
| 5.2.2 | Extension matches content | 1 | n/a | No uploads |
| 5.2.3 | Compressed-file limits | 2 | n/a | Nothing is decompressed |
| 5.2.4 | Per-user file quota | 3 | n/a | No uploads |
| 5.2.5 | Reject symlinks in archives | 3 | n/a | No archives |
| 5.2.6 | Reject over-large images | 3 | n/a | No images |
| 5.3.1 | Untrusted files in a public folder are not executed | 1 | n/a | Nothing untrusted is written to a served directory |
| 5.3.2 | File paths built from trusted data, not user filenames | 1 | met | GET /ui/static/{file} is a two-arm match, not a filesystem lookup — tower-http’s fs feature is deliberately off (crates/admin/src/webadmin/pages/assets.rs). The http-01 responder looks a token up in an in-memory store and touches no path |
| 5.3.3 | Ignore user path information when decompressing | 3 | n/a | Nothing is decompressed |
| 5.4.1 | Validate or ignore user filenames; set Content-Disposition | 2 | met | The chain download names the file from the stored order id, not the path segment, and sets attachment; filename="…" (crates/admin/src/webadmin/pages/orders.rs) |
| 5.4.2 | Served filenames are encoded or sanitized | 2 | met | Same: a generated identifier, so there is nothing to encode |
| 5.4.3 | Antivirus scanning of files from untrusted sources | 2 | n/a | No files are accepted from untrusted sources |
V6 Authentication
Applies to the web admin and to the CLI commands that mint and rotate operator credentials. The ACME listener authenticates keys, not people; that is assessed under V9.
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 6.1.1 | Documented anti-automation and lockout behaviour | 1 | met | Web Admin and the login_max_attempts / login_window_seconds entries in Configuration Reference. The limiter is keyed on the peer address, never the username, so no attacker can lock an operator out by guessing at them |
| 6.1.2 | Documented list of context-specific words barred from passwords | 2 | met | Derived and documented: Password policy |
| 6.1.3 | Multiple authentication pathways documented together | 2 | met | There are two — a panel session and a shell on the host — and Users & Sessions states which operations belong to which and why create and passwd stay on the host |
| 6.2.1 | Passwords at least 8 characters | 1 | met | MIN_PASSWORD_LEN = 12, counted in characters rather than bytes (crates/admin/src/admin/password.rs) |
| 6.2.2 | Users can change their password | 1 | met | The panel’s own account page carries a password card (POST /ui/account/password, POST /api/account/password) beside acme-proxy admin user passwd, so an operator with no shell can rotate their own → Users & Sessions |
| 6.2.3 | Password change requires current and new password | 1 | met | handlers::mfa::verify_current_password (crates/admin/src/webadmin/handlers/mfa.rs) checks the current password before admin::users::change_own_password writes a new one — unconditionally, unlike the second-factor step-up it was split out of, since this requirement has no “nothing yet to protect” exemption. admin user passwd still takes only the new one, answering as it does to a process that can already rewrite the row |
| 6.2.4 | Check against the top 3000 passwords | 1 | met | 13 918 entries compiled in (crates/admin/src/admin/corpus/); every shorter entry is already refused on length |
| 6.2.5 | No composition rules | 1 | met | Deliberately none. The three rules are about length, this deployment’s own words and known-common passwords — none dictates shape (crates/admin/src/admin/password.rs) |
| 6.2.6 | Password fields use type=password | 1 | met | templates/login.html, the step-up fields in templates/account/_card.html, templates/account/_contact.html and templates/operators/_card.html, and the two fields in templates/account/_password_card.html |
| 6.2.7 | Paste and password managers permitted | 1 | met | Standard inputs with autocomplete="username" / "current-password"; nothing blocks paste |
| 6.2.8 | Password verified exactly as received | 1 | met | verify_password hashes the bytes as received: no trimming, no case folding, and an over-long password is rejected rather than truncated. The policy check folds a copy to compare against the corpus and the word list, and never touches what is stored |
| 6.2.9 | Passwords of at least 64 characters permitted | 2 | met | MAX_PASSWORD_LEN = 1024 bytes, a denial-of-service bound rather than a policy |
| 6.2.10 | No forced periodic rotation | 2 | met | Nothing expires a password. The stored form is self-describing, so raising the KDF cost re-encodes a row on its owner’s next login instead of forcing a change |
| 6.2.11 | Context-specific word list used | 2 | met | PasswordContext in crates/admin/src/admin/password.rs, matched as a substring |
| 6.2.12 | Check against breached passwords | 2 | met | Same corpus: breach-derived (xato-net), filtered to the reachable length range |
| 6.3.1 | Credential-stuffing and brute-force controls | 1 | met | LoginLimiter refuses over the limit before the 600 000-iteration KDF runs, which makes it an availability control as much as a credential one; an attempt counts from the moment it starts, so a parallel burst cannot outrun it (crates/admin/src/webadmin/session.rs) |
| 6.3.2 | No default accounts | 1 | met | The admin_users migration seeds no rows and there is no sign-up page; the first operator is created by admin user create on the host |
| 6.3.3 | MFA or a combination of single factors | 2 | partial | TOTP with recovery codes is implemented and admin.require_mfa enforces it for every operator — but it defaults to false, so a stock deployment is single-factor. Hardening tells operators to turn it on. For L3 this would need a hardware factor; see Documented deviations |
| 6.3.4 | No undocumented pathways; consistent strength | 2 | met | The panel and API share one session layer, and every mutating route passes through AuthenticatedWrite, PageSessionWrite or EnrolWrite. The host CLI is the second pathway and is documented as such |
| 6.3.5 | Notify users of suspicious authentication attempts | 3 | met | A completed sign-in from an address not among the operator’s recent ones (admin_users.known_login_ips, last five), a correct password then a refused second factor, and a per-session second-factor lockout each send an admin_sign_in notification to the operator’s own contact_email, through [admin.notify] (crates/admin/src/webadmin/handlers/session.rs, crates/jobs/src/notify/) |
| 6.3.6 | Email not used as an authentication factor | 3 | met | It is not |
| 6.3.7 | Notify after changes to authentication details | 3 | met | A password change, a second-factor enrol/disable, a recovery-code regeneration, a notification-address change and a colleague-admin second-factor reset send an admin_credential_changed notification — from the panel and from the host CLI alike (admin user passwd, contact, totp reset, totp recovery-codes), the CLI queuing the delivery for the running server’s worker |
| 6.3.8 | Valid users not deducible from failed challenges | 3 | met | An unknown username still pays the KDF, against password::dummy_hash(), and every failure returns one invalid_credentials whatever the real cause (crates/admin/src/admin/users.rs) |
| 6.4.1 | Initial passwords and activation codes are random, policy-compliant and short-lived | 1 | n/a | Nothing generates an initial password; the operator supplies one on stdin or in --password-file |
| 6.4.2 | No password hints or secret questions | 1 | met | Neither exists |
| 6.4.3 | Secure forgotten-password reset that does not bypass MFA | 2 | met | Reset is admin user passwd on the host. It revokes every session the operator held and leaves the enrolled factor untouched, so the next sign-in still needs it |
| 6.4.4 | Lost MFA factor requires enrolment-level identity proofing | 2 | met | Either a single-use recovery code, or admin user totp reset on the host — the second being a strictly stronger proof than the live session plus password that enrolment took |
| 6.4.5 | Renewal reminders before an authenticator expires | 3 | n/a | No authentication factor expires |
| 6.4.6 | Administrators can reset but not choose a user’s password | 3 | gap | admin user passwd sets the password, so whoever runs it knows it. This is a host-root operation on a machine that already holds the hashes |
| 6.5.1 | Lookup secrets and TOTPs usable only once | 2 | met | AdminUser::claim_totp_step is an UPDATE … WHERE totp_last_step IS NULL OR totp_last_step < ? decided by rows_affected, so a code resubmitted inside its own 30-second window is refused (crates/store/src/admin_user.rs); recovery codes are consumed by UPDATE … WHERE id = ? AND used_at IS NULL |
| 6.5.2 | Sub-112-bit lookup secrets hashed with an approved KDF and a 32-bit salt | 2 | met | Recovery codes carry 50 bits and are stored through admin::password — PBKDF2-HMAC-SHA256, 600 000 iterations, a 128-bit per-row salt |
| 6.5.3 | Seeds and codes from a CSPRNG | 2 | met | ring::rand::SystemRandom for the TOTP secret (crates/admin/src/admin/totp.rs) and every recovery code (crates/admin/src/admin/recovery.rs) |
| 6.5.4 | Lookup secrets have at least 20 bits of entropy | 2 | met | Ten characters from a 32-symbol alphabet: 50 bits, with zero modulo bias because 256 is a multiple of 32 |
| 6.5.5 | Defined lifetime for codes and TOTPs | 2 | met | 30-second step with RFC 6238 §5.2’s one step of permitted skew either way (SKEW_STEPS = 1). A half-authenticated session additionally dies after PENDING_MFA_TTL, five minutes |
| 6.5.6 | Any factor can be revoked | 3 | met | admin user totp reset, admin user disable, admin session revoke [--user <u> [--session <id>] | --all], and recovery codes are superseded as a set on re-enrolment |
| 6.5.7 | Biometrics only as a secondary factor | 3 | n/a | No biometrics |
| 6.5.8 | TOTP checked against a trusted time source | 3 | met | The server’s own clock; no client-supplied time reaches totp::verify |
| 6.6.1 | PSTN OTP restrictions | 2 | n/a | No SMS or voice factor |
| 6.6.2 | Out-of-band codes bound to their originating request | 2 | n/a | No out-of-band factor. The equivalent binding for TOTP is that the code is only accepted against the pending_mfa session that the password created |
| 6.6.3 | Rate-limit code-based out-of-band mechanisms | 2 | n/a | No out-of-band factor. TOTP guessing is bounded twice — mfa_attempts on the pending row and the five-minute PENDING_MFA_TTL |
| 6.6.4 | Rate-limit push notifications | 3 | n/a | No push factor |
| 6.7.1 | Certificates verifying authentication assertions protected from modification | 3 | met | Account public keys live in accounts under the database’s file mode; a modified key is a key that no longer verifies its own account’s requests |
| 6.7.2 | Challenge nonce at least 64 bits and unique | 3 | met | 256 bits from ring::rand::SystemRandom, unique by primary key and single-use by rows_affected (crates/store/src/nonce.rs) |
| 6.8.1 | Identity cannot be spoofed across identity providers | 2 | n/a | No identity provider |
| 6.8.2 | Signatures on authentication assertions validated | 2 | n/a | No external assertions. The equivalent for ACME JWS is V9.1.1 |
| 6.8.3 | SAML assertions processed once | 2 | n/a | No SAML |
| 6.8.4 | Authentication strength verified from the IdP | 2 | n/a | No identity provider |
V7 Session Management
Applies to the web admin. The ACME listener holds no sessions: every request carries its own signature and its own nonce.
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 7.1.1 | Documented inactivity timeout and absolute lifetime | 2 | met | session_ttl_seconds (12 h, never extended by activity) and session_idle_timeout_seconds (1 h) in Configuration Reference, restated in Web Admin |
| 7.1.2 | Documented concurrent-session policy | 2 | partial | The behaviour is definite — sessions are unlimited per operator, and admin session revoke (--all, one operator’s, or one session with --session <id>) is the lever — but no page states the limit as a policy |
| 7.1.3 | Federated session coordination documented | 2 | n/a | No federation |
| 7.2.1 | Session verification at a trusted backend | 1 | met | Every request resolves hex(SHA-256(token)) against admin_sessions and re-checks state, expiry, idleness and the owner’s status (crates/admin/src/webadmin/session.rs) |
| 7.2.2 | Dynamically generated tokens, not static secrets | 1 | met | mint_token per sign-in; there are no API keys on this listener |
| 7.2.3 | Reference tokens unique, CSPRNG, ≥ 128 bits | 1 | met | 256 bits from ring::rand::SystemRandom, base64url-encoded |
| 7.2.4 | New token on authentication, old one terminated | 1 | met | Sign-in deletes whatever session the request carried; completing MFA is a rotation — AdminSession::promote deletes the pending_mfa row and inserts a new one with a new token and a new CSRF token, in one transaction (crates/store/src/admin_session.rs) |
| 7.3.1 | Inactivity timeout | 2 | met | session_idle_timeout_seconds, checked per request and swept by the reaper |
| 7.3.2 | Absolute maximum session lifetime | 2 | met | expires_at is set at creation and never advanced |
| 7.4.1 | Terminated sessions cannot be reused | 1 | met | Sessions are reference tokens in a table; sign-out deletes the row |
| 7.4.2 | All sessions terminated when an account is disabled or deleted | 1 | met | set_status("disabled") and set_password both call AdminSession::delete_for_user; the liveness check also refuses a session whose owner is no longer active |
| 7.4.3 | Option to terminate other sessions after a factor changes | 2 | met | confirm_totp_enrolment and disable_totp both call revoke_other_sessions; a password change revokes every session unconditionally |
| 7.4.4 | Visible logout on every authenticated page | 2 | met | A “Sign out” control in templates/layout.html, which every page extends |
| 7.4.5 | Administrators can terminate sessions individually or globally | 2 | met | admin session list/revoke on the host terminates globally (--all), one operator’s (--user <u>), or one session (--user <u> --session <id>, the id being the fingerprint the listing prints); the panel’s Operators page does the individual form over HTTP — GET /ui/operators/{username} lists another operator’s sessions and POST /ui/operators/{username}/sessions/{id}/revoke ends one, gated by verify_current_password |
| 7.5.1 | Full re-authentication before changing authentication attributes | 2 | met | check_step_up demands the password again before any change to an existing second factor, and the module doc explains the blast radius that makes it necessary (crates/admin/src/webadmin/handlers/mfa.rs) |
| 7.5.2 | Users can view and terminate their own sessions | 2 | met | The account page’s Sessions card (GET /api/account/sessions, /ui/account) lists every one of the caller’s own live sessions and terminates one individually (POST /api/account/sessions/{id}/revoke) or all at once (“Sign out everywhere”) — closing the gap between nothing and everything the panel used to leave → Sessions |
| 7.5.3 | Further authentication before highly sensitive operations | 3 | partial | Second-factor changes are gated by check_step_up, and the whole /operators colleague-management surface by verify_current_password, which re-prompts even for an operator with no factor. Certificate revocation and account deletion require at least the operator role (admin_users.role), but for an operator holding it a live session is still sufficient authority — no password re-prompt on the CA mutations |
| 7.6.1 | Federated re-authentication behaviour | 2 | n/a | No federation |
| 7.6.2 | Session creation requires explicit user action | 2 | met | A session exists only after a submitted sign-in form |
V8 Authorization
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 8.1.1 | Documented function-level and data-specific rules | 1 | met | Security Model states the four issuance gates; Filters specifies the policy engine; Web Admin states that every route needs a session |
| 8.1.2 | Documented field-level rules | 2 | met | The ACME object shapes are RFC 8555’s, and Audit Trail states which fields are recorded and that none of them reach an ACME object |
| 8.1.3 | Documented environmental and contextual attributes | 3 | met | Address, reverse name, IPAM ownership, request path and EAB identity are each documented under Filters, and the trust placed in a forwarded address under Allowed IP |
| 8.1.4 | Documented use of contextual factors in decisions | 3 | met | The policy expression language, including how an or over an address check weakens a conjunction, is written out in Hardening |
| 8.2.1 | Function-level access restricted to explicit permissions | 1 | met | ACME: POST-as-GET with a kid resolving to the owning account. Admin: three extractors every mutating route passes through |
| 8.2.2 | Data-specific access restricted (IDOR/BOLA) | 1 | met | Order, authorization and certificate reads check the requesting account owns the object; accounts and orders are additionally isolated per profile, so a kid naming another profile does not resolve (crates/protocol/src/extractors/acme.rs) |
| 8.2.3 | Field-level access restricted (BOPLA) | 2 | met | Responses are built from explicit serializer functions, never by serializing a row |
| 8.2.4 | Adaptive controls from contextual attributes | 3 | met | The filter chain evaluates per request, not per session, so a change of address is re-evaluated on the next call |
| 8.3.1 | Authorization enforced at a trusted service layer | 1 | met | Extractors and middleware, server-side. No decision depends on anything the client sends unsigned |
| 8.3.2 | Authorization changes applied immediately | 3 | met | Sessions are reference tokens read from the database each request, so a disabled operator or a revoked session stops working on the next call. filter reload applies policy without a restart |
| 8.3.3 | Access based on the originating subject | 3 | partial | With the relay signer, one upstream account is deliberately multiplexed across every local client — that is the feature. The local gates decide, and the upstream sees only this server. See Documented deviations |
| 8.4.1 | Cross-tenant controls | 2 | met | Profiles are the tenancy boundary: accounts, orders, nonces and EAB credentials are scoped to one, and a kid from another profile fails the prefix check |
| 8.4.2 | Administrative access uses more than network location | 3 | partial | Password plus optional TOTP plus a session, with the bind address and TLS as further layers, and a per-operator role (admin/operator/viewer) scoping what a session may do. There is no device posture assessment and no contextual risk analysis |
V9 Self-contained Tokens
The self-contained token in this system is the ACME JWS on every state-changing request (RFC 8555 §6.2), not a session JWT — the admin session is a reference token and is assessed under V7.
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 9.1.1 | Signature validated before the contents are accepted | 1 | met | verify_jws verifies the signature over the protected header and payload before any handler sees the body (crates/protocol/src/extractors/acme.rs, crates/core/src/jws/signature.rs) |
| 9.1.2 | Algorithm allowlist, no none | 1 | met | Exactly ES256 on P-256 and RS256 are accepted; anything else is Unsupported algorithm. The alg must additionally agree with the key type and the named curve, so alg alone never selects the verifier (crates/core/src/jws/signature.rs) |
| 9.1.3 | Key material from trusted pre-configured sources | 1 | met | A kid resolves to a stored account key whose URL prefix must match this profile’s base_url; a jwk is the key being registered and is only ever trusted for newAccount/revokeCert as RFC 8555 §6.2 defines. jwk and kid together are refused, and a crit header is refused outright |
| 9.2.1 | Validity time span honoured | 1 | met | The equivalent is the nonce: single-use, and refused past nonce.ttl_seconds. Unknown, consumed and expired are made indistinguishable on purpose (crates/store/src/nonce.rs) |
| 9.2.2 | Token type checked against the intended purpose | 2 | met | The protected header must carry exactly the fields RFC 8555 §6.2 defines for the request kind; newAccount requires a jwk, revokeCert takes either, and everything else a kid — an embedded jwk elsewhere is 400 malformed |
| 9.2.3 | Audience restriction | 2 | met | The JWS url must equal profile.base_url plus the request path, byte for byte (RFC 8555 §6.4). A signature captured from one profile does not verify against another |
| 9.2.4 | Same key across audiences carries an audience restriction | 2 | met | Same mechanism: the audience is in the signed url, and the kid prefix pins the profile |
V11 Cryptography
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 11.1.1 | Documented key management policy and lifecycle | 2 | met | Security Model names every secret, what its compromise buys and how it is stored; Secret Rotation is the lifecycle half — a recommended interval and the early-rotation triggers for each — and Hardening covers the CA key specifically |
| 11.1.2 | Cryptographic inventory maintained | 2 | met | The table in Security Model, plus the per-algorithm rationale carried in the module docs of crates/admin/src/admin/password.rs, crates/admin/src/admin/totp.rs and crates/core/src/eab.rs |
| 11.1.3 | Cryptographic discovery mechanisms | 3 | met | One backend: ring, plus rustls for TLS and rcgen for certificate construction. grep -rn ring src/ is the discovery mechanism, and cargo deny fails the build on an unlisted crypto dependency |
| 11.1.4 | Inventory includes a post-quantum migration path | 3 | gap | No PQC migration plan. The ACME wire algorithms are RFC 8555’s to change first |
| 11.2.1 | Industry-validated implementations | 2 | met | ring (BoringSSL-derived) for hashing, HMAC, PBKDF2, signature verification and the RNG; rustls for TLS. Nothing hand-rolls a primitive — crates/admin/src/admin/totp.rs composes ring::hmac per RFC 4226 and is checked against the RFC’s published test vectors |
| 11.2.2 | Crypto agility | 2 | met | Password hashes are stored self-describing (pbkdf2-sha256$600000$…), so the algorithm or cost can change with a new branch in verify_password and needs_rehash re-encodes each row at its owner’s next login — no migration. The signer backend, the CA key type and the key source (file or PKCS#11) are all configuration |
| 11.2.3 | Minimum 128 bits of security | 2 | partial | ECDSA P-256, SHA-256, HMAC-SHA-256 and 256-bit secrets are all at or above the bar. RSA is accepted from 2048 bits (RSA_PKCS1_2048_8192_SHA256), which is about 112 — see Documented deviations |
| 11.2.4 | Constant-time cryptographic operations | 3 | met | ring::constant_time::verify_slices_are_equal under the hood, and subtle::ConstantTimeEq for the TOTP comparison (crates/admin/src/admin/totp.rs) |
| 11.2.5 | Cryptographic modules fail securely | 3 | met | A verification failure is a refusal, never a fallback. A corrupt stored password hash is deliberately not folded into “wrong password” — it refuses and logs admin_password_hash_unreadable (crates/admin/src/admin/users.rs) |
| 11.3.1 | No insecure block modes or weak padding | 1 | partial | Nothing in the tree encrypts. RS256 is RSASSA-PKCS1-v1_5, which RFC 8555 requires — a signature scheme, not the padding oracle this requirement targets. See Documented deviations |
| 11.3.2 | Only approved ciphers and modes | 1 | met | Transport encryption is rustls with safe defaults; the application encrypts nothing itself |
| 11.3.3 | Encrypted data protected against modification | 2 | n/a | No application-layer encryption |
| 11.3.4 | Single-use numbers not reused across key/data pairs | 3 | n/a | No application-layer encryption. ACME nonces are single-use by construction |
| 11.3.5 | Encrypt-then-MAC | 3 | n/a | No application-layer encryption |
| 11.4.1 | Approved hash functions | 1 | met | SHA-256 throughout. The one SHA-1 is HMAC_SHA1_FOR_LEGACY_USE_ONLY inside TOTP, which RFC 6238 §1.2 specifies and which every authenticator app assumes — the module doc in crates/admin/src/admin/totp.rs is the argument for not “fixing” it |
| 11.4.2 | Passwords stored with an approved, expensive KDF | 2 | met | PBKDF2-HMAC-SHA256 at 600 000 iterations with a 128-bit per-row salt — OWASP’s current recommendation for the non-Argon2 case. See Documented deviations for why not Argon2id |
| 11.4.3 | Collision-resistant hashes of adequate length in signatures | 2 | met | SHA-256 for every signature and every integrity use. HMAC-SHA-1’s security rests on the PRF property, not collision resistance |
| 11.4.4 | Approved KDF with key-stretching for password-derived keys | 2 | met | Same PBKDF2 parameters; recovery codes go through the identical path |
| 11.5.1 | Non-guessable values from a CSPRNG with ≥ 128 bits | 2 | met | Session tokens, CSRF tokens, EAB secrets, challenge tokens and ACME replay nonces are all 256 bits from ring::rand::SystemRandom, base64url-encoded, through the one crates/core/src/random.rs. The nonce was a UUID v4 until 0.2.0 — 122 bits, and a form this requirement names explicitly |
| 11.5.2 | RNG works securely under heavy demand | 3 | met | SystemRandom draws from the OS CSPRNG; there is no userspace pool to exhaust |
| 11.6.1 | Approved algorithms for key generation and signatures | 2 | met | rcgen generates ECDSA P-256 by default; the accepted account-key algorithms are the two RFC 8555 defines. Key generation can be delegated to a PKCS#11 token, where the key never leaves the device |
| 11.6.2 | Approved key exchange with secure parameters | 3 | met | rustls with with_safe_default_protocol_versions(): TLS 1.2 and 1.3 only, and only its own vetted groups |
| 11.7.1 | Full memory encryption for data in use | 3 | n/a | A property of the host, not of this process |
| 11.7.2 | Data minimization during processing | 3 | partial | The CA key can live in a PKCS#11 token and never enter this process at all. EAB and TOTP secrets are necessarily readable, because both are verified by recomputing an HMAC — file mode is the boundary, and that is stated in Security Model |
V12 Secure Communication
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 12.1.1 | Only current TLS versions, newest preferred | 1 | met | with_safe_default_protocol_versions() on the rustls ring provider — TLS 1.3 and 1.2 only (crates/net/src/tls.rs) |
| 12.1.2 | Recommended cipher suites, forward secrecy for L3 | 2 | met | rustls ships no suite without forward secrecy and none that is not current; there is no knob to weaken it |
| 12.1.3 | mTLS client certificates validated before use | 2 | n/a | No mTLS. The one place a client certificate is inspected is tls-alpn-01 validation, where the certificate is the challenge response and is checked for the RFC 8737 acmeIdentifier extension rather than for trust |
| 12.1.4 | Certificate revocation such as OCSP stapling | 3 | partial | As a CA, the server publishes a CRL signed over the revocations recorded in its database (Revocation & CRL). As a TLS server it does not staple |
| 12.1.5 | Encrypted Client Hello | 3 | gap | Not offered by rustls in a form this could adopt today |
| 12.2.1 | TLS for all client connectivity, no fallback | 1 | met | With server.tls.enabled the socket speaks TLS instead of cleartext; there is no downgrade path. HTTPS is on the hardening checklist for deployments that terminate elsewhere |
| 12.2.2 | Publicly trusted certificates on external services | 1 | n/a | This is an internal service by design; its clients trust the CA the operator installed |
| 12.3.1 | Encrypted protocols for all inbound and outbound connections | 2 | partial | The relay upstream, webhooks and IPAM are HTTPS. http-01 validation is HTTP because RFC 8555 §8.3 defines it that way, SQLite is a local file, not a connection, and a PostgreSQL connection is TLS only when database.url asks for it with sslmode — the server does not enforce it |
| 12.3.2 | TLS clients validate certificates | 2 | met | The relay client validates against webpki-roots — there, the certificate is the only thing identifying the CA being handed your CSRs. The IPAM clients validate too; insecure_skip_verify exists, defaults off, and warns on every startup while on |
| 12.3.3 | TLS between internal HTTP services | 2 | met | Same set. The http-01 exception above is the protocol’s |
| 12.3.4 | Internal TLS uses trusted certificates | 2 | met | The IPAM clients take a ca_bundle so a NetBox behind an internal PKI is trusted specifically rather than by disabling verification (crates/core/src/config/types/ipam.rs) |
| 12.3.5 | Strong mutual authentication between internal services | 3 | n/a | Single process; there are no intra-service hops |
V13 Configuration
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 13.1.1 | All communication needs documented, including user-supplied destinations | 2 | met | Security Model names all three outbound surfaces and says which of them a client can steer |
| 13.1.2 | Documented connection limits and behaviour at the limit | 3 | met | The database pool size, the admission limiter’s slots, its queue budget and its deadline are all in Configuration Reference, and shedding at the limit is a 503 problem document |
| 13.1.3 | Documented resource-management strategy per external system | 3 | partial | Timeouts are documented per subsystem and every outbound call has one. Retry policy is documented for the job runner but not stated as a policy for the IPAM and webhook clients |
| 13.1.4 | Documented critical secrets and a rotation schedule | 3 | met | The secrets are named and classified in Security Model; Secret Rotation gives a recommended interval and the early-rotation triggers for each, with the CA key called out as structural rather than scheduled |
| 13.2.1 | Authenticated backend communication with non-shared credentials | 2 | partial | The relay upstream authenticates by account key and the IPAM clients by API token, both per-deployment, and PostgreSQL by the role and password in database.url. SQLite is a local file governed by file mode, not by a credential |
| 13.2.2 | Least privilege for backend accounts | 2 | met | custom hooks run with env_clear(), a minimal PATH, a timeout and kill_on_drop (crates/core/src/script_hook.rs); the systemd unit in Deployment runs as a dedicated acme-proxy user and the repository Containerfile runs as a non-root acme-proxy user (uid 1000) owning only /data; the IPAM token needs read access only |
| 13.2.3 | No default service credentials | 2 | met | Nothing ships with a credential. Every secret is either operator-supplied or generated on first start |
| 13.2.4 | Allowlist of external systems the application may contact | 2 | partial | The relay upstream, the IPAM host and the webhook URL are each a single configured destination — an allowlist of one. The http-01 validator is the exception, and deliberately so |
| 13.2.5 | Server-level allowlist of destinations | 2 | partial | Same. The containment for http-01 is scheme, port and hop count rather than destination |
| 13.2.6 | Documented per-connection configuration followed | 3 | met | Each client is built from its own configuration block at startup, so a broken setting stops the server rather than failing every later call |
| 13.3.1 | A secrets management solution; no secrets in source or artifacts | 2 | partial | No secret is in the source tree or the image. Every secret can come from the environment rather than the file, and the CA key can live in a PKCS#11 token — which is the L3 hardware-backed form. There is no vault integration, and the database necessarily holds EAB and TOTP secrets in retrievable form |
| 13.3.2 | Least privilege for secret access | 2 | met | Keys are created 0600 with create_new rather than chmod’ed afterwards (crates/core/src/pemfile.rs); the database file mode is the documented boundary |
| 13.3.3 | Cryptographic operations inside an isolated security module | 3 | partial | Available but not required: --features hsm puts the issuing key in a PKCS#11 token, where it can be used and not copied (Hardware Keys) |
| 13.3.4 | Secrets expire and rotate as documented | 3 | partial | Rotation is now documented per secret in Secret Rotation. Sessions expire on their own and EAB credentials are revocable live; the CA key, the TSIG key and the API tokens rotate on an operator-run cadence, not a timer |
| 13.4.1 | No source-control metadata deployed | 1 | met | .dockerignore is an allowlist — * then !Cargo.toml, !Cargo.lock, !src/, !crates/store/migrations/ — so .git never enters the build context, and the final stage copies only the compiled binary |
| 13.4.2 | Debug modes disabled in production | 2 | met | Log level is configuration and defaults to info; there is no debug endpoint and no development mode. challenge.bypass, the one setting that genuinely weakens the server, is off by default and warns on every startup while on |
| 13.4.3 | No directory listings | 2 | met | Nothing is served from a directory. tower-http’s fs feature is off and static assets are a two-arm match |
| 13.4.4 | HTTP TRACE unsupported | 2 | met | Never routed; axum answers 405 |
| 13.4.5 | Documentation and monitoring endpoints not exposed unless intended | 2 | met | /metrics is a separate listener, off by default; /health is deliberately outside the filter chain and the hardening checklist tells operators not to forward it (Monitoring) |
| 13.4.6 | No detailed version information exposed | 3 | met | No Server header is set and no version appears in any response body |
| 13.4.7 | Web tier serves only specific extensions | 3 | met | The static allowlist is two filenames; everything else is a 404 |
V14 Data Protection
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 14.1.1 | Sensitive data identified and classified | 2 | met | Security Model classifies every secret by what its compromise buys; Database Schema classifies each by the form it is stored in — one-way, retrievable, or never stored |
| 14.1.2 | Documented protection requirements per level | 2 | met | The same two pages, plus Audit Trail for retention |
| 14.2.1 | No sensitive data in URLs or query strings | 1 | met | The session token is in a cookie, the CSRF token in a header, the EAB secret in a response body shown once. No credential is ever a path or query parameter |
| 14.2.2 | Sensitive data not cached in server components | 2 | met | Cache-Control: no-store on every admin response — account contacts and a freshly minted EAB secret must not sit in a disk cache after the tab closes (crates/admin/src/webadmin/mod.rs) |
| 14.2.3 | Sensitive data not sent to untrusted parties | 2 | met | The only outbound payloads are the relay’s own ACME traffic, a webhook to an operator-configured URL and IPAM lookups. No analytics, no third-party asset, no CDN |
| 14.2.4 | Documented controls implemented | 2 | met | Nonces and session tokens reach logs only as fingerprints (crates/store/src/nonce.rs, crates/store/src/admin_session.rs); proxy URLs are redacted before they are logged or Debug-formatted, pinned by neither_debug_nor_redacted_leaks_the_password (crates/net/src/proxy.rs); the database URL’s password is redacted wherever it is printed (redact_url, crates/core/src/logfields.rs), and an EAB row’s Debug omits its HMAC secret (crates/store/src/eab.rs) |
| 14.2.5 | Caching only for expected content types (web cache deception) | 3 | met | no-store on the whole admin listener, and an unknown path returns a 404, never a different valid file |
| 14.2.6 | Return the minimum sensitive data | 3 | met | An EAB secret is shown exactly once at creation; a session is displayed by the fingerprint of its token hash, never by the hash; render_admin_session_json is the one serializer |
| 14.2.7 | Retention classification and scheduled deletion | 3 | partial | audit.retention_days sweeps the trail and the job runner reaps nonces, expired sessions and stale orders. The default is 0 — keep everything — which is the right default for a trail whose value is that it is complete, and is a decision the operator is asked to make |
| 14.2.8 | Strip metadata from user-submitted files | 3 | n/a | No file uploads |
| 14.3.1 | Authenticated data cleared from client storage on termination | 1 | met | The panel keeps nothing in localStorage or sessionStorage; sign-out clears the cookie with a Max-Age=0 Set-Cookie carrying the same attributes |
| 14.3.2 | Anti-caching response header fields | 2 | met | Cache-Control: no-store |
| 14.3.3 | No sensitive data in browser storage beyond session tokens | 2 | met | Only the session cookie exists |
V15 Secure Coding and Architecture
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 15.1.1 | Documented remediation time frames for vulnerable components | 1 | partial | Security Policy states that fixes land on main and in the next release, and cargo deny runs advisories on every CI run and on a schedule. No numeric time frame is committed to |
| 15.1.2 | An SBOM or equivalent inventory is maintained | 2 | met | sbom.cdx.json is a committed CycloneDX 1.5 inventory of the shipped dependency closure (--all-features --target all, dev-dependencies excluded), regenerated and diffed by the sbom CI job and carried in the published crate; cargo deny check gates that same closure |
| 15.1.3 | Documented resource-demanding functionality | 2 | met | The expensive paths are named and bounded: http-01 and dns-01 validation have timeouts, the PBKDF2 cost is documented as a denial-of-service lever with the limiter placed before it, and the admission limiter’s shed-versus-queue reasoning is written out in crates/protocol/src/middlewares/admission.rs |
| 15.1.4 | Risky third-party libraries highlighted | 3 | met | deny.toml is the allow list, run with all-features = true, and the rationale for refusing dependencies is recorded where the refusal was made — crates/admin/src/admin/password.rs on Argon2id, issue #5 on webauthn-rs |
| 15.1.5 | Dangerous functionality highlighted | 3 | met | Security Model names the three request-forgery surfaces, and Security Policy lists the behaviour that looks alarming and is deliberate |
| 15.2.1 | No components past the documented remediation window | 1 | met | The Advisories, licenses & sources CI job fails the build on a RUSTSEC advisory |
| 15.2.2 | Implemented defenses against availability loss | 2 | met | Admission limiter with a queue budget and a request deadline, body limits on both listeners, a login limiter ahead of the KDF, per-call timeouts on every outbound subsystem, and kill_on_drop on script hooks |
| 15.2.3 | Production contains no test or development functionality | 2 | met | Test helpers are #[cfg(test)] or behind testutil; there is no sample data, no seeded account and no development route |
| 15.2.4 | Dependencies from expected repositories | 3 | met | cargo deny check sources restricts registries, and Cargo.lock pins every transitive dependency by hash |
| 15.2.5 | Extra protection around dangerous functionality | 3 | met | Script hooks are the dangerous surface and run in a cleared environment with a minimal PATH, a deadline and kill_on_drop; the CA key can be moved into a PKCS#11 token; the Containerfile is the network-isolation story |
| 15.3.1 | Return only the required subset of fields | 1 | met | Explicit serializers per resource; no row is serialized wholesale |
| 15.3.2 | Do not follow redirects unless intended | 2 | met | Intended, bounded and switchable: follow_redirects and max_redirects on the http-01 validator, with scheme and port checked on every hop (crates/net/src/challenge/http_01.rs) |
| 15.3.3 | Countermeasures against mass assignment | 2 | met | Request bodies deserialize into per-route structs holding only the fields that route accepts; nothing constructs a row from client JSON |
| 15.3.4 | Original client IP transferred correctly and used for decisions | 2 | met | filter.trusted_proxies is the allowlist of hops whose forwarded header is believed; empty means the header is ignored. The admin listener has its own list, admin.filter.trusted_proxies, and with it empty — the default — no caller can choose its own rate-limiter key (crates/core/src/client.rs, crates/admin/src/webadmin/filter.rs, crates/admin/src/webadmin/session.rs) |
| 15.3.5 | Explicit types and strict comparisons | 2 | met | Rust’s type system; there is no coercion to juggle |
| 15.3.6 | JavaScript written to prevent prototype pollution | 2 | n/a | The panel ships no application JavaScript |
| 15.3.7 | Defenses against HTTP parameter pollution | 2 | met | axum extractors read from one named source per parameter — a path segment, a typed query struct, or a JSON body — never from a merged bag |
| 15.4.1 | Thread-safe access to shared objects | 3 | met | Send/Sync are checked at compile time; shared mutable state is behind Mutex or Semaphore (LoginLimiter, Admission) |
| 15.4.2 | State checks and dependent actions are atomic | 3 | met | The single-use idiom is one statement: UPDATE … WHERE <still unused> decided by rows_affected, never a read followed by a write. Key files are created with create_new, which is the atomic form of “exists?” then “create” |
| 15.4.3 | Consistent locking, contained in the owning code | 3 | met | Locks are held inside the type that owns the resource and never across an await |
| 15.4.4 | Resource allocation prevents starvation | 3 | met | The admission limiter refuses past its queue budget rather than queueing without bound — the reasoning for not using GlobalConcurrencyLimitLayer is written out in crates/protocol/src/middlewares/admission.rs |
V16 Security Logging and Error Handling
| # | Requirement | L | Status | Evidence |
|---|---|---|---|---|
| 16.1.1 | A logging inventory exists | 2 | met | Monitoring enumerates every event = "…" name; Audit Trail states what the trail records, where it lives, who can read it and how retention works |
| 16.2.1 | Log entries carry when, where, who, what | 2 | met | The access line carries method, URI, status, latency, client address, profile and a request id (crates/protocol/src/middlewares/access.rs); an audit row carries actor, address, reverse name, identifiers, User-Agent and the same request id |
| 16.2.2 | Synchronized time sources; UTC or explicit offset | 2 | met | Timestamps come from the host clock as Unix seconds in the database and RFC 3339 in the log; host time sync is the operator’s |
| 16.2.3 | Logs only go to documented destinations | 2 | met | One tracing subscriber built in one place — prepare_logging in crates/server/src/logging.rs — with logging.target naming the sink, so a reload cannot drift from startup |
| 16.2.4 | Logs readable by the log processor | 2 | met | logging.json_format produces one JSON object per line, with flatten_event for pipelines that want fields at the top level |
| 16.2.5 | Sensitive data logged according to its protection level | 2 | met | Nonces and session tokens appear only as fingerprints; proxy credentials and the database URL’s password are redacted; a password never enters a log or argv — admin user passwd reads from stdin or --password-file |
| 16.3.1 | All authentication operations logged | 2 | met | admin_login_*, admin_mfa_verified, admin_mfa_attempts_exhausted, admin_logout and admin_password_hash_unreadable, each with the outcome and the method used |
| 16.3.2 | Failed authorization attempts logged | 2 | met | Filter denials, jws_url_mismatch, jws_jwk_and_kid_both_present, nonce_replayed and the *_failed audit rows. certificate_revoke_failed is written specifically so a run of them is visible as somebody enumerating serials |
| 16.3.3 | Security events and control-bypass attempts logged | 2 | met | challenge_validation_bypassed and two other weakened-configuration warnings repeat on every startup so they cannot become background noise (Hardening) |
| 16.3.4 | Unexpected errors and control failures logged | 2 | met | Backend, signer, DNS and IPAM failures each log with outcome = "failure" and their own event name |
| 16.4.1 | Logging components encode data to prevent log injection | 2 | met | JSON mode escapes structurally. In text mode the only client-controlled fields are the request URI, which http::Uri renders percent-encoded, and header values, which HeaderValue::to_str accepts only as visible ASCII — so a User-Agent carrying a control byte is dropped before it can be stored, let alone printed |
| 16.4.2 | Logs protected from unauthorized access and modification | 2 | met | The audit trail has no foreign keys, so deleting an account does not take its history; nothing in the panel can erase it — the audit surface is read-only and pruning is a host command (Audit Trail). The log stream itself is the operator’s to protect |
| 16.4.3 | Logs transmitted to a logically separate system | 2 | partial | The server writes to stdout or a file in a format built for shipping, and Monitoring shows the pipeline — but shipping them is the deployment’s job, not this process’s |
| 16.5.1 | Generic message to the consumer on unexpected errors | 2 | met | Every ACME refusal is an RFC 8555 problem document with a fixed type; internal detail goes to the log and not the body. The http-01 validator’s fetched body is never echoed into a client-visible error, precisely because it is attacker-chosen |
| 16.5.2 | Secure operation when external resources fail | 2 | met | A check that cannot reach its authority answers undecided rather than “allow”, so an IPAM outage degrades to a retryable 500 instead of failing open — the property Filters is built around |
| 16.5.3 | Fail gracefully and securely; no fail-open | 2 | met | Startup refuses rather than degrading: a non-loopback admin.bind_address without TLS, an unknown challenge type, a deadline below signer.custom.timeout_ms. tests/security.rs is the regression set for the request-path equivalents |
| 16.5.4 | A last-resort handler for unhandled exceptions | 3 | met | A CatchPanicLayer on each listener catches a handler panic, logs request_handler_panicked, and answers with that listener’s own error shape — an ACME problem document, or the admin JSON (/api) / HTML (/ui) error — instead of the aborted connection it used to be. panic = "abort" stays unset so the layer can unwind; the panic message goes to the log only |
Documented deviations
Places where this project has knowingly chosen differently from what ASVS asks. Each was argued before this assessment existed; the assessment’s job is to surface them against the standard, not to reverse them.
http-01 validation does not block private addresses — V1.3.6, V13.2.4,
V13.2.5. RFC 8555 requires following redirects, and Boulder’s mitigation —
refusing RFC 1918 targets — cannot apply to a server whose entire purpose is
serving private networks. What contains it instead: only http and https,
only the two configured ports, at most max_redirects hops, a shared timeout,
an off switch, and the fetched body never being echoed into a client-visible
error. →
HTTP-01
RS256 and RSA from 2048 bits — V11.2.3, V11.3.1. RFC 8555 §6.2 names
RS256 as an algorithm an ACME server must accept, and RS256 is
RSASSA-PKCS1-v1_5. Refusing it would refuse conforming clients. Two things
soften it: this is a signature scheme, not the encryption padding V11.3.1
targets, and the key in question is a client’s own account key, whose
compromise costs that client its account rather than costing the CA anything.
Raising the accepted floor to 3072 bits is a protocol-compatibility decision,
not a code change.
PBKDF2-HMAC-SHA256 rather than Argon2id — V11.4.2. Argon2id is the
stronger primitive. Adopting it would add four crates to a certificate
authority’s dependency graph — all audited on every cargo deny check, which
runs with all-features = true — for a subsystem that is disabled by default
and whose password is the bootstrap credential in a design that ends in a
second factor. 600 000 iterations is OWASP’s current recommendation for the
non-Argon2 case, and the stored form is self-describing so the trade can be
revisited without a migration. The argument is in the module doc of
crates/admin/src/admin/password.rs.
admin.require_mfa defaults to false — V6.3.3. Defaulting it on would
brick a panel whose first operator has not enrolled yet, with no way in to fix
it. The hardening checklist tells operators to turn it on, and the panel
supports a bootstrap flow where enrolment is the only thing a session can do.
An L2 claim for the web admin depends on the operator setting it. →
Hardening
No hardware-based authentication factor — V6.3.3 at L3. WebAuthn was
investigated and deferred, and both blocking checks were actually run:
webauthn-rs 0.5.5 is MPL-2.0, which deny.toml’s allow list does not carry,
and webauthn-rs-core hard-depends on openssl, which this tree has avoided
at every turn. Nothing in the design precludes it — another factor kind is
another MfaStep variant, not a change to the state machine. It stays open as
issue #5.
The relay backend multiplexes one upstream account — V8.3.3. One
upstream ACME account, and one centrally held RFC 2136 TSIG key, standing in
for every local client. That is the whole point of the backend: not
distributing a scarce credential is what it exists to do. Every authorization
decision is made locally, before the upstream is ever asked. →
Relay
Two listeners, usually two ports on one host — V3.5.4. They are separate
sockets with separate TLS, separate authentication and separate defaults, and
the admin one binds loopback unless TLS is on. They are not separate
hostnames, and cookies are not port-scoped — which is exactly why the panel
does not rely on SameSite for CSRF and carries a per-session token plus an
Origin check instead. →
Web Admin
The audit trail is a record, not a control — V8.2.4. Nothing in the server compares a live request against the trail. Pinning an identity to an address breaks CGNAT and mobile clients, and that is a deliberate non-feature. It is stated as such in the Security Model.
Gaps
Open shortfalls, worst first. What remains is all L3, recorded only here.
Lower-priority L3 items, recorded here only and with no issue open: no CSP
violation-report endpoint (V3.4.7), no Cross-Origin-Opener-Policy (V3.4.8),
no documented behaviour for browsers lacking security features (V3.1.1,
V3.7.5), no OCSP stapling as a TLS server (V12.1.4), no Encrypted Client Hello
(V12.1.5), no post-quantum migration plan (V11.1.4), no multi-user approval for
issuance (V2.3.5), and admin user passwd letting the resetter learn the
password (V6.4.6).
Re-running this
The requirement text is vendored at rfc/asvs-5.0/, so this page can be
re-derived against a later ASVS release by diffing the chapter files and
revisiting only the rows whose requirement text moved. The per-chapter tables
enumerate every in-scope requirement rather than only the failures for
exactly that reason: a list of gaps alone cannot be compared against anything.
Architecture
This page is about how the pieces fit, and about the handful of decisions that are load-bearing enough that changing them would break something non-obvious. The organising ideas are three: RFC 8555’s checks are hoisted into an extractor so no route can forget them, an ACME endpoint is a profile, and signing, filtering and notifying are each a trait with several implementations.
The workspace
One binary over a Cargo workspace. The root package, acme-proxy, holds the
clap command tree (src/cli/), main.rs and the integration tests; nine
library crates under crates/ hold everything else, each naming only the
crates beneath it:
| Crate | What it holds | Depends on |
|---|---|---|
acme-proxy-core | configuration, the ACME wire types, certificate parsing, the audit vocabulary | — |
acme-proxy-store | the storage layer over SQLite or PostgreSQL, one module per table, and both migration sets | core |
acme-proxy-net | DNS, outbound HTTP and proxies, TLS, listeners, the challenge validators | core |
acme-proxy-policy | the filter engine and the IPAM inventories | core, net |
acme-proxy-jobs | the job queue, notifications, the audit writer, metrics | core, net, store |
acme-proxy-signer | the signing backends and their read side | core, jobs, net, store |
acme-proxy-protocol | the ACME services, extractors, handlers and routers | all of the above |
acme-proxy-admin | the operation layer and the web admin panel | protocol and below |
acme-proxy-server | the runtime: roles, listeners, reload, logging | all of the above |
The crate edges are the layering: a handler cannot reach the runtime that
serves it, and a job queue cannot reach the filters, because the compiler
refuses the import. tests/layering.rs pins each crate’s dependencies to this
table, so an edge across a layer — which Cargo would accept — is a deliberate
change rather than a drive-by one.
Request flow and extractors
Nearly every ACME endpoint is signed by the client using JSON Web Signatures
(JWS). Rather than parsing this manually in each handler, the server leverages
an AcmeRequest<T> Extractor.
The JWS extractor core
Eight checks run in a fixed order, and five of them have their own way out. The
shape matters more than the list: the jwk/kid branch in the middle is where
the two security properties below live, and a linear numbered list hides it.
graph TD
REQ["Signed POST"] --> CT{"Content-Type is<br/>application/jose+json?"}
CT -->|no| E415["415 — body never read,<br/>so no nonce is burned"]
CT -->|yes| DEC["Decode the flattened JWS<br/>and its protected header"]
DEC -->|unparsable| EMAL["malformed (400)"]
DEC --> CRIT{"crit header present?"}
CRIT -->|"yes — this server<br/>implements none"| EMAL
CRIT -->|no| URL{"JWS url equals<br/>the route reached?"}
URL -->|"no — §6.4"| EMAL
URL -->|yes| AUTH{"jwk or kid?"}
AUTH -->|"both, or neither"| EMAL
AUTH -->|jwk| JWK["Re-encode the key as DER SPKI"]
AUTH -->|kid| KID["Load the account, then check the<br/>stored SPKI's own OID against alg"]
KID -->|"unknown kid"| EACC["accountDoesNotExist (400)"]
KID -->|"OID does not match alg"| EALG["badSignatureAlgorithm (400)"]
JWK --> SIG{"Signature verifies?<br/>ES256 or RS256, via ring"}
KID --> SIG
SIG -->|no| E401["unauthorized (401)"]
SIG -->|yes| NONCE{"Nonce fresh and unused?"}
NONCE -->|"no — §6.5"| EBAD["badNonce (400)"]
NONCE -->|yes| H["Handler"]
Only once all of this succeeds is the request handed to the axum handler. Hoisting these checks into the extractor makes them structural: a new signed route cannot forget them, and no handler repeats a four-line preamble.
Three extractors build on that core: AcmeRequest<T> (decode and deserialize
the payload), AcmePostAsGet (require an empty payload, else malformed), and
AcmeOptionalPayload<T> — the last exists for the authorization resource, where
one URL serves both a POST-as-GET read and a §7.5.2 deactivation.
Two security properties worth not breaking
jwk and kid are mutually exclusive (RFC 8555 §6.2) — the branch in the
middle of the diagram. Both present, or neither, is malformed. An embedded
jwk is verified and re-encoded as DER SPKI; a kid is resolved to its account
and verified against the account’s stored SPKI.
The verification algorithm never rests on alg alone. On the kid path,
the stored SPKI’s own AlgorithmIdentifier OID is checked against the client’s
claimed alg before verification — that is the badSignatureAlgorithm exit. EC
coordinates must additionally be exactly 32 octets (RFC 7518 §6.2.1.2) — a short
or long one parses as a different point, which would register one key as two
accounts.
Database & persistence
The server uses sqlx over SQLite or PostgreSQL, chosen by database.url’s
scheme (ADR 0014). The SQL is written
once, in crates/store/src/sql.rs, the only module that names either driver.
The connection pool is private to crates/store/src/. Everything else reaches
the database through a table module, Database::transaction() (a Tx whose
conn() is the connection the table methods take) or
Database::pool_stats() (the metrics gauge), so SQL and its dialect stay in one
module tree. Database::raw_pool() exists only for test fixtures, and
tests/layering.rs fails the build when production code calls it.
Migrations
Migrations are embedded with sqlx::migrate!(), frozen once committed, and
applied only by acme-proxy migrate/init and a process running the worker
role; every other entry point refuses a database whose schema is behind. The
rules and the reasons are in
ADR 0003, and the how-to in
Contributing. The schema is
also the only frozen surface before 1.0.0
(ADR 0001).
Schema details
The tables, their constraints and the reasoning behind each — profile isolation,
the CHECKed state machines, the audit trail’s deliberate lack of foreign keys,
and the three different ways a secret is stored — have their own page:
Database Schema.
Revocation is the one piece worth naming here, because it constrains the request
path rather than the schema: it writes a reason and a timestamp and deliberately
does not touch the order status, since RFC 8555 defines no “revoked” order
status.
Two front ends, one operation layer
src/cli/ and crates/admin/src/webadmin/ are two front ends;
crates/admin/src/admin/ is the operation layer both dispatch to and neither
owns.
src/cli/ crates/admin/src/webadmin/
(clap) (axum)
\ /
\ /
crates/admin/src/admin/ops.rs — delete_account, revoke_order, load_order_detail…
crates/admin/src/admin/users.rs — create_user, authenticate, set_password…
crates/admin/src/admin/changes.rs — an operator change: write, audit row, sessions, notification
crates/admin/src/admin/render.rs — render_*_json, the one JSON shape of each object
crates/admin/src/admin/password.rs — the KDF, shared by both
A handler in crates/admin/src/webadmin/handlers/ is a few lines over an
admin::ops call and a render_*_json, the same way a src/cli/ command body
is a few lines over the same call and either the same render_*_json or a
render_*_line of its own, in src/cli/render.rs. That is what keeps the
password policy, the duplicate check and the rehash-on-login identical between
them.
Two consequences worth knowing:
- The destructive operations come in pairs.
delete_account(id, db)simply deletes;confirm_delete_account(id, yes, reader, db)asks first. The split exists becauseassume_yes: bool+reader: &mut impl BufReadare a terminal’s concerns — a caller with no terminal was passingtrueand an empty reader, asserting a confirmation that never happened. The CLI calls the wrapper; the web calls the bare form. crates/admin/src/webadmin/is notcrates/admin/src/admin/web/. That would invert the dependency, putting an HTTP server inside the operation layer.
The admin listener is assembled by webadmin::build_admin_app, which takes
&[Arc<Profile>] as a slice — build_app consumes the Vec, so the admin
side must be built first, and the signature is where that ordering is stated
rather than a borrow error to rediscover. Its state is AdminState, not
AppState: the latter holds exactly one Profile, and this listener is
cross-profile by nature (revoking an order needs that order’s own profile’s
revocation route, which may name a different CA from any default).
Order lifecycle
sequenceDiagram
participant Client
participant Axum Router
participant Filters
participant Order Manager
participant Job Queue
participant Challenge Validator
participant Signer Backend
Client->>Axum Router: POST /newOrder
Axum Router->>Filters: Validate Client IP & Identifiers
Filters-->>Axum Router: Allow/Deny
Axum Router->>Order Manager: Create Order + Authorizations + Challenges
Order Manager-->>Axum Router: Order Object (status: pending)
Axum Router-->>Client: 201 Created
Client->>Axum Router: POST /chall/{id} (trigger)
Axum Router->>Job Queue: claim + enqueue challenge_validate
Axum Router-->>Client: 200 OK + challenge object (processing)
Job Queue->>Challenge Validator: Validate domain control
Challenge Validator-->>Job Queue: Pass/Fail
Job Queue->>Order Manager: Commit challenge + authz + order
Note over Order Manager: Order -> "ready" once every<br/>authorization is valid
Client->>Axum Router: POST /chall/{id} (poll)
Axum Router-->>Client: 200 OK + challenge object (either way)
Note over Client,Challenge Validator: This exchange in detail:<br/>Challenge Validation
Client->>Axum Router: POST /finalize (with CSR)
Axum Router->>Filters: Re-validate identifiers from CSR
Filters-->>Axum Router: Allow/Deny
Axum Router->>Job Queue: claim + enqueue signer_issue
Axum Router-->>Client: 200 OK + order (processing)
Job Queue->>Signer Backend: Request Signature (worker only)
Signer Backend-->>Job Queue: Signed Certificate
Job Queue->>Order Manager: Record certificate (order -> valid)
Client->>Axum Router: POST-as-GET order (poll)
Axum Router-->>Client: 200 OK (Certificate URL)
Four properties hold this together. Three are transactional:
- Order creation inserts the order, its authorizations and their challenges in one transaction. A half-written order would be finalizable for names that were never authorized.
- A validation outcome commits the challenge, the authorization and the order
together, and the “is every authorization valid?” read happens inside
that transaction. From the pool, two concurrent validations of one order could
each read before the other’s write landed, and neither would promote the order
to
ready. Every one of those writes is also guarded on the state it leaves, since the verdict is computed by a job long after the request that claimed the challenge read those rows. finalizeclaims the order (ready → processing) and queues itssigner_issuejob in one transaction, so no crash can leave an orderprocessingwith nothing coming to settle it. The job runs in theworkerrole — the only one that builds a signing backend — which is what keeps the CA key out of the process parsing client requests.
The fourth is about the answer rather than the write: challenge validation
returns 200 plus the challenge object whether it passed or failed
(§7.5.1). A 4xx would surface as a transport failure to certbot’s acme
library rather than as a failed challenge.
Pluggable signing keys
The SignerBackend trait is the seam for how a
certificate is obtained — locally, from an upstream ACME server, or from a
script. Inside the local_ca backend there is a second, narrower seam for
where the private key lives, and it is worth knowing that it is rcgen’s own
trait, not one this project invented.
graph LR
ISSUE["LocalCa::issue<br/>LocalCa::revoke"] --> SB["spawn_blocking<br/>— unconditionally"]
SB --> ISS["Issuer<'static, CaSigningKey>"]
ISS --> SW["Software(KeyPair)<br/>a PEM file on disk"]
ISS --> PK["Pkcs11(...)<br/>behind --features hsm"]
PK --> MOD["the PKCS#11 module (.so)"]
MOD --> TOK[("token / HSM")]
The spawn_blocking sits before the branch, not inside one arm of it: see
the first consequence below.
rcgen::SigningKey is public, and every signing entry point local_ca uses is
generic over it:
| call | signature |
|---|---|
csr.signed_by(&issuer) | signed_by(&self, issuer: &Issuer<impl SigningKey>) |
CertificateRevocationListParams::signed_by(&issuer) | signed_by(&self, issuer: &Issuer<'_, impl SigningKey>) |
params.self_signed(&key) | self_signed(&self, signing_key: &impl SigningKey) |
So LocalCa holds an Issuer<'static, CaSigningKey> — a small enum in
crates/signer/src/local_ca/key.rs with a Software(KeyPair) variant and,
behind the hsm feature, a Pkcs11(..) one. Adding a key source (a cloud KMS,
a remote signer daemon) means adding a variant that implements two rcgen trait
methods: sign, der_bytes/algorithm. Nothing in issue, revoke,
crl_der or the CSR sanitisation changes, because none of it ever names the key
type.
Two consequences worth preserving:
- Signing runs on the blocking pool.
rcgen::SigningKey::signis synchronous and called from deep insidesigned_by, so there is nothing to await through. A key source that talks to hardware or a network would otherwise stall a runtime worker for its whole round trip, soLocalCa::issueand the CRL rebuild inrevokeboth go throughspawn_blockingunconditionally — not gated on the variant, which would be a branch someone eventually gets wrong. - The key type must be
Send + Sync. It lives inside anArc<dyn SignerBackend>. Where the underlying handle is not (cryptoki’sSessionisSendbut notSync), astd::sync::Mutexis the right wrapper: the signing call never awaits, so an async mutex would buy nothing.
Architecture Decision Records
An architecture decision record (ADR) holds the why behind a design: the
problem it answers, the choice made, and what that choice rules out. The rest of
the documentation says what the server does. The module docs (//!) say what
each piece of code guarantees. The ADRs are where the argument lives, so that
the argument is not restated wherever the decision is mentioned.
Read the relevant ADR before changing something it covers. If a change reverses a decision, write a new ADR that supersedes the old one and change the old one’s status. Do not rewrite its argument after the fact.
| ADR | Decision | Status |
|---|---|---|
| 0001 | Before 1.0.0, only the database schema is a compatibility promise | Accepted |
| 0002 | One binary over a layered workspace of lockstep crates | Accepted |
| 0003 | Migrations are append-only and applied only by the schema owners | Accepted |
| 0004 | Row ids are UUID v7 stored as BLOBs, and their type says where they came from | Accepted |
| 0005 | A Rust enum owns each vocabulary, and SQL checks only the closed ones | Accepted |
| 0006 | A request does no slow or privileged work; it queues it | Accepted |
| 0007 | One binary runs as role processes, and only the worker holds the CA key | Accepted |
| 0008 | State that more than one process can see lives in the database | Accepted. One exception stands, the web admin’s login limiter, for as long as |
| 0009 | Dependencies are pure Rust on ring, add no global state, and earn their place | Accepted; metrics clause superseded by 0011, crypto clause narrowed by 0014 |
| 0010 | Errors derive thiserror, carry their whole message, and panic only at startup | Accepted. This was re-argued more than once before it was written down, which |
| 0011 | Metrics are built on prometheus-client, one registry per scrape | Accepted |
| 0012 | Images are built natively per architecture, uncached, and published only past a guard | Accepted |
| 0013 | main is the trunk, and a patch line is a release branch cut when a fix needs one | Accepted |
| 0014 | PostgreSQL is chosen by the URL’s scheme, over one set of queries | Accepted |
Decisions argued elsewhere
Some decisions are explained where their subject lives, and have no record here, because a second copy would drift from the first.
In the book:
- Hoisting the JWS checks into an extractor, and the two security properties it must keep.
- Pluggable signing keys:
rcgen’s ownSigningKeyas the seam, and signing on the blocking pool. - Secrets are stored three different ways, on purpose
- Columns nothing ever compares against: no identity is pinned to an address or a User-Agent.
- Evidence has no foreign keys: the audit trail and the revocation ledger outlive what they describe.
- Profiles: an endpoint is a profile, its path is derived from its name, and it is a database boundary.
- Bypass is not a shortcut: why validation is on by default.
- When a check cannot decide: the filter’s three-valued answers.
- Reloading the configuration: a reload is a rebuild and a swap, all or nothing.
- Why a second listener, and the web admin’s CSRF, roles and read-only audit view.
- Why the audit trail survives deletion.
- Delivery semantics of notifications.
- Paging: every listing is paged and says so.
In a module’s own documentation (//!), where the decision concerns that
module alone:
crates/jobs/src/jobs/mod.rs: why background work is a durable queue and not atokio::spawn, and theRetry/Failedsplit every handler must honour.crates/signer/src/local_ca/crl.rs: how the CRL number stays monotonic across processes.crates/policy/src/filter/policy.rs: Kleene logic, and why a rule’s stages are an intersection.crates/policy/src/ipam/mod.rs: why an inventory never denies, so an outage cannot fail open.crates/core/src/script_hook.rs: the one contract everycustomscript runs under.crates/core/src/templating.rs:.htmlescapes and.j2does not, decided by the name.crates/jobs/src/metrics.rs: bounded cardinality, families declared even when empty, and counters that survive a reload.crates/jobs/src/notify/expiry.rs: why the expiry notice is a digest.crates/server/src/reload.rs: the swap’s mechanics and what stays frozen.
Writing an ADR
Name the file NNNN-kebab-title.md, taking the next free number, and list it
both in the table above and in SUMMARY.md under this page. doc/lint.py
refuses an ADR that is missing from either list, or that lacks one of the five
sections below.
The title is # ADR NNNN: <decision>, and the sections come in this order:
## Status:Accepted, orSuperseded by ADR NNNN. Name a proposal to replace it if one is on record.## Context: the problem, and the history that makes the rule necessary. History with no rule behind it does not belong here — git andCHANGELOG.mdalready keep it.## Decision: what was chosen, stated as rules.## Consequences: what the decision costs and what it rules out.## Enforced by: the tests, startup refusals and type-level constructs that hold the decision in place, named so that a grep finds them. Write “Review only” when nothing does.
Link to the page that owns a configuration key rather than restating its default, since the book documents each key in exactly one file.
ADR 0001: Before 1.0.0, only the database schema is a compatibility promise
Status
Accepted.
Context
A project that promises stability on every surface from its first release ends up carrying a compatibility layer for every shape it has ever had: aliases for renamed keys, dual syntaxes, lowering passes that translate an old section into a new one. Each of these is code that must be tested, documented and reasoned about for as long as the promise lasts, and each makes the next redesign harder.
acme-proxy is still finding its design. Several subsystems were replaced
wholesale rather than extended: the filter chain became a policy engine
(Policy), the Mattermost notifier became a generic
webhook (Webhook), the acme_proxy signer
became relay, and filter.netbox became [ipam]. None of these would have
been worth doing if each had to keep reading the old shape.
One surface cannot be treated this way. The database holds accounts, orders and
issued certificates that clients and relying parties depend on, and sqlx
checksums every applied migration. An upgrade that cannot open the existing
database is not an upgrade.
Decision
Before 1.0.0 the database schema is the only compatibility guarantee:
crates/store/migrations/ is append-only (ADR
0003), so an upgrade is starting the
new binary against the existing database.
Everything else may be renamed or removed in any release: configuration keys, profile names and the ACME URLs derived from them, the admin JSON API, log event names and the CLI. Such a change owes exactly three things:
- An entry under the release’s
### Breakingheading inCHANGELOG.md, naming the old spelling and the new one. The changelog’s Compatibility section is the canonical statement of this policy; the README,SECURITY.mdand the book each carry one sentence linking to it. - A startup refusal naming the removed key wherever practical, so an unmigrated configuration stops the server instead of coming up looking configured. This is a one-line error message, not a compatibility path.
- Never an alias, a dual syntax or a legacy lowering. The old shape is deleted; the new code does not learn to read it.
The refusals themselves are removed at 1.0.0.
Consequences
- A redesign costs one changelog entry and one diagnostic, which is what makes replacing a subsystem cheaper than accreting around it.
- A key must still parse to be refused by name. The removed
[filter]fields therefore survive incrates/core/src/config/types/filter.rs; a field that is gone would fail as an opaque serde error instead of a named one. - Operators must read the Breaking section before an upgrade.
acme-proxy filter showbuilds a policy exactly as startup does, so it checks a migrated configuration before a restart. - Frozen means frozen: the comments inside committed migrations still say
acme_proxywhere the code saysrelay, because editing them would change their checksum (ADR 0003).
The how-to for a rename is in Contributing.
Enforced by
refuse_removed_keysincrates/policy/src/filter/build.rs, guarded byevery_removed_key_is_refused_by_name.the_old_acme_proxy_backend_name_is_refused_by_its_new_one(crates/signer/src/lib.rs) andthe_removed_mattermost_backend_is_refused_by_name(crates/jobs/src/notify/mod.rs).the_netbox_type_is_refused_by_name(crates/policy/src/filter/build.rs).- The Breaking entry and the “no alias” rule are review only.
ADR 0002: One binary over a layered workspace of lockstep crates
Status
Accepted.
Context
The server was one crate of about 106,000 lines. Every change recompiled all of it, and nothing but review stopped the web admin from issuing SQL or the notifier from reaching into a signer. Layering held by convention, and a convention is broken by the first drive-by import that compiles.
The code was not meant as a library for anyone else. It exists so the binary and the integration tests can reach it. That shapes what a split has to buy: compiler-enforced edges and faster rebuilds, not a public API.
Decision
- One binary, nine library crates. The
acme-proxypackage at the repository root holds theclapcommand tree (src/cli/),main.rsand every suite intests/. The libraries undercrates/are, bottom-up:core,store,net,policy,jobs,signer,protocol,admin,server. The crate map is in Architecture. - A crate names only crates beneath it. Cargo refuses a cycle but accepts an
edge that skips across a layer, so the intended edges are pinned in a table
(
CRATE_DEPSintests/layering.rs) and a test compares every member’s[dependencies]against it. Adding an edge means editing that table on purpose. - Lockstep, with no semver promise. Every member takes
version.workspace, and every internal dependency is pinned with=in[workspace.dependencies]. An item becomespubonly because another crate needs it, not because it is an API. - Test scaffolding lives per crate, behind a
test-utilfeature (#[cfg(any(test, feature = "test-util"))] pub mod testutil). Only other crates’[dev-dependencies]turn it on, so no normal build ships it. A fixture belongs to the lowest crate its types allow. - Every cargo command takes
--workspace. At a root that is also a package, a bare command acts on that package alone, and the library crates would drop out of lint, tests and coverage with nothing going red.
Consequences
- A handler cannot reach the runtime that serves it, and the job queue cannot reach the filters. The compiler refuses the import.
- A doc link cannot point up the graph, since a crate cannot link to its dependants. Such a mention is a plain code span instead.
- A member’s unit tests run in the member’s own directory. A test that reads a
repository file anchors on
env!("CARGO_MANIFEST_DIR"). tracingtargets are the crate paths (acme_proxy_store::db, …). The default filteracme_proxy=infostill covers them all becauseEnvFiltermatches targets by string prefix.- All ten packages are published together. A release bumps the workspace version
once, and every
=pin with it.
Enforced by
- Cargo, for cycles.
crate_dependencies_follow_the_layersintests/layering.rs, for edges across a layer.- The CI jobs, each of which passes
--workspace(.github/workflows/ci.yml).
ADR 0003: Migrations are append-only and applied only by the schema owners
Status
Accepted.
Context
The schema is the one surface this project promises not to break (ADR
0001), and sqlx enforces that promise
mechanically: it records a checksum for every migration it applies. Editing a
committed file does not quietly diverge a deployment’s schema. It makes the next
startup fail with a checksum mismatch. Before the first release, a schema change
meant editing the migration in place and deleting sqlite.db. That stopped
being possible once a database outside the repository had run the files.
SQLite adds its own constraints:
- It cannot add a
CHECK,UNIQUEor foreign key to an existing table. - It gives
VARCHAR(n)TEXT affinity and enforces no length. - Under
foreign_keys = ON,DROP TABLEperforms an implicitDELETE FROM, which firesON DELETE CASCADEinto every child table. - It gives
sqlxno migration lock.
Migrations used to run inside Database::connect. That made every subcommand an
upgrade step. Once the server could run as several processes (ADR
0007), two processes starting together raced
MIGRATOR::run with nothing to serialise them.
Decision
Append-only. Every file in crates/store/migrations/ is frozen, comments
included — and, since ADR 0014, so is every
file in crates/store/migrations-postgres/. A schema change is a new file in
each set (sqlx migrate add <name>). The rules below describe SQLite’s
constraints; PostgreSQL needs no rebuild for a CHECK or a width, so its file
is usually the one-line ALTER TABLE the change actually is:
- A new column is a new
ALTER TABLE … ADD COLUMNfile, even when it plainly belongs to an existing table. - A new or dropped
CHECK,UNIQUEor foreign key is a table rebuild, written in a new migration. - A wrong declared width is also a rebuild. Where a width follows a constant in the code, a test pins the two together.
- A rebuild re-creates the table’s indexes, since
DROP TABLEtakes them with it. It also names every column in itsINSERT … SELECT, since a forgotten column is dropped silently. Rows in a table with children are parked in constraint-freeCREATE TABLE … AS SELECTcopies before anything is dropped, then put back parent-first. Otherwise the cascade empties the children.
Applied explicitly. Database::open connects and never migrates.
Database::migrate has two callers:
acme-proxy migrateandacme-proxy init;- a
serveprocess that runs theworkerrole (server::apply_or_require_schema).
Every other entry point calls pending_migrations and refuses by name, pointing
at acme-proxy migrate. A default single-process serve includes the worker
role, so it still migrates a fresh database on its own.
The pool is private to crates/store/. Everything else goes through a table
module, Database::transaction() (a Tx whose conn() hands out the
connection) or Database::pool_stats(). Database::raw_pool() exists for test
fixtures only. On SQLite the connection pins two pragmas:
foreign_keys(true), because the schema’sON DELETE CASCADEdepends on it;journal_mode = WAL, because every ACME response writes a nonce row, and the default rollback journal takes a database-wide exclusive lock on each write.
PostgreSQL needs neither — foreign keys are always enforced and there is no journal mode to choose — so its pool pins a connection limit instead, which is the resource several role processes there actually share.
Runtime queries. Queries use sqlx::query(...), not the compile-time
query! macros, so building needs no DATABASE_URL.
Consequences
- An upgrade is starting the new binary. There is no dump/restore procedure.
- Several constraints were declared before anything wrote to them:
admin_users.totp_*andadmin_sessions.state’s'pending_mfa'. Adding them afterwards would have cost a rebuild each. - Stale comments stay stale. Three migrations still call the relay backend
acme_proxy. Treat grep results incrates/store/migrations/as read-only. sqlx::migrate!()embeds the migration set at compile time. Adding a file does not invalidate the build on its own; touchcrates/store/src/db.rs.- SQL, and the dialect it is written in, lives in one crate. That is what kept a second backend a contained change when PostgreSQL arrived in ADR 0014.
- PostgreSQL gives
sqlxan advisory migration lock where SQLite gives it none, so the race above cannot happen there. The one-owner rule still holds on both: it is about which process may own the schema, which is a deployment property, not only about the race that exposed it.
Enforced by
sqlx’s checksums, at startup.only_the_schema_owners_apply_migrationsandproduction_code_never_reaches_the_raw_poolintests/layering.rs.- Rebuild guards in
crates/store/src/db.rs:the_blob_migration_preserves_every_row;the_audit_log_rebuild_keeps_every_row_and_relaxes_the_event_check.
- Width pins in the same file:
declared_token_widths_match_random_token;declared_issuer_widths_match_the_issuer_id;every_id_column_is_declared_a_blob.
ADR 0004: Row ids are UUID v7 stored as BLOBs, and their type says where they came from
Status
Accepted.
Context
Row ids started as UUID v4 strings. Two costs followed:
- A random id is written into an index at a random leaf on every insert. SQLite feels that mildly, and PostgreSQL would pay a page split and a full-page WAL write per row. That backend is issue #4.
- The paged listings break ties on a whole-second
created_atwithORDER BY created_at, id. With v4 ids, rows created in the same second come back in a different random order for each pair.
Converting ids from text to BLOBs added a third problem. SQLite never compares a
bound String equal to a BLOB, so an internal lookup that kept a &str
parameter would match nothing, silently, for ever. The case that exposed this
was a job-retention test. It inserted a job with a hand-written id and asserted
it was gone afterwards. Had Job::find_by_id taken a &str and parsed it
internally, that assertion would have passed whether or not the sweep had run.
Decision
-
Every id is minted in one place (
mint()incrates/store/src/id.rs), as a UUID version 7. The leading 48-bit millisecond timestamp makes ids created close together share a prefix, and makes them sort by creation. -
Ids are stored as the 16 bytes themselves.
sqlx’suuidfeature encodes aUuidas a SQLite BLOB, and maps the same type to PostgreSQL’s nativeuuid. -
Existing rows keep their v4 ids. An id is a foreign key, a
kidis a credential a client stored, and an order id is part of a URL a client polls. A table can therefore hold both versions, and only the v7 ids sort by creation. -
The parameter type records provenance. In
crates/store/src/:- an id typed
&strmay be junk from outside the process, so it is parsed and an unparseable value answers “absent”; - an id typed
Uuidwas read out of a row.
There is no third case, and no parallel
_uuidvariant of any function. - an id typed
-
parse()is deliberately narrower thanUuid::try_parse. It accepts only the 36-character hyphenated form thatUuid::to_string()produces, so an id written in another spelling answers “not found” rather than resolving. -
Columns that look like ids but do not point at a row stay text. The list, and how to query a BLOB id by hand, are in Database Schema.
Consequences
- An internal lookup handed a stale string fixture is a compile error rather than a silently empty result.
- Only a few functions keep
&str, where the value genuinely arrives from outside the process (the request path and job payloads). They are listed incrates/store/src/id.rs. - The conversion migration (
20260827120000_uuid_ids_as_blobs.sql) is the worked example of the DROP-cascade hazard in ADR 0003. - Some fresh values are deliberately not row ids and do not go through
mint(): thex-request-idfallback, the job lease owner and a notification’sdelivery_id.
Enforced by
- The type system, for
&strversusUuid. every_id_column_is_declared_a_blobandthe_blob_migration_preserves_every_row(crates/store/src/db.rs).- The unit tests of
parseincrates/store/src/id.rs.
ADR 0005: A Rust enum owns each vocabulary, and SQL checks only the closed ones
Status
Accepted.
Context
Many columns hold a word from a fixed list: an order’s status, an audit event,
an operator’s role, a job’s kind. There are two places the list can be enforced:
a SQL CHECK constraint, or a Rust type.
A CHECK catches a typo before it parks a row in a state nothing can reach. But
SQLite cannot alter one without rebuilding the table (ADR
0003). Once a vocabulary grows, its
constraint becomes a migration per new word, and a rolling upgrade breaks: an
older binary refuses to write a word only a newer one knows.
String literals in Rust are worse than either. Before
crates/store/src/status.rs existed, about 30 comparisons against literals were
spread across handlers, storage and the relay flow. A typo compiled, and
silently changed policy.
Decision
- Every vocabulary is a Rust enum, and the enum is the authority. Values are
written only through its
as_str(), and read back through a parse that handles an unknown word explicitly. - Closed vocabularies also keep a
CHECK. Everystatuscolumn (accounts,orders,authorizations,challenges,jobs,upstream_orders,eab_keys,admin_users) changes only with its state machine. The enums instatus.rsreproduce their columns’ existing strings byte for byte, so the frozen constraints stay valid. - Open vocabularies carry no
CHECK. These are vocabularies that grow with features:audit_log.event.20260909120000rebuilt the table to drop thatCHECKwhen the administrative actions widened it.AuditEventis the authority, andAuditEvent::outcomeis an exhaustive match, so a new name must classify itself as a success or a failure.admin_users.role.NULLreads asadmin, so an operator created before the column existed keeps full authority. Any other unknown value folds toviewer, the least privilege.jobs.kind, which is an open set by design.
- Operator surfaces refuse an unknown value by name (Admin
CLI) wherever the vocabulary is closed, since “no
rows” reads exactly like “nothing is in that state”.
--kindis the exception, because kinds are an open set. Job::statusandUpstreamOrder::statusstayStringin their models. An older binary must still render a row a newer one wrote. Their enums exist for the operator surface.
Consequences
- A new audit event or role is a Rust change and a changelog line, with no migration.
- A value written by hand into the database is handled by the parse rule, never
trusted: a garbage role reads as
viewer. - The database alone no longer guarantees that
audit_log.eventholds a known word. The insert path bindsAuditEvent::as_str(), never a free string.
Enforced by
- The
CHECK (status IN (…))constraints in the migrations. EVENT_COUNTandALL_AUDIT_EVENTSincrates/core/src/audit/mod.rs, a compile-time assertion that no variant is missing from the list.AdminRole::from_storageincrates/store/src/admin_user.rs.the_audit_log_rebuild_keeps_every_row_and_relaxes_the_event_check(crates/store/src/db.rs).
ADR 0006: A request does no slow or privileged work; it queues it
Status
Accepted.
Context
Three ACME operations used to do their real work inside the HTTP request:
POST /chall/{id}awaited the challenge validation. That reached out to an address the client named, over DNS, HTTP or TLS, to a host that may never answer. It held an admission permit for the whole ofchallenge.timeout_ms. Up toserver.max_concurrent_requestsclients pointing at black-holed addresses could pin every permit. It also forced a startup rule that the request timeout exceed the validation timeout.finalizecalled the signing backend. The process parsing untrusted JWS and CSRs from the internet therefore heldca.key, a PKCS#11 login or a relay’s upstream account.- Revocation called the backend too, from
POST /revokeCert, from the admin panel, and fromorder revokeon the host. Run beside a live server, the last of these could silently drop a revocation from the CRL.
RFC 8555 already has the states that queued work needs. A challenge is
processing while “the server is working on it” (§7.1.6), an order is
processing while “the certificate is being issued” (§7.4), and §8.2 pairs both
with Retry-After. certbot, acme.sh and lego all poll.
Decision
A request claims, queues and answers. The worker does the work.
- Validation.
POST /chall/{id}claims the challenge (pending → processing, a compare-and-swap), writes achallenge_validatejob, and answers200with the challenge readingprocessingplusRetry-After. Seecrates/protocol/src/acme/validate.rs. - Issuance.
finalizechecks the CSR and runs the filter synchronously. It then claims the order and queuessigner_issuein one transaction, and answersprocessingfor every backend. Seecrates/protocol/src/acme/issue.rs. - Revocation. A local CA’s revocation is a
revocationsrow and the order’s stamp in one transaction; the worker then signs the CRL. A relay orcustomrevocation is asigner_revokejob that the request waits on, for up toserver.request_timeout_msless a second, answering503+Retry-Afterif the job is still running. Seecrates/protocol/src/acme/revoke.rs. - A verdict is terminal. A check that ran records its answer, pass or fail.
A job is retried only when the attempt could not happen at all: the database
is unreachable, or this process does not mount the profile. A backend’s own
BadCsrmakes the orderinvalid, notreadyagain, because the client was already toldprocessingand polls for a terminal state. - No stranded claim. A failed enqueue releases its claim.
abandonrecords a failure rather than leaving the client polling forever.recoverre-queues a challenge leftprocessingwith no live job. - Each kind has one handler, covering every profile. A row names its subject, the subject names its profile, and the profile names its validators or backend. The job registry refuses a second handler for a kind anyway.
Consequences
challenge.timeout_msbounds a job attempt, not an HTTP request. The only timeout rule left at startup concernssigner.custom’s read hooks, the one backend call still made inline.- A client sees
processingon every finalize, including against a local CA. This shipped under### Breaking. - The request path never names a signing backend, which is what lets the CA key leave the ACME process entirely (ADR 0007).
- A request no longer holds a permit during outbound I/O to a host the client chose.
- The integration harness runs a real worker. A test that triggers or finalizes must poll for the outcome rather than read it from the response.
Enforced by
the_request_path_never_holds_a_signerandthe_cli_never_builds_a_signerintests/layering.rs.check_request_timeout, the remaining startup refusal.tests/challenges.rs,tests/orders.rsandtests/revoke_cert.rs, which drive each operation through the queue.
ADR 0007: One binary runs as role processes, and only the worker holds the CA key
Status
Accepted.
Context
A certificate authority has three very different kinds of work, and they carry different risks:
- Serving ACME means parsing untrusted JWS and CSRs from the internet.
- Serving the web admin means holding operator sessions that can revoke certificates and mint EAB credentials.
- Doing background work means dialling hosts the clients chose, talking to an upstream CA, sending mail, and signing with the CA key.
In one process, a bug in the first kind of work reaches the key used by the third. The obvious answer, separate binaries, would cost a second configuration surface and a second release artefact. It would also create a protocol between the binaries, where one database and one job queue already do the job.
Decision
-
One binary, one configuration, several processes.
acme-proxy serve --role acme,admin,workerselects what a process does. With no--role, a process does all three, exactly as before the flag existed. -
The three roles:
acmeserves ACME and the root router;adminserves the panel;workerdrains the job queue and owns the schema (ADR 0003) and the first-run material.
Each role only enqueues work it does not do itself. Every role may serve its own
/metrics. -
Only the worker builds a signing backend. The signer is split in two (
crates/signer/):SignerInfo, the read side: the CA certificate, the stored CRL, a relay’s lazily discovered directory andhttp-01store, and acustomscript’s read hooks. Every role builds it, from public material only.SignerBackend, the write side: issue and revoke. Only the worker builds it, inAssembly::build_parts.
Profilehas no backend field at all. The job handlers take their backend fromGenerationParts::signers. -
The worker stores each CA’s first CRL before it serves (
store_first_crls, at startup and before publishing a reload), since the read side never signs. -
Roles are a flag, not a configuration key, so they cannot change under
SIGHUP. That is what lets a reload compare the same role set on both sides. -
ProcessRoleis notsockets::Role. The first names the three jobs a process does. The second names the three listeners it holds (acme,admin,metrics).workerholds no socket, andmetricsis a socket any role may serve, so neither set fits inside the other.
Consequences
- An
acmeoradminprocess never readsca.key, never logs in to a token, and never registers upstream. It starts even when it cannot read the key. - An
acmeoradminprocess that finds no CA certificate refuses to start, namingacme-proxy init, because the first-run material belongs to the worker. - Without a worker in the same process, queued work waits for one elsewhere. The
process logs the advisory
server_role_no_worker. A process that does not own the schema and finds it behind refuses withserver_schema_behind. - The wake-up is in-process. A worker in another process picks up a new row
within
jobs.poll_interval_ms, whereas the claim itself is race-free across processes. - Counters live per process, so each process needs its own
metrics.bind_address. Arolelabel is added when the metrics are rendered. --roleis parsed by clap, so an unknown name is refused beforeConfig::loadand beforeDatabase::open, which creates the database file.- State that several processes share cannot live in memory or in a file one of them owns (ADR 0008).
The supported topologies are in Deployment.
Enforced by
the_request_path_never_holds_a_signerandthe_cli_never_builds_a_signer(tests/layering.rs).tests/roles.rs, which starts acme and admin processes against akey_paththat does not exist and drives an order tovalidacross three processes over one file-backed database.an_unknown_role_is_refused_by_name(crates/server/src/roles.rs).info_from_config_agrees_with_the_backends_own_info(crates/signer/src/lib.rs).
ADR 0008: State that more than one process can see lives in the database
Status
Accepted. One exception stands, the web admin’s login limiter, for as long as one admin process is the supported topology.
Context
Several pieces of state used to live in process memory or in files beside the binary. Each was correct only while exactly one process existed:
- A local CA’s revocations. They lived in an in-memory ledger plus a JSON
sidecar and a CRL file, and
GET /crlserved the in-memory DER. Each process rewrote both files at startup, andorder revokeon the host wrote them beside a running server. Updates were lost,crl_numberwas duplicated, and the server served a stale CRL. A revocation could silently vanish from the CRL. - The relay’s published
http-01key authorizations, held in an in-memory token store. The relay job publishes them, and the root router serves them, which may be in another process. - Reload. Both of the above had to be handed from the outgoing
configuration generation to the incoming one through an in-memory handover
(
CarriedState). Otherwise a reload would empty the CRL ledger or the token store under a live upstream fetch. Changing the profile set,[signer],[dns]or[proxy]was therefore refused onSIGHUP. - Notification delivery. A delivery queued by one process for a profile
another process had not yet reloaded was retired as
Failedfor good.
Role processes (ADR 0007) turn every one of these from a latent bug into a routine one.
Decision
- Revocation state is two tables, keyed on the issuer: the hex SHA-256 of
the CA certificate’s SPKI, so two profiles over one CA are one issuer.
revocationsholds one row per serial. A repeat does nothing, so the first revocation’s time and reason stand.crlsholds the current CRL for each issuer. It is replaced only through a compare-and-swap oncrl_number(StoredCrl::replace_if_number), which keeps the number monotonic across processes.GET /crlserves the stored row.- The old sidecar is imported once, in the same transaction as the first
crlsrow, and never written again. signer.local_ca.crl_pathbecomes an export that nothing reads back.
http01_tokensholds the relay’s key authorizations, keyed on the upstream’s token and reaped by an hourly sweep.- A backend whose configuration did not change is reused on reload; one whose
configuration changed is rebuilt. The outgoing and incoming instances share
the tables, so there is nothing to hand over.
CarriedStateis gone. - An unknown profile or backend is a bounded
Retry, notFailed, so a delivery survives a rolling reload across processes. - Files stay files where that is a security property.
ca.key(or its PKCS#11 token) and the upstream account key remain files. The database is readable by every role and by every backup, whereas a key file can be readable only by the worker’s uid.
Consequences
order revokebeside a runningserveappears in that server’sGET /crlonce its worker has signed, with no restart and no lost update.- The profile set, each profile’s signer,
[dns]and[proxy]all reload. Of everything in the configuration, onlydatabase.urlis still refused onSIGHUP. crl_numbernever goes backwards (RFC 5280 §5.2.3). A client that meets a lower number than it has cached keeps the cached CRL, which means it keeps trusting what was revoked.admin.login_max_attemptsis still counted in memory, per process. A second admin process would get its own brute-force budget, which is why one admin process is the supported topology.
The CRL store’s concurrency rules are in crates/signer/src/local_ca/crl.rs.
Enforced by
concurrent_revocations_through_two_instances_are_all_keptandtwo_cas_over_one_database_keep_separate_crls(crates/signer/src/local_ca/mod.rs).tests/crl.rsandtests/roles.rs.tests/http01_responder.rs.an_unknown_profile_or_backend_is_retried_within_its_budget(crates/jobs/src/notify/job.rs).
ADR 0009: Dependencies are pure Rust on ring, add no global state, and earn their place
Status
Accepted. The metrics clause is superseded by ADR 0011: latency histograms were the case where a metrics library would earn its place, and they arrived. The crypto clause is narrowed by ADR 0014: it governs the crypto this project calls, not what a driver carries for a wire protocol of its own.
Context
A certificate authority’s dependency graph is part of its attack surface, and
this one is audited by cargo deny at all-features = true and inventoried in
a committed SBOM. Every crate added is one more crate to audit, one more licence
to allow, and one more release to track.
Two kinds of cost are easy to miss when choosing a dependency:
- A second crypto stack or a C toolchain.
aws-lc-rsor OpenSSL makes the build depend on a C compiler, and it means two implementations of the same primitives to keep patched. - Process-global state. A library that installs a global (a rustls
CryptoProvider::install_default, a metrics recorder, a global HTTP client) makes behaviour depend on which code ran first. In a test suite that means two tests sharing counters or a provider, with the outcome depending on test order.
Decision
-
ringis the only crypto backend this project calls.rcgen,rustls,tokio-rustls,hickory-proto’s TSIG signing andlettre’s SMTP TLS are all built on it, as is the PostgreSQL driver’s TLS (tls-rustls-ring). There is noaws-lc-rs, nonative-tlsand no OpenSSL.subtlesupplies the one constant-time comparisonringno longer offers.A driver may carry its own for a wire protocol of its own.
sqlx-postgresbrings RustCrypto (sha2,hmac,md-5,stringprep) because PostgreSQL authenticates with SCRAM-SHA-256, and no configuration of the driver avoids it. The narrowing is deliberate and bounded: that code is reachable only from the driver’s own handshake, never from anything here, and the alternative was to have no second backend at all. A dependency that wanted a second stack for work this project does — signing, hashing, comparing — is still refused. -
No global installs. The rustls provider is passed to every config builder explicitly and never installed as the process default. The metrics registry is a value held by the assembly, never a global recorder.
-
Prefer an edge to a crate. Before a new crate is added, check whether the capability is already in the graph through something else.
hyper(viaaxum),hickory-proto(via the resolver),x509-parser(viarcgen),percent-encoding(viaurl) andsubtle(viarustls) were each promoted to direct dependencies rather than replaced by a fresh one. -
Hand-roll a small, stable protocol rather than import a stack. Examples:
- RFC 6238 TOTP on
ring; - password hashing with
ring’s PBKDF2 rather than four crates for Argon2; - the relay’s ACME client on
hyper; - the Prometheus text format.
Each is a page of code against a frozen specification, with its own tests.
- RFC 6238 TOTP on
-
Choose by maintenance and safety as well as features. For example,
cryptokirather thanpkcs11, because it wraps sessions in safe types and tracks the OASIS specification. -
Each dependency’s own reason stays beside it in
Cargo.toml, as a one-line comment.
Consequences
cargo buildneeds only a Rust toolchain.- A test can build two registries or two TLS configurations without them interfering with each other.
- A hand-rolled protocol is code this project owns. Each one carries its own RFC test vectors (HOTP and TOTP) or an end-to-end test against a real counterpart (the relay client against the e2e lab’s ACME server).
- The validators deliberately trust no root store (RFC 8555 §8.3, RFC 8737 §3:
the responder’s certificate is the proof, not an identity). The relay client
is the opposite case, a client of a real CA, and uses
webpki-roots. - When a library would genuinely earn its place, as histograms would for metrics, adopting it is a new decision that supersedes this one for that subsystem.
Enforced by
cargo deny checkwithall-features = true(deny.toml), and the SBOM drift check (thesupply-chainandsbomCI jobs).- Otherwise review only.
ADR 0010: Errors derive thiserror, carry their whole message, and panic only at startup
Status
Accepted. This was re-argued more than once before it was written down, which is why it is written down.
Context
The workspace has a couple of dozen error types, and each needs a Display.
Written by hand, that is roughly 250 lines of match and write! in which each
message sits several screens from the variant it describes. That distance is
what let messages drift from their variants, more than once.
anyhow carries startup errors up to the CLI, where
CliError::failed(error.to_string()) prints them. to_string() renders only an
anyhow::Error’s outermost message. An error wrapped with .context() would
therefore print its context and silently lose the underlying cause.
The ACME error type, Problem (RFC 8555 §6.7), is the Err of nearly every
request-path function. clippy’s result_large_err lint flags a large Err in
every one of them.
Decision
- Every error type derives
thiserror::Error. No hand-writtenDisplayorstd::error::Errorimpl remains, and a new one should not add any. Three shapes need care:- A field literally named
sourceis taken as the#[source]and must implementError.ScriptError::Spawn’s is a formattedString, so it is nameddetailinstead. - A variant whose wording depends on an inner enum needs a function, not a
format string. For example,
RevokeError::Signergoes throughsigner_detail. - A variant that renders through a method of its own takes an expression:
#[error("{}", self.reason())].
- A field literally named
anyhowwithout.context(). Everyanyhowerror is built withanyhow!("… {error}"), so its message already contains its cause. That is what keepsCliError::failed(error.to_string())lossless.Problemstays small. Every constructor goes through a privateProblem::build. The optional members (identifier,subproblems, and type-specific extras such asbadSignatureAlgorithm’salgorithms) live behind anOption<Box<_>>.identifierrenders only insidesubproblems, because RFC 8555 §6.7.1 forbids it at the top level, and the serializer enforces that rather than trusting callers.- Error handling splits by phase. A startup path (database connect,
migrations, reading the configuration) may
panic!orunwrapto fail fast. A request path returnsResultand degrades:- a database error becomes a
500Problem; - a nonce that fails to save means the response goes out without
Replay-Nonce, and the failure is logged.
- a database error becomes a
Consequences
- A variant and its message sit on adjacent lines.
- Error types add one proc-macro crate to the audited graph.
syn,quoteandproc-macro2were already there, viaserde_deriveandasync-trait. - A startup error prints its whole chain on one line. Adding a
.context()anywhere would silently start truncating that line. - A handler that needs a dynamic status or header returns
Result<Response, Problem>and builds the response itself. That is howkeyChange’s409gets itsLocationheader.
Enforced by
- clippy’s
result_large_err, run with-D warningsin CI. Problem::to_value, for whereidentifiermay appear.- Otherwise review only: nothing mechanically forbids
.context()or a hand-writtenDisplay.
ADR 0011: Metrics are built on prometheus-client, one registry per scrape
Status
Accepted. Supersedes the metrics clause of ADR 0009, which listed the Prometheus text format among the protocols this project writes by hand.
Context
The exporter was written by hand: a BTreeMap of counters per family and a
write! per series. ADR 0009 accepted that for counters and a gauge, and named
latency histograms as the point where a library would earn its place.
A histogram is where hand-writing stops being small. Each series needs its
buckets, a running sum and a count, and every bucket is rendered cumulatively
with a final +Inf bucket. A mistake in any of that produces a scrape the
collector reads wrongly, or rejects outright.
Of the Rust clients, most keep a process-wide registry or recorder: the
metrics facade installs a global recorder, and the prometheus crate has a
default registry. ADR 0009 refuses global state, because tests sharing one
registry would count each other’s requests. prometheus-client has no global
state at all: a Registry is an ordinary value. It adds dtoa and a derive
macro to the graph; itoa and parking_lot were already there. Its licence is
Apache-2.0 OR MIT.
Decision
crates/jobs/src/metrics.rsis built onprometheus-client. The families are fields ofMetrics, whichserver::Assemblyholds across reloads, as before.- The registry is built per scrape.
Metrics::renderbuilds aRegistryover clones of the families (a clone shares its series) with therolelabel as a registry label. Nothing registers into shared state, and the roles stay a builder step onMetrics. - Every family is declared even when empty. The library leaves an empty family out of the exposition; a wrapper keeps it in, so a dashboard can tell “has not happened yet” from a misspelled name.
- The output is OpenMetrics, the library’s only text format. Series names
are unchanged; a counter’s
# TYPEline drops_total, and the body ends with# EOF. - Label values are escaped before they reach the library, which writes them verbatim.
- Two histograms: request latency by profile and route, and issuance latency
(finalize accepted to certificate stored) by profile. Neither carries
statusorreason, since a histogram multiplies every label by its buckets.
Consequences
- Histograms and their buckets are the library’s to get right, not this project’s.
- A new family is a field and a
registercall inrender. It appears on an empty registry at once, whichtests/grafana_dashboard.rsrelies on. - Series order within a family is no longer sorted, so tests match series individually rather than comparing the whole output.
- A scraper that only reads the old text format and ignores the content type
would misread
# EOFand the_total-less# TYPElines. Prometheus negotiates OpenMetrics and reads both.
Enforced by
- The unit tests in
crates/jobs/src/metrics.rs, which assert the exposition as text, including an empty registry’s families and the closing# EOF. tests/grafana_dashboard.rs, both directions, per family.cargo deny checkand the SBOM drift check, as for every dependency.
ADR 0012: Images are built natively per architecture, uncached, and published only past a guard
Status
Accepted.
Context
The repository shipped a Containerfile and documented a container deployment,
but published no image, so every container user compiled the crate. A
contribution (#3) added a
workflow that published one on a release tag. Its approach was right: a
tag-only trigger, GHCR with the workflow’s own token, and actions pinned by SHA.
Four of its details were not.
- It built the lab’s binary. The
Containerfilecompiled--profile e2e, which is release without fat LTO, tuned for the e2e lab’s inner loop. Nothing outside the binary tells the two profiles apart, so a published image of the wrong one would not have been noticed. - It emulated arm64 with QEMU. The release profile is fat LTO with one codegen unit. Under emulation that is the slowest build this project has, with no timeout.
- It configured a build cache that could not help.
type=ghais scoped to the ref, so a tag never reads another tag’s entries. The build’s real cache is twoRUN --mount=type=cachemounts, which no cache exporter preserves. AndCOPY . .sits directly above the one expensiveRUN, so layer reuse buys nothing. The cost was real, though:mode=maxwrites gigabytes into the repository’s shared 10 GB Actions cache, and evicts therust-cacheentries every CI job depends on. - Nothing checked the tag. CI runs on pushes to
main, not on tags. A tag that did not match the manifest’s version, or that pointed at a commit CI had never passed, would have published.
An image is also the one prebuilt artifact this project distributes, and its users run it as their certificate authority. That calls for provenance an operator can check.
Decision
- The
Containerfiletakes the cargo profile as a build argument,CARGO_PROFILE, defaulting torelease. The lab passese2eexplicitly. The default is the distribution build, so a hand-runpodman build .reproduces the published image instead of a near miss. - Each architecture builds on a native runner,
ubuntu-latestandubuntu-24.04-arm, as a matrix. Each leg pushes one single-architecture image by digest, with no tag. A final job joins the two digests into one manifest list and tags that, after checking there are exactly two. - No build cache, and the workflow says why, since an absent cache is the first thing a reader would add.
- A guard job runs before any build. It refuses a tag that differs from
[workspace.package].version, or from any crate’s=x.y.zpin. It refuses a tag off its release line,mainforX.Y.0andrelease/X.Yfor a patch (ADR 0013). It also refuses a tag whose commit has no successfulpushrun ofci.ymlon that branch. It fails rather than waits: the release procedure tags only once CI is green. - Build provenance is attested once, on the manifest list’s digest, and pushed to the registry. BuildKit’s own per-image attestations are off: with them on, each leg pushes an index instead of an image, and joining those indexes would carry attestation manifests that nothing references.
- A release tag publishes
X.Y.Z,X.Yandlatest. The floatingX.Yis the newest release of its line. It never crosses a minor, so it never picks up a breaking change, and an operator following it gets patch releases unattended. There is no floatingX: before 1.0 a minor release is where breaking changes land. Both floating tags move only on a tag push, so a manual republish of an older tag moves neither.latestalso needs the tag to be the highest release, so a patch to an older line leaves it alone. - Every push to
mainpublishesedgeandsha-<commit>, once all ofci.ymlhas passed on it:ci.ymlcalls this workflow as its last job. The build is the same release build, attested the same way; only the tags differ. - The image carries the default feature set, the same binary
cargo install acme-proxyproduces.hsmneeds a build of one’s own.
Consequences
- An uncached release build takes tens of minutes per architecture, and now
runs on every merge to
mainas well as on every release. A public repository’s runners are not billed, andtimeout-minutesis set to stop a wedged builder, not as an estimate. Merges tomainqueue rather than cancel each other, since a cancelled publish leaves orphaned manifests. edgeand thesha-tags accumulate a version per merge in GHCR. Pruning them is a registry policy, and it must keep untagged versions (below).- A hand-run
podman build .is now as slow as a release build. Contributors building the lab image by hand pass--build-arg CARGO_PROFILE=e2e, astests/e2e/common.rsdoes. - The attestation covers the manifest list. Verifying by tag finds it, since a tag resolves to the list. Verifying the digest of one architecture’s image finds nothing. If that ever matters, the fix is a second attestation per leg, not moving this one.
- The per-architecture images show in GHCR as untagged versions. The manifest list references them, so a cleanup policy must never prune untagged versions.
- A package that the workflow’s token creates under an organisation starts private. The first release needs a one-time change to inherit the repository’s visibility, and the workflow’s run summary says so.
- A failed architecture fails the release with no partial publish. The two legs
are independent (
fail-fast: false), so the healthy one still shows whether the fault is the architecture or the change.
Enforced by
guardin.github/workflows/release.yml: the version and pin check, the release-line check, the CI check, and the highest-release check that gateslatest.- The
imagejob in.github/workflows/ci.yml, whichneeds:every other job before it publishesedge. - The digest-count check in that workflow’s
publishjob. tests/e2e/common.rs, whose image build namesCARGO_PROFILE=e2e; nothing else selects the lab’s profile.
ADR 0013: main is the trunk, and a patch line is a release branch cut when a fix needs one
Status
Accepted.
Context
Every change landed on main, and every release was tagged there. That left no
way to ship a fix alone. Once main held work towards the next minor, some of
it breaking under ADR 0001, a bug in the last
release could only be fixed by releasing the next minor with it. An operator
who needed the fix had to migrate their configuration to get it.
The only image was a release’s. Nothing let an operator try the next release before it was cut, short of building it themselves.
Decision
mainis the trunk. It is the default branch. Every pull request, feature or fix, targets it, and it is releasable at every commit: CI is green before a merge.- A minor release,
X.Y.0, is a tag onmain. - A patch line is the branch
release/X.Y, cut from theX.Y.0tag when the first fix needs to ship on it, not at release time. A minor that never needs a patch never gets a branch. - A fix lands upstream first. It merges on
main, then is cherry-picked withgit cherry-pick -xonto a topic branch and merged intorelease/X.Ythrough a pull request. A fix is never made on a release branch alone, so the next minor cannot regress it. - A patch release,
X.Y.ZwithZabove 0, is a tag onrelease/X.Y. Its version bump, pins, SBOM and changelog section are made on that branch. The changelog section is then cherry-picked tomain, so the trunk’s changelog lists every release. - Images follow
ADR 0012.
A release tag publishes
X.Y.Z,X.Yandlatest. Every merge tomainpublishesedge. A push to a release branch publishes nothing, because its image is its next patch tag’s. mainand everyrelease/*branch are protected by a repository ruleset: changes arrive by pull request with the CI checks passing, and the branch can be neither force-pushed nor deleted.
Consequences
- A fix reaches operators without the next minor’s changes, and
X.Ygives them patch releases without a configuration change. - Each backport is a second pull request, and a cherry-pick can conflict once
mainhas moved away from the release. The-xtrailer records where each one came from. - Only the newest release line is maintained as a rule. An older one gets a branch only if a fix is worth the backport, and the nightly advisory scan covers the newest release branch alone.
- The book tracks
main, so between releases it can describe behaviour no release has yet. - A patch tag on
main, or a tag on a commit its branch does not contain, is refused before anything builds. The version inmain’sCargo.tomlstays at the last minor until the next one is cut, soedgereports that version.
Enforced by
guardin.github/workflows/release.yml: “The tag is on its release line”, and “CI passed on the tagged commit” against that branch.- The
pushtrigger in.github/workflows/ci.yml, which coversmainandrelease/[0-9]+.[0-9]+so that the guard has a run to read. - The repository rulesets on
mainandrelease/*, configured in the repository settings rather than in the tree. - The procedures in Contributing: review only.
ADR 0014: PostgreSQL is chosen by the URL’s scheme, over one set of queries
Status
Accepted.
Context
SQLite across processes is safe on one local disk and not across hosts. The
role split (ADR 0007) lets acme, admin and
worker run as separate processes, but only as separate processes on one
filesystem, so the deployment page promised PostgreSQL for as long as it did
not exist.
Two earlier decisions had already paid for most of it. ADR
0003 keeps the pool private to
crates/store/, so SQL and the dialect it is written in live in one crate.
ADR 0004 chose UUID v7 partly because a v4 primary
key costs PostgreSQL a page split and a full-page WAL write per row. What was
left was a real fork: two pool types, two row types, two parameter syntaxes,
and around 350 bind sites.
The obvious answers were both bad. sqlx::Any cannot carry a Uuid at all —
its type set is Null/Bool/SmallInt/Integer/BigInt/Real/Double/
Text/Blob — and it does not translate SQL. Writing every statement twice
doubles the SQL and guarantees the two copies drift, which is the one failure
mode nothing here would catch: a query used by an operator listing can be wrong
for months.
What made a single set of queries possible is that almost nothing in this
schema is dialect-specific to begin with. Every timestamp is an epoch-second
integer, so there is no strftime, julianday, datetime() or date type
anywhere. There is no CAST, no || concatenation, no LIKE, no
GROUP_CONCAT, no IFNULL, no CTE and no window function. RETURNING,
ON CONFLICT … DO NOTHING, DO UPDATE … excluded.*, partial unique indexes
and bound LIMIT/OFFSET are already spelled the way both accept.
Decision
- The scheme of
database.urlpicks the backend, atDatabase::open.sqlite:creates the file;postgres:/postgresql:expects the database to exist, because creating one is an operator’s act and not something a server does to a cluster it was pointed at. Any other scheme is refused by name. - One set of queries, in
crates/store/src/sql.rs. A statement is asql::Querycarrying its SQL and aVec<Value>until asql::Execsays which driver is on the other end.sql::Rowhides which row came back andsql::Builderreplacessqlx::QueryBuilder. Nothing outside that module names either driver. - Statements keep
?and the seam rewrites to$1…$n. Writing$nin the source would have worked on both — sqlx’s SQLite driver parses a$Nmarker and binds argumentN— but three things here build SQL by concatenation: thelive_certificate!predicate spliced into the middle of three statements, theIN (?, ?, …)lists expanded per element, and theformat!ed fragments injob::claim_nextandjob::settle. Every number would then be a hand-maintained constant. One rewrite at the edge cannot drift. - An absent value carries the type it would have had. SQLite has no typed
null; PostgreSQL sends a type OID per parameter and refuses
column "eab_kid" is of type uuid but expression is of type bigint. Every bind site knows the type statically, soValue::Null(NullKind)costs nothing and is declared nowhere twice. - Three things fork, and each asks
Dialect. The identifier search (json_each/json_extract/instragainstjsonb_array_elements/->>/strpos); the unique-violation matchers; and the_sqlx_migrationsprobe, which asked by reading the table and swallowing the error, where on PostgreSQL a failed statement aborts the surrounding transaction. Nothing else may fork without a line here. strpos, neverposition(needle in haystack). It takes its arguments the other way round, so onepush_bindsequence would bind the two dialects in different orders — a wrong answer rather than an error.- Two migration sets, both append-only.
migrations/for SQLite, frozen since 0.1.0;migrations-postgres/from its own first release. The PostgreSQL set is not a transcription: the SQLite files carry three table rebuilds that exist only because SQLite cannot add aCHECK, aUNIQUEor a foreign key to an existing table, plus a text-to-blob id conversion, and no PostgreSQL deployment has that history to replay. Every declared width is transcribed literally. - A unique constraint PostgreSQL must name is named in the migration. SQLite reports the offending columns and gives sqlx no constraint name; PostgreSQL reports the constraint and never the columns. A matcher passes both spellings, so the index name is part of the schema rather than whatever the server happened to generate.
- A database is one backend or the other, and
acme-proxy transferis the way across. Not a dual-write mode and not a sync: an offline copy of every row, refused unless the target is migrated and empty. It exists because the order row is a certificate’s only record — a deployment that moved to PostgreSQL by starting empty would leave every certificate it had issued impossible to revoke, which is the outcomelive_certificates_refusalexists to prevent. The copy is driven by a declared column manifest rather than by reading the source’s shape, because the seam decodes into a known Rust type and “read this column as whatever it is” would mean deciding at runtime whether SQLite’s untyped BLOB is abyteaor auuid. The manifest’s own hazard — a column added to the schema and forgotten here — is answered the way ADR 0003 answers it for a table rebuild: by introspecting the live schema and refusing a manifest that has drifted. - The database URL is redacted wherever it is printed. A DSN carries
user:password@; the startup log line and theSIGHUPrefusal both go throughlogfields::redact_url.
Consequences
- One binary and one container image serve both, and a deployment moves from
SQLite to PostgreSQL by changing one key and running
acme-proxy transfer. What that command cannot check is that the source is stopped, so it says so in its prompt: a copy taken while a worker is issuing is a torn snapshot that looks exactly like a good one. Txno longer derefs toSqliteConnection.tx.conn()is what&mut *txwas, andJobQueue::enqueue_intakes asql::Exec— the one SQLite-typed signature that had leaked outsidecrates/store/.- A declared
VARCHAR(n)is now enforced. SQLite ignores the width, which is hownonces.valuestayedVARCHAR(36)after the nonce became a 43-character token; on PostgreSQL that would have rejected every nonce the server mints. The width pins indb.rsare what keep the two honest. sqlx/postgresbrings RustCrypto (sha2,hmac,md-5,stringprep) for SCRAM-SHA-256. That is a second crypto stack in the graph, which ADR 0009 argues against; its clause is narrowed rather than worked around, because there is no configuration of the driver that avoids it. TLS stays onring(tls-rustls-ring), and sqlx builds itsClientConfigwithbuilder_with_providerrather thaninstall_default, so the “nothing installs process-global state” rule is untouched.- PostgreSQL gives sqlx an advisory migration lock, which SQLite does not. The one-owner rule in ADR 0003 is therefore belt and braces there rather than load-bearing — it stays, because the rule is about which process may own the schema, not only about the race.
- The coverage floor cannot see this backend: a dialect arm not taken is not an
uncovered line. That is what the CI job and its
REQUIREguard are for.
Enforced by
- The whole
acme-proxy-storesuite, on both backends. Every test there callsDatabase::connect_for_test(), which is PostgreSQL whenTEST_POSTGRES_URLnames one — the coverage that foundMAX(x, 0), SQLite’s scalar two-argument max, in a path no dialect-specific test would have singled out.connect_in_memory()means SQLite, and is the opt-out for the seven tests that are about SQLite. tests/postgres.rs, which runs the dialect-sensitive paths against both backends, andpostgres_is_available_when_it_is_required, which fails rather than skips whenACME_PROXY_REQUIRE_POSTGRESis set.transfer::tests::the_manifest_names_every_columnand…every_table, against the live schema on whichever backend is running, plusa_database_survives_a_round_trip_through_the_other_backend, which seeds all fifteen tables and compares values after a copy out and back.- The
postgresjob in.github/workflows/ci.yml, which sets that variable and also runsrolesandreload— several processes over one database, which is the deployment this exists for. sql::tests::numbering_is_contiguous_from_oneand the literal-skipping cases beside it.production_code_never_reaches_the_raw_poolandonly_the_schema_owners_apply_migrations(tests/layering.rs), unchanged.declared_token_widths_match_random_token,declared_issuer_widths_match_the_issuer_idandevery_id_column_is_declared_a_blob(crates/store/src/db.rs).logfields::tests, for the redaction, andreload::tests::a_refusal_over_a_dsn_keeps_the_host_and_drops_the_password.
Database Schema
acme-proxy stores everything in one database — accounts, orders, the audit
trail and the web admin’s own operators. There is no second datastore and no
cache. This page describes what is in it and why, for anyone reading the
database directly, writing a migration, or trying to understand what a delete
cascades to.
Two backends, one schema. SQLite is the default and PostgreSQL is what a
multi-node deployment needs; the scheme of
database.url picks between
them. Everything below describes both — the tables, the
constraints, the cascades and the reasoning are the same either way. Where the
two differ it is noted inline, and the differences are three: the column types
(BLOB/uuid, INTEGER/bigint), the AUTOINCREMENT spelling, and the
recipes for reading it by hand at the bottom of this page. The SQL the server
issues is written once; crates/store/src/sql.rs is the seam and its //!
says what had to fork.
A database is one backend or the other; there is no dual-write mode. Moving
between them is
acme-proxy transfer, which
copies every row of every table and is guarded by
crates/store/src/transfer.rs’s column manifest — a migration that adds a
column adds it there too, or the copy would silently leave it behind.
There are two migration sets, one per dialect, and both are frozen and
append-only — SQLite’s as of 0.1.0, PostgreSQL’s from its first release. A
schema change is a new sqlx migrate add file in each, never an edit to a
committed one. The PostgreSQL set is deliberately not a transcription of the
SQLite one: those files carry table rebuilds that exist only because SQLite
cannot add a CHECK, a UNIQUE or a foreign key to an existing table, and no
PostgreSQL deployment has that history to replay. The schema is also the only
surface frozen before 1.0.0 — the freeze says nothing about configuration keys,
which may still be renamed. See
Contributing for the three
consequences that catch people out.
The tables at a glance
erDiagram
accounts ||--o{ orders : "account_id"
orders ||--o{ authorizations : "order_id"
authorizations ||--o{ challenges : "authz_id"
orders ||--o| upstream_orders : "order_id"
admin_users ||--o{ admin_sessions : "user_id"
admin_users ||--o{ admin_recovery_codes : "user_id"
accounts {
blob id PK
text profile "UNIQUE(profile, pubkey)"
blob pubkey
text status "CHECK valid|deactivated|revoked"
blob eab_kid "no FK - see below"
text created_ip
text last_seen_ip
}
orders {
blob id PK
text profile
blob account_id FK
text status "CHECK pending|ready|processing|valid|invalid"
text identifiers "JSON array"
text replaces "RFC 9773 certID"
text certificate "PEM chain"
text cert_serial
integer cert_not_after "what the leaf says - see below"
integer revoked_at
}
authorizations {
blob id PK
blob order_id FK
text identifier "JSON, UNIQUE(order_id, identifier)"
text status "CHECK pending|valid|invalid|deactivated|expired|revoked"
}
challenges {
blob id PK
blob authz_id FK
text type "CHECK http-01|dns-01|tls-alpn-01, UNIQUE(authz_id, type)"
text token
text status "CHECK pending|processing|valid|invalid"
}
upstream_orders {
blob order_id PK "also the concurrency guard"
text upstream_order_url
blob csr_der
text client_ip "parked request context"
}
eab_keys {
blob kid PK
blob secret "retrievable on purpose"
text profile "NULL = every endpoint"
text status "CHECK active|revoked"
}
nonces {
text value PK
integer created_at
}
audit_log {
integer id PK "AUTOINCREMENT"
text event "no CHECK - the Rust enum"
text outcome "CHECK success|failure"
text account_id "no FK, deliberately"
text order_id "no FK, deliberately"
text identifiers "frozen into the row"
}
jobs {
blob id PK
text kind "no CHECK - see below"
text dedup_key "partial UNIQUE(kind, dedup_key)"
text payload "JSON, the subject's identity"
text status "CHECK ready|running|done|failed|cancelled"
integer run_at "the durable schedule"
integer attempts "incremented at claim"
integer deadline "give up after this"
integer lease_until "when a dead runner's row is reclaimed"
text lease_owner "which runner holds it"
}
admin_users {
blob id PK
text username UK
text password_hash "one-way"
blob totp_secret
text status "CHECK active|disabled"
text role "no CHECK - NULL reads as admin"
}
admin_sessions {
text token_hash PK "SHA-256 of the token"
blob user_id FK
text state "CHECK pending_mfa|active"
integer mfa_attempts
}
admin_recovery_codes {
blob id PK
blob user_id FK
text code_hash
integer used_at "stamped, not deleted"
}
revocations {
text issuer PK "SHA-256 of the CA's SPKI"
text serial PK
integer revoked_at "the first one stands"
integer not_after "NULL is never pruned"
}
crls {
text issuer PK
integer crl_number "moves only by compare-and-swap"
blob der "what GET /crl serves"
integer next_update
}
http01_tokens {
text token PK "the upstream's own token"
text key_authorization
integer expires_at "a backstop, swept hourly"
}
The diagram has three clusters, and the two things worth noticing are the edges that are not drawn:
- The ACME graph —
accounts → orders → authorizations → challenges, withupstream_ordershanging off an order andeab_keysandnoncesstanding alone. audit_log,jobsand a local CA’srevocations,crls, plus the relay’shttp01_tokens, all connected to nothing. That is policy in each case, not an omission. The audit trail and the revocation ledger must outlive what they describe, the job queue is generic (both below), and the last three are state every role process shares (ADR 0008).- The admin island —
admin_usersand its two children — which never joins toaccounts. Anadmin_usersrow is an operator of this server; anaccountsrow is a client key that asks it for certificates. They are different populations and the schema says so.
Profiles are a database boundary
accounts.profile and orders.profile are NOT NULL, and accounts is keyed
UNIQUE(profile, pubkey). One client key presenting itself at two endpoints is
two independent accounts with separate orders and separate authorizations —
see Profiles & Routing.
eab_keys.profile is the one nullable member of the set, and NULL means
“valid at every endpoint” rather than “unknown”.
Request-path lookups always take the profile. The admin CLI deliberately uses
unscoped lookups (find_any_by_id, find_any_by_kid), because an operator
holding an id wants the row, not a reminder about which endpoint it belongs to.
Every foreign key is indexed and cascades
SQLite indexes primary keys and UNIQUE constraints and nothing else, so before
20260727120000_indexes_and_constraints.sql every order read and every
challenge trigger was a full table scan. That migration rebuilt the four ACME
tables to add both halves at once:
ON DELETE CASCADEon every foreign key, so an account or an order can genuinely be deleted. On SQLite this depends on theforeign_keyspragma, whichcrates/store/src/db.rspins on for every connection; PostgreSQL always enforces them.- An index on every foreign key:
idx_orders_account_id,idx_authorizations_order,idx_challenges_authz.
Two later indexes serve one query each: idx_orders_cert_serial on (profile, cert_serial), which POST /revokeCert uses on every request, and
idx_orders_created_at / idx_orders_status_created_at, which the web admin’s
newest-first cross-account listing needs and the ACME path never did.
idx_orders_replaces_claim is different — it is a partial unique index on
(profile, replaces) where replaces IS NOT NULL AND status != 'invalid'. It
is not a lookup index at all; it is RFC 9773 §5’s “already replaced?” rule
enforced in SQL, which is what makes 409 alreadyReplaced race-free and what
lets an order that fails release its claim. See
Renewal Information.
CHECK constraints hold the state machines
Every status column carries a CHECK (status IN (…)) — accounts, orders,
authorizations, challenges, jobs, upstream_orders, eab_keys and
admin_users — as do challenges.type, admin_sessions.state, and
audit_log’s outcome and actor_kind.
They are there because a typo in a status would otherwise park a row in a state nothing can read back, and the row would look fine. With the constraint it is a failed write at the moment of the mistake.
Open vocabularies carry none: audit_log.event (its CHECK was dropped by
20260909120000), admin_users.role and jobs.kind are validated by a Rust
enum instead, because each grows with features and a CHECK would cost a table
rebuild per new word. See ADR
0005.
This is also why a new CHECK is expensive: SQLite cannot add one to an
existing table, so it needs a full table rebuild in a new migration. Several
constraints were therefore declared before anything wrote them —
admin_users.totp_secret/totp_pending_secret/totp_last_step and
admin_sessions.state’s 'pending_mfa' value are the worked example, added in
20260808120000 and only used once the second factor shipped.
The audit trail has no foreign keys, deliberately
audit_log names an account_id and an order_id with no constraint behind
either. An audit row has to survive the account or order it describes being
deleted — a CASCADE there would destroy the evidence along with its subject,
which is the one thing an audit trail may not do. The identifiers are frozen
into the row for the same reason, rather than being read back through a join
that may no longer resolve.
Two more consequences of that decision:
idisINTEGER PRIMARY KEY AUTOINCREMENT, not a plain rowid. An operator types this id, andAUTOINCREMENTis what stops SQLite handing out the rowid of a purged row a second time.outcomeis denormalized fromeventand written from the single definition inAuditEvent::outcome, so “show me everything that was refused” is an index lookup rather thanevent LIKE '%_failed'written out in three front ends.
Rows are only ever INSERTed. There is no setter and no UPDATE against this
table anywhere in the crate; the only statement that removes anything is the
retention sweep. See Audit Trail.
revocations follows the same rule for the same reason. A local CA’s
revocation must outlive an order an operator deletes, or the serial would drop
off the CRL, so it has no foreign key to orders either.
accounts.eab_kid is a similar deliberate non-key: it records which credential
was used at registration, but an EAB credential is revocable and the account
outlives it, so there is no constraint tying the two together.
The job queue is generic, and its schema says so
jobs is the second table with no foreign key, for a different reason than
audit_log’s. A queue is generic: payload names whatever kind means — a
local order today, a certificate serial or nothing at all tomorrow — so a typed
foreign key would either be wrong for every other kind or force one nullable
column per kind. A job whose subject was deleted is retired by its handler
(“the order no longer exists”), which is a terminal outcome recorded in
last_error, not an orphan nothing sweeps.
Three more shapes worth knowing before touching it:
kindcarries noCHECK, unlike every other enum-ish column here. A kind is registered in code by whichever subsystem owns it, and SQLite cannot alter aCHECKwithout a table rebuild, so every future kind would cost one. The runner claims only the kinds its registry holds, so an unrecognised one is left alone rather than mis-run — which is also what lets an older binary meet a row a newer one wrote.- The identity index is partial:
UNIQUE(kind, dedup_key) WHERE status IN ('ready', 'running'). Only a live job holds an identity. A plainUNIQUEwould let one finished job block its own key for ever, which is fatal for a periodic kind whose key is a constant and wrong for an order retried after a failure. attemptsincrements when the row is claimed, not when it completes, so a job that reliably kills the process still exhausts its budget instead of crash-looping. The same reasoning is why the reclaim sweep leaves the counter alone.
status = 'cancelled' is written by jobs cancel and its panel and API
twins. It was declared before anything wrote it — the
admin_sessions.state = 'pending_mfa' treatment, where a CHECK was written
before anything filled it precisely so no rebuild would be needed later.
Two expiry columns on orders, and neither is the order’s
orders.not_after is the validity the client asked for in newOrder (RFC
8555 §7.4) — usually NULL, and clamped by the signer when set.
orders.cert_not_after is what the issued leaf actually says, stamped by
Order::finalize from the same DER that cert_serial and cert_pubkey come
from. The order object’s own expires is a third thing again, and is not a
column here.
cert_not_after has three meaningful states:
- an epoch second;
NULL— issued before the column existed; the expiry sweep backfills it;- a negative sentinel — the sweep looked and the chain would not parse.
Writing
NULLback would have it re-parsed on every pass for ever.
It is optional where cert_serial is not: a chain whose serial cannot be read
cannot be revoked, so it is a failed issuance, while an unreadable validity is
only housekeeping. Its index is partial on certificate IS NOT NULL AND revoked_at IS NULL, which is the expiry digest’s own predicate.
An order holding a live certificate is never deleted
The order row is a certificate’s only record: revokeCert and order revoke
find it by serial, the expiry digest lists it, and renewal information (RFC
9773) is derived from it. Deleting it would make a certificate that is still
trusted impossible to revoke.
So account delete, order delete and eab delete --delete-accounts are
refused, on every surface, while any order they would remove holds a live
certificate — issued, not revoked, and not yet expired (cert_not_after
NULL, negative or in the future). The check runs before the confirmation
prompt and again inside the DELETE itself (live_certificate! in
crates/store/src/order.rs), so a certificate issued in between is not lost.
There is no override flag. The daily order_sweep follows the same rule: a
valid order is never swept, whatever its age.
Secrets are stored three different ways, on purpose
The three storage shapes in this schema are not an inconsistency — each one follows from what the server has to do with the value later.
| Column | Shape | Why |
|---|---|---|
eab_keys.secret | Raw bytes, retrievable | HMAC verification needs the same secret back on every request. A lost one is replaced, never recovered — eab create prints it once. |
admin_users.password_hash | One-way KDF (PBKDF2-HMAC-SHA256), unreadable | A password is only ever compared. No code path can read it out. |
admin_sessions.token_hash | hex(SHA-256(token)), no KDF | A 256-bit CSPRNG token has no dictionary to slow down. The hash exists solely so a database read yields nothing replayable. |
admin_recovery_codes.code_hash follows the password shape — a recovery code is
only ever compared. It is a table rather than a JSON column so that consuming
one is UPDATE … WHERE id = ? AND used_at IS NULL with rows_affected deciding
a race, the same primitive nonces uses. used_at is stamped rather than
deleted, so “7 of 10 remaining” is a count and a spent code leaves a trail.
admin_users.totp_secret and totp_pending_secret are plaintext BLOBs on
purpose: verification recomputes the HMAC, so the server needs the same bytes
back every attempt. That is eab_keys.secret’s situation, not a password’s, and
any wrapping key would live in the same directory as the database.
Columns nothing ever compares against
accounts.created_ip/created_ptr/last_seen_ip/last_seen_ptr,
orders.created_ip/created_ptr, admin_sessions.created_ip/user_agent, and
audit_log’s client_ip/client_ptr/user_agent are forensics only. No
code path compares a live request against any of them.
That is a decision, not an oversight. Pinning an identity to an address breaks CGNAT and mobile clients; pinning it to a User-Agent breaks on the next browser update. They answer “who asked for this, and from where” after the fact, and nothing else.
One column is compared, and only to decide whether to send a message:
admin_users.known_login_ips, the operator’s last five distinct sign-in
addresses, decides whether a sign-in is reported as coming from a new address.
It never allows or denies anything.
None of them reaches an ACME object either — the wire format is RFC 8555’s and stays that way. They surface through the admin CLI and the web admin only.
Ids are UUID v7, stored as bytes
Every row this server creates is keyed by a UUID version 7 (RFC 9562 §5.7),
minted in one place, acme_proxy_store::id::mint. Ids created close together
share a prefix, and they sort by creation; why that matters, and the rule that
an id’s Rust type says where it came from, are in
ADR 0004.
On SQLite the column holds the sixteen bytes, not the thirty-six characters
of the rendering; on PostgreSQL it is a native uuid, and sqlx maps the same
Rust type to both. Nothing on the wire changes either way — an id is still
rendered by Uuid::to_string, so account URLs, kids, order URLs and every
admin API member are the same lower-case hyphenated form they always were.
What changes is an ad-hoc query, and only on SQLite, where an id column prints
as a blob and wants hex():
# Readable ids.
sqlite3 sqlite.db "SELECT lower(hex(id)), profile, status FROM accounts;"
# Looking one up by the id from a URL or a log line.
sqlite3 sqlite.db "SELECT status FROM orders
WHERE id = unhex(replace('6ba7b810-9dad-41d1-80b4-00c04fd430c8','-',''));"
PostgreSQL needs neither, since it prints and parses the hyphenated form:
psql -c "SELECT id, profile, status FROM accounts;"
psql -c "SELECT status FROM orders WHERE id = '6ba7b810-9dad-41d1-80b4-00c04fd430c8';"
A few columns look like ids and are not, so they stay text: orders.replaces is
an RFC 9773 certID, audit_log.actor_id may be an account id or an admin
username, audit_log.account_id and order_id name a row that may already be
gone, and request_id is whatever the caller sent.
Rows created before this changed were converted in place and keep their v4 ids, so a table holds both versions and only the v7s sort by creation.
Reading it directly
On SQLite, the file is sqlite.db by default and is opened in WAL mode,
so there are normally sqlite.db-wal and sqlite.db-shm beside it. Copying
only sqlite.db gives you a database missing every recent write; back up all
three, or use sqlite3 sqlite.db ".backup backup.db", which is consistent by
construction. On PostgreSQL, back it up the way you back up any other
database — pg_dump is consistent by construction and there are no sidecar
files to miss.
Ids are stored as bytes on SQLite, so they need hex() on the way out and
unhex() on the way in; on PostgreSQL they are a native uuid and need
neither. See
Ids are UUID v7, stored as bytes.
The recipes below are SQLite’s. On PostgreSQL the ids need no wrapping and the
epoch-second columns read with to_timestamp(created_at) in place of
datetime(created_at,'unixepoch').
# What has this account been issued?
sqlite3 sqlite.db "SELECT lower(hex(id)), status, identifiers,
datetime(created_at,'unixepoch')
FROM orders
WHERE account_id = unhex(replace('…','-',''))
ORDER BY created_at DESC;"
# Everything refused in the last day.
sqlite3 sqlite.db "SELECT datetime(created_at,'unixepoch'), event, profile, client_ip, reason
FROM audit_log WHERE outcome = 'failure'
AND created_at > strftime('%s','now','-1 day');"
# Which migrations have run.
sqlite3 sqlite.db "SELECT version, description, success FROM _sqlx_migrations;"
Read-only inspection of a running server is safe under WAL, and on PostgreSQL
by its own MVCC. Writing to the database behind the server’s back is not — the
CHECK constraints will catch a bad status, but nothing will re-sync the
in-memory state a handler is holding.
Use the Admin CLI instead.
Custom Plugins Examples
This section provides complete examples of custom plugin scripts that can be
integrated into acme-proxy. These scripts must be marked as executable (chmod +x).
Custom signer script
A custom signer that passes the CSR to a fictional internal API to obtain a certificate.
#!/bin/bash
# /etc/acme-proxy/signer/internal-pki.sh
set -e
# We only handle the "issue" hook in this example
if [ "$ACME_SIGNER_HOOK" = "issue" ]; then
# Read the JSON payload from stdin
PAYLOAD=$(cat)
# Extract the base64 encoded CSR
CSR_B64=$(echo "$PAYLOAD" | jq -r '.csr_der_base64')
# Call internal PKI API
# The API is expected to return a JSON with a 'certificate_pem' field.
RESPONSE=$(curl -s -X POST https://pki.internal.company.com/api/sign \
-H "Content-Type: application/json" \
-d "{\"csr\": \"$CSR_B64\", \"order_id\": \"$ACME_SIGNER_ORDER_ID\"}")
# Extract PEM from response
PEM=$(echo "$RESPONSE" | jq -r '.certificate_pem')
if [ -n "$PEM" ] && [ "$PEM" != "null" ]; then
# Output the PEM chain to stdout (leaf first, then issuers)
echo "$PEM"
exit 0
else
# Exit 1 (any non-zero other than 3) = internal failure -> the client
# gets a 500 and the order is marked invalid.
#
# Exit 3 is RESERVED for "this CSR is bad" -> the client gets a 400
# badCSR and the order stays "ready" so it can retry with a corrected
# CSR. Only use 3 when the PKI rejected the CSR itself, never for an
# API outage like this one.
echo "internal PKI API did not return a certificate" >&2
exit 1
fi
fi
# Hooks this example does not implement. `revoke` must succeed or the proxy
# leaves the order un-revoked, so returning non-zero here would be wrong for a
# real deployment;
# implement it, or set supports_crl/supports_renewal_info = false (the default)
# so those hooks are never invoked at all.
exit 1
Custom filter script
A custom filter that checks the client IP against a threat intelligence feed before allowing the connection.
#!/bin/bash
# /etc/acme-proxy/filters/threat-intel.sh
if [ "$ACME_FILTER_HOOK" = "connection" ]; then
# Skip checking local IPs
if [[ "$ACME_FILTER_CLIENT_IP" == 10.* ]] || [[ "$ACME_FILTER_CLIENT_IP" == 192.168.* ]]; then
exit 0
fi
# Query threat intel API
STATUS=$(curl -s -o /dev/null -w "%{http_code}" "https://threat.internal/api/check?ip=$ACME_FILTER_CLIENT_IP")
if [ "$STATUS" = "200" ]; then
# IP is clean
exit 0
else
# IP is flagged, deny the request
echo "Client IP $ACME_FILTER_CLIENT_IP is flagged in Threat Intel"
exit 1
fi
fi
exit 0
Custom IPAM script
A custom IPAM backend that reads the permitted names for an address out of a
CSV the estate already maintains, one address,name,name,... row per machine.
#!/bin/bash
# /etc/acme-proxy/ipam/lookup.sh
set -u
INVENTORY="/etc/acme-proxy/ipam/inventory.csv"
# The address arrives twice — in the environment and in the JSON on stdin.
# This script uses the environment, so it never has to read stdin at all; a
# script that exits without reading it is fine and is not an error.
ROW=$(grep -m1 "^${ACME_IPAM_CLIENT_IP}," "$INVENTORY")
if [ -z "$ROW" ]; then
# 3 is RESERVED: "this inventory holds no record of that address". The
# `ipam` check words its own refusal for it, distinct from the one below.
exit 3
fi
# Exit 0 with the permitted names, one per line. They are lowercased and
# stripped of a trailing dot for you, so print whatever form the file holds.
# Printing nothing here would mean "recorded, and entitled to nothing" — a
# different answer from exit 3, and also a refusal.
echo "$ROW" | cut -d, -f2- | tr ',' '\n'
exit 0
Every other non-zero exit — a missing inventory file, a
grepthat could not run, a timeout — is reported as a retryable500, never as a denial. That is deliberate: an inventory this server cannot reach has decided nothing, so issuance stops rather than failing open. Do not use a non-zero exit to refuse a client; refuse by not printing the name.
acme-proxy filter explainreally runs the policy, so it executes this script too.
Custom notification script
A custom notification script that sends a Slack message when a certificate is revoked.
#!/bin/bash
# /etc/acme-proxy/notify/slack.sh
WEBHOOK_URL="https://hooks.slack.com/services/AAAAAAAAA/BBBBBBBBB/XXXXXXXXXXXXXXXXXXXXXXXX"
# The JSON payload is always on stdin; read it before anything else. For
# certificate_revoked it carries the reason code, which has no env var.
PAYLOAD=$(cat)
if [ "$ACME_NOTIFY_HOOK" = "certificate_revoked" ]; then
# ACME_NOTIFY_IDENTIFIERS is only populated for certificate_issued, so it
# would render empty here — take the order id from the environment and the
# reason from the payload instead.
REASON=$(echo "$PAYLOAD" | jq -r '.reason // "unspecified"')
MESSAGE="🚨 Certificate Revoked! Serial: \`$ACME_NOTIFY_CERT_SERIAL\` | Order: \`$ACME_NOTIFY_ORDER_ID\` | Reason: \`$REASON\`"
curl -s -X POST -H 'Content-type: application/json' \
--data "{\"text\": \"$MESSAGE\"}" \
"$WEBHOOK_URL"
fi
exit 0
A notification script’s exit code is only logged — it can never fail the ACME request that triggered it.
Testing & Coverage
acme-proxy relies on a multi-layered testing strategy combining lightning-fast
unit/integration tests with real-world End-to-End (E2E) scenarios.
Prerequisites
- cargo-nextest: The project requires
cargo nextestto execute the integration suite.nextestruns each test in its own isolated process. This is load-bearing because tests involving thecustomscripts exec generated bash files. Under standardcargo test(which runs in threads), file descriptor sharing causes intermittentETXTBSYfailures. - llvm-cov: For coverage reporting.
- Podman / Docker: Required for running the E2E suite.
Install the required Rust tools:
cargo install cargo-nextest cargo-llvm-cov
rustup component add llvm-tools-preview
Running the unit & integration suite
To run the complete in-memory test suite:
cargo nextest run --workspace
These tests use an in-memory SQLite database and an in-memory local CA, and
nothing reaches a real network. A few suites write to a temporary directory or
bind a loopback socket, each because the thing under test needs one: roles
and reload (real processes, ports and a config.toml), filters and
custom_signer (scripts, and the IPAM mocks), and revoke_cert (a CA on disk
that two processes share).
PostgreSQL. Set TEST_POSTGRES_URL to a server’s URL and every
crates/store/ test that calls Database::connect_for_test() runs against it
instead of SQLite; tests/postgres.rs runs the dialect-sensitive paths against
both, and skips without it. CI’s postgres job sets
ACME_PROXY_REQUIRE_POSTGRES=1 as well, which turns that skip into a failure.
A test that calls Config::load() holds ENV_LOCK
(acme_proxy_core::config::ENV_LOCK, or testutil::EnvGuard, which holds it
for you). ACME_PROXY_* and ACME_PROXY_CONFIG are process state: a test
setting one while another loads makes the second read the first’s variables.
There is one lock for every crate on purpose, since a per-module lock would
serialise a module against itself and nothing else.
A test that triggers a challenge or finalizes an order must poll for the
result. Validation and issuance run in the job queue, and every test app runs
a real worker, so the response only says processing. await_order and
await_challenge in tests/common/ are the helpers.
The hsm feature (PKCS#11)
crates/signer/src/local_ca/pkcs11.rs is behind the hsm feature, so the
command above neither compiles nor lints it — --all-targets does not enable
features. Run it explicitly:
cargo nextest run --workspace --features acme-proxy-signer/hsm
cargo clippy --workspace --all-targets --features acme-proxy-signer/hsm -- -D warnings
The PKCS#11 tests create a SoftHSM2 token in a temporary directory, generate
a P-256 key inside it, self-sign a CA certificate through the token, and then
drive the real LocalCa end to end — issuing a leaf that must verify against
that CA, and a CRL that must too. The key is generated through cryptoki
itself, so softhsm2 is the only prerequisite; opensc/pkcs11-tool is not
needed.
# Debian/Ubuntu
sudo apt install softhsm2
# Arch
sudo pacman -S softhsm
When no SoftHSM2 module is found the PKCS#11 tests skip with a message
rather than failing, so --features hsm stays green without it. CI has a
dedicated hsm job — separate from test so the coverage floor, which a
feature-gated file sits outside of entirely, does not fight the feature.
cargo nextestmatters more than usual here:SOFTHSM2_CONFis process-global and read atC_Initialize, and the PKCS#11 context is cached per module for the life of the process. Process-per-test isolation is what keeps those from leaking between tests.
Code coverage
CI enforces a hard floor of 97% of lines, over every package in the
workspace (main.rs is excluded — it is pure socket and exit wiring). The
shortest way to see the same number locally:
cargo llvm-cov nextest --workspace --summary-only
CI splits that in two, because it wants several views of one test run: the run
itself with --no-report, then lcov.info, an HTML tree and the summary that
gates, each generated from the profiles left on disk.
--workspacehas to reach the report, and thereportsubcommand cannot take it.cargo llvm-cov reportrejects the flag, and with no package selection it measures the package cargo picks — at a root that is also a package, the root package alone. Reporting from saved profiles at workspace scope iscargo llvm-cov --no-run --workspace, which is what CI uses:cargo llvm-cov --no-run --workspace --summary-only \ --ignore-filename-regex 'src/main\.rs' --fail-under-lines 97Not
-ponce per member either: a crate built twice under different features contributes two coverage maps that way, and its lines are counted twice.
Gotcha: a handler annotated with
#[instrument]reports far lower coverage than it actually has. The attribute moves the body into a generatedasyncblock, so the signature lines show zero hits and the body lines carry no region at all —handlers/authz.rssits around 40% whiletests/challenges.rsdrives nearly every branch in it. Checkcargo llvm-cov report --textfor the file before writing tests against the percentage. (Installing atracingsubscriber in tests does not fix this; measured, it moves the total by 0.03 points.)
Which is why
crates/admin/src/webadmin/carries no#[instrument]at all. It is a rule for that module, not a preference: the access middleware already opens the request span, so the attribute would buy nothing and cost the module’s reported coverage.
The password KDF is slow on purpose
admin::password runs PBKDF2-HMAC-SHA256 at 600 000 iterations — roughly 85 ms
in a release build. Unoptimised, ring takes ~1.1 s for the same hash, and
the admin suites pay it at least twice per test (the harness creates an
operator and signs in). With twenty of them in parallel that was most of the
suite’s CPU time and a 40-second critical path. So the workspace Cargo.toml
builds ring at opt-level = 3 in the dev and test profiles
([profile.dev.package.ring]), which brings a debug build to ~90 ms per hash.
Only ring is raised: the loop is compiled entirely inside it, and optimising
the workspace crates instead measurably changes nothing.
The override lives in Cargo.toml because nothing else reaches the build
nextest runs: cargo --config … nextest run and
CARGO_PROFILE_DEV_PACKAGE_RING_OPT_LEVEL are both silently ignored there.
The admin::password unit tests still mostly go through a private
hash_with_iterations at a cheap setting — the same code path, the same salt
generation and encoding, without 600 000 rounds dozens of times over. Two
deliberately pay the real cost: the encoding must reflect the real constants,
and the dummy hash must cost what a real row costs, or an unknown username would
answer faster and enumerate the operator table.
If you add a test that signs in, expect it to cost one real hash.
Testing the web admin
tests/admin_api.rs drives the real build_admin_app through
tower::ServiceExt::oneshot, the same way tests/orders.rs drives the ACME
side. The harness helpers live in tests/common/mod.rs:
| Helper | |
|---|---|
admin_config() | a Config with [admin] enabled |
test_admin_app(config) | the admin router + its database |
test_admin_app_with_signer(config) | also returns the signer, for tests that must issue before revoking |
test_admin_app_logged_in(config) | creates one operator, signs in, returns an AdminSessionHandle |
admin_request(app, method, path, session, body) | one request, optionally authenticated |
admin_login, session_cookie_token, json_body |
test_admin_app and test_app_full share one_profile, so the two cannot
drift into mounting subtly different endpoints.
The CSRF table is the regression suite. mutating_endpoints() in
tests/admin_api.rs lists every unsafe method and path, and two tests assert
each of them refuses a missing, wrong, and foreign token. AuthenticatedWrite
already makes the check structural — a mutating handler cannot reach a session
without it — but the residual risk is a new handler taking Authenticated by
mistake, and that table is what catches it. A new endpoint under /api that
is not in that list is a review catch.
E2E testing (real clients)
The E2E suite spins up complete environments using testcontainers-rs to run
real ACME clients (certbot, acme.sh, lego) against the proxy.
The E2E suite is #[ignore]d by default to keep the main test cycle fast. You
must have Podman or Docker running.
Run the E2E suite with:
cargo nextest run -E 'binary(e2e)' --run-ignored all
# or, with plain cargo:
cargo test --test e2e -- --ignored
Do not run
cargo nextest run e2e. nextest’s bare positional filter matches against test names, not binary ids, and none of this suite’s test names contain the substring “e2e” — so that command silently matches nothing and reports0 tests runrather than failing. The-E 'binary(e2e)'expression is what selects the binary.
Rootless Podman is auto-detected: the harness points DOCKER_HOST at the user’s
podman socket if unset, and fails with a clear message naming systemctl --user start podman.socket rather than starting it itself.
The tests/e2e/common.rs harness automatically builds the necessary container
images from the Containerfiles in the repository, provisions a dedicated
podman network, and asserts on the container logs. It tests complex scenarios
like Key Rollover (via lego), NetBox filter mocks, and full TLS-ALPN-01
responses.
Contributing to acme-proxy
Thank you for your interest in contributing to acme-proxy! Whether you’re
fixing a bug, adding a new feature, or improving documentation, your help is
welcome.
Development environment
To start developing, ensure you have the following installed:
- Rust (latest stable version)
sqlite3(for database inspection, thoughsqlxhandles migrations)- mdBook (if you want to build this documentation locally)
Initial setup
Clone the repository and build the project:
git clone https://github.com/acme-proxy/acme-proxy.git
cd acme-proxy
cargo build
Testing
The suite is what holds RFC 8555 compliance in place, and CI enforces a coverage floor, so a change that adds a branch generally has to add a test for it.
Before submitting a pull request, run the full suite with nextest:
cargo nextest run --workspace
--workspace is not optional. The repository root is both the acme-proxy
package and the root of a workspace of library crates under crates/, and a
bare cargo command at such a root acts on the root package alone — the unit
tests of every library crate would simply not run.
Use
cargo nextest run --workspace, notcargo test. This is a requirement, not a preference: several tests execute a script file they have just written, and undercargo test— which runs tests as threads of a single process — another thread’sCommand::spawncan fork while the file’s write descriptor is still open, failing withETXTBSYroughly one run in three. nextest’s process-per-test isolation removes the race entirely. See Testing & Coverage.
What CI will check
Your pull request has to pass all of these:
cargo fmt --all --check
cargo clippy --workspace --all-targets -- -D warnings
cargo llvm-cov nextest --workspace --summary-only --fail-under-lines 97
cargo test --workspace --doc # llvm-cov skips doc-tests
cargo deny check # supply-chain audit, against deny.toml
RUSTDOCFLAGS="-D warnings -A rustdoc::private_intra_doc_links" \
cargo doc --workspace --no-deps --all-features # every intra-doc link
mdbook build doc/ && python3 doc/lint.py # this book
cargo test --doc compiles the doc examples but not the intra-doc links, of
which the workspace has a great many; cargo doc -D warnings is what catches a
link a rename broke. Private intra-doc links are allowed on purpose: the library
exists for the binary and the tests, and a public item explaining itself by
naming the private thing it delegates to is the good outcome.
doc/lint.py holds the book to its own conventions: 80-column prose, no
numbered headings, every fence tagged, every relative link and anchor
resolving, every ADR listed, and no configuration key documented in two
files — two copies of a default drift silently.
Four more jobs check what the ones above cannot:
msrvreadsrust-versionout ofCargo.tomland runscargo check --locked --workspace --all-targets --all-featureson exactly that toolchain, so the minimum stated there is one CI has verified.hsmruns clippy and the suite with--features acme-proxy-signer/hsmagainst SoftHSM2.--all-targetsenables no features, so without this job the PKCS#11 code would be neither linted nor tested; it is not folded into the coverage job, whose floor a feature-gated file sits outside of.postgresruns the wholeacme-proxy-storesuite,tests/postgres.rs,rolesandreloadagainst a real PostgreSQL server, withACME_PROXY_REQUIRE_POSTGRES=1so a skipped test is a failure. It is separate from the coverage job forhsm’s reason.e2eruns nightly, not on a push: a subset of the container lab intests/e2e/, with real certbot, acme.sh and lego clients.
The sbom job additionally regenerates sbom.cdx.json and fails if it differs
from the commit — see Changing dependencies.
Note the coverage floor is enforced, so new code generally needs new tests.
cargo test --doc is the only thing that compiles the startup example in
src/lib.rs.
Writing tests
- Unit Tests: Keep them close to the code (in the same file, in a
mod tests). - Integration Tests: Located in the
tests/directory. These tests spin up a full in-memory axum router and SQLite database to test the entire ACME flow.
See the Testing & Coverage page for more details.
Code style
- Format your code using
cargo fmt. - Ensure all lints pass by running
cargo clippy --workspace --all-targets -- -D warnings. - Document public APIs using rustdoc comments (
///). - Comments, doc comments and error-message strings are written in English, as are identifiers and log messages.
- Every
tracingcall carriesevent = "<subsystem>_<object>_<outcome>"as its first field, as a string literal rather than a computed value, so the name stays greppable. Several are asserted by the end-to-end suite — grep before renaming one. - The crate is edition 2024; see
rust-versioninCargo.tomlfor the minimum toolchain.
Changing the database schema
Both migration directories are append-only: crates/store/migrations/
(SQLite, since 0.1.0) and crates/store/migrations-postgres/ (PostgreSQL). Add
a migration to each; never edit a committed one:
sqlx migrate add --source crates/store/migrations add_widget_table
sqlx migrate add --source crates/store/migrations-postgres add_widget_table
sqlx tracks each migration by a checksum, so editing a file that has already
run turns every existing deployment into a startup failure. One build-system
trap while you work: sqlx::migrate!() embeds the set at compile time and
adding or removing a file under either directory does not on its own
invalidate the build, so a test can be run against the previous set — touch crates/store/src/db.rs after changing the directory. This reverses the rule
that held before the first release, when the server had never been deployed and
a schema change meant editing the migration and running rm -f sqlite.db*.
Three consequences:
- A new column is a new file, even when it plainly belongs to an existing
table.
ALTER TABLE ADD COLUMNis cheap; putting it in the originalCREATE TABLEis what breaks. Name it incrates/store/src/transfer.rs’s manifest too, oracme-proxy transferdrops it;the_manifest_names_every_columnrefuses a manifest that has drifted. - A new
CHECK,UNIQUEor foreign key needs a table rebuild in the SQLite set, because SQLite cannot add one to an existing table; PostgreSQL’sALTER TABLE … ADD CONSTRAINTneeds none. Write the rebuild in the new migration, and remember the two things a rebuild loses silently: anINSERT … SELECTdrops any column you forget to name, andDROP TABLEtakes the table’s indexes with it — including ones declared in an earlier migration, which will not run again to put them back. - A wrong declared width is a rebuild too. SQLite gives
VARCHAR(n)TEXT affinity and enforces no length, so a width that no longer matches its data costs nothing at runtime and is wrong everywhere else — in what.schematells an operator, and in any port to a dialect that does check.20260826120000_declared_widths_for_random_tokens.sqlis the worked example. Where the width follows a constant insrc/, pin the two together with a test; that file’sVARCHAR(43)isTOKEN_BYTESand nothing else, so a change to the constant has to reach the schema.
Adding a configuration key
A key is a field on one of the section structs under
crates/core/src/config/types/, with a #[serde(default)] that makes the whole
section optional. Beyond the field itself, a new key owes:
- Documentation in exactly one book page, as a
### Referenceentry naming its environment variable, plus an entry inconfig.toml.example(which a test deserializes, so it cannot rot into invalid TOML).doc/lint.pyrefuses a key documented in two pages. - A decision about scope. A section listed in
PROFILE_SECTIONSis per-profile and inherited key by key (see Profiles); one describing the process —[jobs],[audit],[metrics],[proxy],[admin]— is not. - A decision about reload. A reload rebuilds everything from the new
configuration, so a key reloads unless something snapshots it at startup.
Only
database.urlis refused onSIGHUP(FROZENincrates/server/src/reload.rs); a new key joins it only with a reason.
A list-valued key has one more obligation, and one thing to know:
#[serde(deserialize_with = "string_list")]on the field. An environment variable can only carry a string, and this is what splitsa,binto a list, at any depth: inside a profile or inside a named table ([filter.check.<name>],[notify.webhook.<name>], …) alike. It also reads a variable set to the empty string (a${VAR:-}shell default) as[]. Without it the key loads from a file and fails from the environment;every_list_field_reads_a_comma_separated_stringrefuses the omission.- A value containing a literal comma, such as a regex with
{2,3}, can only be set from the file, since the comma is the separator.
The environment source pins prefix_separator("_"). Without it, config
reuses the nested separator __ after the prefix and silently ignores every
ACME_PROXY_* variable.
Changing a configuration key
The schema is the only frozen surface. Before 1.0.0, renaming or removing a configuration key is a normal change rather than one to design around — that is what keeps the code free of a compatibility layer for every shape a section has ever had. What such a change owes:
- An entry in the changelog under the release’s
### Breakingheading, naming the old spelling and the new one. See Compatibility. - A startup error naming the replacement, where practical, so an unmigrated
configuration stops the server instead of coming up looking configured and
doing nothing.
crates/policy/src/filter/build.rs’srefuse_removed_keysand thesigner.backend = "acme_proxy"arm incrates/signer/src/lib.rsare the worked examples. A key must still parse to be refused by name, which is why the removed[filter]fields survive incrates/core/src/config/types/filter.rs; a field that is gone fails as an opaque serde error instead. - No alias, no dual syntax, no legacy lowering. Delete the old shape. The refusals themselves are one-line diagnostics and go away at 1.0.0.
Changing dependencies
sbom.cdx.json at the repository root is a committed CycloneDX
1.5 inventory of the dependency closure that ships in
the binary — the artifact ASVS 5.0 V15.1.2 asks for, alongside the cargo deny
gate. It is scoped --all-features --target all, so the hsm/cryptoki path
and every platform-gated crate are covered; dev-dependencies are excluded, since
they cannot reach a released build.
Regenerate it after any change to Cargo.toml or Cargo.lock, and when cutting
a release (it records the crate version). The sbom CI job runs the same recipe
and fails on any difference:
export SOURCE_DATE_EPOCH=0
cargo metadata --locked --format-version 1 >/dev/null
cargo cyclonedx --all-features --target all --spec-version 1.5 \
--format json --override-filename sbom.cdx -q
jq --arg from "path+file://$PWD" --arg to "path+file:///acme-proxy" \
'walk(if type == "string" then ((if startswith($from) then $to + .[($from | length):] else . end) | gsub("path\\+file:///acme-proxy#acme-proxy@"; "path+file:///acme-proxy#")) else . end) | del(.metadata.timestamp)' \
sbom.cdx.json > sbom.cdx.json.tmp
mv sbom.cdx.json.tmp sbom.cdx.json
rm -f crates/*/sbom.cdx.json
The tool writes one document per workspace member. The committed one is the binary’s, whose closure already names every library crate, so the per-member copies are deleted rather than committed.
cargo install cargo-cyclonedx@0.5.9 --locked provides the generator; keep the
version in step with the pin in .github/workflows/ci.yml, since it is written
into the document. SOURCE_DATE_EPOCH makes the output reproducible (it also
suppresses the otherwise-random serialNumber); the jq pass drops the
wall-clock timestamp and rewrites the bom-ref values the tool derives from
the checkout path — both the absolute directory it embeds and the name@
segment it drops when that directory’s basename happens to equal the crate
name, so the file is identical whether it was regenerated in a worktree named
acme-proxy or anything else.
Cutting a release
main is the trunk: every pull request targets it, and a minor release is a
tag on it. A patch release comes from a release/X.Y branch, cut from the
X.Y.0 tag when the first fix needs to ship.
ADR 0013 argues the model.
Every crate of the workspace is published to crates.io together, at the
binary’s version: the library crates are internal, with no semver promise of
their own, and exist on crates.io only so cargo install acme-proxy can build.
A minor release
-
On
main, bumpversionin[workspace.package]of the rootCargo.tomland every=x.y.zpin on anacme-proxy-*crate in[workspace.dependencies], together. The exact pins are what keep the crates in step. -
Regenerate
sbom.cdx.json(above); it records the version. -
Check the whole set packages and builds from its own archives, then publish it, in dependency order:
cargo publish --workspace --dry-run cargo publish --workspacePublishing a workspace in one command needs cargo 1.90 or later, below the minimum supported Rust version.
-
Once that commit is on
mainand its CI run is green, push the bare version as a tag:git tag -a 0.6.0 -m 0.6.0 && git push origin 0.6.0The tag triggers
.github/workflows/release.yml, which builds the image on an amd64 and an arm64 runner and publishesghcr.io/acme-proxy/acme-proxyas0.6.0,0.6andlatest, with a provenance attestation. ADR 0012 explains its shape. Itsguardjob stops the release, with nothing published, in four cases:- The tag is not the workspace version, or a crate pin is stale. The tag is most likely mistyped: delete it and push the right one. If step 1 was incomplete, fix the manifest first.
- The tag is off its line.
X.Y.0must be onmain, andX.Y.Zonrelease/X.Y. Move the tag. - There is no CI run on that branch for the commit. The tag is on a commit that was never pushed to the branch. Move the tag.
- CI is pending or failed. Wait for it, or fix it, then use “Re-run all jobs” on the release run for the same tag.
To rehearse the workflow without publishing, run it from a branch under “Run workflow”. It builds both architectures and pushes nothing.
The first time the package is published, it is private, even though the repository is public. In the package’s settings, make it inherit access from the repository, then check that
podman pullworks with no credentials.
A patch release
-
If
release/X.Ydoes not exist yet, cut it from the minor’s tag and push it. The branch ruleset lets a new branch be created without a pull request:git switch -c release/0.6 0.6.0 && git push origin release/0.6 -
Backport every fix it ships, as below.
-
On a topic branch off
release/X.Y, bump the version and the pins, regenerate the SBOM, and give the changelog its## [X.Y.Z]section, holding the backported entries. Merge it intorelease/X.Yby pull request. -
Publish the crates from that commit, as in step 3 above.
-
Once its CI run on
release/X.Yis green, tag it:git tag -a 0.6.1 -m 0.6.1on the branch’s head, then push the tag. The image is published as0.6.1and0.6, and aslatestonly when no higher release exists. -
Cherry-pick the changelog section onto
main, so the trunk’s changelog lists every release, and take the fixed entries out of[Unreleased]there.
Backporting a fix
A fix is merged on main first, and reaches a release branch as a
cherry-pick, never the other way around:
git switch -c backport/0.6/fix-name origin/release/0.6
git cherry-pick -x <commit on main>
Open the pull request against release/0.6. The -x trailer names the commit
on main the fix came from. When the cherry-pick conflicts, resolve it on the
topic branch and say in the pull request what differs from the original.
Trying the next release
Every merge to main publishes ghcr.io/acme-proxy/acme-proxy:edge, once
the whole of ci.yml has passed on it. The image job at the end of ci.yml
calls release.yml for that.
Submitting a pull request
- Fork the repository and create your branch from
main, fixes included. A fix is backported to a release branch after it merges (see above). - Write clear, descriptive commit messages.
- If you’ve added code that should be tested, add tests.
- If you’ve changed APIs, update the documentation in this
mdBook. - Open a PR, describing the problem you’re solving and how you fixed it.
Architecture guidelines
If you are proposing a large feature (like a new Signer or Filter), please review the Architecture & Design documentation first. It’s often best to open an Issue to discuss the design before writing extensive code.
Open work
Planned and deferred work lives in the
issue tracker, one issue per
item, labelled by subsystem (server, store, webadmin, signer, ipam,
notify). Several of them record why something was investigated and not
built, which is worth reading before proposing it again.