Troubleshooting
This section covers common issues when operating acme-proxy in production.
The server refuses to start
Several configuration mistakes are deliberately fatal at startup rather than silently degrading at runtime. The log line names the problem in each case.
| Message | Cause | Fix |
|---|---|---|
no enabled [profiles] | No profile is defined, or all are enabled = false. The server serves ACME only through profiles. | Add [profiles.default], or set ACME_PROXY_PROFILES__DEFAULT__ENABLED=true. |
Unknown or empty challenge.enabled | The list is empty or names a type that does not exist. This is fatal even with challenge.bypass = true — a bypassing server still has to advertise a challenge type. | Use one or more of http-01, dns-01, tls-alpn-01. |
filter.check.<name> has neither allow nor deny entries | An allowed_ip or path check with both lists empty permits everything, so it is a no-op and almost certainly a mistake. | Populate filter.check.<name>.allow or .deny, or delete the check and any rule naming it. |
request_timeout_ms too small | It must exceed signer.custom.timeout_ms when that backend is configured with its crl or renewal_info hook on — those hooks run inline inside a request, so a smaller whole-request deadline would cut them off every time. challenge.timeout_ms and the script’s issue hook are not checked against it: both run in the job queue. | Raise server.request_timeout_ms. |
Two profiles sharing ca.key but differing elsewhere | One CA key under two [signer] configurations would issue under two policies from one identity. | Make the [signer] sections identical (they then share one backend), or give each profile its own CA files. |
| An entry name with an underscore | [filter.check.*], [filter.rule.*], [notify.custom.*], [notify.webhook.*] and profile names must match ^[a-z0-9-]+$. The name is also an environment-variable segment, which the config crate lowercases. | Use hyphens: threat-intel, not threat_intel. |
A custom notify backend enabled with no entries | "custom" is listed in notify.enabled while notify.custom_enabled is empty. | Populate notify.custom_enabled, or drop "custom". |
filter.enabled is no longer a setting | A 0.1.x [filter] section. The flat all-must-pass chain was replaced wholesale by named checks and rules; filter.exempt_paths and filter.custom_enabled are refused the same way. | Declare each filter as a [filter.check.<name>] with a type, write a [filter.rule.<name>] whose when names them, and list the rules in filter.rules. See Filters. |
[filter.rule.<name>] is configured but filter.rules is empty | A rule is declared but nothing selects it, so the policy it describes would never run. Inherited rules count: a profile that leaves filter.rules empty meets this too. | List the rule in filter.rules, or delete it. |
Upstream requires EAB, no .kid sidecar | The relay backend has never registered with the upstream. | Run acme-proxy upstream register --profile <name> --eab-kid …, or set signer.relay.eab.kid/hmac_key in configuration. |
key_source = "pkcs11" … built without | signer.local_ca.key_source = "pkcs11" on a binary with no PKCS#11 support. Deliberately fatal rather than falling back to the file key, which would silently leave the CA key on disk. | Rebuild with cargo build --release --features hsm. |
| A PKCS#11 key that is not the certified one | The token key’s public key does not match cert_path — almost always a wrong key_label. | See Hardware Keys. |
admin.bind_address … is not loopback while admin.tls.enabled is false | The session cookie is always sent Secure, which a browser will not store over plain HTTP on anything but localhost. Refused rather than warned about, because the symptom is otherwise “signing in succeeds and then immediately signs you out” with nothing in any log to explain it. | Set admin.tls.enabled = true, or bind 127.0.0.1 and reach the panel through an SSH tunnel. |
metrics.bind_address and … are both | The metrics endpoint is a listener of its own and cannot share a socket with the ACME or admin one. Checked at startup and on every reload. | Give it its own port. |
notify.webhook.<entry>.body is not a valid template (likewise .url, .method, .headers) | Every webhook entry is compiled and validated at startup: the body template, the URL, a verb outside POST/PUT/PATCH, and a header name or value a wire format will not carry. Deliberately fatal — a broken entry would otherwise become a permanent delivery failure discovered on the first event that mattered. | Fix the entry. See Webhook. |
admin.template_dir … is not a directory, or a template that does not compile | The panel’s per-file overrides are compiled at startup for the same reason, so a broken one refuses to start rather than serving a 500 later. | Point it at a directory, or leave it empty for the compiled-in defaults. See Customizing the Panel. |
A [proxy] URL that is not http://, or a no_proxy entry with a port | proxy.http_url / proxy.https_url name the proxy to dial, which is spelled http://host:port even for HTTPS targets; there is no SOCKS support. A no_proxy entry matches a host or a network, never a port. | Drop the scheme to http://, and the port from the no_proxy entry. |
ipam.phpipam.sources naming vip or fhrp | phpIPAM records neither roles nor redundancy groups, so those two sources exist only for NetBox. Refused by name rather than silently returning nothing. | Use dns_name, custom_field or device. See phpIPAM. |
Failures specific to a hardware CA key — PIN, token, slot and mechanism problems — have their own table in Hardware Keys (PKCS#11).
A wildcard order is rejected
- Symptoms —
newOrderfor*.example.comreturnsrejectedIdentifiernamingdns-01. - Cause — Wildcards can only be proven with a DNS challenge, so
acme-proxyaccepts them only whendns-01is amongchallenge.enabled. - Fix — Add
dns-01tochallenge.enabled. See Challenge Validation.
A client suddenly cannot order anything
- Symptoms — Every order-side request from one client returns
unauthorized. - Cause — The account has been deactivated — by the client itself, or by
acme-proxy account deactivate. Deactivation is permanent and blocks all issuance. - Fix — The client must register a new account.
An order sits in processing and never finishes
- Symptoms — A client polls an order that reached
processingand stays there. Only therelaybackend defers like this;local_caandcustomanswerfinalizeinline. - Cause — The issuance is a background job, and it is either waiting out a backoff after a retryable upstream failure or has been retired for good. The order object itself will not say which.
- Fix — Ask the queue:
# What is the job doing, and how many attempts has it spent?
acme-proxy jobs list --kind signer_relay_issue --limit 20
# The upstream's own error text, and the URLs it was talking to.
acme-proxy jobs show <job-id>
A job still ready is waiting for its next attempt at the run_at it
prints — acme-proxy jobs run-now <id> pulls that forward. A failed one has
spent its budget: run-now grants exactly one more attempt, and
acme-proxy jobs cancel <id> abandons the ACME order so the client stops
polling and can order again. See Job queue, and
job_run_abandoned in Monitoring for the
log line that says it happened.
SQLite database locks
- Symptoms — The server logs show
database is lockederrors during high concurrency order creation. - Cause — The database may not be utilizing Write-Ahead Logging (WAL) or your filesystem does not support proper locking mechanisms (e.g., NFS).
- Fix — Ensure
acme-proxyis running on a local filesystem andjournal_mode = WALis applied. (The server automatically attempts to enable WAL on startup).
Reading the database directly
When the CLI cannot answer a question — usually “what does the row actually say?” — the database is readable while the server runs, because WAL allows a reader alongside the writer.
# Which migrations have run, and did they all succeed?
sqlite3 sqlite.db "SELECT version, description, success FROM _sqlx_migrations;"
# Everything this server refused in the last day, and why.
sqlite3 sqlite.db "SELECT datetime(created_at,'unixepoch'), event, profile, client_ip, reason
FROM audit_log WHERE outcome = 'failure'
AND created_at > strftime('%s','now','-1 day');"
# An order that will not progress: its status and its authorizations'.
sqlite3 sqlite.db "SELECT o.status, a.identifier, a.status, c.type, c.status, c.error
FROM orders o JOIN authorizations a ON a.order_id = o.id
JOIN challenges c ON c.authz_id = a.id WHERE o.id = '…';"
Read, do not write. The
CHECKconstraints will catch an impossible status, but nothing re-syncs the in-memory state a running handler is holding. Use the Admin CLI to change anything.
Backups must include the WAL. Copying sqlite.db on its own gives you a
database missing every recent write. Either take all three files (sqlite.db,
-wal, -shm) with the server stopped, or run sqlite3 sqlite.db ".backup backup.db", which is consistent by construction and safe against a running
server.
The schema itself — every table, constraint and index, and why each is shaped the way it is — is documented in Database Schema.
Upstream Let’s Encrypt rate limits
- Symptoms — Order finalization fails with HTTP 429 Too Many Requests from the upstream CA.
- Cause — When using the
relaysigner backend, all internal clients share a single external ACME account. Let’s Encrypt applies rate limits (e.g., 50 certificates per registered domain per week). - Fix — Request a rate limit increase for the root domain you are relaying, or implement careful caching mechanisms on your internal servers to prevent excessive renewals.
Network challenge blockages
- Symptoms — Order stays in
pendingstate, or challenge verification fails. - Cause — If
challenge.bypass = false, the server must reach the internal client over HTTP (port 80) or DNS to verify domain ownership. Firewalls might be blocking this internal callback. - Fix — Ensure the host running
acme-proxyhas egress network access to reach the internal servers requesting certificates.
EAB registration fails
- Symptoms — Client receives an error stating
External Account Binding is required. - Cause — The client is connecting to a profile that requires EAB, but did
not provide the
kidandhmaccredentials. - Fix — Create EAB credentials using the
acme-proxy eab createCLI command and configure the client to use them.
Order finalization fails (413 payload too large)
- Symptoms — Finalizing the order returns a 413 error.
- Cause — The encoded Certificate Signing Request (CSR) exceeds the
server.max_body_byteslimit (default 128 KiB). - Fix — Ensure your client is not generating excessively large CSRs or increase the limit in your configuration.
Order finalization fails (badCSR: identifier mismatch)
- Symptoms — Finalizing the order returns
400 badCSRcomplaining about the requested identifiers. - Cause — The DNS Subject Alternative Names in your CSR are not exactly the set of names the order authorized. The comparison is set equality, so an extra name fails just as an omitted one does.
- Fix — Configure your ACME client to put every ordered domain — and nothing else — into the CSR’s SANs.
Order finalization fails (badCSR: common name)
- Symptoms —
400 badCSR, “CSR common name is a domain the order does not cover”. - Cause — The CSR’s Subject Common Name looks like a DNS name that the order
does not authorize.
acme-proxydoes not ignore the CN: a domain-shaped CN must be covered by the order, precisely so a name cannot be smuggled past the identifier filters by moving it out of the SANs. - Fix — Either add that name to the order, or drop the CN. A CN that is not
domain-shaped — a human label like
rcgen self signed cert— is tolerated and ignored. Note that thelocal_casigner empties the subject entirely on the issued certificate, so a CN is never carried through to the leaf.
Order finalization fails (badCSR: non-DNS SAN)
- Symptoms —
400 badCSRfrom thelocal_casigner for a CSR containing an IP, email or URI SAN. - Cause —
local_caaccepts DNS SANs only, and rejects anything else outright rather than stripping it. - Fix — Remove the non-DNS SANs from the CSR.
Connection refused or 503 service unavailable
- Symptoms — High-throughput ACME clients receive HTTP 503 errors during bursts.
- Cause — The server’s load shedder activated because the number of
concurrent in-flight requests exceeded
server.max_concurrent_requests. - Fix — Increase
max_concurrent_requestsandadmission_wait_ms, or configure your client to retry with exponential backoff.