A renewal job can save a new certificate while the website continues serving the old one. DNS credentials can stop working after a team handover, and expiry alerts can reach someone who no longer maintains the site. To catch these failures, track each certificate, assign an owner, and verify what users receive after renewal.
This checklist covers renewal schedules, DNS access, service reloads, monitoring, and recovery for websites, APIs, dashboards, and reverse proxies. On a self-managed Virtarix VPS, your team maintains the certificate tools and the software that serves each certificate.
Check all four stages of renewal
Confirm each stage separately:
- Validation succeeds. The certificate authority can verify control through the configured HTTP, DNS, or other supported challenge.
- Issuance succeeds. A new certificate and its chain are written to the expected location.
- Deployment succeeds. The correct service loads the new certificate without breaking its configuration.
- External verification succeeds. A client connecting to every public hostname receives the intended certificate, chain, and validity dates.
Check that DNS validation works, the service can read the renewed files, and the reload succeeds. Then connect from outside the server to confirm that the public endpoint serves the new certificate. A timer marked successful does not establish all of those results.
Name the certificate owner
Assign a primary and backup owner for each responsibility below. One person may cover several roles in a small team, but someone else must be able to access the tools and respond when that person is unavailable.
| Role | Accountable for | Evidence to retain |
|---|---|---|
| Service owner | Business impact, maintenance window, and escalation priority | Current contact and service criticality |
| Renewal owner | Renewal client, schedule, challenge method, logs, and remediation | Last successful test and last successful renewal |
| DNS-account owner | Zone access and DNS API credential recovery for DNS validation | Account owner, backup administrator, and recovery route |
| Incident owner | Coordinating response before expiry or after a failed deployment | Paging route, backup contact, and escalation deadline |
Use team-controlled accounts for DNS and certificate-authority access where the provider supports them. Remove stale personal accounts after ownership is transferred and tested. Store recovery codes or equivalent recovery material in the team's approved secret system; do not place credentials or private keys in the certificate inventory.
Ownership must also survive absence. The backup owner should be able to locate the inventory, read the renewal logs, access the required DNS or hosting account, run the documented test, and contact the incident owner without relying on the primary owner's device or memory.
Map every certificate
List certificate files and renewal records on the server. Separately list public hostnames and listeners from DNS, proxies, load balancers, and application settings. Compare the lists to find unused certificates and endpoints missing from renewal automation.
Keep the issuance and expiry fields together:
| Domain/SAN | Service | Issuer | Renewal method | Expiry |
|---|---|---|---|---|
app.example.com
|
Public reverse proxy | Certificate authority name | ACME HTTP-01, DNS-01, or documented manual method | UTC date and remaining days |
Link that row to its operational ownership and deployment fields:
| Domain/SAN | Owner | DNS owner | Reload target | Last test |
|---|---|---|---|---|
app.example.com
|
Primary and backup renewal owners | Team account or Not required
|
Exact service and listener | UTC date, result, and evidence link |
For a multi-domain certificate, record every Subject Alternative Name rather than only the common name. If the same certificate is installed on several endpoints, list every reload target and test every endpoint. Also record whether a service terminates TLS locally or receives traffic through another proxy; otherwise a team may renew the wrong layer.
Review the inventory after adding or removing a domain, changing DNS providers, moving a listener, replacing the renewal client, changing a service account, or handing the system to another team. A scheduled inventory review is useful, but configuration changes are the events most likely to make renewal data stale.
Confirm renewal visibility
Use two independent signals:
- External expiry monitoring connects to each public hostname with the correct Server Name Indication and measures the certificate actually served to clients.
- Renewal-job monitoring records whether the scheduled job ran, whether it attempted a renewal, its result, the certificate name, a safe error summary, and where its detailed logs can be found.
Send expiry alerts early enough for the team to investigate access, validation, and rate-limit problems before the certificate becomes urgent. Use more than one threshold, route the later threshold with higher urgency, and alert separately when the renewal job has not run within its expected interval. The inventory should identify who receives each alert and who is paged if the first owner does not acknowledge it.
Do not infer renewal success only from a zero process exit or the presence of new files. Some clients exit successfully when no certificate is due, and a renewed file can remain undeployed. Store the old and new certificate serial number or SHA-256 fingerprint, validity window, deployment result, and external observation. Do not put private-key material in logs.
Monitor the prerequisites as well: DNS API authentication, filesystem permissions, free space, time synchronisation, challenge reachability, renewal-client scheduling, and access to the certificate authority. A failure in any one of them can surface first as an expiry alert.
Certbot example: inspect and test renewal
The commands in this example are specific to Certbot and should be run only on a host where Certbot is the documented renewal client. Confirm the installed version and read the Certbot renewal documentation before changing options.
sudo certbot certificates
sudo certbot renew --dry-run
Certbot documents renew --dry-run as a test against the Let's Encrypt staging service by default. It can run pre- and post-hooks and may temporarily change and roll back web-server configuration; deploy hooks do not run during a dry run unless explicitly enabled. Run it in a controlled window, inspect the complete result, and test the real service reload path separately. Do not add --force-renewal to a routine test: it requests live renewals regardless of certificate age and can consume certificate-authority limits.
Test the reload path
Issuance and deployment should be separate observable steps. After a new certificate is issued:
- Confirm that the new certificate covers the intended hostname set and that its chain, validity dates, serial number, and fingerprint differ as expected.
- Confirm the service account can read the certificate and private key without broadening permissions.
- Run the service's native configuration test before a reload.
- Reload the exact web server, application gateway, or reverse proxy that terminates TLS. Use a restart only when the software requires it and the maintenance plan permits the interruption.
- Connect to every public hostname and relevant port from outside the service, using the correct SNI value.
- Compare the served certificate's names, issuer, serial or fingerprint, validity dates, and chain with the certificate that was issued.
- Run an application smoke test through TLS and record the result.
The following OpenSSL command is a generic inspection example for an HTTPS endpoint. Replace both instances of app.example.com; do not treat the output as proof until it has been compared with the intended certificate record.
openssl s_client -connect app.example.com:443 -servername app.example.com </dev/null 2>/dev/null \
| openssl x509 -noout -subject -issuer -dates -serial -fingerprint -sha256
Test every terminating endpoint. Round-robin DNS, multiple proxies, containers, or copied certificate files can leave one node serving an old certificate while another is current. If the renewal client supports a deploy hook, make the hook run only after successful issuance and have it fail visibly when the configuration test, reload, or post-deployment verification fails.
Keep DNS and account access aligned
DNS validation depends on more than a valid API token. Record the authoritative DNS provider, account owner, zone, credential location, token scope, renewal client or plugin that consumes it, and the recovery route if the primary owner is unavailable. Verify that the credential can update only the required zone or records where the provider supports scoped access.
When transferring ownership, create and test the replacement access before removing the old credential. Then revoke the old token, confirm that renewal uses the replacement, and update the inventory. Do not leave a personal DNS credential in a root-owned configuration file simply because automation still works.
Test account recovery without exposing secrets. Confirm that the team controls the account email or identity-provider group, backup administrators have access, and MFA recovery is documented. Make sure the incident owner can contact the DNS provider without depending on a former employee's device or personal account.
Emergency procedure for a failed renewal
Use a dated incident record and keep the old, currently served certificate in place while it remains valid and the service is healthy. Do not replace working files with an unverified manual result.
- Confirm impact and time remaining. Inspect the certificate served externally, note the UTC expiry time, affected hostnames, service owner, and incident owner.
- Identify the failed stage. Separate validation, issuance, file deployment, service reload, and external verification. Preserve the exact safe error and relevant log location.
- Protect service continuity. Pause repeated automated attempts if they are causing a retry storm or certificate-authority limits. Do not delete the certificate files, keys, or renewal configuration already serving the endpoint.
- Repair the narrow dependency. Restore challenge reachability, DNS access, filesystem permission, disk capacity, renewal configuration, or service configuration as appropriate. Test the repair before issuing a live certificate.
- Perform a temporary manual renewal only when supported. Follow the current renewal-client and certificate-authority documentation, use the existing certificate identity where appropriate, and record every changed option. Avoid creating an unrelated duplicate certificate as an undocumented workaround.
- Validate before deployment. Check hostname coverage, chain, validity dates, serial or fingerprint, file ownership, and the service's native configuration test.
- Reload and verify. Reload the exact TLS-terminating service, inspect the externally served certificate for every hostname and endpoint, and complete an application smoke test.
- Escalate while time remains. If the failure is not resolved by the documented internal deadline, contact the incident owner, DNS or certificate-authority account owner, and the relevant software vendor or provider. Virtarix does not manage the customer's certificate software or renewal incident.
If deployment breaks the service, revert only the deployment change: restore the last-known-good service configuration or certificate reference, run the native configuration test, reload, and verify externally. A rollback to an already expired certificate does not restore trustworthy service, so continue the incident and escalation path until a valid certificate is served. After recovery, remove temporary credentials, restore automation, run a safe renewal test, and document the root cause and prevention action.