Benchmarking Infrastructure Resilience with the Certificate Management Maturity Model
Public Key Infrastructure (PKI) and certificate management have fundamentally shifted from routine IT administrative tasks to critical cybersecurity imperatives. The explosion of machine identities—containers, microservices, APIs, and IoT devices—means that non-human identities now vastly outnumber human ones in any modern enterprise environment.
The primary catalyst driving organizations to mature their certificate management practices is Google’s impending proposal to reduce the maximum public TLS certificate validity from 398 days to just 90 days. Concurrently, the finalization of the National Institute of Standards and Technology (NIST) Post-Quantum Cryptography (PQC) standards requires organizations to rethink how cryptographic algorithms are deployed and rotated.
Relying on manual processes to manage cryptographic identities is no longer mathematically or operationally viable. To understand where your infrastructure stands and how to prepare for these industry shifts, organizations rely on the Certificate Management Maturity Model (CMMM). This framework categorizes PKI operations into four distinct phases, providing a clear roadmap from manual chaos to automated crypto-agility.
Level 1: The Reactive Phase
The lowest tier of the maturity model is characterized by decentralized purchasing, manual tracking, and a severe lack of visibility. In this phase, certificate management is often referred to as the "spreadsheet era."
Different departments purchase their own certificates from various Certificate Authorities (CAs) using corporate credit cards. Tracking is handled through static Excel files, Wiki pages, or disjointed Jira tickets. Because there is no centralized inventory, IT and security teams are completely blind to the cryptographic footprint of their organization.
The Business Impact
The inevitable result of Level 1 maturity is the frequent, unexpected outage. When a spreadsheet is inevitably forgotten or an employee leaves the company, a critical certificate expires silently.
According to reports from the Ponemon Institute, over 80% of organizations have experienced at least one disruptive outage due to an expired certificate in the past two years. The consequences are highly public and financially devastating. In 2023, Starlink suffered a global outage caused by a single expired ground station certificate. Similarly, Cisco issued field notices in 2024 for multiple products failing due to hardcoded, expiring root certificates.
Furthermore, the Reactive Phase breeds shadow IT. Developers, frustrated by slow IT procurement processes, often spin up free, unmonitored certificates or use self-signed certificates in production. These rogue certificates create massive blind spots that attackers can exploit for lateral movement or man-in-the-middle (MitM) attacks.
Level 2: The Managed Phase
Organizations usually enter the Managed Phase immediately following a catastrophic outage. The defining characteristic of Level 2 is the pursuit of visibility.
At this stage, organizations deploy automated discovery tools that scan internal and external networks on standard ports (like 443) to build a centralized inventory of all active certificates. IT teams configure alerting systems to trigger email or Slack notifications 30, 60, or 90 days before a certificate is due to expire.
The Alert Fatigue Problem
While Level 2 significantly reduces the frequency of unexpected outages, it introduces a new operational bottleneck: alert fatigue.
Security teams now know exactly when a certificate will expire, but the actual processes of generating a Certificate Signing Request (CSR), submitting it to the CA, validating domain ownership, downloading the new certificate, and installing it on the endpoint remain entirely manual.
As the volume of microservices and machine identities grows, IT administrators find themselves drowning in renewal tickets. If Google enforces a 90-day maximum lifespan for public TLS certificates, the workload for a Level 2 organization will quadruple overnight.
To bridge the gap between manual tracking and full automation, organizations need highly reliable, noise-free monitoring. This is where Expiring.at provides critical infrastructure support. By offering precise expiration tracking and customizable webhooks, Expiring.at ensures that infrastructure teams receive actionable alerts exactly when they need them, integrating cleanly into existing incident management workflows without overwhelming the team with unactionable noise.
Level 3: The Automated Phase
Level 3 represents the "DevOps Era" of certificate management and is the minimum required maturity level to survive the transition to 90-day public certificates and short-lived internal certificates.
In the Automated Phase, manual CSR generation and installation are entirely eliminated. Instead, organizations adopt standardized protocols like the Automated Certificate Management Environment (ACME) defined in RFC 8555. Certificate provisioning and renewal become zero-touch processes integrated directly into CI/CD pipelines and infrastructure-as-code (IaC) deployments.
Implementing Automation with Kubernetes cert-manager
For cloud-native environments, moving to Level 3 usually involves deploying cert-manager, the de facto standard for Kubernetes certificate automation. cert-manager allows developers to request certificates natively using YAML manifests, completely abstracting the CA interaction.
Here is a practical example of how a Level 3 organization automates TLS for a web application using Let's Encrypt and an ACME ClusterIssuer:
# 1. Define the ACME ClusterIssuer
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-production
spec:
acme:
server: https://acme-v02.api.letsencrypt.org/directory
email: security@yourdomain.com
privateKeySecretRef:
name: letsencrypt-production-account-key
solvers:
- http01:
ingress:
class: nginx
---
# 2. Request the Certificate
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: api-gateway-cert
namespace: production
spec:
secretName: api-gateway-tls
duration: 2160h # 90 days
renewBefore: 360h # 15 days
issuerRef:
name: letsencrypt-production
kind: ClusterIssuer
dnsNames:
- api.yourdomain.com
In this setup, cert-manager automatically monitors the Certificate resource. When the renewBefore threshold is crossed, it automatically negotiates with the ACME server, completes the HTTP-01 challenge, downloads the new certificate, updates the Kubernetes Secret, and reloads the Ingress controller—all without human intervention.
For mobile device management (MDM) and IoT environments, Level 3 organizations utilize protocols like Simple Certificate Enrollment Protocol (SCEP) or Enrollment over Secure Transport (EST) to securely provision identities to headless devices at scale.
Level 4: The Crypto-Agile Phase
The pinnacle of the Certificate Management Maturity Model is Level 4: Crypto-Agility. While Level 3 focuses on operational efficiency (stopping outages and saving time), Level 4 focuses on long-term cryptographic resilience and Zero Trust Architecture (ZTA).
Crypto-agility is the ability of an organization to completely swap out cryptographic algorithms, CAs, or trust stores across the entire enterprise rapidly and without operational downtime.
Preparing for Post-Quantum Cryptography
The urgency for Level 4 maturity is driven by the advent of quantum computing. In August 2024, NIST released the finalized Post-Quantum Cryptography standards:
* FIPS 203 (ML-KEM): For general encryption and key encapsulation.
* FIPS 204 (ML-DSA) & FIPS 205 (SLH-DSA): For digital signatures.
Current algorithms like RSA and ECC will eventually be broken by cryptographically relevant quantum computers (CRQCs). Organizations at Level 1 or 2 will take years to find and manually replace every RSA key in their infrastructure. A Level 4 organization can update a central policy-as-code repository and automatically rotate the entire fleet of certificates to ML-DSA signatures in a matter of hours.
Dynamic Issuance and Service Meshes
Level 4 organizations enforce strict policy-as-code using Role-Based Access Control (RBAC) to dictate exactly which workloads can request specific types of certificates.
Internal PKI lifespans shrink from years to days, or even hours, to limit the blast radius if a private key is compromised. This is achieved using frameworks like SPIFFE and SPIRE, which issue cryptographic identities based on workload attributes rather than static IP addresses or hostnames.
Furthermore, mutual TLS (mTLS) between microservices is handled transparently by service meshes like Istio or Linkerd. Application developers do not write any code to handle encryption; the service mesh sidecar proxies intercept traffic, validate the peer's certificate, and establish a secure tunnel automatically.
# Example Istio PeerAuthentication policy enforcing strict mTLS
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: default-strict-mtls
namespace: istio-system
spec:
mtls:
mode: STRICT
With a single policy deployment, a Level 4 organization guarantees that all inter-service communication is encrypted using automatically rotated, short-lived certificates.
Compliance and Regulatory Pressures
Maturing your certificate management is no longer just an engineering best practice; it is increasingly a strict legal and regulatory requirement.
- PCI-DSS v4.0: The latest iteration of the Payment Card Industry Data Security Standard enforces much stricter requirements on cryptography. Organizations are now explicitly required