CRL vs. OCSP vs. OCSP Stapling
Issuing a TLS certificate is a solved problem. Thanks to the ACME protocol and automated authorities, provisioning cryptographic trust takes seconds. Revoking that trust, however, has historically been the most broken component of the Public Key Infrastructure (PKI) ecosystem.
For decades, the industry has wrestled with mechanisms that were either too bloated to be practical, too slow for modern web performance, or fundamentally compromised user privacy.
As the industry prepares for Google's "Moving Forward, Together" initiative—which aims to reduce the maximum validity of public TLS certificates from 398 days to just 90 days—the conversation around revocation is shifting. Shorter lifespans drastically reduce the window of vulnerability for a compromised key. But if a private key is exposed or a CA mandates a mass revocation on day two of a 90-day lifecycle, you still need a reliable way to tell the world not to trust that certificate for the remaining 88 days.
Understanding how we got here, why legacy protocols fail, and how modern infrastructure handles revocation is critical for DevOps engineers and security professionals tasked with maintaining resilient systems.
The Heavyweight: Certificate Revocation Lists (CRL)
The earliest attempt at solving the revocation problem was the Certificate Revocation List (CRL). The concept is straightforward: the Certificate Authority (CA) publishes a cryptographically signed file containing the serial numbers of all revoked, unexpired certificates they have issued.
When a client connects to a server, it downloads this list, verifies the CA's signature, and checks if the server's certificate serial number is on it.
The Problem with CRLs
The fatal flaw of the CRL is unbounded growth. Certificates are only removed from a CRL when they naturally expire. For a massive public CA, a CRL can easily grow to tens or hundreds of megabytes.
Forcing a client—especially a mobile device on a slow network—to download a 50MB file and parse millions of serial numbers during a TLS handshake causes unacceptable latency. Because of this bloat, traditional CRLs are effectively dead for public web browsing.
Where CRL is Still King
Despite its failure on the public web, CRL remains heavily utilized in closed enterprise environments. In Zero Trust Architectures (ZTA), mutual TLS (mTLS), and internal IoT networks, the number of issued certificates is tightly controlled.
For a corporate VPN or an internal Kubernetes cluster issuing certificates via an internal CA, a CRL might only contain a few dozen entries. In these environments, downloading a tiny CRL is fast, reliable, and entirely offline, making it a perfectly viable solution for internal compliance and access control.
The Privacy and Performance Nightmare: Live OCSP
To solve the bandwidth problem of CRLs, the industry introduced the Online Certificate Status Protocol (OCSP). Instead of downloading the entire list of revoked certificates, the client simply asks the CA about the specific certificate it just encountered.
During the TLS handshake, the client pauses, extracts the OCSP responder URL from the certificate's Authority Information Access (AIA) extension, and sends an HTTP request to the CA: "Is certificate serial #12345 valid?" The CA responds with a signed "good," "revoked," or "unknown" status.
Why Live OCSP Failed
Live OCSP solved the bandwidth problem but introduced three severe new issues:
- The Privacy Leak: By querying the CA every time a user visits a website, the CA receives a real-time log of the user's IP address and the exact domain they are visiting.
- The Latency Penalty: Live OCSP requires the client to perform a DNS lookup, establish a TCP connection, and complete an HTTP request to a third-party server before it can finish the TLS handshake with the original website.
- The "Soft-Fail" Flaw: Because CA OCSP responders frequently went offline under massive global load, browsers had to make a choice: if the OCSP server doesn't respond, do we block the user from the website (hard-fail), or do we assume the certificate is valid and let them through (soft-fail)?
Browsers chose soft-fail to prevent massive internet outages. However, this rendered OCSP useless against active Man-in-the-Middle (MITM) attacks. An attacker intercepting a connection simply blocks the outbound OCSP request. The browser times out, soft-fails, and accepts the revoked, compromised certificate anyway.
The danger of live OCSP was perfectly illustrated during a major GlobalSign outage. The CA accidentally revoked an intermediate certificate. Because millions of clients were performing live OCSP and CRL checks, major websites like Wikipedia and the Financial Times became completely inaccessible to users, proving that relying on live CA infrastructure for client handshakes creates a catastrophic single point of failure.
The Server-Side Savior: OCSP Stapling
If clients shouldn't download massive lists, and they shouldn't query the CA in real-time, who should handle the revocation check? The answer is the web server itself.
OCSP Stapling (formally known as the TLS Certificate Status Request extension) shifts the burden of proof from the client to the server.
Instead of the client querying the CA, the web server periodically queries the CA's OCSP responder in the background. The server caches the CA's cryptographically signed, time-stamped response. When a client initiates a TLS handshake, the server "staples" this cached OCSP response directly to the certificate payload.
The Benefits of Stapling
- Zero Client Latency: The client receives the certificate and the revocation status in the exact same round-trip. No external DNS or HTTP requests are required.
- Privacy Preserved: The CA only sees the IP address of the web server requesting the OCSP response, not the end-users visiting the site.
- High Availability: If the CA's OCSP responder goes offline, the web server continues to serve the cached, valid response until its natural expiration (usually 3 to 7 days). Cloudflare uses this edge architecture aggressively, shielding end-users from CA infrastructure instability.
OCSP Must-Staple
To fix the soft-fail problem, the industry introduced the "OCSP Must-Staple" X.509 v3 extension. When you request a certificate, you can ask the CA to embed this extension. It explicitly tells the browser: "Reject this connection if I do not provide a valid, stapled OCSP response."
While Must-Staple fixes the security gap, adoption remains incredibly low (under 1%). Server administrators are terrified of bricking their own websites. If a server fails to fetch a fresh OCSP response and the cache expires, a Must-Staple certificate will cause browsers to hard-fail the connection, resulting in downtime.
Implementing OCSP Stapling in Production
Enabling OCSP stapling is a mandatory best practice for any public-facing infrastructure. Modern ingress controllers and web servers like Caddy and Traefik handle OCSP stapling automatically by default, requiring zero configuration.
However, if you are running Nginx or Apache, you must configure it manually.
Nginx Configuration
In Nginx, simply turning stapling on isn't enough. You must provide a DNS resolver so Nginx can resolve the CA's OCSP responder hostname, and you must provide the chain of trust to verify the response.
```nginx
server {
listen 443 ssl;
server_name example.com;
ssl_certificate /path/to/fullchain.pem;
ssl_certificate_key /path/to/privkey.pem;
# Enable OCSP Stapling
ssl_stapling on;
ssl_stapling_verify on;
# Point