Skip to content

Certificates are not issued

This page helps you when cert-manager doesn’t issue the certificates for a Kubernetes deployment set up with Envoy Gateway and cert-manager.

kubectl get certificates -n goiabada shows goiabada-tls-auth and goiabada-tls-admin with READY at False, or lists nothing at all. Meanwhile, https://auth.example.com fails in the browser with a certificate or connection error, since the Gateway has no certificate to serve.

The certificates come from a chain of four pieces, and any of them can stop it:

  1. The generated Gateway carries the annotation cert-manager.io/cluster-issuer: "letsencrypt-prod". cert-manager reads it, and creates one Certificate for each HTTPS listener, named after the listener’s certificateRefs: goiabada-tls-auth and goiabada-tls-admin. It does that only with its Gateway API support turned on.
  2. The ClusterIssuer named letsencrypt-prod asks Let’s Encrypt for each certificate.
  3. Let’s Encrypt checks you control each host name with an HTTP-01 challenge: it fetches a token from http://<host>/.well-known/acme-challenge/ on port 80.
  4. cert-manager answers that request through a temporary HTTPRoute it attaches to the Gateway named in the ClusterIssuer’s solver, and a solver pod it starts in the goiabada namespace. The Gateway’s http listener on port 80 carries it.

So the usual causes are:

  • No Certificates at all: cert-manager runs without --enable-gateway-api, so it never reads the Gateway. Step 2 of Gateway and certificates turns it on, after Envoy Gateway has brought the Gateway API’s resources.
  • The issuer doesn’t match: the ClusterIssuer isn’t named letsencrypt-prod, isn’t ready, or its solver’s parentRefs names another Gateway or namespace than the one you deployed into.
  • DNS doesn’t point at the Gateway yet, so Let’s Encrypt fetches the token from somewhere else.
  • The names were looked up before their records existed. Before it asks Let’s Encrypt, cert-manager fetches the token itself, and a resolver that answered “no such name” then keeps that answer for as long as your zone allows, often 30 minutes, even once the records exist. The Challenge’s reason ends in no such host, while nslookup from your own machine finds the address. It clears by itself; creating the ClusterIssuer only once both names resolve, as Gateway and certificates does, avoids it.
  • Port 80 doesn’t reach Envoy: a firewall or security group in front of the load balancer, or the Local traffic policy with nodes that have no Envoy pod.
  • The solver pod is refused, because the namespace enforces the restricted Pod Security standard and the pod doesn’t meet it. The generated namespace only warns, and cert-manager’s own solver pod meets restricted, but a podTemplate in the ClusterIssuer’s solver that sets its own security context, or cert-manager run with --acme-http01-solver-run-as-non-root=false, can break that. kubectl get events -n goiabada shows a refused pod.
  1. Follow the chain to where it stops:

    Terminal window
    kubectl get certificates,certificaterequests,orders,challenges -n goiabada
    kubectl describe challenges -n goiabada

    A Challenge’s status and events say what Let’s Encrypt or cert-manager’s own check got back, such as a host that doesn’t resolve or a connection that timed out.

  2. Check the issuer is ready and named as the Gateway expects:

    Terminal window
    kubectl get clusterissuer letsencrypt-prod
  3. Check that the Gateway has an address and that DNS points at it:

    Terminal window
    kubectl get gateway goiabada -n goiabada
    nslookup auth.example.com

    PROGRAMMED should be True, and both host names should resolve to the ADDRESS it shows. Check what the cluster resolves too, since it can remember a name as missing after your machine finds it:

    Terminal window
    kubectl run dnscheck --rm -i --restart=Never --image=busybox -- nslookup auth.example.com
  4. Under the Local traffic policy, check that Envoy runs a pod on every node:

    Terminal window
    kubectl get daemonset -n envoy-gateway-system

    If your load balancer still can’t reach Envoy, switch to the Cluster traffic policy, whose gatewayclass.yaml has no envoyDaemonSet. Goiabada then sees a node’s address for every client. If your eg GatewayClass has no parametersRef to the EnvoyProxy at all, apply the gatewayclass.yaml for your traffic policy from Gateway and certificates.

  5. Once the cause is fixed, a challenge still pending passes at cert-manager’s next check. One Let’s Encrypt has already refused ends that attempt, and cert-manager waits before the next one, an hour after the first failure and longer after each later one. To try again at once, renew with cmctl, cert-manager’s command-line tool:

    Terminal window
    cmctl renew --namespace goiabada --all

How the certificates are served and renewed

Section titled “How the certificates are served and renewed”

The Gateway terminates TLS for both host names, with the certificate in each listener’s Secret. On the http listener on port 80, Goiabada’s one HTTPRoute, goiabada-https-redirect, answers every request with a 301 to HTTPS. cert-manager’s challenge route matches the exact path /.well-known/acme-challenge/<token>, which takes precedence over it, so the challenges still reach the solver. cert-manager renews each certificate before it expires the same way, so port 80 has to stay reachable after the first issue too.