Skip to content

Monitoring

This page helps you watch a running Goiabada through its Prometheus metrics, and alert when something goes wrong.

The metrics say how many requests each route answers and how fast, how close the database pool is to its cap, how often the rate limiter refuses, how many tokens are issued and refused, whether the cleanup run succeeds, and, from the admin console, how the auth server answers it. For which request failed and why, read the logs.

Each server has a metrics listener of its own, off until you turn it on:

Server Turned on by Port
Auth server GOIABADA_AUTHSERVER_METRICS_ENABLED=true 9190, set by GOIABADA_AUTHSERVER_LISTEN_PORT_METRICS
Admin console GOIABADA_ADMINCONSOLE_METRICS_ENABLED=true 9191, set by GOIABADA_ADMINCONSOLE_LISTEN_PORT_METRICS
  1. Turn the listeners on. On Kubernetes, answer the setup wizard’s metrics question, as Scrape on Kubernetes describes. With Docker Compose or native binaries, add the two variables above to each server’s environment: the Compose files and the env file the wizard writes leave them off.

  2. Restart both servers.

  3. Check what the auth server serves, from beside it. On Kubernetes:

    Terminal window
    kubectl port-forward -n goiabada deploy/goiabada-authserver 9190:9190
    curl -s http://localhost:9190/metrics | grep goiabada_build_info

    With Docker Compose, from a container on the Compose network: curl -s http://goiabada-authserver:9190/metrics.

  4. Point your scraper at both servers, on Kubernetes or outside it.

Each listener listens on every address by default, because a scraper in another container or pod reaches it over the network. When the scraper runs on the same host, set GOIABADA_AUTHSERVER_LISTEN_HOST_METRICS and GOIABADA_ADMINCONSOLE_LISTEN_HOST_METRICS to 127.0.0.1. The environment variables list every setting and its flag.

The setup wizard asks Expose Prometheus metrics?, with three answers:

Answer Flag What the manifest gets
None (default) --metrics=none Nothing: the metrics listeners stay off
Pod annotations --metrics=annotations Both listeners on, a container port named metrics on each server, 9190 and 9191, and the prometheus.io/scrape, port and path annotations on both pod templates
PodMonitor --metrics=podmonitor Both listeners on, the metrics container ports, and one monitoring.coreos.com/v1 PodMonitor selecting both servers by that port

Pick the answer that matches what scrapes your cluster:

Scraper Answer Finds Goiabada by
The prometheus-community prometheus Helm chart Pod annotations The annotations
kube-prometheus-stack, or any Prometheus the Prometheus Operator runs PodMonitor A PodMonitor
Grafana Alloy Pod annotations The metrics container port
The OpenTelemetry Collector Pod annotations The metrics container port

Alloy and the Collector don’t read the annotations; that answer is simply the one that turns the listeners on and names the port without writing a PodMonitor, which kubectl apply refuses on a cluster without the Operator.

With the NetworkPolicies on, the wizard also asks which namespace the scraper runs in (--metrics-namespace, monitoring by default). Each policy then gains a second ingress rule admitting that namespace to its server’s metrics port alone. Give it the namespace your scraper’s pods actually run in, which kubectl get pods -A shows: a policy naming any other admits nothing. To admit a scraper in another namespace too, add its namespaceSelector to that second rule’s from. Without the NetworkPolicies, any pod in the cluster can read the metrics.

The wizard writes these on each server’s pod template, here the auth server’s:

annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9190"
prometheus.io/path: "/metrics"

They’re a convention, not part of Kubernetes or Prometheus. The prometheus-community prometheus chart’s default configuration reads them, and kube-prometheus-stack ignores them. For the NetworkPolicy, the namespace is the one the chart’s Prometheus runs in.

The Prometheus Operator, which kube-prometheus-stack installs, scrapes only what a PodMonitor or a ServiceMonitor names. The wizard writes one PodMonitor for both servers:

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: goiabada
namespace: goiabada
labels:
release: kube-prometheus-stack
spec:
selector:
matchExpressions:
- key: app
operator: In
values:
- goiabada-authserver
- goiabada-adminconsole
podMetricsEndpoints:
- port: metrics
path: /metrics

A Prometheus the Operator runs selects only the PodMonitors its podMonitorSelector matches, in the namespaces its podMonitorNamespaceSelector matches. kube-prometheus-stack, by default, selects only those labeled release: <its release name>. The wizard asks for the labels to give the PodMonitor (--podmonitor-labels), such as release=kube-prometheus-stack, blank for none. Read what your Prometheus selects by with:

Terminal window
kubectl get prometheus -A -o jsonpath='{..podMonitorSelector}'

A PodMonitor no Prometheus selects is accepted and scraped by nothing. PodMonitor is one of the Operator’s CRDs: on a cluster without them, kubectl apply exits 1 after applying everything else in the file. For the NetworkPolicy, the namespace is the one Prometheus runs in, monitoring in the usual kube-prometheus-stack install.

Alloy discovers pods itself. Keep the targets on the metrics container port, and forward them to the prometheus.remote_write component your configuration already has:

discovery.kubernetes "goiabada" {
role = "pod"
namespaces {
names = ["goiabada"]
}
}
discovery.relabel "goiabada" {
targets = discovery.kubernetes.goiabada.targets
rule {
source_labels = ["__meta_kubernetes_pod_container_port_name"]
regex = "metrics"
action = "keep"
}
rule {
source_labels = ["__meta_kubernetes_pod_name"]
target_label = "pod"
}
rule {
source_labels = ["__meta_kubernetes_pod_container_name"]
target_label = "container"
}
}
prometheus.scrape "goiabada" {
targets = discovery.relabel.goiabada.output
forward_to = [prometheus.remote_write.default.receiver]
}

Alloy’s service account needs to list and watch pods in the goiabada namespace, which the Alloy Helm chart grants. For the NetworkPolicy, the namespace is the one Alloy runs in.

The Collector’s prometheus receiver takes a Prometheus scrape configuration, so it finds the pods the same way:

receivers:
prometheus:
config:
scrape_configs:
- job_name: goiabada
kubernetes_sd_configs:
- role: pod
namespaces:
names: [goiabada]
relabel_configs:
- source_labels: [__meta_kubernetes_pod_container_port_name]
regex: metrics
action: keep
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod
- source_labels: [__meta_kubernetes_pod_container_name]
target_label: container

Add the receiver to a metrics pipeline, and give the Collector’s service account permission to list and watch pods in the goiabada namespace. For the NetworkPolicy, the namespace is the one the Collector runs in.

Run the wizard again with the answer you want and compare. Or set GOIABADA_AUTHSERVER_METRICS_ENABLED and GOIABADA_ADMINCONSOLE_METRICS_ENABLED to "true" in the two servers’ ConfigMaps, and add the metrics port to each container yourself, here the auth server’s, with 9191 for the admin console:

ports:
- containerPort: 9090
- name: metrics
containerPort: 9190

A Prometheus on the same Compose network as Goiabada scrapes both servers by their service names, with the ports published nowhere:

scrape_configs:
- job_name: goiabada
static_configs:
- targets:
- goiabada-authserver:9190
- goiabada-adminconsole:9191

For native binaries, list the hosts the servers run on, and set the listen hosts to 127.0.0.1 when Prometheus runs on the same host.

A starting set, as a Prometheus rules file. The thresholds are where to begin, not measurements: tune them once you have a few weeks of traffic. With the Prometheus Operator, put the groups under the spec of a PrometheusRule carrying the labels your Prometheus selects rules by, as for the PodMonitor.

groups:
- name: goiabada
rules:
- alert: GoiabadaDatabasePoolWaits
expr: sum by (instance) (rate(goiabada_db_wait_duration_seconds_total[5m])) > 0.1
for: 10m
annotations:
summary: Requests are waiting for a database connection.
- alert: GoiabadaServerErrors
expr: |
sum(rate(goiabada_http_requests_total{status=~"5.."}[5m]))
/ sum(rate(goiabada_http_requests_total[5m])) > 0.02
for: 10m
annotations:
summary: More than 2% of requests are answered with a server error.
- alert: GoiabadaRateLimitRefusals
expr: sum by (limiter) (increase(goiabada_rate_limit_refusals_total[10m])) > 50
annotations:
summary: The rate limiter is refusing a burst of requests.
- alert: GoiabadaNoSuccessfulCleanup
expr: |
time() - max(max_over_time(goiabada_cleanup_last_success_timestamp_seconds[1d])) > 86400
and on () count(goiabada_db_max_open_connections offset 1d) > 0
annotations:
summary: No cleanup run has completed in 24 hours.
- alert: GoiabadaAfterResponseJobsDropped
expr: sum by (class) (increase(goiabada_after_response_jobs_dropped_total[10m])) > 0
annotations:
summary: Mail was not sent because too many jobs were in flight.
- alert: GoiabadaAuthServerFailingTheConsole
expr: |
sum(rate(goiabada_upstream_requests_total{status=~"error|5.."}[5m]))
/ sum(rate(goiabada_upstream_requests_total[5m])) > 0.05
for: 5m
annotations:
summary: The admin console's calls to the auth server are failing.

What each one is for:

  • Pool waits. The database pool is capped at GOIABADA_DB_MAX_OPEN_CONNS connections per pod, and a request that finds every connection in use waits for one. The expression is the time spent waiting, per second, on one pod: above 0.1, requests wait a tenth of a second for every second that passes. Compare goiabada_db_connections{state="in_use"} with goiabada_db_max_open_connections: a pool that sits at its cap needs a larger cap, if the database’s connection limit has room for it across every pod, or more pods. The average wait is rate(goiabada_db_wait_duration_seconds_total[5m]) / rate(goiabada_db_wait_count_total[5m]).
  • Server errors. A 5xx is a fault on Goiabada’s side or in what it depends on, the database or the mail server, rather than a client’s mistake. Find which route with sum by (route, status) (rate(goiabada_http_requests_total{status=~"5.."}[5m])), and the cause in the ERROR records the same requests wrote to the logs.
  • Rate-limit refusals. Every refusal is counted, so a credential-stuffing run or a flood shows as a spike on the limiter it hit, as it happens. The rate limits give each limiter’s budget, and the audit log’s rate_limit_exceeded events say which accounts and addresses the refusals were about. The limits counted by a user or an email always apply. Those counted by an IP address apply, and show here, only once GOIABADA_AUTHSERVER_RATELIMITER_ENABLED is true. The setup wizard asks, and turns it on by default, except on Kubernetes under the Cluster traffic policy, where every client shares a node’s address.
  • No successful cleanup. One auth server pod claims a cleanup run every 12 hours, deleting expired sessions, tokens, codes and audit records. Each pod reports the last run it completed, so the deployment’s latest is the maximum across pods, kept for a day so a rollout doesn’t lose it. The second line holds the alert back until the metrics go back a day, since a new deployment has had no run yet. A run that keeps failing writes an ERROR record for the step that failed.
  • Dropped jobs. Forgot-password, self-registration and the notice of an email change send their mail after the response, at most 64 of each at once. A dropped job’s mail is never sent: the user waits for an email that doesn’t arrive. A sustained goiabada_after_response_jobs_in_flight near 64 is the same thing about to happen, usually a slow mail server.
  • The auth server failing the console. The admin console’s failure mode is the auth server. A status of error is a call that got no response at all: refused, reset or timed out.

Alert as well on the scrape itself failing, Prometheus’s up series at 0 for Goiabada’s targets, and on pods restarting, which Kubernetes reports through kube-state-metrics rather than Goiabada. When a scrape fails, see Metrics are not scraped.

The metrics listener answers GET /metrics in the Prometheus text exposition format, and any other path with a 404. It isn’t part of the application’s routes: a scrape is neither logged nor counted in the request metrics.

A label only ever takes a value from a set declared in the code when the metric is registered, and a value outside that set is recorded as other. So no metric can produce more series than the product of its label sets, whatever requests arrive, and nothing taken from a request or from the database reaches a label: no user, client, session, email address, IP address, path, query, user agent or error description. The route label is the server’s own route table, read at startup, and a request no route answered is recorded under unmatched.

Prometheus adds the job, instance and, on Kubernetes, pod and container labels itself. A metric both servers expose has the same name, labels and meaning on each, so one query covers both, and those labels tell them apart.

Every metric Goiabada exposes, with each label and the values it can take:

metric type labels server meaning
goiabada_http_requests_total counter route: the route table, unmatched; method: GET, HEAD, POST, PUT, PATCH, DELETE, CONNECT, OPTIONS, TRACE; status: the response’s status code both HTTP requests answered, by the route that answered them, the method and the exact status code. Every request the server’s main listeners answer is counted, /health and the static files included; a scrape of the metrics listener is not.
goiabada_http_request_duration_seconds histogram route: the route table, unmatched; method: GET, HEAD, POST, PUT, PATCH, DELETE, CONNECT, OPTIONS, TRACE both How long requests took to answer, in seconds. The buckets are 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30 and 60: the slowest requests send mail and can take up to 40 seconds.
goiabada_db_max_open_connections gauge none auth server The most connections the database pool may hold open at once, GOIABADA_DB_MAX_OPEN_CONNS on PostgreSQL, MySQL and SQL Server, and 1 on SQLite.
goiabada_db_connections gauge state: in_use, idle auth server Connections the database pool holds, by whether a request is using them or they are idle. in_use reaching the cap is the pool fully taken.
goiabada_db_wait_count_total counter none auth server Requests that waited for a database connection because the pool was at its cap. Its rate rising is the pool becoming too small for the load.
goiabada_db_wait_duration_seconds_total counter none auth server Time spent waiting for a database connection, in seconds, summed over every wait. Its rate divided by the wait count’s is the average wait.
goiabada_db_connections_closed_total counter reason: max_idle, max_idle_time, max_lifetime auth server Connections the database pool closed, by the limit that closed them: the idle cap, the idle time or the lifetime.
goiabada_rate_limit_refusals_total counter limiter: pwd_account_net, pwd_account, pwd_ip, otp, email_verification, email_verification_send, account_password, activate, register, register_email, reset_pwd, forgot_pwd_email, forgot_pwd_ip, dcr, ropc_ip auth server Requests the rate limiter refused with 429, by the limiter that refused them. Every refusal is counted, where the audit log records one per key per window, so a spike in its rate is a credential-stuffing or flooding attempt as it happens.
goiabada_tokens_issued_total counter grant_type: authorization_code, refresh_token, client_credentials, password, implicit auth server Token responses the server answered with, by the grant that issued them: the token endpoint’s four grants, and the implicit grant’s tokens from the authorization endpoint.
goiabada_token_requests_refused_total counter grant_type: authorization_code, refresh_token, client_credentials, password; error: invalid_request, invalid_client, invalid_grant, unauthorized_client, unsupported_grant_type, invalid_scope, server_error auth server Token requests the token endpoint refused, by the grant_type the request named and the error code it was answered with. A grant type the endpoint does not redeem is recorded as other. A request the rate limiter refused is counted in goiabada_rate_limit_refusals_total instead.
goiabada_cleanup_runs_total counter outcome: completed, failed, interrupted auth server Cleanup runs this instance performed, by how they ended: every step succeeded, a step failed, or shutdown cut the run short. The run is claimed by one instance every 12 hours, so on a deployment of several pods each counts only the runs it won.
goiabada_cleanup_last_run_duration_seconds gauge none auth server How long this instance’s last cleanup run took, in seconds, however it ended. The worker’s worker task completed log record carries the same duration.
goiabada_cleanup_last_success_timestamp_seconds gauge none auth server When this instance’s last cleanup run completed, as a Unix timestamp in seconds, or 0 when none has since it started. Across a deployment the latest success is the maximum over its pods.
goiabada_after_response_jobs_in_flight gauge class: recovery, registration, account_notice auth server Work handed off to run after a response that is running now, by class: forgot-password’s code and mail, self-registration’s mail, and the notice of an email change. Each class runs at most 64 at once.
goiabada_after_response_jobs_dropped_total counter class: recovery, registration, account_notice auth server Work handed off to run after a response that was dropped because its class already had 64 running, by class. A dropped job’s mail is never sent.
goiabada_upstream_requests_total counter target: admin_api, settings, token, jwks, sessions; status: the response’s status code, or error admin console Calls the admin console made to the auth server, by the client that made them and the status code the auth server answered, or error when no response arrived: a connection refused or reset, or a timeout. admin_api is every call to the admin and account APIs, settings the public settings, token the sign-in’s code exchange, the refresh and the client credentials grant, jwks the signing keys, and sessions the administrators’ browser sessions, which the auth server stores; a retry is a call of its own. A rising rate of error or 5xx is the auth server failing as the console sees it.
goiabada_upstream_request_duration_seconds histogram target: admin_api, settings, token, jwks, sessions admin console How long the admin console’s calls to the auth server took until the response headers arrived, in seconds, by the client that made them, with the buckets of goiabada_http_request_duration_seconds.
goiabada_settings_cache_requests_total counter result: hit, miss admin console Lookups of the auth server’s public settings, which every console page needs, by whether a cached value answered them. A miss found no fresh value, whether it started the call to the auth server or waited on one already under way; the value is kept 30 seconds.
goiabada_build_info gauge version: this binary’s version; commit: this binary’s commit both Always 1. The labels say which release is running.
go_goroutines gauge none both Goroutines that currently exist.
go_memstats_heap_inuse_bytes gauge none both Heap bytes in use.

The values of the limiter label, and what each one refuses:

limiter Refuses
pwd_account_net Failed sign-ins on one account from one address, at the sign-in form and the password grant together
pwd_account Failed sign-ins on one account from any address, at the sign-in form and the password grant together
pwd_ip Every sign-in form submission from one address
otp Failed OTP codes for one user
email_verification Failed email verification codes for one account
email_verification_send Requests to send a verification email for one account
account_password Failed passwords on one account’s password, OTP and email changes together
activate Account activation requests from one address
register Self-registrations from one address
register_email Self-registrations for one email address
reset_pwd Password reset requests from one address
forgot_pwd_email Forgot-password requests for one email address
forgot_pwd_ip Forgot-password requests from one address
dcr Dynamic client registrations from one address
ropc_ip Every password grant request from one address