Monitoring
This page helps you watch a running Goiabada through its Prometheus metrics, and alert when something goes wrong.
The metrics say how many requests each route answers and how fast, how close the database pool is to its cap, how often the rate limiter refuses, how many tokens are issued and refused, whether the cleanup run succeeds, and, from the admin console, how the auth server answers it. For which request failed and why, read the logs.
Turn the metrics on
Section titled “Turn the metrics on”Each server has a metrics listener of its own, off until you turn it on:
| Server | Turned on by | Port |
|---|---|---|
| Auth server | GOIABADA_AUTHSERVER_METRICS_ENABLED=true |
9190, set by GOIABADA_AUTHSERVER_LISTEN_PORT_METRICS |
| Admin console | GOIABADA_ADMINCONSOLE_METRICS_ENABLED=true |
9191, set by GOIABADA_ADMINCONSOLE_LISTEN_PORT_METRICS |
-
Turn the listeners on. On Kubernetes, answer the setup wizard’s metrics question, as Scrape on Kubernetes describes. With Docker Compose or native binaries, add the two variables above to each server’s environment: the Compose files and the env file the wizard writes leave them off.
-
Restart both servers.
-
Check what the auth server serves, from beside it. On Kubernetes:
Terminal window kubectl port-forward -n goiabada deploy/goiabada-authserver 9190:9190curl -s http://localhost:9190/metrics | grep goiabada_build_infoWith Docker Compose, from a container on the Compose network:
curl -s http://goiabada-authserver:9190/metrics. -
Point your scraper at both servers, on Kubernetes or outside it.
Each listener listens on every address by default, because a scraper in another container or pod reaches it over the network. When the scraper runs on the same host, set GOIABADA_AUTHSERVER_LISTEN_HOST_METRICS and GOIABADA_ADMINCONSOLE_LISTEN_HOST_METRICS to 127.0.0.1. The environment variables list every setting and its flag.
Scrape on Kubernetes
Section titled “Scrape on Kubernetes”The setup wizard asks Expose Prometheus metrics?, with three answers:
| Answer | Flag | What the manifest gets |
|---|---|---|
| None (default) | --metrics=none |
Nothing: the metrics listeners stay off |
| Pod annotations | --metrics=annotations |
Both listeners on, a container port named metrics on each server, 9190 and 9191, and the prometheus.io/scrape, port and path annotations on both pod templates |
| PodMonitor | --metrics=podmonitor |
Both listeners on, the metrics container ports, and one monitoring.coreos.com/v1 PodMonitor selecting both servers by that port |
Pick the answer that matches what scrapes your cluster:
| Scraper | Answer | Finds Goiabada by |
|---|---|---|
The prometheus-community prometheus Helm chart |
Pod annotations | The annotations |
| kube-prometheus-stack, or any Prometheus the Prometheus Operator runs | PodMonitor | A PodMonitor |
| Grafana Alloy | Pod annotations | The metrics container port |
| The OpenTelemetry Collector | Pod annotations | The metrics container port |
Alloy and the Collector don’t read the annotations; that answer is simply the one that turns the listeners on and names the port without writing a PodMonitor, which kubectl apply refuses on a cluster without the Operator.
With the NetworkPolicies on, the wizard also asks which namespace the scraper runs in (--metrics-namespace, monitoring by default). Each policy then gains a second ingress rule admitting that namespace to its server’s metrics port alone. Give it the namespace your scraper’s pods actually run in, which kubectl get pods -A shows: a policy naming any other admits nothing. To admit a scraper in another namespace too, add its namespaceSelector to that second rule’s from. Without the NetworkPolicies, any pod in the cluster can read the metrics.
Pod annotations
Section titled “Pod annotations”The wizard writes these on each server’s pod template, here the auth server’s:
annotations: prometheus.io/scrape: "true" prometheus.io/port: "9190" prometheus.io/path: "/metrics"They’re a convention, not part of Kubernetes or Prometheus. The prometheus-community prometheus chart’s default configuration reads them, and kube-prometheus-stack ignores them. For the NetworkPolicy, the namespace is the one the chart’s Prometheus runs in.
A PodMonitor
Section titled “A PodMonitor”The Prometheus Operator, which kube-prometheus-stack installs, scrapes only what a PodMonitor or a ServiceMonitor names. The wizard writes one PodMonitor for both servers:
apiVersion: monitoring.coreos.com/v1kind: PodMonitormetadata: name: goiabada namespace: goiabada labels: release: kube-prometheus-stackspec: selector: matchExpressions: - key: app operator: In values: - goiabada-authserver - goiabada-adminconsole podMetricsEndpoints: - port: metrics path: /metricsA Prometheus the Operator runs selects only the PodMonitors its podMonitorSelector matches, in the namespaces its podMonitorNamespaceSelector matches. kube-prometheus-stack, by default, selects only those labeled release: <its release name>. The wizard asks for the labels to give the PodMonitor (--podmonitor-labels), such as release=kube-prometheus-stack, blank for none. Read what your Prometheus selects by with:
kubectl get prometheus -A -o jsonpath='{..podMonitorSelector}'A PodMonitor no Prometheus selects is accepted and scraped by nothing. PodMonitor is one of the Operator’s CRDs: on a cluster without them, kubectl apply exits 1 after applying everything else in the file. For the NetworkPolicy, the namespace is the one Prometheus runs in, monitoring in the usual kube-prometheus-stack install.
Grafana Alloy
Section titled “Grafana Alloy”Alloy discovers pods itself. Keep the targets on the metrics container port, and forward them to the prometheus.remote_write component your configuration already has:
discovery.kubernetes "goiabada" { role = "pod" namespaces { names = ["goiabada"] }}
discovery.relabel "goiabada" { targets = discovery.kubernetes.goiabada.targets
rule { source_labels = ["__meta_kubernetes_pod_container_port_name"] regex = "metrics" action = "keep" } rule { source_labels = ["__meta_kubernetes_pod_name"] target_label = "pod" } rule { source_labels = ["__meta_kubernetes_pod_container_name"] target_label = "container" }}
prometheus.scrape "goiabada" { targets = discovery.relabel.goiabada.output forward_to = [prometheus.remote_write.default.receiver]}Alloy’s service account needs to list and watch pods in the goiabada namespace, which the Alloy Helm chart grants. For the NetworkPolicy, the namespace is the one Alloy runs in.
The OpenTelemetry Collector
Section titled “The OpenTelemetry Collector”The Collector’s prometheus receiver takes a Prometheus scrape configuration, so it finds the pods the same way:
receivers: prometheus: config: scrape_configs: - job_name: goiabada kubernetes_sd_configs: - role: pod namespaces: names: [goiabada] relabel_configs: - source_labels: [__meta_kubernetes_pod_container_port_name] regex: metrics action: keep - source_labels: [__meta_kubernetes_pod_name] target_label: pod - source_labels: [__meta_kubernetes_pod_container_name] target_label: containerAdd the receiver to a metrics pipeline, and give the Collector’s service account permission to list and watch pods in the goiabada namespace. For the NetworkPolicy, the namespace is the one the Collector runs in.
A manifest generated without metrics
Section titled “A manifest generated without metrics”Run the wizard again with the answer you want and compare. Or set GOIABADA_AUTHSERVER_METRICS_ENABLED and GOIABADA_ADMINCONSOLE_METRICS_ENABLED to "true" in the two servers’ ConfigMaps, and add the metrics port to each container yourself, here the auth server’s, with 9191 for the admin console:
ports:- containerPort: 9090- name: metrics containerPort: 9190Scrape outside Kubernetes
Section titled “Scrape outside Kubernetes”A Prometheus on the same Compose network as Goiabada scrapes both servers by their service names, with the ports published nowhere:
scrape_configs:- job_name: goiabada static_configs: - targets: - goiabada-authserver:9190 - goiabada-adminconsole:9191For native binaries, list the hosts the servers run on, and set the listen hosts to 127.0.0.1 when Prometheus runs on the same host.
Set up alerts
Section titled “Set up alerts”A starting set, as a Prometheus rules file. The thresholds are where to begin, not measurements: tune them once you have a few weeks of traffic. With the Prometheus Operator, put the groups under the spec of a PrometheusRule carrying the labels your Prometheus selects rules by, as for the PodMonitor.
groups:- name: goiabada rules: - alert: GoiabadaDatabasePoolWaits expr: sum by (instance) (rate(goiabada_db_wait_duration_seconds_total[5m])) > 0.1 for: 10m annotations: summary: Requests are waiting for a database connection. - alert: GoiabadaServerErrors expr: | sum(rate(goiabada_http_requests_total{status=~"5.."}[5m])) / sum(rate(goiabada_http_requests_total[5m])) > 0.02 for: 10m annotations: summary: More than 2% of requests are answered with a server error. - alert: GoiabadaRateLimitRefusals expr: sum by (limiter) (increase(goiabada_rate_limit_refusals_total[10m])) > 50 annotations: summary: The rate limiter is refusing a burst of requests. - alert: GoiabadaNoSuccessfulCleanup expr: | time() - max(max_over_time(goiabada_cleanup_last_success_timestamp_seconds[1d])) > 86400 and on () count(goiabada_db_max_open_connections offset 1d) > 0 annotations: summary: No cleanup run has completed in 24 hours. - alert: GoiabadaAfterResponseJobsDropped expr: sum by (class) (increase(goiabada_after_response_jobs_dropped_total[10m])) > 0 annotations: summary: Mail was not sent because too many jobs were in flight. - alert: GoiabadaAuthServerFailingTheConsole expr: | sum(rate(goiabada_upstream_requests_total{status=~"error|5.."}[5m])) / sum(rate(goiabada_upstream_requests_total[5m])) > 0.05 for: 5m annotations: summary: The admin console's calls to the auth server are failing.What each one is for:
- Pool waits. The database pool is capped at
GOIABADA_DB_MAX_OPEN_CONNSconnections per pod, and a request that finds every connection in use waits for one. The expression is the time spent waiting, per second, on one pod: above 0.1, requests wait a tenth of a second for every second that passes. Comparegoiabada_db_connections{state="in_use"}withgoiabada_db_max_open_connections: a pool that sits at its cap needs a larger cap, if the database’s connection limit has room for it across every pod, or more pods. The average wait israte(goiabada_db_wait_duration_seconds_total[5m]) / rate(goiabada_db_wait_count_total[5m]). - Server errors. A 5xx is a fault on Goiabada’s side or in what it depends on, the database or the mail server, rather than a client’s mistake. Find which route with
sum by (route, status) (rate(goiabada_http_requests_total{status=~"5.."}[5m])), and the cause in theERRORrecords the same requests wrote to the logs. - Rate-limit refusals. Every refusal is counted, so a credential-stuffing run or a flood shows as a spike on the limiter it hit, as it happens. The rate limits give each limiter’s budget, and the audit log’s
rate_limit_exceededevents say which accounts and addresses the refusals were about. The limits counted by a user or an email always apply. Those counted by an IP address apply, and show here, only onceGOIABADA_AUTHSERVER_RATELIMITER_ENABLEDistrue. The setup wizard asks, and turns it on by default, except on Kubernetes under the Cluster traffic policy, where every client shares a node’s address. - No successful cleanup. One auth server pod claims a cleanup run every 12 hours, deleting expired sessions, tokens, codes and audit records. Each pod reports the last run it completed, so the deployment’s latest is the maximum across pods, kept for a day so a rollout doesn’t lose it. The second line holds the alert back until the metrics go back a day, since a new deployment has had no run yet. A run that keeps failing writes an
ERRORrecord for the step that failed. - Dropped jobs. Forgot-password, self-registration and the notice of an email change send their mail after the response, at most 64 of each at once. A dropped job’s mail is never sent: the user waits for an email that doesn’t arrive. A sustained
goiabada_after_response_jobs_in_flightnear 64 is the same thing about to happen, usually a slow mail server. - The auth server failing the console. The admin console’s failure mode is the auth server. A
statusoferroris a call that got no response at all: refused, reset or timed out.
Alert as well on the scrape itself failing, Prometheus’s up series at 0 for Goiabada’s targets, and on pods restarting, which Kubernetes reports through kube-state-metrics rather than Goiabada. When a scrape fails, see Metrics are not scraped.
What the listener serves
Section titled “What the listener serves”The metrics listener answers GET /metrics in the Prometheus text exposition format, and any other path with a 404. It isn’t part of the application’s routes: a scrape is neither logged nor counted in the request metrics.
A label only ever takes a value from a set declared in the code when the metric is registered, and a value outside that set is recorded as other. So no metric can produce more series than the product of its label sets, whatever requests arrive, and nothing taken from a request or from the database reaches a label: no user, client, session, email address, IP address, path, query, user agent or error description. The route label is the server’s own route table, read at startup, and a request no route answered is recorded under unmatched.
Prometheus adds the job, instance and, on Kubernetes, pod and container labels itself. A metric both servers expose has the same name, labels and meaning on each, so one query covers both, and those labels tell them apart.
Metrics catalog
Section titled “Metrics catalog”Every metric Goiabada exposes, with each label and the values it can take:
| metric | type | labels | server | meaning |
|---|---|---|---|---|
goiabada_http_requests_total |
counter | route: the route table, unmatched; method: GET, HEAD, POST, PUT, PATCH, DELETE, CONNECT, OPTIONS, TRACE; status: the response’s status code |
both | HTTP requests answered, by the route that answered them, the method and the exact status code. Every request the server’s main listeners answer is counted, /health and the static files included; a scrape of the metrics listener is not. |
goiabada_http_request_duration_seconds |
histogram | route: the route table, unmatched; method: GET, HEAD, POST, PUT, PATCH, DELETE, CONNECT, OPTIONS, TRACE |
both | How long requests took to answer, in seconds. The buckets are 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30 and 60: the slowest requests send mail and can take up to 40 seconds. |
goiabada_db_max_open_connections |
gauge | none | auth server | The most connections the database pool may hold open at once, GOIABADA_DB_MAX_OPEN_CONNS on PostgreSQL, MySQL and SQL Server, and 1 on SQLite. |
goiabada_db_connections |
gauge | state: in_use, idle |
auth server | Connections the database pool holds, by whether a request is using them or they are idle. in_use reaching the cap is the pool fully taken. |
goiabada_db_wait_count_total |
counter | none | auth server | Requests that waited for a database connection because the pool was at its cap. Its rate rising is the pool becoming too small for the load. |
goiabada_db_wait_duration_seconds_total |
counter | none | auth server | Time spent waiting for a database connection, in seconds, summed over every wait. Its rate divided by the wait count’s is the average wait. |
goiabada_db_connections_closed_total |
counter | reason: max_idle, max_idle_time, max_lifetime |
auth server | Connections the database pool closed, by the limit that closed them: the idle cap, the idle time or the lifetime. |
goiabada_rate_limit_refusals_total |
counter | limiter: pwd_account_net, pwd_account, pwd_ip, otp, email_verification, email_verification_send, account_password, activate, register, register_email, reset_pwd, forgot_pwd_email, forgot_pwd_ip, dcr, ropc_ip |
auth server | Requests the rate limiter refused with 429, by the limiter that refused them. Every refusal is counted, where the audit log records one per key per window, so a spike in its rate is a credential-stuffing or flooding attempt as it happens. |
goiabada_tokens_issued_total |
counter | grant_type: authorization_code, refresh_token, client_credentials, password, implicit |
auth server | Token responses the server answered with, by the grant that issued them: the token endpoint’s four grants, and the implicit grant’s tokens from the authorization endpoint. |
goiabada_token_requests_refused_total |
counter | grant_type: authorization_code, refresh_token, client_credentials, password; error: invalid_request, invalid_client, invalid_grant, unauthorized_client, unsupported_grant_type, invalid_scope, server_error |
auth server | Token requests the token endpoint refused, by the grant_type the request named and the error code it was answered with. A grant type the endpoint does not redeem is recorded as other. A request the rate limiter refused is counted in goiabada_rate_limit_refusals_total instead. |
goiabada_cleanup_runs_total |
counter | outcome: completed, failed, interrupted |
auth server | Cleanup runs this instance performed, by how they ended: every step succeeded, a step failed, or shutdown cut the run short. The run is claimed by one instance every 12 hours, so on a deployment of several pods each counts only the runs it won. |
goiabada_cleanup_last_run_duration_seconds |
gauge | none | auth server | How long this instance’s last cleanup run took, in seconds, however it ended. The worker’s worker task completed log record carries the same duration. |
goiabada_cleanup_last_success_timestamp_seconds |
gauge | none | auth server | When this instance’s last cleanup run completed, as a Unix timestamp in seconds, or 0 when none has since it started. Across a deployment the latest success is the maximum over its pods. |
goiabada_after_response_jobs_in_flight |
gauge | class: recovery, registration, account_notice |
auth server | Work handed off to run after a response that is running now, by class: forgot-password’s code and mail, self-registration’s mail, and the notice of an email change. Each class runs at most 64 at once. |
goiabada_after_response_jobs_dropped_total |
counter | class: recovery, registration, account_notice |
auth server | Work handed off to run after a response that was dropped because its class already had 64 running, by class. A dropped job’s mail is never sent. |
goiabada_upstream_requests_total |
counter | target: admin_api, settings, token, jwks, sessions; status: the response’s status code, or error |
admin console | Calls the admin console made to the auth server, by the client that made them and the status code the auth server answered, or error when no response arrived: a connection refused or reset, or a timeout. admin_api is every call to the admin and account APIs, settings the public settings, token the sign-in’s code exchange, the refresh and the client credentials grant, jwks the signing keys, and sessions the administrators’ browser sessions, which the auth server stores; a retry is a call of its own. A rising rate of error or 5xx is the auth server failing as the console sees it. |
goiabada_upstream_request_duration_seconds |
histogram | target: admin_api, settings, token, jwks, sessions |
admin console | How long the admin console’s calls to the auth server took until the response headers arrived, in seconds, by the client that made them, with the buckets of goiabada_http_request_duration_seconds. |
goiabada_settings_cache_requests_total |
counter | result: hit, miss |
admin console | Lookups of the auth server’s public settings, which every console page needs, by whether a cached value answered them. A miss found no fresh value, whether it started the call to the auth server or waited on one already under way; the value is kept 30 seconds. |
goiabada_build_info |
gauge | version: this binary’s version; commit: this binary’s commit |
both | Always 1. The labels say which release is running. |
go_goroutines |
gauge | none | both | Goroutines that currently exist. |
go_memstats_heap_inuse_bytes |
gauge | none | both | Heap bytes in use. |
Rate limiters
Section titled “Rate limiters”The values of the limiter label, and what each one refuses:
limiter |
Refuses |
|---|---|
pwd_account_net |
Failed sign-ins on one account from one address, at the sign-in form and the password grant together |
pwd_account |
Failed sign-ins on one account from any address, at the sign-in form and the password grant together |
pwd_ip |
Every sign-in form submission from one address |
otp |
Failed OTP codes for one user |
email_verification |
Failed email verification codes for one account |
email_verification_send |
Requests to send a verification email for one account |
account_password |
Failed passwords on one account’s password, OTP and email changes together |
activate |
Account activation requests from one address |
register |
Self-registrations from one address |
register_email |
Self-registrations for one email address |
reset_pwd |
Password reset requests from one address |
forgot_pwd_email |
Forgot-password requests for one email address |
forgot_pwd_ip |
Forgot-password requests from one address |
dcr |
Dynamic client registrations from one address |
ropc_ip |
Every password grant request from one address |