Documentation → Monitoring

Monitoring

This guide covers the essentials. Read the full upstream guide on GitHub → For a specific release, select its tag in the repository.

Enable a private management listener

[health]
bind_address = "127.0.0.1"
bind_port = 8080

These endpoints use plain HTTP and have no authentication. Keep them on loopback or a private management network.

Health probes

EndpointMeaning
/livezThe process can respond, including during loading and drain.
/readyzHTTP 200 when the runtime is running and at least one served zone is ACTIVE; otherwise 503.
/healthzAlias for readiness.
/metricsPrometheus scrape; may return 429 when rate-limited.

Readiness is not an all-zones check. Probe every expected zone over UDP and TCP and compare its SOA serial with the primary. A primary outage alone is not a reason to restart a responsive process.

Prometheus

Start by scraping the configured private management address. Useful metric families include:

ConcernMetric
Per-zone availabilityborondns_secondary_zone_state
Initial loadingborondns_secondary_zone_loading_seconds
Served serialborondns_secondary_zone_soa_serial
Refresh failuresborondns_secondary_zone_refresh_failures
Failed transfersborondns_transfer_sessions_failed_total
Query trafficborondns_queries_received_total
RRL dropsborondns_rrl_responses_dropped_total

Retain history externally: counters reset on restart. The upstream health and metrics reference defines additional metrics, labels, scrape limits, and response bodies.

Logs and investigation

journalctl -u borondns --since '10 minutes ago'

Investigate zones that remain LOADING, expired zones, transfer failures, and authentication errors. Check primary reachability, credentials, certificate validity, and transfer limits.

Keep metrics.hot_path_detail = "full" for complete query-path instrumentation. Reduced and off modes suppress measurements and cannot be interpreted as complete traffic counts.

For JSON status snapshots, see the optional observability API. Its bearer token does not protect the probes or /metrics. Use the operational SLO guide to choose deployment-specific alert thresholds.