Every component of Stellar Control writes structured logs and serves Prometheus metrics and a health
endpoint. The services of the core configure them in the observability section of the
global configuration; components built with the Rust SDK take
STELLAR_METRICS_ADDR.
Logs#
- Format. One JSON object per line on standard output (
observability.log_format: json, the default), flat, with the fields of the event and of the current span; or human-readable text (text) for development. - Filter.
observability.log_level, in the syntax oftracingdirectives:info,debug,stellar_reconciler=debug,info.RUST_LOG, when set, takes precedence. - Context. Every line of a service carries the
componentandinstance. The handling of a NATS message runs in amessagespan with itsoperation,subject,correlation(the telecommand or run identifier) andconfig(the configuration revision), so every log line about a telecommand can be found by its identifier. - At start, a service logs its effective configuration (
effective configuration), the address of its metrics endpoint, andcomponent started.
The language server logs to standard error, at warn unless RUST_LOG says otherwise.
Endpoints#
| Path | Answer |
|---|---|
/metrics | Prometheus text format |
/healthz | ok while the process runs |
A service serves them on observability.metrics_addr (0.0.0.0:9100 by default); give each
instance on a host its own port with STELLAR__OBSERVABILITY__METRICS_ADDR. The all-in-one
binary serves one endpoint for all its services. A driver, transport, gateway or connector of the
Rust SDK serves them on STELLAR_METRICS_ADDR, and none when it is unset.
Every metric carries the labels component (the kind: reconciler, driver…) and instance.
Durations are histograms with buckets from 0.5 ms to 5 s.
Metrics#
Every component#
| Metric | Type | Labels | Meaning |
|---|---|---|---|
stellar_component_info | gauge | version | 1 while the component runs |
stellar_component_start_time_seconds | gauge | Start time, in seconds since the epoch | |
stellar_messages_total | counter | operation, outcome | NATS messages handled |
stellar_message_duration_seconds | histogram | operation | Time to handle a message |
Services of the core#
| Metric | Type | Labels | Service | Meaning |
|---|---|---|---|---|
stellar_reconciler_leader | gauge | reconciler | 1 on the leader, 0 on a standby instance | |
stellar_reconciler_registrations_total | counter | outcome (accepted, rejected) | reconciler | Registrations of instances |
stellar_connector_pending | gauge | connector, kind | reconciler | Messages of a connector consumer not delivered or not acknowledged |
stellar_connector_lag_seconds | gauge | connector, kind | reconciler | Age of its oldest message not acknowledged |
stellar_connector_at_risk | gauge | connector, kind | reconciler | 1 when the retention is about to remove messages it has not read |
stellar_compute_messages_total | counter | outcome | compute | Decoded telemetry messages handled |
stellar_compute_samples_total | counter | target | compute | Samples published |
stellar_values_samples_total | counter | outcome | compute (value table) | Samples applied to the current value table |
stellar_executor_telecommands_total | counter | state | executor | Telecommand transitions |
stellar_executor_runs_total | counter | outcome (succeeded, failed) | executor | Runs finished |
stellar_alarms_transitions_total | counter | state | alarms | Alarm transitions |
stellar_scheduler_runs_total | counter | outcome (fired, missed) | scheduler | Scheduled runs fired or missed |
stellar_transfers_bytes_total | counter | target, outcome (asked, received, written) | transfers | Bytes of file transfers |
Drivers, transports and gateways (Rust SDK)#
| Metric | Type | Labels | Meaning |
|---|---|---|---|
stellar_driver_frames_total | counter | target, outcome (decoded, failed) | Frames decoded |
stellar_driver_telecommands_total | counter | target, outcome (encoded, failed, expired) | Telecommands handled |
stellar_driver_cfdp_pdus_total | counter | target, direction (sent, received) | CFDP PDUs of the entity of the driver |
stellar_transport_units_total | counter | target, outcome (framed, failed, expired) | Units to frame |
stellar_transport_frames_total | counter | target, outcome (unwrapped, failed) | Frames received |
stellar_transport_cop1_frames_total | counter | target, link, kind (first, retransmitted) | Frames sent by the COP-1 FOP |
stellar_transport_cop1_alerts_total | counter | target, link | COP-1 alerts (lockout, transmission limit): the FOP stops |
stellar_gateway_uplink_frames_total | counter | target, outcome (sent, failed, dropped, expired) | Frames to uplink |
stellar_gateway_throughput_bps | gauge | direction (uplink, downlink), target | Mean throughput over the last report window |
stellar_gateway_bytes_total | counter | direction, target | Bytes carried since the start |
Gateways also publish their throughput on NATS, stellar.metrics.<gateway>.<target>, every
metrics_period (5 s): the transfer manager uses it, and the web console shows it (see
Writing a Gateway).
Connector lag#
Every 15 s, the reconciler leader measures each consumer of each
output connector: messages not delivered, messages not
acknowledged, age of the oldest of them. It publishes them on stellar.metrics.connector.<name>
and as the stellar_connector_* gauges; GET /v1/connectors gives the same on demand.
A consumer is at risk when the retention of its stream is about to remove messages it has not
read: its oldest unread message has reached 80 % of the max_age of the stream, or the stream,
bound by max_bytes, is 90 % full while its oldest message is still unread. The reconciler then
logs a warning. Alert on it:
groups:
- name: stellar-connectors
rules:
- alert: ConnectorAboutToLoseData
expr: stellar_connector_at_risk == 1
labels: {severity: critical}
annotations:
summary: "Connector {{ $labels.connector }} ({{ $labels.kind }}) is about to lose data"Suggested alerts#
| Condition | Expression |
|---|---|
| A component is down | up == 0 on its scrape target |
| No reconciler leader | max(stellar_reconciler_leader) < 1 |
| Registrations rejected | increase(stellar_reconciler_registrations_total{outcome="rejected"}[10m]) > 0 |
| Runs failing | increase(stellar_executor_runs_total{outcome="failed"}[1h]) > 0 |
| Scheduled runs missed | increase(stellar_scheduler_runs_total{outcome="missed"}[1h]) > 0 |
| Frames not decoded | rate(stellar_driver_frames_total{outcome="failed"}[5m]) > 0 |
| COP-1 alerts | increase(stellar_transport_cop1_alerts_total[10m]) > 0 |
| Connector about to lose data | stellar_connector_at_risk == 1 |
Readiness in NATS#
Health of the instances is also visible without Prometheus:
- every instance publishes a heartbeat on
stellar.ctl.hb.<kind>.<instance>(healthy, ordegradedwith a reason; a gateway adds whether its link is available). Its key instellar_instancesexpires after three heartbeat periods: an expired instance is unbound and its targets are no longer ready; - the readiness of each target, with a reason per link, is in
stellar_readiness; GET /v1/topology,GET /v1/instancesandGET /v1/instances/{kind}/{instance}show them, as do the Topology and Instances views of the web console.
NATS itself serves its monitoring on its HTTP port (8222 in the development setup):
/healthz, /varz, /jsz.