Stellar ControlMission control · by Stellar Systems v0.1.0

Deployment

Observability

Logs, Prometheus metrics, health endpoints and alerts of the services, drivers, transports, gateways and connectors.

Every component of Stellar Control writes structured logs and serves Prometheus metrics and a health endpoint. The services of the core configure them in the observability section of the global configuration; components built with the Rust SDK take STELLAR_METRICS_ADDR.

Logs#

  • Format. One JSON object per line on standard output (observability.log_format: json, the default), flat, with the fields of the event and of the current span; or human-readable text (text) for development.
  • Filter. observability.log_level, in the syntax of tracing directives: info, debug, stellar_reconciler=debug,info. RUST_LOG, when set, takes precedence.
  • Context. Every line of a service carries the component and instance. The handling of a NATS message runs in a message span with its operation, subject, correlation (the telecommand or run identifier) and config (the configuration revision), so every log line about a telecommand can be found by its identifier.
  • At start, a service logs its effective configuration (effective configuration), the address of its metrics endpoint, and component started.

The language server logs to standard error, at warn unless RUST_LOG says otherwise.

Endpoints#

PathAnswer
/metricsPrometheus text format
/healthzok while the process runs

A service serves them on observability.metrics_addr (0.0.0.0:9100 by default); give each instance on a host its own port with STELLAR__OBSERVABILITY__METRICS_ADDR. The all-in-one binary serves one endpoint for all its services. A driver, transport, gateway or connector of the Rust SDK serves them on STELLAR_METRICS_ADDR, and none when it is unset.

Every metric carries the labels component (the kind: reconciler, driver…) and instance. Durations are histograms with buckets from 0.5 ms to 5 s.

Metrics#

Every component#

MetricTypeLabelsMeaning
stellar_component_infogaugeversion1 while the component runs
stellar_component_start_time_secondsgaugeStart time, in seconds since the epoch
stellar_messages_totalcounteroperation, outcomeNATS messages handled
stellar_message_duration_secondshistogramoperationTime to handle a message

Services of the core#

MetricTypeLabelsServiceMeaning
stellar_reconciler_leadergaugereconciler1 on the leader, 0 on a standby instance
stellar_reconciler_registrations_totalcounteroutcome (accepted, rejected)reconcilerRegistrations of instances
stellar_connector_pendinggaugeconnector, kindreconcilerMessages of a connector consumer not delivered or not acknowledged
stellar_connector_lag_secondsgaugeconnector, kindreconcilerAge of its oldest message not acknowledged
stellar_connector_at_riskgaugeconnector, kindreconciler1 when the retention is about to remove messages it has not read
stellar_compute_messages_totalcounteroutcomecomputeDecoded telemetry messages handled
stellar_compute_samples_totalcountertargetcomputeSamples published
stellar_values_samples_totalcounteroutcomecompute (value table)Samples applied to the current value table
stellar_executor_telecommands_totalcounterstateexecutorTelecommand transitions
stellar_executor_runs_totalcounteroutcome (succeeded, failed)executorRuns finished
stellar_alarms_transitions_totalcounterstatealarmsAlarm transitions
stellar_scheduler_runs_totalcounteroutcome (fired, missed)schedulerScheduled runs fired or missed
stellar_transfers_bytes_totalcountertarget, outcome (asked, received, written)transfersBytes of file transfers

Drivers, transports and gateways (Rust SDK)#

MetricTypeLabelsMeaning
stellar_driver_frames_totalcountertarget, outcome (decoded, failed)Frames decoded
stellar_driver_telecommands_totalcountertarget, outcome (encoded, failed, expired)Telecommands handled
stellar_driver_cfdp_pdus_totalcountertarget, direction (sent, received)CFDP PDUs of the entity of the driver
stellar_transport_units_totalcountertarget, outcome (framed, failed, expired)Units to frame
stellar_transport_frames_totalcountertarget, outcome (unwrapped, failed)Frames received
stellar_transport_cop1_frames_totalcountertarget, link, kind (first, retransmitted)Frames sent by the COP-1 FOP
stellar_transport_cop1_alerts_totalcountertarget, linkCOP-1 alerts (lockout, transmission limit): the FOP stops
stellar_gateway_uplink_frames_totalcountertarget, outcome (sent, failed, dropped, expired)Frames to uplink
stellar_gateway_throughput_bpsgaugedirection (uplink, downlink), targetMean throughput over the last report window
stellar_gateway_bytes_totalcounterdirection, targetBytes carried since the start

Gateways also publish their throughput on NATS, stellar.metrics.<gateway>.<target>, every metrics_period (5 s): the transfer manager uses it, and the web console shows it (see Writing a Gateway).

Connector lag#

Every 15 s, the reconciler leader measures each consumer of each output connector: messages not delivered, messages not acknowledged, age of the oldest of them. It publishes them on stellar.metrics.connector.<name> and as the stellar_connector_* gauges; GET /v1/connectors gives the same on demand.

A consumer is at risk when the retention of its stream is about to remove messages it has not read: its oldest unread message has reached 80 % of the max_age of the stream, or the stream, bound by max_bytes, is 90 % full while its oldest message is still unread. The reconciler then logs a warning. Alert on it:

YAML
groups:
  - name: stellar-connectors
    rules:
      - alert: ConnectorAboutToLoseData
        expr: stellar_connector_at_risk == 1
        labels: {severity: critical}
        annotations:
          summary: "Connector {{ $labels.connector }} ({{ $labels.kind }}) is about to lose data"

Suggested alerts#

ConditionExpression
A component is downup == 0 on its scrape target
No reconciler leadermax(stellar_reconciler_leader) < 1
Registrations rejectedincrease(stellar_reconciler_registrations_total{outcome="rejected"}[10m]) > 0
Runs failingincrease(stellar_executor_runs_total{outcome="failed"}[1h]) > 0
Scheduled runs missedincrease(stellar_scheduler_runs_total{outcome="missed"}[1h]) > 0
Frames not decodedrate(stellar_driver_frames_total{outcome="failed"}[5m]) > 0
COP-1 alertsincrease(stellar_transport_cop1_alerts_total[10m]) > 0
Connector about to lose datastellar_connector_at_risk == 1

Readiness in NATS#

Health of the instances is also visible without Prometheus:

  • every instance publishes a heartbeat on stellar.ctl.hb.<kind>.<instance> (healthy, or degraded with a reason; a gateway adds whether its link is available). Its key in stellar_instances expires after three heartbeat periods: an expired instance is unbound and its targets are no longer ready;
  • the readiness of each target, with a reason per link, is in stellar_readiness;
  • GET /v1/topology, GET /v1/instances and GET /v1/instances/{kind}/{instance} show them, as do the Topology and Instances views of the web console.

NATS itself serves its monitoring on its HTTP port (8222 in the development setup): /healthz, /varz, /jsz.

Stellar Control · v0.1.0

↑↓ to moveEnter to open