Stellar ControlMission control · by Stellar Systems v0.1.0

Operations

Alarm Handling

The lifecycle of alarms, acknowledgement and shelving by operators, automatic reactions and notifications.

The alarm service (stellar-alarms) watches the limits and the conditions the catalogue declares, manages the lifecycle of each alarm, and publishes every transition on NATS, where the CLI, the web console and external systems follow them. It keeps no local state: the current state of every alarm lives in the stellar_alarms KV bucket, and its transitions in the ALARMS stream.

Alarms are declared in the catalogue — limits of a measure, and named alarms with a condition, a severity and an optional reaction: see Alarms. This page covers what happens at run time.

Sources and naming#

SourceName of the alarmSeverity
Limits of a measure or a derived measureThe measure: tcu[TCU1].anode_voltagecritical outside the hard limits, else warning outside the soft limits
Named alarm (when)Its name: tcu[TCU2].cathode_overheatIts severity

On a multi-instance component, an alarm exists per instance. A component without instances writes its alarms component.alarm (eps.undervoltage). Measures, derived measures, telecommands and alarms of a component share one namespace, so the alarm of a limit and a named alarm never clash.

Evaluation#

The service reads the real-time samples of PARAMS — measures and derived measures published by the compute stage — and evaluates the conditions the way the compute stage does, with the same history and deadlines. cathode_overheat: stable(cathode_temperature > 1200 degC, 5 s) is raised 5 seconds after the threshold is crossed, even without a new sample.

  • Real time only. Deferred samples (telemetry stored on board and dumped later) never raise an alarm: they are archived and analysed afterwards.
  • Unknown conditions change nothing. A condition on a measure without a value, or on an empty window (during a loss of signal), leaves the alarm as it is.

Lifecycle#

stateDiagram-v2
    [*] --> NORMAL
    NORMAL --> ACTIVE_UNACK: condition true
    ACTIVE_UNACK --> ACTIVE_ACK: acknowledged
    ACTIVE_UNACK --> CLEARED_UNACK: back to normal
    CLEARED_UNACK --> NORMAL: acknowledged
    CLEARED_UNACK --> ACTIVE_UNACK: condition true again
    ACTIVE_ACK --> NORMAL: back to normal
StateMeaning
NORMALCondition false, nothing to acknowledge
ACTIVE_UNACKCondition true, not acknowledged
ACTIVE_ACKCondition true, acknowledged
CLEARED_UNACKCondition false again, but its activation was not acknowledged
  • An activation must be acknowledged, even once it has cleared.
  • While active, a more severe severity (warning to critical) asks for a new acknowledgement; a less severe one keeps the acknowledgement.
  • Every transition increments seq and becomes an event, published with Nats-Msg-Id = <target>.<component>.<instance>.<alarm>-<seq>.

The current state of an alarm, in stellar_alarms under <target>.<component>.<instance>.<alarm>:

FieldMeaning
state, previousState, and state before the last transition
severitySeverity of the activation, kept while cleared and unacknowledged
active_sinceStart of the activation
changed_atLast transition
valueValue of the measure at the last transition, for the alarm of its limits
seqNumber of the last transition, from 1
causecondition, acknowledged, shelved or unshelved
byOperator of the last transition, when not the condition
shelved_untilEnd of the shelving in progress

Recovery without loss nor duplicate#

A transition is first written in the KV bucket by a conditional write, then published. When the service meets a target for the first time (at start-up, or after a failover), it reads PARAMS and the telecommand events again, like the compute stage, and republishes any transition of the KV whose last event in the stream carries a lower seq. It then evaluates the conditions and publishes only the differences with the recorded state.

A standby instance of the service waits on a lease in stellar_leader; it takes over within reconciler.leader_ttl if the active one stops.

Operators: acknowledge, shelve, unshelve#

Shell
stellar alarms                                           # alarms not NORMAL, or shelved
stellar alarms --target sat1-fm
stellar alarms --all                                     # every alarm recorded
stellar alarms ack sat1-fm 'tcu[TCU1].anode_voltage'
stellar alarms shelve sat1-fm 'tcu[TCU1].anode_voltage' --for '30 min'
stellar alarms unshelve sat1-fm 'tcu[TCU1].anode_voltage'
stellar alarms watch                                     # every transition from now
stellar alarms --target sat1-fm watch
text
sat1-fm  tcu[TCU1].anode_voltage  ACTIVE_UNACK critical (310.2)
sat1-fm  tcu[TCU2].cathode_overheat  ACTIVE_ACK critical, shelved until 2026-10-02T11:00:00Z

The options of stellar alarms (--target, --all, and the API options --api, --as, --role, --token) come before its subcommand. Every command is traced with the identity of the operator: a command without identity is refused (400 api::anonymous).

Acknowledge#

ACTIVE_UNACK becomes ACTIVE_ACK, and CLEARED_UNACK becomes NORMAL. Acknowledging an alarm that is NORMAL or already ACTIVE_ACK is refused ("nothing to acknowledge").

Shelve#

Shelving an alarm silences its reactions for a bounded duration, at most alarms.max_shelve (8 h by default), written as a duration: 30 min, 2h. A new shelving replaces the one in progress.

While shelved, the alarm goes on with its lifecycle, and its transitions are still published, with shelved_until: operators still see it. It only triggers no reaction. At the end of the shelving, the service unshelves it by itself (cause unshelved, without operator); it reads the shelvings in progress again when it starts.

Unshelve#

Ends a shelving before its term. Unshelving an alarm that is not shelved is refused.

Through the API#

RequestEffect
GET /v1/alarms (?target=, ?all=true)Alarms not NORMAL or shelved; every alarm with all
POST /v1/alarms/{target}/{alarm}/acknowledgeAcknowledge
POST /v1/alarms/{target}/{alarm}/shelve with {"for": "30 min"}Shelve
POST /v1/alarms/{target}/{alarm}/unshelveUnshelve
GET /v1/alarms/watch (?target=)Transitions from now, over WebSocket

The alarm is written in the path as operators write it: tcu[TCU1].anode_voltage. The API sends the command to the alarm service, the only writer of the alarms of its targets, by request-reply on stellar.alarm.control.<target>.

StatusCodeCause
202—Applied
400api::invalid-body, api::anonymousMalformed shelving, or no identity
404api::invalid-target, api::invalid-alarm, api::unknown-commandTarget or alarm not written as expected, unknown action
409alarm::refusedNothing to acknowledge, alarm not shelved, never recorded, or duration out of bounds
503alarm::no-serviceNo alarm service is running
504alarm::timeoutThe alarm service did not answer

Notifications#

Every transition is published on stellar.alarm.evt.<target>.<component>.<instance>.<alarm> (_ for a component without instances), in the replayable ALARMS stream (90 days by default):

JSON
{
  "target": "sat1-fm", "component": "tcu", "instance": "TCU1", "alarm": "anode_voltage",
  "previous": "NORMAL", "state": "ACTIVE_UNACK", "severity": "critical", "value": 310.2,
  "at": "2026-10-02T10:18:03Z", "seq": 7, "cause": "condition"
}

Ways to follow them:

  • stellar alarms watch, or GET /v1/alarms/watch over WebSocket;
  • the alarms view of the web console;
  • NATS directly, with follow-up credentials that may subscribe to stellar.alarm.evt.> (stellar auth nats-creds, see NATS Accounts and Credentials);
  • an output connector taking alarms, to store them or relay them (to a paging system, for instance).

Automatic reactions#

A named alarm can declare on_raise: run "<procedure>". The service then launches the procedure on the target of the alarm, once per activation.

YAML
alarms:
  cathode_overheat:
    severity: critical
    when: stable(cathode_temperature > 1200 degC, 5 s)
    on_raise: run "Safe TCU"
  • Trigger. Only a new activation triggers the reaction: a transition to ACTIVE_UNACK from NORMAL or CLEARED_UNACK, caused by the condition. A more severe severity does not, a shelved alarm does not, and the end of a shelving does not.
  • Binding. The compiler checks the procedure: found in the libraries compiled against the platform, exactly one role of the platform (bound to the target of the alarm) and, for a component with instances, at most one input of the instance enum (bound to the instance of the alarm). It sends no hazardous telecommand, directly or through its calls.
  • Environment. That of the run in progress on the target; without a run, the last environment declared for the target, its current phase. The request is resolved like any launch: a refusal (procedure not allowed in the environment, or in in_orbit a target not ready) is logged and nothing is launched.
sequenceDiagram
    participant A as Alarm service
    participant E as Executor
    participant R as Run holding the target
    A->>A: NORMAL → ACTIVE_UNACK (condition)
    A->>E: submit the reaction run (reaction, Nats-Msg-Id = <key>-<seq>)
    E->>E: take the lease of the target from the run
    R->>R: next renewal: suspended (by = the reaction)
    E->>A: the reaction ends: passed, or failed → alarm <alarm>-reaction
  • The reaction passes first. A reaction is a safing: the compiler guarantees it sends nothing hazardous. The service submits it at once with by = alarm <key>, reaction = alarm <key>: reaction "<procedure>" and Nats-Msg-Id = <key>-<seq>, so an activation launches one reaction only. The executor gives it the lease of the target whoever holds it: a run in progress, or waiting for a decision, an answer or the confirmation of a hazardous telecommand. It waits only for the reaction of another alarm, 60 seconds at most.
  • The run it takes the target from is suspended within a second, whatever it waits for: suspended by alarm sim-1.tcu.TCU2.cathode_overheat: reaction "Safe TCU" (run …), which took sim-1``. The instruction under way may end (a wait, a question), but no telecommand leaves any more: a telecommand about to be sent waits there. What the run was waiting for is answered after the reaction, on the state it left.
  • The suspended run stays suspended until an operator decides: resume it (replay, skip, or fail, so that its if failed blocks run), once the reaction is over and its leases free, or abort it. The run ends the instruction under way before it can be resumed. See Questions, Decisions and Control.
  • A reaction that fails raises an alarm. When the reaction cannot be launched, or its run does not pass, the service raises <alarm>-reaction on the same component and instance (sim-1 · tcu[TCU2] · cathode_overheat-reaction), critical, with the reason as its value. It is cleared at once and stays to be acknowledged (CLEARED_UNACK).

A dry run plays the alarms and their reactions too, with faults scheduled by its request to raise them: see Dry runs.

Metrics#

The service counts the transitions in the Prometheus counter stellar_alarms_transitions_total. See Observability.

See also#

Stellar Control · v0.1.0

↑↓ to moveEnter to open