The alarm service (stellar-alarms) watches the limits and the conditions the catalogue declares,
manages the lifecycle of each alarm, and publishes every transition on NATS, where the CLI, the web
console and external systems follow them. It keeps no local state: the current state of every
alarm lives in the stellar_alarms KV bucket, and its transitions in the ALARMS stream.
Alarms are declared in the catalogue — limits of a measure, and named alarms with a condition, a severity and an optional reaction: see Alarms. This page covers what happens at run time.
Sources and naming#
| Source | Name of the alarm | Severity |
|---|---|---|
| Limits of a measure or a derived measure | The measure: tcu[TCU1].anode_voltage | critical outside the hard limits, else warning outside the soft limits |
Named alarm (when) | Its name: tcu[TCU2].cathode_overheat | Its severity |
On a multi-instance component, an alarm exists per instance. A component without instances writes
its alarms component.alarm (eps.undervoltage). Measures, derived measures, telecommands and
alarms of a component share one namespace, so the alarm of a limit and a named alarm never clash.
Evaluation#
The service reads the real-time samples of PARAMS — measures and derived measures published by
the compute stage — and evaluates the conditions the way the compute stage does, with the same
history and deadlines. cathode_overheat: stable(cathode_temperature > 1200 degC, 5 s) is raised
5 seconds after the threshold is crossed, even without a new sample.
- Real time only. Deferred samples (telemetry stored on board and dumped later) never raise an alarm: they are archived and analysed afterwards.
- Unknown conditions change nothing. A condition on a measure without a value, or on an empty window (during a loss of signal), leaves the alarm as it is.
Lifecycle#
stateDiagram-v2
[*] --> NORMAL
NORMAL --> ACTIVE_UNACK: condition true
ACTIVE_UNACK --> ACTIVE_ACK: acknowledged
ACTIVE_UNACK --> CLEARED_UNACK: back to normal
CLEARED_UNACK --> NORMAL: acknowledged
CLEARED_UNACK --> ACTIVE_UNACK: condition true again
ACTIVE_ACK --> NORMAL: back to normal| State | Meaning |
|---|---|
NORMAL | Condition false, nothing to acknowledge |
ACTIVE_UNACK | Condition true, not acknowledged |
ACTIVE_ACK | Condition true, acknowledged |
CLEARED_UNACK | Condition false again, but its activation was not acknowledged |
- An activation must be acknowledged, even once it has cleared.
- While active, a more severe severity (warning to critical) asks for a new acknowledgement; a less severe one keeps the acknowledgement.
- Every transition increments
seqand becomes an event, published withNats-Msg-Id=<target>.<component>.<instance>.<alarm>-<seq>.
The current state of an alarm, in stellar_alarms under <target>.<component>.<instance>.<alarm>:
| Field | Meaning |
|---|---|
state, previous | State, and state before the last transition |
severity | Severity of the activation, kept while cleared and unacknowledged |
active_since | Start of the activation |
changed_at | Last transition |
value | Value of the measure at the last transition, for the alarm of its limits |
seq | Number of the last transition, from 1 |
cause | condition, acknowledged, shelved or unshelved |
by | Operator of the last transition, when not the condition |
shelved_until | End of the shelving in progress |
Recovery without loss nor duplicate#
A transition is first written in the KV bucket by a conditional write, then published. When the
service meets a target for the first time (at start-up, or after a failover), it reads PARAMS
and the telecommand events again, like the compute stage, and republishes any transition of the
KV whose last event in the stream carries a lower seq. It then evaluates the conditions and
publishes only the differences with the recorded state.
A standby instance of the service waits on a lease in stellar_leader; it takes over within
reconciler.leader_ttl if the active one stops.
Operators: acknowledge, shelve, unshelve#
stellar alarms # alarms not NORMAL, or shelved
stellar alarms --target sat1-fm
stellar alarms --all # every alarm recorded
stellar alarms ack sat1-fm 'tcu[TCU1].anode_voltage'
stellar alarms shelve sat1-fm 'tcu[TCU1].anode_voltage' --for '30 min'
stellar alarms unshelve sat1-fm 'tcu[TCU1].anode_voltage'
stellar alarms watch # every transition from now
stellar alarms --target sat1-fm watchsat1-fm tcu[TCU1].anode_voltage ACTIVE_UNACK critical (310.2)
sat1-fm tcu[TCU2].cathode_overheat ACTIVE_ACK critical, shelved until 2026-10-02T11:00:00ZThe options of stellar alarms (--target, --all, and the API options --api, --as,
--role, --token) come before its subcommand. Every command is traced with the identity of the
operator: a command without identity is refused (400 api::anonymous).
Acknowledge#
ACTIVE_UNACK becomes ACTIVE_ACK, and CLEARED_UNACK becomes NORMAL. Acknowledging an alarm
that is NORMAL or already ACTIVE_ACK is refused ("nothing to acknowledge").
Shelve#
Shelving an alarm silences its reactions for a bounded duration, at most alarms.max_shelve
(8 h by default), written as a duration: 30 min, 2h. A new shelving replaces the one in
progress.
While shelved, the alarm goes on with its lifecycle, and its transitions are still published,
with shelved_until: operators still see it. It only triggers no reaction. At the end of the
shelving, the service unshelves it by itself (cause unshelved, without operator); it reads the
shelvings in progress again when it starts.
Unshelve#
Ends a shelving before its term. Unshelving an alarm that is not shelved is refused.
Through the API#
| Request | Effect |
|---|---|
GET /v1/alarms (?target=, ?all=true) | Alarms not NORMAL or shelved; every alarm with all |
POST /v1/alarms/{target}/{alarm}/acknowledge | Acknowledge |
POST /v1/alarms/{target}/{alarm}/shelve with {"for": "30 min"} | Shelve |
POST /v1/alarms/{target}/{alarm}/unshelve | Unshelve |
GET /v1/alarms/watch (?target=) | Transitions from now, over WebSocket |
The alarm is written in the path as operators write it: tcu[TCU1].anode_voltage. The API sends
the command to the alarm service, the only writer of the alarms of its targets, by request-reply
on stellar.alarm.control.<target>.
| Status | Code | Cause |
|---|---|---|
| 202 | — | Applied |
| 400 | api::invalid-body, api::anonymous | Malformed shelving, or no identity |
| 404 | api::invalid-target, api::invalid-alarm, api::unknown-command | Target or alarm not written as expected, unknown action |
| 409 | alarm::refused | Nothing to acknowledge, alarm not shelved, never recorded, or duration out of bounds |
| 503 | alarm::no-service | No alarm service is running |
| 504 | alarm::timeout | The alarm service did not answer |
Notifications#
Every transition is published on stellar.alarm.evt.<target>.<component>.<instance>.<alarm>
(_ for a component without instances), in the replayable ALARMS stream (90 days by default):
{
"target": "sat1-fm", "component": "tcu", "instance": "TCU1", "alarm": "anode_voltage",
"previous": "NORMAL", "state": "ACTIVE_UNACK", "severity": "critical", "value": 310.2,
"at": "2026-10-02T10:18:03Z", "seq": 7, "cause": "condition"
}Ways to follow them:
stellar alarms watch, orGET /v1/alarms/watchover WebSocket;- the alarms view of the web console;
- NATS directly, with follow-up credentials that may subscribe to
stellar.alarm.evt.>(stellar auth nats-creds, see NATS Accounts and Credentials); - an output connector taking
alarms, to store them or relay them (to a paging system, for instance).
Automatic reactions#
A named alarm can declare on_raise: run "<procedure>". The service then launches the procedure
on the target of the alarm, once per activation.
alarms:
cathode_overheat:
severity: critical
when: stable(cathode_temperature > 1200 degC, 5 s)
on_raise: run "Safe TCU"- Trigger. Only a new activation triggers the reaction: a transition to
ACTIVE_UNACKfromNORMALorCLEARED_UNACK, caused by the condition. A more severe severity does not, a shelved alarm does not, and the end of a shelving does not. - Binding. The compiler checks the procedure: found in the libraries compiled against the platform, exactly one role of the platform (bound to the target of the alarm) and, for a component with instances, at most one input of the instance enum (bound to the instance of the alarm). It sends no hazardous telecommand, directly or through its calls.
- Environment. That of the run in progress on the target; without a run, the last environment
declared for the target, its current phase. The request is resolved like any launch: a refusal
(procedure not allowed in the environment, or in
in_orbita target not ready) is logged and nothing is launched.
sequenceDiagram
participant A as Alarm service
participant E as Executor
participant R as Run holding the target
A->>A: NORMAL → ACTIVE_UNACK (condition)
A->>E: submit the reaction run (reaction, Nats-Msg-Id = <key>-<seq>)
E->>E: take the lease of the target from the run
R->>R: next renewal: suspended (by = the reaction)
E->>A: the reaction ends: passed, or failed → alarm <alarm>-reaction- The reaction passes first. A reaction is a safing: the compiler guarantees it sends nothing
hazardous. The service submits it at once with
by=alarm <key>,reaction=alarm <key>: reaction "<procedure>"andNats-Msg-Id=<key>-<seq>, so an activation launches one reaction only. The executor gives it the lease of the target whoever holds it: a run in progress, or waiting for a decision, an answer or the confirmation of a hazardous telecommand. It waits only for the reaction of another alarm, 60 seconds at most. - The run it takes the target from is suspended within a second, whatever it waits for:
suspended by alarm sim-1.tcu.TCU2.cathode_overheat: reaction "Safe TCU" (run …), which tooksim-1``. The instruction under way may end (a wait, a question), but no telecommand leaves any more: a telecommand about to be sent waits there. What the run was waiting for is answered after the reaction, on the state it left. - The suspended run stays suspended until an operator decides: resume it (replay, skip, or
fail, so that its
if failedblocks run), once the reaction is over and its leases free, or abort it. The run ends the instruction under way before it can be resumed. See Questions, Decisions and Control. - A reaction that fails raises an alarm. When the reaction cannot be launched, or its run does
not pass, the service raises
<alarm>-reactionon the same component and instance (sim-1 · tcu[TCU2] · cathode_overheat-reaction), critical, with the reason as its value. It is cleared at once and stays to be acknowledged (CLEARED_UNACK).
A dry run plays the alarms and their reactions too, with faults scheduled by its request to raise them: see Dry runs.
Metrics#
The service counts the transitions in the Prometheus counter stellar_alarms_transitions_total.
See Observability.
See also#
- Alarms: declaring limits, named alarms and reactions.
- API reference: alarms.