Stellar ControlMission control · by Stellar Systems v0.1.0

Operations

Leases and Crash Recovery

How runs hold their targets with leases, and how a run survives the restart of its executor.

Two mechanisms make runs safe to operate on shared targets and robust to failures: leases on targets, so that two runs never command the same target at once, and event sourcing with executor leases, so that a run resumes after its executor stops, on the same instance or another one, without sending a telecommand twice.

Leases on targets#

A run holds a lease on each of its targets while it runs, in the KV bucket stellar_leases, one key per target:

JSON
{"runs": ["01K6A3D3Y4Q8K1P9W2N6R5T7XZ"], "shared": false, "renewed_at": "2026-10-02T10:15:03Z"}
  • All at start. The executor takes every lease when the run starts, in the sorted order of the target names, with conditional writes.
  • Never waiting. A target held by another run fails the start of the run at once: the leases already taken are released together, and the reason names the holder (`flatsat-1` is held by run 01K6…). Since nothing waits, there is no deadlock; only the reaction of an alarm waits for the reaction of another one, 60 s at most.
  • Renewed. The run renews its leases every second. A lease not renewed for 5 s is free: a crashed run does not hold its targets forever.
  • Released at the end of the run and while it is suspended; taken again on resumption.
  • Taken by a reaction. The reaction of an alarm (on_raise) takes the lease from any ordinary run, whatever that run waits for. The run notes it at its next renewal (a mark left in the bucket for it, under preempted.<run_id>, even when the reaction is already over): it is suspended at once, sends nothing more, and takes its leases again when an operator resumes it. The holder of a lease taken by a reaction is named after it: `sim-1` is held by the alarm …: reaction "Safe TCU" (run …). See Automatic reactions.
  • Shared. A run that sends only telecommands declared changes_state: false takes shared leases: several such runs hold a target together. A run that may change the on-board state holds its targets alone.

Direct telecommands on a held target#

A direct telecommand (stellar send, POST /v1/tc) on a target held by a run is refused: 409 tc::target-held by the API, REJECTED by the executor. The one exception is a telecommand with changes_state: false on a target under a shared lease.

The telecommands of a file transfer started by a run carry the header Stellar-For-Run and pass the lease of that run; other direct telecommands remain refused.

GET /v1/topology and the web console show the lease of each target, with its runs.

Event sourcing#

Every transition of a run is an event of its log (stellar.run.evt.<run_id>, stream RUNS), published with the message id <run_id>-<seq>. Every telecommand identifier is logged (telecommand_sent) before the telecommand leaves, and the telecommand is published with that identifier as message id: JetStream deduplicates a telecommand sent again with the same identifier.

Executor leases and takeover#

Each run has a record in the KV bucket stellar_runs:

FieldMeaning
irHash of the resolved run
byWho launched it
ownerThe executor instance holding it
renewed_atLast renewal
finishedWhether it is over

The owner renews the record every second, with a conditional write; if the write fails, it lost the run and stops interpreting it. A record not renewed for 5 s is taken over by an executor instance, possibly the same one restarted:

  1. It logs taken_over with its instance name.
  2. It reads the log back and replays the run without publishing again the events already logged: steps finished keep their outcome, answers already given are reused.
  3. It resumes where the log ends.
flowchart TD
    A[Executor stops] --> B[Record not renewed for 5 s]
    B --> C[Another instance takes the run: taken_over]
    C --> D[Replay the log]
    D --> E{Step interrupted?}
    E -->|no| F[Go on]
    E -->|yes, only changes_state: false| G[Replay the step: same identifiers, deduplicated]
    E -->|yes, after a state-changing telecommand| H[decision_required: replay, skip, fail or abort]

An interrupted step#

  • If the step had sent only telecommands with changes_state: false, it is replayed: its telecommands go out again with their identifiers, deduplicated by JetStream, and their follow-up restarts from the events and samples already archived.
  • If it had sent a telecommand that changes the on-board state, the run waits for a decision (decision_required: "interrupted by a restart after a telecommand changing the on-board state"). replay starts a new attempt with new identifiers, skip counts the step as passed, abort ends the run. See Questions, Decisions and Control.
  • A run suspended when its executor stopped stays suspended.

Waits are not replayed: a wait or a retry interval already elapsed in the log does not wait again.

A loss of link during a run is not a special case: the telecommand fails (the gateway refuses the frame, SEND_FAILED) or its verification times out (VERIFY_TIMEOUT), and the step fails, retried when it is eligible. See Retries.

Clock#

The whole chain works in one time scale, UTC by default (time.scale). within windows and freshness are measured on the ground reception time of the samples, never on the on-board time. See Time and Freshness.

Stellar Control · v0.1.0

↑↓ to moveEnter to open