Two mechanisms make runs safe to operate on shared targets and robust to failures: leases on targets, so that two runs never command the same target at once, and event sourcing with executor leases, so that a run resumes after its executor stops, on the same instance or another one, without sending a telecommand twice.
Leases on targets#
A run holds a lease on each of its targets while it runs, in the KV bucket stellar_leases, one
key per target:
{"runs": ["01K6A3D3Y4Q8K1P9W2N6R5T7XZ"], "shared": false, "renewed_at": "2026-10-02T10:15:03Z"}- All at start. The executor takes every lease when the run starts, in the sorted order of the target names, with conditional writes.
- Never waiting. A target held by another run fails the start of the run at once: the leases
already taken are released together, and the reason names the holder
(
`flatsat-1` is held by run 01K6…). Since nothing waits, there is no deadlock; only the reaction of an alarm waits for the reaction of another one, 60 s at most. - Renewed. The run renews its leases every second. A lease not renewed for 5 s is free: a crashed run does not hold its targets forever.
- Released at the end of the run and while it is suspended; taken again on resumption.
- Taken by a reaction. The reaction of an alarm (
on_raise) takes the lease from any ordinary run, whatever that run waits for. The run notes it at its next renewal (a mark left in the bucket for it, underpreempted.<run_id>, even when the reaction is already over): it is suspended at once, sends nothing more, and takes its leases again when an operator resumes it. The holder of a lease taken by a reaction is named after it:`sim-1` is held by the alarm …: reaction "Safe TCU" (run …). See Automatic reactions. - Shared. A run that sends only telecommands declared
changes_state: falsetakes shared leases: several such runs hold a target together. A run that may change the on-board state holds its targets alone.
Direct telecommands on a held target#
A direct telecommand (stellar send, POST /v1/tc) on a target held by a run is refused:
409 tc::target-held by the API, REJECTED by the executor. The one exception is a telecommand
with changes_state: false on a target under a shared lease.
The telecommands of a file transfer started by a run carry the header Stellar-For-Run and pass
the lease of that run; other direct telecommands remain refused.
GET /v1/topology and the web console show the lease of each target, with its runs.
Event sourcing#
Every transition of a run is an event of its log (stellar.run.evt.<run_id>, stream RUNS),
published with the message id <run_id>-<seq>. Every telecommand identifier is logged
(telecommand_sent) before the telecommand leaves, and the telecommand is published with
that identifier as message id: JetStream deduplicates a telecommand sent again with the same
identifier.
Executor leases and takeover#
Each run has a record in the KV bucket stellar_runs:
| Field | Meaning |
|---|---|
ir | Hash of the resolved run |
by | Who launched it |
owner | The executor instance holding it |
renewed_at | Last renewal |
finished | Whether it is over |
The owner renews the record every second, with a conditional write; if the write fails, it lost the run and stops interpreting it. A record not renewed for 5 s is taken over by an executor instance, possibly the same one restarted:
- It logs
taken_overwith its instance name. - It reads the log back and replays the run without publishing again the events already logged: steps finished keep their outcome, answers already given are reused.
- It resumes where the log ends.
flowchart TD
A[Executor stops] --> B[Record not renewed for 5 s]
B --> C[Another instance takes the run: taken_over]
C --> D[Replay the log]
D --> E{Step interrupted?}
E -->|no| F[Go on]
E -->|yes, only changes_state: false| G[Replay the step: same identifiers, deduplicated]
E -->|yes, after a state-changing telecommand| H[decision_required: replay, skip, fail or abort]An interrupted step#
- If the step had sent only telecommands with
changes_state: false, it is replayed: its telecommands go out again with their identifiers, deduplicated by JetStream, and their follow-up restarts from the events and samples already archived. - If it had sent a telecommand that changes the on-board state, the run waits for a decision
(
decision_required: "interrupted by a restart after a telecommand changing the on-board state").replaystarts a new attempt with new identifiers,skipcounts the step as passed,abortends the run. See Questions, Decisions and Control. - A run suspended when its executor stopped stays suspended.
Waits are not replayed: a wait or a retry interval already elapsed in the log does not wait
again.
Losing the link#
A loss of link during a run is not a special case: the telecommand fails (the gateway refuses the
frame, SEND_FAILED) or its verification times out (VERIFY_TIMEOUT), and the step fails,
retried when it is eligible. See Retries.
Clock#
The whole chain works in one time scale, UTC by default (time.scale). within windows and
freshness are measured on the ground reception time of the samples, never on the on-board time.
See Time and Freshness.