Stellar Control is a set of stateless services around a NATS server with JetStream. There is no central data plane: each component takes its input and output subjects from the compiled configuration and from the bindings the reconciler gives it. This page describes the services, the state they share, the subjects they use, and how the system scales and survives failures.
The big picture#
flowchart TB
repo[(Git: configuration repository)] -- "stellar compile --publish" --> ir[(stellar_ir / stellar_config)]
subgraph core[Services of the core]
api[API]
rec[Reconciler]
exe[Executor]
comp[Compute stage + current value table]
alm[Alarm service]
sch[Scheduler]
xfer[Transfer manager]
end
subgraph edge[Components at the edge]
drv[Driver] --> tr[Transport] --> gw[Gateway]
conn[Output connectors]
end
ir --> rec
cli[CLI / web console] --> api
api -- "tc.submit, run.submit" --> exe
exe -- "tc.encode" --> drv
gw -- "tm.raw" --> drv
drv -- "tm.decoded" --> comp
comp -- "param" --> exe
comp -- "param" --> alm
rec -- "bindings, credentials" --> edge
gw <--> target((Target))
core -. "JetStream streams" .-> conn- A telecommand flows from the API or the executor to the driver of the link (ACK 1), through an optional transport, to the gateway (ACK 2); its verification (ACK 3) is a predicate on the measures that come back. See Telecommand Lifecycle.
- Telemetry flows from the gateway (raw frames) to the driver (raw values), to the compute stage (calibrated samples and derived measures), to the current value table, the executor, the alarm service and the output connectors. See Measures and Current Values.
State model#
The state is split into three kinds, each with one owner and one storage.
| State | Content | Source of truth | Storage |
|---|---|---|---|
| Desired | Catalogues, targets, links, environments, steps, procedures | People, through Git | Git, then a compiled snapshot in the stellar_ir object store, named by stellar_config |
| Observed | Live instances, versions, capabilities, health | The components themselves | NATS KV with a TTL (stellar_instances), fed by registrations and heartbeats |
| Operational | Telecommands in flight, runs, current values, leases, alarms, passes, transfers | The component that owns the partition | JetStream streams and KV buckets |
The reconciler replaces a control plane: it compares desired and observed state, binds the registered instances to the declared needs, and publishes the readiness of each target. There is no relational configuration database.
Every compiled snapshot has a hash. This hash is stamped on every message (Stellar-Config
header) and on every run: any value, event or report can be traced back to the exact
configuration that produced it.
Services of the core#
| Service | Binary | Role | Scaling |
|---|---|---|---|
| API | stellar-api | HTTP and WebSocket; resolves requests against the current snapshot; gives ULIDs to direct telecommands and runs; serves the web console | Stateless: any number of active instances, sharing a port on one host (SO_REUSEPORT) |
| Reconciler | stellar-reconciler | Desired vs observed state, bindings, readiness, NATS credentials of bound components, consumers of output connectors | One leader; others stand by |
| Executor | stellar-executor | Direct telecommands and runs of procedures; evaluates requires, expect, check; carries ACK 3 | Several active instances: runs by identifier, direct telecommands one at a time through a shared consumer |
| Compute stage | stellar-compute | Calibration, pure and temporal derived measures, current value table (stellar_values) | One active instance per consumer; others stand by |
| Alarm service | stellar-alarms | Limits and alarms, their lifecycle, notifications, automatic reactions | One active instance; others stand by |
| Scheduler | stellar-scheduler | Passes and runs scheduled on them | Several active instances (each run submitted once) |
| Transfer manager | stellar-transfers | File transfers by chunks or CFDP, quality of continuous streams | One active instance; others stand by |
stellar-mcs runs all of them in one process, for small cells; see
Running the Services.
Next to them:
- Drivers, transports and gateways are components outside the core, one process each, written with the SDKs or in any language speaking the NATS contract. See Drivers, Transports and Gateways.
- Output connectors subscribe read-only to the data of the MCS to store it or relay it. See Output Connectors.
- The web editor (
stellar-editor) is a service of its own, relayed by the API under/v1/editor. See Web Editor.
NATS subjects#
The control plane is addressed by instance, the data plane by logical role and by target. Every
variable part is a NATS token: ASCII letters, digits, - and _. A single-instance component
uses _ as instance.
| Subject | Content |
|---|---|
stellar.ctl.register | Registrations (request/reply) |
stellar.ctl.hb.<kind>.<instance> | Heartbeats |
stellar.ctl.rpc.<kind>.<instance>.<verb> | Control commands to one instance (status, bindings, credentials, …) |
stellar.tc.submit.<target> | Direct telecommands submitted by the API |
stellar.tc.encode.<codec>.<target> | Semantic telecommands to encode |
stellar.tc.wrap.<transport>.<target> | Units of a driver to frame, when the link has a transport |
stellar.tc.uplink.<gateway>.<target> | Frames to send |
stellar.tc.evt.<target>.<tc_id> | Telecommand lifecycle events |
stellar.tm.raw.<gateway>.<target> | Raw frames received |
stellar.tm.unit.<transport>.<target> | Units extracted by a transport, for the driver |
stellar.tm.decoded.<target> | Raw values decoded by a driver |
stellar.param.<target>.<component>.<instance>.<measure> | Calibrated samples |
stellar.alarm.evt.<target>.<component>.<instance>.<alarm> | Alarm transitions |
stellar.alarm.control.<target> | Commands of operators to the alarm service |
stellar.file.chunk.<target>.<file_id> | File chunks received |
stellar.file.decode.<target> | Requests to drivers to decode a verified file |
stellar.stream.<target>.<stream_id> | Segments of continuous streams |
stellar.sim.evt.<target> | Fault changes of a simulated target |
stellar.metrics.<gateway>.<target> | Throughput of a gateway |
stellar.metrics.connector.<name> | Lag of an output connector |
stellar.run.submit | Runs to start |
stellar.run.evt.<run_id> | Run log |
stellar.run.approval.<run_id> | Answers to questions and decisions |
stellar.run.control.<run_id> | Suspend, resume, abort |
The full reference is in NATS Contract Reference.
Streams, buckets and object stores#
Durable subjects are captured by JetStream streams; everything under stellar.ctl.> and
stellar.metrics.> stays on core NATS.
| Stream | Subjects | Default retention | Purpose |
|---|---|---|---|
TC_COMMANDS | stellar.tc.submit.>, encode.>, wrap.>, uplink.> | 7 days | Telecommands at each stage, deduplicated by identifier |
TC_EVENTS | stellar.tc.evt.> | 30 days | Lifecycle of every telecommand, replayable |
TM_RAW | stellar.tm.raw.> | 30 days | Opaque frames, re-decodable |
TM_DECODED | stellar.tm.decoded.> | 1 day | Raw values before calibration |
PARAMS | stellar.param.> | 30 days | Samples, real-time and deferred |
RUNS | stellar.run.submit, evt.>, approval.> | 365 days | Event sourcing of runs; source of reports |
ALARMS | stellar.alarm.evt.> | 90 days | Alarm transitions |
FILE_CHUNKS | stellar.file.chunk.> | 7 days | Chunks until stored |
STREAMS | stellar.stream.> | archive.streams | Segments of continuous streams |
State read by key lives in KV buckets (stellar_config, stellar_instances, stellar_leader,
stellar_bindings, stellar_readiness, stellar_runs, stellar_leases, stellar_values,
stellar_passes, stellar_pass_gateways, stellar_schedule, stellar_schedule_rules,
stellar_alarms, stellar_transfers, stellar_cop1, stellar_modes), large or content-addressed blobs in
object stores (stellar_ir, stellar_files, stellar_uploads, stellar_reports). Values are
JSON, so that components in any language can read them; the compiled IR is the exception
(postcard). See NATS and JetStream.
Long-term archiving is the job of output connectors: the API only reads NATS and never queries the database of a connector.
The reconciler#
The reconciler is the only component that knows both what should run and what does run.
- Desired state. Its leader watches the
currentkey ofstellar_config, loads the snapshot fromstellar_ir, and gives each target a revision. - Observed state. Every component registers on
stellar.ctl.registerand sends heartbeats; the reconciler records them instellar_instanceswith a TTL of three heartbeat periods. An expired key means the instance is gone: its links are unbound, and the targets concerned are no longer ready. - Bindings. For each link of each target, it binds a registered, healthy and compatible
driver, transport and gateway, and delivers the bindings to the instances with the
bindingsverb (stellar_bindings). See Links and Bindings. A binding also carries the current mode of its target (stellar_modes) and its parameters in that mode; the reconciler delivers the bindings again at each change of mode. See Modes and Parameters. - Readiness. A target is ready when all its links are bound; otherwise each link gets a
reason (
stellar_readiness). - Credentials. It issues to each bound driver, transport, gateway and connector a short-lived NATS JWT carrying exactly its permissions. See NATS Accounts and Credentials.
Safe points#
A new configuration applies at safe points only. A target held by a run (a lease in
stellar_leases) keeps its revision until the lease is released; the others move to the new
snapshot at once. The components bound to a target stamp its revision on their messages.
Leadership and failover#
The reconciler is not on the critical path of a pass: bindings, readiness and permissions are persisted in NATS KV. While it fails over, runs, telecommands and telemetry continue; only new bindings wait for the new leader.
- Election by KV lease. The key
reconcilerof the bucketstellar_leader, whose TTL isreconciler.leader_ttl(5 s), is created by a conditional write: one instance succeeds. - Renewal. The leader renews the key every
reconciler.leader_renew(1 s), requiring the revision it knows. A failed renewal means the lease is lost: it stops writing and becomes a candidate again. - Detection. Standby instances watch the key and retry the creation when it expires. The TTL is enforced by the NATS server, independent of the clocks of the instances.
- Epoch. The revision of the creation of the lease is its epoch. Every value the reconciler writes carries it, and readers ignore values of an older epoch: a frozen former leader cannot do harm.
- Takeover without handover. The new leader recomputes everything from the snapshot and the instances, and rewrites the values of older epochs. Its writes are idempotent and respect the safe points.
The same lease mechanism, with the component name as key, gives high availability to the
compute stage, the alarm service and the transfer manager: one instance consumes, the others
take over within leader_ttl if it stops or freezes.
Scaling and ordering#
The target is the unit of partitioning: moving work between instances moves whole targets, and the order of the telecommands of a target is preserved.
- Drivers and gateways are partitioned by their bindings. A binding is kept while its instance fits; a new link goes to the eligible instance serving the fewest links. An instance joining only takes new links; an instance leaving hands its links, whole, to the others.
- One message at a time per subject. Each encoding and uplink subject has one durable
consumer on
TC_COMMANDS(encode_<codec>_<target>,uplink_<gateway>_<target>) delivering one message at a time: a telecommand is encoded and sent once, in order, even while a link moves. A message redelivered after a failover is deduplicated by itsNats-Msg-Id. - The executor takes direct telecommands one at a time from a consumer shared by its instances, and publishes each telecommand to encode before acknowledging its submission, which keeps the order of submission per target. Runs are spread by run identifier.
- The compute stage and the current value table run one active instance for now.
nats.partitions(64) fixes the number of partitions of their future partitioning; it must exceed the maximum number of instances of a component.
Configuration and identity of components#
Operating parameters (NATS URL, timeouts, listen addresses, identity provider…) come from the global configuration file, distinct from the desired state: it describes the deployment, not the mission. See Global Configuration.
Each component starts with a bootstrap identity limited to registration and heartbeat; once bound, it receives credentials for exactly its data. People go through the API, with an OIDC token or a declared identity where the environment allows it. See Identity and Roles.
Repository layout#
The source repository has three parts that depend on each other in one direction only, plus
examples built on the SDK alone; scripts/check-layers.py checks it in CI.
sdk/ the contract (stellar-contract), the common types (stellar-common),
the Rust SDK (stellar-sdk) and the Python SDK (stellar-mcs)
core/ the configuration model and its compiler, the services, the web application
tooling/ the CLI, the language server, the web editor, the VS Code extension,
the simulator, the all-in-one binary, stellar-admin
examples/ drivers, transports, gateways and connectors on the SDK; the example
configuration repository (examples/config); fictitious ICDs (examples/icd)sdk ← core ← tooling: of sdk/, the core only uses the contract and the common types,
never the SDK. A driver thus embeds only the contract and the SDK, and the core ships without
them.