The core of Stellar Control is a set of stateless services connected to NATS. Each one is a binary
of its own, and they all share the same command line, logging, metrics and shutdown behaviour.
For small cells, the stellar-mcs binary runs all of them in one process.
The binaries#
| Binary | Kind | Role | Streams it creates or updates |
|---|---|---|---|
stellar-reconciler | reconciler | Desired vs observed state, bindings, readiness, NATS credentials of the bound instances, connector consumers and their lag | TM_RAW, STREAMS |
stellar-compute | compute | Calibration, derived measures, current value table | TM_DECODED, PARAMS, TC_EVENTS |
stellar-executor | executor | Direct telecommands and runs of procedures | TC_COMMANDS, TC_EVENTS, PARAMS, RUNS |
stellar-api | api | HTTP and WebSocket API, web console, relay of the editor | TC_COMMANDS, TC_EVENTS, RUNS, ALARMS |
stellar-alarms | alarms | Limits and alarms, their lifecycle and reactions | PARAMS, TC_EVENTS, ALARMS |
stellar-scheduler | scheduler | Runs bound to passes, recurring rules | RUNS |
stellar-transfers | transfers | File transfers and quality of continuous streams | FILE_CHUNKS, TC_COMMANDS, PARAMS |
stellar-editor | editor | Web editor of the configuration repository (optional) | — |
stellar-mcs | all of the above | The services of the core in one process | all of the above |
Streams are created at start with an idempotent create or update, so the services can start in any order once NATS is up. The KV buckets and object stores are created by the service that first needs them. See NATS and JetStream for what each one holds.
Container images#
The binaries are also delivered as two container images, built from the repository by
deployment/docker/build.sh and tagged by commit:
| Image | Content |
|---|---|
stellar-mcs | Every binary above, stellar-lsp, the CLI stellar, stellar-simulator, the web console in /usr/share/stellar/web and the example configuration in /usr/share/stellar/examples/config; runs stellar-mcs by default |
stellar-admin | stellar-admin (provisioning of the cells), with the nomad and stellar command lines |
Both run as the user stellar (uid 10001), under tini, and work with a read-only root file
system. A container chooses its service by its command and reads the global configuration from
STELLAR_CONFIG:
docker run -e STELLAR_CONFIG=/etc/stellar/stellar.yaml \
-v ./stellar.yaml:/etc/stellar/stellar.yaml:ro stellar-mcs:<commit> stellar-apiNATS runs from its official image (nats:2.12). deployment/docker/smoke.sh <commit> checks a
pair of images: NATS and a cell in one process, the example configuration published, the API and
the web console answering.
Outside the core, the simulator (stellar-simulator, see
Simulated Targets) and the drivers, transports, gateways and
connectors built with the SDKs take their settings from environment
variables instead.
Before the first start#
The services work on a compiled configuration. Publish one before the first start, and after each change of the configuration repository:
stellar compile path/to/config --publish nats://nats.example.org:4222Until a configuration is published, the API answers 503 on what needs it and the web console
shows no configuration. See Compilation, Snapshots and Locks.
Command line#
stellar-reconciler --config /etc/stellar/stellar.yaml --instance rec-1| Option | Environment | Default | Meaning |
|---|---|---|---|
--config <path> | STELLAR_CONFIG | built-in defaults | Global configuration |
--instance <name> | STELLAR_INSTANCE | <kind>-<8 random characters> | Instance name, unique per kind; letters, digits, - and _ |
Every key of the configuration can also be overridden with STELLAR__<SECTION>__<KEY>
variables.
Lifecycle#
- The configuration is loaded and validated; the effective configuration is logged.
- The metrics endpoint starts on
observability.metrics_addr. - The NATS client connects to
nats.url, withnats.credentialsandnats.tls. When the server does not answer within 5 s, the refusal is logged with its reason and the connection is retried in the background: a service can start before NATS. - The service runs until SIGINT or SIGTERM, then stops cleanly, waiting at most 2 s for its pending NATS messages.
Exit codes#
| Code | Meaning |
|---|---|
| 0 | Stopped on a signal |
| 1 | The service failed (logged as component failed) |
| 2 | Invalid command line, configuration, log filter or instance name |
A supervisor (systemd, Nomad) should restart a service that exits with 1.
Logs and metrics#
- Logs go to standard output, one JSON object per line (
observability.log_format: json), or as text for development.observability.log_levelfilters them;RUST_LOGoverrides it. - Metrics are served in the Prometheus format on
http://<metrics_addr>/metrics, and/healthzanswersokwhile the process runs. Give each instance on a host its own port, for instanceSTELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9101.
See Observability for the metrics of each service.
Scaling and high availability#
The target is the unit of partitioning: adding or removing an instance moves whole targets, and the order of the telecommands of a target is kept. Each service follows one of two patterns.
| Service | Instances | How they share the work |
|---|---|---|
api | several, all active | Stateless; instances on one host share api.listen (SO_REUSEPORT) and the kernel spreads the connections |
executor | several, all active | Runs spread by their identifier; each target is held by a lease; direct telecommands taken one at a time from a shared consumer |
scheduler | several, all active | Each run is submitted once: Nats-Msg-Id = schedule + pass, deduplicated by JetStream |
reconciler | one leader, the others stand by | Lease reconciler in stellar_leader |
compute, alarms, transfers | one active, the others stand by | Leases compute, alarms, transfers in stellar_leader |
editor | one | Its drafts are working copies on the local disk |
Leader leases#
A leader holds a key of the stellar_leader bucket whose TTL is reconciler.leader_ttl (5 s),
created by a conditional write that only one instance wins. It renews the key every
reconciler.leader_renew (1 s), requiring the revision it knows: a failed renewal means the
lease is lost, and the instance stops writing at once and becomes a candidate again. Standby
instances watch the key and try again when it expires; the TTL is enforced by the NATS server,
not by the clocks of the instances.
The revision obtained when the lease is created is the epoch of the leader. The reconciler writes it in every value it writes, and readers ignore values of an older epoch: a frozen former leader cannot overwrite the new one.
The standby instances of the compute stage, alarm service and transfer manager give high
availability, not load sharing: one of them takes over within leader_ttl when the active one
stops or freezes, and starts again from the streams. Spreading them by target
(nats.partitions) will come with load tests.
Configuration changes#
A new snapshot applies to a target only at a safe point: a target with a run holding a lease, or
a pass in progress, keeps its revision until it ends. The reconciler watches the current key
of stellar_config and publishes the revision of each target in stellar_readiness.
All in one: stellar-mcs#
stellar-mcs runs the reconciler, the compute stage, the executor, the API, the alarm service,
the scheduler, the transfer manager and, when editor.repository is set, the editor, in one
process. Each service behaves as if it ran alone, with its own instance name; they share one
configuration, one log, one metrics endpoint and one NATS client. It takes about a quarter of
the memory of the separate services, which suits demonstrations, integration tests,
development and small cells.
stellar-mcs --config stellar.yaml
stellar-mcs --without editor,scheduler| Option | Environment | Default | Meaning |
|---|---|---|---|
--config <path> | STELLAR_CONFIG | built-in defaults | Global configuration |
--instance <suffix> | STELLAR_INSTANCE | random | Suffix of the instance names: api-<suffix>, executor-<suffix>… |
--without <kinds> | STELLAR_WITHOUT | none | Services not to run, comma-separated: reconciler, compute, executor, api, alarms, scheduler, transfers, editor |
An unknown name in --without exits with code 2. The first service to fail stops the process,
so that the supervisor restarts it whole.
A minimal deployment#
# NATS with JetStream
docker run -d --name nats -p 4222:4222 docker.io/library/nats:2.12.15-alpine -js
# The configuration, then the services
stellar compile examples/config --publish nats://localhost:4222
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9101 stellar-reconciler &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9102 stellar-compute &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9103 stellar-executor &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9104 stellar-api &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9105 stellar-alarms &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9106 stellar-scheduler &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9107 stellar-transfers &For a complete orchestrated stack, see Nomad; for one stack per tenant, see Cells.