Stellar ControlMission control · by Stellar Systems v0.1.0

Deployment

Running the Services

Command line, lifecycle, logs, metrics, high availability and scaling of the services of the core, and the all-in-one binary.

The core of Stellar Control is a set of stateless services connected to NATS. Each one is a binary of its own, and they all share the same command line, logging, metrics and shutdown behaviour. For small cells, the stellar-mcs binary runs all of them in one process.

The binaries#

BinaryKindRoleStreams it creates or updates
stellar-reconcilerreconcilerDesired vs observed state, bindings, readiness, NATS credentials of the bound instances, connector consumers and their lagTM_RAW, STREAMS
stellar-computecomputeCalibration, derived measures, current value tableTM_DECODED, PARAMS, TC_EVENTS
stellar-executorexecutorDirect telecommands and runs of proceduresTC_COMMANDS, TC_EVENTS, PARAMS, RUNS
stellar-apiapiHTTP and WebSocket API, web console, relay of the editorTC_COMMANDS, TC_EVENTS, RUNS, ALARMS
stellar-alarmsalarmsLimits and alarms, their lifecycle and reactionsPARAMS, TC_EVENTS, ALARMS
stellar-schedulerschedulerRuns bound to passes, recurring rulesRUNS
stellar-transferstransfersFile transfers and quality of continuous streamsFILE_CHUNKS, TC_COMMANDS, PARAMS
stellar-editoreditorWeb editor of the configuration repository (optional)—
stellar-mcsall of the aboveThe services of the core in one processall of the above

Streams are created at start with an idempotent create or update, so the services can start in any order once NATS is up. The KV buckets and object stores are created by the service that first needs them. See NATS and JetStream for what each one holds.

Container images#

The binaries are also delivered as two container images, built from the repository by deployment/docker/build.sh and tagged by commit:

ImageContent
stellar-mcsEvery binary above, stellar-lsp, the CLI stellar, stellar-simulator, the web console in /usr/share/stellar/web and the example configuration in /usr/share/stellar/examples/config; runs stellar-mcs by default
stellar-adminstellar-admin (provisioning of the cells), with the nomad and stellar command lines

Both run as the user stellar (uid 10001), under tini, and work with a read-only root file system. A container chooses its service by its command and reads the global configuration from STELLAR_CONFIG:

Shell
docker run -e STELLAR_CONFIG=/etc/stellar/stellar.yaml \
    -v ./stellar.yaml:/etc/stellar/stellar.yaml:ro stellar-mcs:<commit> stellar-api

NATS runs from its official image (nats:2.12). deployment/docker/smoke.sh <commit> checks a pair of images: NATS and a cell in one process, the example configuration published, the API and the web console answering.

Outside the core, the simulator (stellar-simulator, see Simulated Targets) and the drivers, transports, gateways and connectors built with the SDKs take their settings from environment variables instead.

Before the first start#

The services work on a compiled configuration. Publish one before the first start, and after each change of the configuration repository:

Shell
stellar compile path/to/config --publish nats://nats.example.org:4222

Until a configuration is published, the API answers 503 on what needs it and the web console shows no configuration. See Compilation, Snapshots and Locks.

Command line#

Shell
stellar-reconciler --config /etc/stellar/stellar.yaml --instance rec-1
OptionEnvironmentDefaultMeaning
--config <path>STELLAR_CONFIGbuilt-in defaultsGlobal configuration
--instance <name>STELLAR_INSTANCE<kind>-<8 random characters>Instance name, unique per kind; letters, digits, - and _

Every key of the configuration can also be overridden with STELLAR__<SECTION>__<KEY> variables.

Lifecycle#

  1. The configuration is loaded and validated; the effective configuration is logged.
  2. The metrics endpoint starts on observability.metrics_addr.
  3. The NATS client connects to nats.url, with nats.credentials and nats.tls. When the server does not answer within 5 s, the refusal is logged with its reason and the connection is retried in the background: a service can start before NATS.
  4. The service runs until SIGINT or SIGTERM, then stops cleanly, waiting at most 2 s for its pending NATS messages.

Exit codes#

CodeMeaning
0Stopped on a signal
1The service failed (logged as component failed)
2Invalid command line, configuration, log filter or instance name

A supervisor (systemd, Nomad) should restart a service that exits with 1.

Logs and metrics#

  • Logs go to standard output, one JSON object per line (observability.log_format: json), or as text for development. observability.log_level filters them; RUST_LOG overrides it.
  • Metrics are served in the Prometheus format on http://<metrics_addr>/metrics, and /healthz answers ok while the process runs. Give each instance on a host its own port, for instance STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9101.

See Observability for the metrics of each service.

Scaling and high availability#

The target is the unit of partitioning: adding or removing an instance moves whole targets, and the order of the telecommands of a target is kept. Each service follows one of two patterns.

ServiceInstancesHow they share the work
apiseveral, all activeStateless; instances on one host share api.listen (SO_REUSEPORT) and the kernel spreads the connections
executorseveral, all activeRuns spread by their identifier; each target is held by a lease; direct telecommands taken one at a time from a shared consumer
schedulerseveral, all activeEach run is submitted once: Nats-Msg-Id = schedule + pass, deduplicated by JetStream
reconcilerone leader, the others stand byLease reconciler in stellar_leader
compute, alarms, transfersone active, the others stand byLeases compute, alarms, transfers in stellar_leader
editoroneIts drafts are working copies on the local disk

Leader leases#

A leader holds a key of the stellar_leader bucket whose TTL is reconciler.leader_ttl (5 s), created by a conditional write that only one instance wins. It renews the key every reconciler.leader_renew (1 s), requiring the revision it knows: a failed renewal means the lease is lost, and the instance stops writing at once and becomes a candidate again. Standby instances watch the key and try again when it expires; the TTL is enforced by the NATS server, not by the clocks of the instances.

The revision obtained when the lease is created is the epoch of the leader. The reconciler writes it in every value it writes, and readers ignore values of an older epoch: a frozen former leader cannot overwrite the new one.

The standby instances of the compute stage, alarm service and transfer manager give high availability, not load sharing: one of them takes over within leader_ttl when the active one stops or freezes, and starts again from the streams. Spreading them by target (nats.partitions) will come with load tests.

Configuration changes#

A new snapshot applies to a target only at a safe point: a target with a run holding a lease, or a pass in progress, keeps its revision until it ends. The reconciler watches the current key of stellar_config and publishes the revision of each target in stellar_readiness.

All in one: stellar-mcs#

stellar-mcs runs the reconciler, the compute stage, the executor, the API, the alarm service, the scheduler, the transfer manager and, when editor.repository is set, the editor, in one process. Each service behaves as if it ran alone, with its own instance name; they share one configuration, one log, one metrics endpoint and one NATS client. It takes about a quarter of the memory of the separate services, which suits demonstrations, integration tests, development and small cells.

Shell
stellar-mcs --config stellar.yaml
stellar-mcs --without editor,scheduler
OptionEnvironmentDefaultMeaning
--config <path>STELLAR_CONFIGbuilt-in defaultsGlobal configuration
--instance <suffix>STELLAR_INSTANCErandomSuffix of the instance names: api-<suffix>, executor-<suffix>…
--without <kinds>STELLAR_WITHOUTnoneServices not to run, comma-separated: reconciler, compute, executor, api, alarms, scheduler, transfers, editor

An unknown name in --without exits with code 2. The first service to fail stops the process, so that the supervisor restarts it whole.

A minimal deployment#

Shell
# NATS with JetStream
docker run -d --name nats -p 4222:4222 docker.io/library/nats:2.12.15-alpine -js

# The configuration, then the services
stellar compile examples/config --publish nats://localhost:4222
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9101 stellar-reconciler &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9102 stellar-compute &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9103 stellar-executor &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9104 stellar-api &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9105 stellar-alarms &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9106 stellar-scheduler &
STELLAR__OBSERVABILITY__METRICS_ADDR=127.0.0.1:9107 stellar-transfers &

For a complete orchestrated stack, see Nomad; for one stack per tenant, see Cells.

Stellar Control · v0.1.0

↑↓ to moveEnter to open