Stellar ControlMission control · by Stellar Systems v0.1.0

Getting Started

Changelog

Notable changes of Stellar Control: the services, the configuration languages (catalogues, topology, steps and procedures), the API, the CLI and the SDKs. The format follows Keep a Changelog.

Every pull request that changes the code, a language, the API, the CLI or an SDK adds its entry under Unreleased, in the section that fits (Added, Changed, Deprecated, Removed, Fixed, Security), and updates the user documentation in docgen/. CI checks both (scripts/check-changelog.py); the labels no-changelog and no-docs exempt a pull request that changes nothing a user can see.

[Unreleased]#

Added#

  • stellar monitor (STE-1322): the MCS live in the terminal (ratatui) — the targets and their readiness, the current values of the target chosen, the runs in progress, the active alarms and the telecommands through their acknowledgements —, through the API like the web console. The API follows every telecommand over GET /v1/tc/watch?target=… (WebSocket, each TcEvent with its target).

  • stellar login (STE-1321): the CLI signs in at the identity provider of a cell by the device authorization grant — a code to approve in a browser —, exchanges its refresh token for the token of the organization of the cell, and keeps the session per API in ~/.config/stellar/credentials (0600). Every command to that API uses it when no --token is given, renewed before it expires; stellar watch takes a cut watch up again with a renewed token; stellar logout forgets it. GET /v1/auth/config names the client of the CLI (auth.cli_client_id), which stellar-admin makes at the provider (a native application with device flow, one for every cell) and gives each cell.

  • Packages of other repositories (STE-1306): a configuration repository declares in stellar.yaml the packages it takes from others — hsc-100: {git: …, tag: v1.0.0, path: packages/hsc-100} —, catalogues and libraries alike. stellar fetch puts them in the cache (STELLAR_CACHE) at the commit of their tag, through git, and records commit and hash in packages.lock (--update to move on); uses and depends then resolve them as local packages, offline. A package altered in the cache (package::external-changed), not the one declared, or a draft is refused. stellar schema repository gives the schema of stellar.yaml.

  • Alarms and their reactions play in a dry run (STE-1315): stellar run --dry starts the alarm service too. Transitions show in the log (alarm dry-chamber.vacuum.vacuum_leak ACTIVE_UNACK (critical)) and the report; a reaction is a run of the dry run, its events prefixed by [reaction] and its report after that of the run; once the reactions are over, the dry run aborts the run they suspended. With a fault scheduled by the request, the recovery of a procedure — leak, alarm, safing, report — is tested without anybody.

  • The templates of the cells from the catalog of Stellar Systems: deployment/docker/build.sh prepares them with templates.sh. With STELLAR_CATALOG, each use case of Stellar Control is its repository cloned at the tag pinned in the catalog (a branch is refused), checked with the CLI of the commit; the images carry the catalog (ONBOARDING_CATALOG set in stellar-onboarding) and a templates.MANIFEST of the repositories, tags and commits. Without it, the examples of the repository, as before. Publish images does the same with the variable STELLAR_CATALOG_REPO and the secret TEMPLATES_TOKEN.

  • The onboarding starts without its catalog when ONBOARDING_CATALOG names a missing file (an image built without a catalog), with a warning.

  • Faults scheduled by a run (STE-1312): a run request lists faults to switch on during the run, then off — {target: chamber, fault: vacuum_leak, at: {step: Plateau, after: 10 s}, off: {after: 60 s}} — on the start of the run or the first entry in a step, so that a dry run, and the CI of a procedure repository, rehearse its recovery (if failed blocks, reactions to alarms) without anybody at a terminal. stellar run --dry --fault ROLE:FAULT[@[STEP+]DELAY][~OFF] adds some. Simulated targets only: refused in in_orbit (run::fault-not-simulated), and checked against the simulated gateway of the target when the run starts. Each switch is in the log of the run and on stellar.sim.evt.<target>.

  • The use cases of the catalog in the online trial: with ONBOARDING_CATALOG (the catalog.yaml of stellar-catalog), the wizard shows the use cases in its order, with their titles, taglines, personas and the public repository of their template.

    • #/new?usecase=<id> opens the wizard on a use case; before signing in, the home page names it and /auth/login?next=… / /auth/signup?next=… go back to it (routes of the interface only).
    • Each environment has its Getting started: install the CLI, clone the repository of its use case, open its console, run its demonstration from the command line.
  • A subsystem reused on the platform that imports it (STE-1305): a platform that imports components of a package serves that package for them.

    • A link carrying only those components binds the package's own driver (catalogue hsc-100@^1.0 for a cubesat-6u that imports the camera), in the version imported.
    • A procedure written for the package (uses cam: hsc-100) runs unchanged on a target of the platform, its lock still valid; one that uses a component not imported is refused (run::not-imported).
    • Renamed imports (imager: {from: hsc-100.camera}) are translated: the binding carries the names (components of the bound link, applied by the Rust and Python SDKs to telecommands and samples), and the executor names telecommands, measures, parameters, files and links as the target does.
    • The alarms of an imported component take their on_raise reaction from a library of the package when the platform's libraries do not define it.
    • The compiled origin of a renamed import names its component in the package (hsc-100.camera@1.0.0); unrenamed imports keep their IR hash.
  • CSP over CAN (STE-1304): the transports csp1-can and csp2-can put CSP packets on a CAN bus in the CAN Fragmentation Protocol of libcsp, CFP 1 and CFP 2, CRC32 included (crc32), routed through a node (via, CFP 1) or sent from the address of the ground's CAN interface (can_address, CFP 2); downlink, packets are put back together per link, interleaved ones included, within reassembly_timeout_ms. In Rust (examples/rust/transports/csp-can, on the new stellar_sdk::cfp) and in Python (examples/python/csp_can.py, on the new stellar_mcs.csp, stellar_mcs.can and stellar_mcs.cfp), both checked bit for bit against frames of libcsp 2.1, both ways. The obc-csp driver plays the same end-to-end scenario over csp1-kiss and csp1-can, unchanged.

  • Extended CAN identifiers: CanFrame takes 29-bit identifiers (CanFrame::extended), on NATS as 4 octets with the high bit set; the gateway bench-can sends and receives both kinds.

  • Decision fail: a step not eligible for retry that failed can be answered fail — it counts as failed and the run goes on as after any failure, its if failed blocks included (abort still skips them). In the console (Fail), stellar answer, the API.

  • A dry run never waits for a decision: stellar run --dry answers it by itself, fail by default, so that the if failed blocks run and the command exits with 1; the answer is logged and in the report as given by the dry run. --on-decision skip|replay|abort|ask changes it (replay three times at most). The web editor and the VS Code extension run the same dry run (STE-1299).

  • What a simulated telecommand does next (then): a telecommand of a simulation can list the steps the target goes through by itself after it — COOLING then READY, IMAGING then back to READY. Each step has after (from the end of the delay) and an optional until condition on the measures and the arguments, which applies it earlier, plus reply and ramp. A telecommand received meanwhile cancels the steps still to come on the same measures; until is evaluated every 100 ms of simulated time, so a seed replays the same run. Checked by the compiler (simulation::then-order, simulation::empty-step), in the JSON Schema, played by the engine, waited for by stellar sim check (STE-1298).

  • Playground in the browser: on the onboarding (#/playground, no account), the configuration repository of a template edited and checked as you type by Stellar Control compiled to WebAssembly — problems underlined, the snapshot (the same hash as stellar compile), the simulation of a platform written from its catalogue and checked. tooling/wasm (crate stellar-wasm, built by tooling/wasm/build.sh) exposes check, generateSimulation, simCheck and schema to JavaScript; the engine of the simulated targets is a crate of its own, stellar-sim-engine, without I/O; GET /api/templates/<name>/files of the onboarding (STE-1295).

  • Simulations written from the catalogues: stellar generate simulation <catalogue> writes simulations/<name>.yaml, a simulated target that gives each telecommand what its verification expects (replies, ramps at half the window or at a rate argument, echoes), readbacks of parameters, initial values within the limits that meet the preconditions and keep the alarms quiet, a noise, a fault per alarm and per hard limit, streams and files; what the catalogue does not say is listed to complete by hand. stellar sim check <simulation> checks a simulation against its catalogue without NATS. A dry run writes on the fly the simulation of a platform that has none (STE-1294).

  • Sign-in with the identity provider, and cells for their users only: the keys of the provider read from its jwks_uri (auth.jwks_url, again every hour); the roles also read from a string of scopes, only those with auth.roles_prefix; tokens signed with ES384 (Logto), RS384 and RS512 accepted; auth.required, a token for every request of the API (a WebSocket gives it as token), but GET /v1/auth/config, which tells the web console how to sign in (auth.client_id, auth.organization). The console signs in by itself (code with PKCE, then the token of the organization of the cell, renewed in memory). With identity in platform.yaml, stellar-admin makes an organization of Logto per cell, its owner a member with the roles operator and supervisor, the address of the cell among the redirect URIs of the application of the consoles, all removed with the cell (STE-1291).

  • Pictures of the templates: scripts/template-screenshots runs each template locally with its demonstration and shoots the Target view of the web console (screenshot.png, framed by screenshot in template.yaml); the wizard of the onboarding shows them. The acceptance test of the test bench soaks 15 s at full load (STE-1292).

  • Onboarding of the online trial (tooling/onboarding, image stellar-onboarding, job onboarding on control.stellar-systems.eu): sign-up and sign-in through the identity provider (OpenID Connect, Logto), a wizard that picks a use case, the environment made in the background (queued when the platform is full) and followed live, My environments (open, extend, sleep, wake up, delete), a sleeping environment woken up by its first visit, emails (ready, about to expire, expired), the journal of every action and measures; delivered and checked by deployment/release.sh, installed by the bootstrap (STE-1292).

  • Templates of the use cases (examples/use-cases): a test bench (DC/DC converter module and its supply), a survey fleet behind satellite windows, a balloon gondola, a laboratory cryostat and a pumping station, beside the example configuration for the space missions; each runs as it is on simulated targets, with a template.yaml (title, description, use case, the procedure to run first) and a README. The image stellar-admin carries them; stellar-admin finds them in templates_dir and reads the simulators of a cell from its topology; CI checks each and runs its demonstration dry (STE-1290).

  • Releases of the production platform: deployment/release.sh builds and pushes the images to ghcr.io/stellar-factory, records the release in Nomad (stellar-admin release set), runs the cells again with it one at a time (stellar-admin cell upgrade, job cells-upgrade), checks a witness cell through Traefik and confirms it, or goes back to the previous confirmed release; --rollback on demand. deployment/bootstrap.sh installs the node once over SSH (Nomad with ACLs, NATS, the operator, the settings, the jobs, the witness cell); the workflow Publish images pushes the images of a commit (STE-1288).

  • Cells ready to use for an online trial (stellar-admin): with cells_dir, each cell gets a configuration repository of its own made from a template, committed as its user and published when it starts; provisioning.yaml gives the offers (life, longest life, sleep when idle, instances, limits), templates with their simulators, quotas and a hook for the events of the cells; cell status, cell extend and cell tick (expiry, warning before it, sleep when idle), run periodically by deployment/nomad/jobs/containers/admin.nomad.hcl; the Python library stellar_admin.provisioning.Provisioner for the onboarding, idempotent and resumed after an interruption (STE-1289).

  • Web editor without review (editor.forge.kind: direct): for a sandbox, Publish compiles the draft, merges it into the target branch and publishes its configuration at once, the user named as its publisher (STE-1289).

  • Last access to the API: written in the bucket stellar_activity (key api, at most once a minute, health checks excluded), what puts an idle cell to sleep (STE-1289).

  • Container images: stellar-mcs (core services, all-in-one, editor, CLI, simulator, web console, example configuration) and stellar-admin (provisioning, with nomad and stellar), built by deployment/docker/build.sh and tagged by commit, checked by deployment/docker/smoke.sh in CI (STE-1286).

  • Cells in containers behind Traefik: deployment/nomad/jobs/containers/cell.nomad.hcl runs a cell in containers of the image stellar-mcs (Docker driver, read only, without capabilities), routed by Traefik from the tags of its API service; jobs/containers/nats.nomad.hcl runs NATS in a container. stellar-admin chooses with runtime: containers and ingress: traefik in platform.yaml, which also takes the image, its tag, the Traefik settings and more variables of the job (STE-1287).

  • Encode and decode with a driver, sending nothing: the control verbs encode and decode of every driver (Rust and Python SDKs), POST /v1/encode and POST /v1/decode, and stellar encode / stellar decode: the frame a driver makes of a telecommand, the samples it makes of a frame calibrated by the catalogue, in the context of the link (STE-1270).

  • Web console: tab Parameters of a target: its current mode (since when, by whom or by which run), a mode declared by the operator, and the parameters of its components with their value in each mode and their readback matched within the tolerance (STE-1280).

  • Web console: redesigned console (overview, targets, runs, alarms, passes, file transfers, instances, topology, editor) and the API routes it uses: history of a measure (GET /v1/targets/{target}/values/history), the journal of the reconciler (GET /v1/reconciler/events and …/watch), and WebSockets following the passes, schedules and transfers (/v1/passes/watch, /v1/schedules/watch, /v1/transfers/watch) (STE-1258, STE-1259).

  • log in steps: log <value> and log "<text with {values}>" write values in the log of the run (event logged), in the report, the web console, stellar watch and the Python client (Logged); never failing, a missing value is unknown, a stale one (stale) (STE-1269).

  • Documentation: user documentation built with stellar-docgen (docgen/), and the OpenAPI document of the API generated from the code, served at GET /v1/openapi.json (STE-1267).

  • Code generators: stellar generate driver, the typed skeleton of a Python driver from a catalogue (STE-1264); stellar generate client, typed Python functions launching the procedures of a library (STE-1265); --check in CI.

  • Python API client: stellar_mcs.Client launches runs, follows their events, answers their questions and decisions, controls them and uploads content (STE-1263).

  • Parameters and modes: component parameters (setter, readback, tolerance), topology modes with per-target values and overrides per environment, configure <role> for <mode> and param <path> [in <mode>] in the procedures, the current mode of a target (/v1/targets/{target}/mode, stellar mode, stellar parameters), mode and parameters in the link context of drivers and gateways (STE-1257).

  • Links: transports between drivers and gateways (STE-1231); link parameters overridden by environment (STE-1230); COP-1 in the driver SDK and the ccsds-tc transport (STE-1207, STE-1233); CSP over KISS, csp1-kiss and csp2-kiss (STE-1234); CFDP class 2 (STE-1195).

  • Example drivers and gateways: platform-v3 over CCSDS (STE-1206), tcu-can and the CAN bench gateway (STE-1208), the satlink RF bench (STE-1229), obc-csp (STE-1234), LTTM files of platform-v3 decoded into deferred values (STE-1243).

  • Cells and security: one NATS account and one set of services per tenant (STE-1236), stellar-admin (STE-1240), the all-in-one stellar-mcs (STE-1241), TLS to NATS and client certificates (STE-1237), HTTPS and WSS in front of the cells (STE-1238), JetStream encrypted at rest and secrets in Nomad variables (STE-1239), OIDC identity and NATS credentials (STE-1186).

  • Tools: language server (STE-1202), VS Code extension (STE-1203), web editor of the configuration repository (STE-1204), web console for monitoring, telecommands and runs; descriptions in catalogues, steps and procedures (STE-1225).

  • SDKs: Python SDK for drivers, gateways, output connectors and pass feeders (STE-1223); drivers declare their catalogue as name@requirement (STE-1227).

  • Output connectors: framework, lag alert and the TimescaleDB reference connector (STE-1198 to STE-1200); continuous streams archived with their retention and quality (STE-1196, STE-1197).

  • File transfers: chunked downloads and their state machine, LTTM processing, uploads with a separate activation (STE-1191 to STE-1194); the on-board directory in the current values (STE-1190).

  • Passes and scheduling: passes in JetStream KV fed by the flight dynamics and the ground station providers, schedules and recurring rules, plans validated by a supervisor, runs bound to a pass (STE-1181 to STE-1185, STE-1188, STE-1189).

  • Alarms: evaluation and lifecycle, acknowledgement, shelving, notifications, automatic reactions on_raise (STE-1178 to STE-1180).

  • Executor: interpreter of resolved runs, event sourcing and takeover after a restart, leases on targets, retries and operator decisions, suspension and abort, evidence and test reports, confirmations of hazardous telecommands (STE-1167 to STE-1172, STE-1187).

  • Compute: calibration, pure and temporal derived measures, the current value table, restart by replay (STE-1163, STE-1164, STE-1176, STE-1177).

  • Simulator: simulation engine, driver and gateway, faults, passes, throughput and on-board files; dry runs on ephemeral simulated targets (STE-1156 to STE-1159, STE-1175).

  • Reconciler: registrations and heartbeats, leader election, snapshots applied at safe points, bindings and readiness of targets, NATS JWTs of the bound instances (STE-1149 to STE-1153).

  • Languages and compiler: catalogues (measures, calibration, derived measures, telecommands, alarms, files, imports), topology (environments, targets, links, connectors), steps and procedures with their safety rules and maximum durations, diagnostics in plain language, typed IR with content addressing, locked versions and revalidation (STE-1130 to STE-1147, STE-1221).

  • CLI: stellar check, lock, compile, send, watch, runs and dry runs (STE-1148, STE-1165, STE-1174, STE-1175).

Changed#

  • Stellar Control is the name of the product, instead of Stellar-MCS: documentation, web console, onboarding and playground, emails, help of the command line, OpenAPI document, titles of the JSON Schemas, VS Code extension, application of the consoles at the identity provider (identity.console, now Stellar Control console). The technical names stay: crates and packages stellar-*, the binary and image stellar-mcs, NATS subjects, variables (STE-1297).

  • stellar-admin: one Nomad namespace for every cell: each cell is the job cell-<name> of the namespace cells, its secrets the variable nomad/jobs/cell-<name> (the secrets_path of jobs/stellar.nomad.hcl), rather than a namespace of its own; deleting a cell also removes the service registrations Nomad kept. Cells created before must be deleted and created again (STE-1287).

  • User documentation: the page of the web console describes the redesigned console (STE-1279).

Fixed#

  • After the reaction of an alarm, a dry run aborted the run it suspended, skipping its if failed blocks, and the run ended « aborted by an operator » (STE-1324): the dry run now resumes it as it answers a decision (--on-decision, fail by default), so the recovery of the procedure runs on the state the reaction made safe; --on-decision abort keeps the abort. A suspended run can be resumed with the new choice fail (stellar resume --choice fail): the step where it stopped fails, and the run goes on as after any failure. An aborted run ends with the reason aborted by <who>, a step failed so with suspended by …, failed by ….
  • A telecommand sent steps after wait until X could still be rejected by its precondition on X (STE-1323): the server hands a sample to the wait until a moment before its stream stores it, and requires read the stream without waiting for it, unlike check and log. The precondition of a telecommand of a run now waits, 2 s at most, for the values of its component to be no older than the newest sample of that component the run received.
  • stellar-admin refused by a provider behind Cloudflare (STE-1321): its requests to Logto carried the agent of Python's urllib, which Cloudflare refuses (403 error code: 1010); they are now named stellar-admin, and those of the CLI stellar-cli/<version>.
  • A reaction to an alarm waited for the run holding its target (STE-1318): a safing (on_raise) waited for the run in progress to release the target, 60 s at most, then failed for good when that run was waiting for a decision, an answer or a confirmation. The reaction now takes the lease from any ordinary run, whatever it waits for (reaction in the run submission, leases::preempt); that run is suspended within a second, its log naming the reaction, and sends nothing more until an operator resumes it. A reaction that cannot be launched, or whose run fails, raises the critical alarm <alarm>-reaction on the same component, to be acknowledged.
  • A telecommand sent right after wait until X, or after the verification of the previous one, was rejected by its precondition on X (STE-1310): requires read the current value table, which follows the samples a moment after wait until, expect and verification see them. Preconditions, check and the start of wait until now read the table together with the samples of the measure it has not taken yet, the latest by time as the table keeps it, without waiting. A rejection now gives the values read and their age (pressure = 101325 Pa (received 0.8 s ago)).
  • A check or a log right after an expect read the value before the one the expect had just seen: the current value table follows the samples with a short delay. They now wait for it to catch up with what the run saw of the component, 2 s at most (STE-1290).
  • Dry runs of procedures with modes: configure found no value of the parameters of the simulated target dry-<role>; it takes those of a target of the same platform (STE-1290).
  • The API builds again after the merge of the web console: its new routes and types are in the OpenAPI document (STE-1259).
  • The maximum duration of a step with configure counts one readback window per parameter read back, not one per setting telecommand (STE-1257).

Stellar Control · v0.1.0

↑↓ to moveEnter to open