Stellar ControlMission control · by Stellar Systems v0.1.0

Operations

Simulated Targets

Simulations described in YAML and played by stellar-simulator, for dry runs, CI of procedures, load tests and operator training.

A simulated target replaces the driver and the gateway of a target by a simulation engine, driven by a YAML description. It serves dry runs, the CI of procedure repositories, load tests and the training of operators — before any hardware exists.

The engine, stellar-simulator, registers as a driver and as a gateway, under the same contract as real ones: the MCS does not tell a simulated target from a real one, except by the environment that allows it (typically SIM). A procedure validated on a simulated target runs unchanged on a flatsat and in orbit.

A simulation#

Simulations live in simulations/<name>.yaml of the configuration repository. stellar check compiles them against the catalogue of their platform, and the engine reads them from the repository. The simulated platform-v3 of the example repository:

YAML
simulation: platform-v3-sim
platform: platform-v3@1.4.0
seed: 42
initial:
  tcu[*]: {responding: true, mode: STANDBY, anode_voltage: 0 V}
telecommands:
  ping:
    reply: {responding: true}
    delay: 200 ms
  set_mode:
    reply: {mode: args.mode}
    delay: 2 s
  standby:
    reply: {mode: STANDBY}
    delay: 1 s
  safe_mode:
    reply: {mode: SAFE_MODE}
    delay: 1 s
  set_anode_voltage:
    ramp: {anode_voltage: {to: args.voltage, rate: args.ramp}}
  activate_file:
    delay: 500 ms
  start_stream: {stream: start}
  stop_stream: {stream: stop}
measures:
  anode_voltage: {noise: 0.5 V}
  cathode_temperature: {model: first_order, target: 900 degC, tau: 60 s}
faults:
  tcu_silent:
    tcu[TCU3].responding: false
  anode_overvoltage:
    tcu[TCU1].anode_voltage: 310 V
  cathode_overheat:
    tcu[TCU2].cathode_temperature: 1300 degC
gateway:
  rate_bps: 1000000
files:
  lttm: {every: 10 min, size: 64 KiB}

The description is compiled against the catalogue: an unknown telecommand, measure, component, instance, argument or file type is an error (simulation::unknown-…), shown in the editor like any diagnostic. stellar schema simulation prints its JSON Schema.

A simulation written from the catalogue#

A platform needs no simulation written by hand to be simulated: stellar generate simulation writes one from its catalogue, a simulated target that keeps the contract of the catalogue — each telecommand gives what its verification (ACK 3) expects — and a starting point to make more realistic.

Shell
stellar generate simulation lab-psu --repository examples/config    # simulations/lab-psu-sim.yaml
stellar generate simulation platform-v3@1.4.0 --name platform-v3-twin -o twin.yaml

What the catalogue says becomes behaviour:

In the catalogueIn the simulation
verify: [{mode: STANDBY, within: 5s}] (a boolean or an enum)reply: {mode: "STANDBY"} after a delay of a fifth of the window, 1 s at most
verify: [{voltage: args.voltage, within: 3s}] (a number)ramp: {voltage: {to: args.voltage, rate: …}}, reached at half the window: the rate covers the range of the argument, else the limits of the measure
A condition without withinreply, at once
An argument that is a rate of the measure (ramp in V/s)ramp at that rate: rate: args.ramp
verify: [echo] aloneThe telecommand described with no effect ({}), which the engine echoes
Conditions in expressionsInverted when they can be: not responding, a and b, mode is not SAFE_MODE, pressure < 50 kPa, stable(…)
A parameter with set and readbackThe readback takes the argument of the setter
Limits of a measureIts initial value in the middle of its soft limits (else hard)
requires of the telecommandsInitial booleans and enums that meet them
AlarmsInitial values that keep them quiet; a fault raising each (<component>_<alarm>)
Hard limitsA fault past the high limit (<component>_<measure>_high)
Float measuresA noise of 0.2 % of their limits; a fifth of the tolerance of a readback; none on a measure compared exactly by a verification
stream componentstream: start and stop for the telecommands named so
files componentA file of each type every 10 min

What the catalogue does not say is listed at the end of the file and on the terminal, to complete by hand: the effect of a telecommand that changes the state without a verification (the current of a load), a condition that cannot be inverted, an alarm without a fault. The telecommands are commented with their description.

Before writing, the command compiles the simulation and checks it: each telecommand is sent alone to a fresh engine, in simulated time, and its verifications are evaluated as the executor would.

text
lab-psu-sim on lab-psu@1.0.0: each telecommand sent alone, its verification evaluated
  psu.set_voltage                  ok       verified after 1500ms (voltage=30.0)
  psu.output_on                    ok       verified after 300ms
  psu.output_off                   ok       verified after 0s

An existing file is kept (--force replaces it). The command exits with 1 when a telecommand is not played as its catalogue verifies it.

Checking a simulation: stellar sim check#

Shell
stellar sim check lab-psu-sim --repository examples/config

The same check, on a simulation of the repository, written by hand or generated and edited: it needs no NATS. Each telecommand gets a verdict:

VerdictMeaning
ok verified after <t>Its verifications held, the last one t after it was sent; its echo came when the catalogue verifies it
ok nothing to verifyThe catalogue verifies nothing of it
skippedNot checkable offline: its precondition does not hold in the initial state, or a verification reads a temporal function
FAILEDA verification did not hold in its window, or the echo did not come

The arguments are their default, else the middle of their range, else the last value of their enum, true or a byte. The command exits with 1 when a telecommand failed.

Reference of the format#

Top level#

KeyRequiredMeaning
simulationyesName of the simulation, and software name of its driver (platform-v3-sim)
platformyesCatalogue simulated, name@version
seednoSeed of the random generator; 0 by default
periodnoPublication period of the measures; 1 s by default
initialnoInitial values
telecommandsnoEffect of each telecommand
measuresnoModels of the measures
faultsnoNamed fault scenarios
gatewaynoThe simulated link: rate, passes, losses, COP-1
filesnoFiles generated in the on-board directory

Names#

A telecommand or a measure is written alone when its name is unique in the platform, else component.name. The measures of a reply or a ramp are those of the component of the telecommand, for the instance the telecommand is sent to.

In initial and faults, a key names a component — tcu[*] for all its instances, tcu[TCU3] for one, psu for a component without instances — or a measure directly (tcu[TCU3].responding).

Values#

Literals with their unit (0 V, 1.2 A), booleans, enum values, or an argument of the telecommand (args.mode). A measure without initial value is 0, false, or the first value of its enum.

Telecommands#

KeyMeaning
replyMeasures set after delay: a value, or args.<argument>
delayDelay before the reply and the ramps start; none by default
rampMeasures moving towards to at rate (unit of the measure per second); both accept args.<argument>
streamstart or stop: the telecommand starts or stops the stream it is sent to (a telecommand of the stream component)
thenWhat the target does next by itself, in order: the transitions an equipment goes through after a command

A telecommand whose catalogue declares verify: [echo] receives a conforming echo when the simulation describes it, with no effect if need be: tcu.blank_firing: {}.

What a telecommand does next: then#

An equipment often goes on by itself after a command: a camera told READY cools down, then is ready once its detector reaches its set point; an acquisition goes back to READY when it ends. then lists these steps, applied in order after the first effect of the telecommand:

YAML
telecommands:
  camera.set_mode:
    reply: {mode: COOLING}
    delay: 200 ms
    ramp:
      detector_temperature: {to: 17 degC, rate: 0.5}
    then:
      - after: 120 s
        until: "detector_temperature <= 18 degC"
        reply: {mode: args.mode}
  camera.start_acquisition:
    reply: {mode: IMAGING}
    delay: 150 ms
    ramp:
      acquisition_progress: {to: 100, rate: 5}
    then:
      - after: 20 s
        reply: {mode: READY, acquisition_progress: 0}
Key of a stepMeaning
afterRequired. When the step applies, from the end of the delay of the telecommand; with until, the latest it applies. The steps are in the order of their after (simulation::then-order)
untilA condition: the step applies as soon as it holds, after at the latest. It reads the measures of the component (faults applied), its derived measures and the arguments of the telecommand (args.mode is READY), as a verify does; no temporal function (a derived measure over time is unknown there, and the step waits for its after)
reply, rampAs for the telecommand, args.<argument> included; at least one of them (simulation::empty-step)
  • A step never applies before the one above it, nor before the end of the delay.
  • until is evaluated every 100 ms of simulated time, on the same instants whatever the speed of the simulation: a seed replays the same run.
  • A command received meanwhile takes over. A telecommand sent to the same instance cancels the steps still to come of the earlier telecommands, when they set or ramp a measure it sets or ramps itself. For example, set_mode STANDBY during the cooling of a set_mode READY leaves the camera in STANDBY.
  • A fault keeps its forced values over the steps.

stellar sim check and stellar generate simulation let the steps run within the window of each verification. verify: [{mode: args.mode, within: 130s}] holds once the step sets the mode.

Measures#

KeyMeaning
noiseStandard deviation of a Gaussian noise, in the unit of the measure (0.5 V)
model: first_orderThe measure tends towards target with the time constant tau
target, tauTarget value and time constant of the first-order model

Faults#

A fault is a named set of forced values. While a fault is active, its values replace those of the model. Faults are switched while a run is in progress, to test if failed blocks, alarms and their reactions (see Driving faults).

KeyDefaultMeaning
rate_bps1 000 000Theoretical rate announced by the gateway at registration, in bit/s
passesnoneFictitious passes: {first, every, duration}, in simulated time
lossnoneShare of the frames lost in each direction, 10 %, up to 90 %
cop1falseTelecommands go up in TC frames sequenced by COP-1
  • Passes. passes: {first: 0 s, every: 90 min, duration: 10 min} — a pass lasts at most its period (simulation::invalid-passes). Outside a pass, the gateway reports the link unavailable, refuses the uplink and publishes no frame: current values age as in orbit. Without passes, the link is always available. The measured throughput is the one the SDK counts.
  • Losses. loss drops that share of the frames each way: a telecommand sent but never received on board, a downlink frame never published. It exercises retries, timeouts and the retransmissions of file transfers.
  • COP-1. With cop1: true, telecommands go up in AD frames of a fixed virtual channel, taken on board by a FARM-1 whose CLCW comes down in the telemetry. The simulated transport sim-cop1@1, which the link declares, runs COP-1 and lets the rest through. See COP-1 and the Link Component.

On-board files: files#

YAML
files:
  lttm: {every: 10 min, size: 64 KiB}

For each file type declared by the files component of the catalogue, a file is born at each multiple of every in simulated time, of the given size.

  • A file has an identifier (numbered from 1 in the order of generation), a type, a size, a generation date and a generation (1 at birth). Its checksum uses the algorithm of the catalogue.
  • Its content is deterministic (seed, identifier, generation), except for an lttm: it records the samples of every measure at its generation, one JSON line per on-board instant, one period apart going back, padded with newlines to its size. The simulated driver decodes the lttm files into deferred samples.
  • The list telecommand of the transfer protocol brings down a directory frame, decoded into measures of files; the read telecommand brings down a chunk frame, at the rate of the link. Requests not served are lost at the end of a pass or when the gateway restarts.
  • The write telecommand writes on board and returns an echo. Its first chunk, at offset 0 and different from the content in place, starts a new generation; an absent file is created, of the last declared type. The checksum of a written file is computed at the next listing.
  • For a platform that transfers its files with CFDP, the simulator runs a CFDP entity on board (entity 2; the ground is entity 1). It sends the files the ground asks by Proxy Put Request, within the rate of the pass, and writes the files it receives. Its PDUs suffer the losses, which exercises NAKs and retransmissions.

Streams#

A telecommand marked stream: start starts the stream it is sent to; stream: stop stops it. A stream produces its segments on board at its nominal rate; outside a pass, they are lost. See Continuous Streams.

Raw values#

For a measure calibrated by the MCS (raw and calibration in the catalogue), the engine publishes the raw value by inverting the calibration: polynomial of degree 1, monotonic table, or enumeration. Another calibration is refused when the simulator starts.

Determinism#

The clock of the engine is simulated and its random generator internal: a seed replays exactly the same samples, which makes CI tests of procedures reproducible. Several engines simulate many targets, for load tests.

Running a simulated target#

Shell
stellar-simulator --repository examples/config --simulation platform-v3-sim \
    --target sim-1 --gateway sim-gw-1
OptionDefaultMeaning
--repository <dir>. (STELLAR_REPOSITORY)Configuration repository holding simulations/<name>.yaml
--simulation <name>—Simulation to play
--target <target>—Target served, as in the topology (sim-1)
--gateway <instance>—Instance name of the gateway, as in the topology (sim-gw-1)
--driver <instance><simulation>-1Instance name of the driver
--transport <instance><simulation>-cop1-1Instance name of the sim-cop1 transport, with cop1: true

It connects to NATS with the variables of the SDK: STELLAR_NATS_URL (nats://localhost:4222 by default) and STELLAR_NATS_CREDENTIALS. Logs follow RUST_LOG.

The driver registers with the name of the simulation as software, version 1.0.0, which a link driver: platform-v3-sim@1 accepts. The topology declares the simulated target like any other:

YAML
targets:
  sim-1:
    platform: platform-v3@1.4.0
    environments: [SIM, AIT]
    links:
      nominal: {driver: platform-v3-sim@1, gateway: sim-gw-1, default: true}
      direct:  {driver: platform-v3-sim@1, gateway: sim-gw-1}

Its frames are an internal JSON of the simulator, opaque to the MCS like any frame.

Driving faults#

Shell
stellar sim faults sim-gw-1                   # the faults and whether each is active
stellar sim fault sim-gw-1 tcu_silent on
stellar sim fault sim-gw-1 tcu_silent off
JSON
{"faults": [{"name": "tcu_silent", "active": true}, {"name": "anode_overvoltage", "active": false}]}

stellar sim talks to NATS directly (--nats, STELLAR_NATS_URL, and the NATS security options --nats-credentials, --nats-ca, --nats-cert, --nats-key), not to the API. The API exposes the same:

RequestEffect
GET /v1/sim/{gateway}/faultsFaults and their state
PUT /v1/sim/{gateway}/faults/{fault} with {"active": true}Switch a fault

A gateway that does not answer gives 404 sim::unknown-gateway; a real gateway, which does not know these verbs, 422 sim::not-simulated; an unknown fault 422 sim::refused.

Each change is published on stellar.sim.evt.<target> as {fault, active, at}, so that the evidence and the reports of a run show the faults active during it.

Faults scheduled by a run#

A run request can switch faults itself, to rehearse the recovery of a procedure (its if failed blocks, the reactions to its alarms) without anybody at a terminal:

YAML
# runs/tvac-leak.yaml
run: TVAC cycle
environment: SIM
targets: {chamber: chamber-sim, psu: psu-sim, dut: cam-sim}
inputs: {hot: 50 degC, cold: -30 degC}
faults:
  - {target: chamber, fault: vacuum_leak, at: {step: Plateau, after: 10 s}, off: {after: 60 s}}
KeyMeaning
targetRole of the target, as in targets
faultFault, as its simulation names it
atWhen it is switched on: after a delay (10 s) from the start of the run, or from the first entry in the step step (its name, Plateau, or its path, TVAC cycle / Plateau); at once by default
offWhen it is switched off, after a delay from when it was switched on (or from the first entry in its step); never during the run without it
  • Checked twice. When the run is resolved, the role must be one of the run (run::fault-target), the step one the run may enter (run::fault-step), and the environment not in_orbit (run::fault-not-simulated). When the run starts, before anything is sent, the target must be served by a simulated gateway whose simulation has that fault: otherwise the run ends failed, fault `vacuum_leak` of `chamber` cannot be scheduled: `chamber-1` is not a simulated target (…).
  • In the log. Each switch is logged in the run (log fault `vacuum_leak` switched on for `chamber-sim` (scheduled by the request)) and published on stellar.sim.evt.<target>, so the report shows it.
  • Delays are those of the executor, like wait: in a dry run, --speed does not shorten them.
  • After a restart of the executor, the faults of the run are scheduled again from its resumption.

Control verbs of the simulated gateway#

The simulated gateway answers these verbs on stellar.ctl.rpc.gateway.<instance>.<verb>, besides status:

VerbRequestEffect
faults{}Fault scenarios and whether each is active
fault{name, active}Switch a fault
files{}The on-board directory of the simulation
generate{file_type, size}Create a file on board right away
rewrite{id}Rewrite a file with its next generation (exercises supersession)
corrupt{id, active}Alter the chunks sent down of a file, without changing its listed checksum (exercises CORRUPTED)
stream_gap{stream, segments}Lose the next segments of a stream (exercises gaps)
Shell
nats request stellar.ctl.rpc.gateway.sim-gw-1.generate '{"file_type": "lttm", "size": 65536}'
nats request stellar.ctl.rpc.gateway.sim-gw-1.corrupt '{"id": 3, "active": true}'

Dry runs#

A dry run plays a procedure on ephemeral simulated targets, without any shared infrastructure:

Shell
stellar run --dry examples/config/runs/hot-standby.yaml --repository examples/config \
    --report report.html

It starts a local nats-server (--nats-server, STELLAR_NATS_SERVER), then, in its own process, the reconciler, the compute stage, the value table, the alarm service and the executor, on the snapshot of the local repository. Each role gets a target dry-<role>, played by the simulation of the repository whose platform is the version of the library, with the links its telecommands use. Its parameters, for configure, are those of the target dry-<role> when the topology declares one, else those of the first target of the same platform and version. The procedure runs in a SIM environment without human orchestration, ignoring allowed in, the environment and the targets of the request.

A platform without a simulation in simulations/ gets one written from its catalogue on the fly (see A simulation written from the catalogue), which the log says: a procedure runs dry as soon as its catalogue exists.

OptionMeaning
--repository <dir>Configuration repository; . by default
--seed <n>Seed of the simulations, instead of that of each simulation
--speed <x>Speed of simulated time (2 runs the simulations twice as fast); not the waits of the executor
--report <file>Write the report of the run
--on-decision <choice>How the dry run answers a decision: fail (default), skip, replay (3 times at most, then fail), abort, or ask on the terminal
--fault <ROLE:FAULT[@[STEP+]DELAY][~OFF]>A fault to switch on, after those of the request (see Faults scheduled by a run): chamber:vacuum_leak@120s, chamber:vacuum_leak@Plateau+10s~60s, sat:tcu_silent@TCU answers. Repeatable
--config <file>Global configuration (compilation and executor parameters)

Decisions. When a step not eligible for retry fails, a run waits for a decision (see Decisions). Nobody watches a dry run: it answers by itself, fail by default, so that the if failed blocks run — the recovery of the procedure is tested too — and the command exits with 1. The answer is logged and in the report as given by the dry run. The web editor and the VS Code extension run the same dry run.

Alarms and reactions. The alarms of the targets are evaluated as in operation, those of limits and the named ones: each transition shows in the log (alarm dry-chamber.vacuum.vacuum_leak ACTIVE_UNACK (critical)) and in a section of the report. A reaction (on_raise) is a run of the dry run: it takes the target from the run (see Automatic reactions), its events show in the log prefixed by [reaction], and its report follows that of the run. Once the reactions are over, the dry run resumes the run they suspended as it answers a decision (--on-decision): by default the step where it stopped fails, and its if failed blocks run on the state the reaction made safe (suspended by alarm …, failed by the dry run); with --on-decision abort, the run is aborted without them. The command exits with 1. Reactions still running when the run ends are followed two minutes at most.

Faults. The faults of the request, and those of --fault, are switched by the dry run like any run (see Faults scheduled by a run): the CI of a procedure repository tests its recovery.

Shell
stellar run --dry runs/hot-standby.yaml --fault "sat:tcu_silent@TCU answers"

The log shows as with stellar watch, and the command succeeds when the procedure succeeds: this is the CI of a procedure repository. See Running Procedures and the CLI Reference. The VS Code extension and the web editor run the same dry run on the procedure under the cursor.

See also#

Stellar Control · v0.1.0

↑↓ to moveEnter to open