A campaign chains whole runs in order: scenarios, actions on the board, and gates that refuse to go further unless the board is in a state worth measuring. It keeps a cursor on persistent storage, can be paused at a step boundary and resumed, and a single step can be re-run.
Campaigns are a sequence, not a timeline: each step starts when the previous one finishes, because the length of a scenario or of a reboot is not known in advance.
Example#
An excerpt of dsl/campaigns/release_healthcheck.yaml:
version: "satlink.campaign/v1"
metadata:
name: "release_healthcheck"
display_name: "Release health check"
description: >
What a new release must survive before it is trusted: the PL loopback byte
for byte, then every modulation, coding, scrambler, symbol rate and frame
length the product offers, then the RF path.
author: "F. du Cray"
tags: ["healthcheck", "release", "coverage"]
schedule:
cron: "0 3 * * *"
skip_if_running: true
steps:
- id: "no_corpse"
description: "Nothing has fed this board, so nothing should have counted."
gate: { check: "counters_clear" }
- id: "pl_reference"
board: { do: "apply_profile", name: "pl_loopback" }
- id: "pl_target"
board: { do: "condition_chain", target: "pl" }
- id: "byte_exact"
gate: { check: "byte_exact_loopback", min_accuracy_pct: 99.0, blocks: 4 }
- id: "mod_qpsk"
scenario: { name: "hc_qpsk", timeout_s: 120 }
- id: "mod_bpsk"
scenario: { name: "hc_bpsk", timeout_s: 120 }
on_failure: "continue"File structure#
| Key | Required | Meaning |
|---|---|---|
version | yes | satlink.campaign/v1 |
metadata.name | yes | Identifier used to run it |
metadata.description | yes | What the campaign establishes |
metadata.display_name, author, tags | no | Shown in listings |
schedule | no | Start automatically (see Schedules). Absent: manual only |
steps | yes | The roadmap |
pass_criteria.tolerated | no | Steps allowed to fail without failing the campaign |
Each step has a unique id, an optional description, exactly one of scenario, board or
gate, and an optional on_failure (stop, the default, or continue).
Step kinds#
scenario#
- id: "fec_conv"
scenario: { name: "hc_qpsk_conv", timeout_s: 120 }
on_failure: "continue"Runs a scenario by name and takes its verdict. timeout_s gives up on the run after that long.
A cancelled scenario run fails the step (neither pass nor accusation: the detail says which).
board#
do | Fields | Effect |
|---|---|---|
apply_profile | name | Apply a profile |
condition_chain | target, profile (optional) | Condition the chain for a target (see Chain Conditioning). Name the profile: the conditioning applies it after tuning, so a separate apply_profile before it would be undone. On a fresh board with no profile applied and none named, it is refused |
recover | — | Pulse the receiver's clear bits (the recovery that is not a reboot) |
wait | seconds | Wait, for a settling time with no better observable |
reboot | wait_s (default 180) | Reboot the board and wait for the API |
power_cycle | wait_s (default 180) | Cut and restore the board's power and wait for it |
gate#
check | Fields | Passes when |
|---|---|---|
byte_exact_loopback | min_accuracy_pct (default 99.0), blocks (default 4) | The PL loopback returns a payload with at least this byte accuracy over blocks blocks |
counters_clear | — | The PL frame counters read zero before anything was fed. A non-zero counter means the fabric carries state from an earlier run, and every later reading would re-measure it |
register | name, equals | A catalogued register reads the given value |
A gate always stops the campaign when it fails: a gate declared with on_failure: continue is
refused at load. 99.0 % rather than 100 % in the byte-exact gate allows for the documented
DMA block-boundary artefact (about 0.1 point), which is not a demodulation error.
The byte-exact gate is the one everything else rests on: a board can come up with a receive path that locks and counts frames while delivering wrong bytes, and nothing measured on such a boot means anything.
Pass criteria#
pass_criteria:
tolerated:
- step: "rate_156k"
reason: "STE-nnn: 156 kbaud arm under investigation"The campaign passes when every step passed, except the tolerated ones. Each tolerated entry needs
a non-empty reason, or the file is refused.
Running and following#
satlinkctl campaign list
satlinkctl campaign get release_healthcheck
satlinkctl campaign run release_healthcheck --follow
satlinkctl campaign active # what is running, and the board provider
satlinkctl campaign runs # every execution, newest first
satlinkctl campaign status <id> # cursor, attempts, verdict
satlinkctl campaign pause <id> # stop at the NEXT step boundary
satlinkctl campaign resume <id> --follow
satlinkctl campaign rerun <id> <step-id> # run one step again, appending an attempt--follow polls the saved state rather than a WebSocket, so it picks the campaign back up when
the board returns from a reboot. The same operations are available under /api/v1/campaigns and
/api/v1/campaign-runs, and ws://<board>:8080/ws/campaign-runs/{id} streams progress.
The cursor is written to [daemon] campaign_state_dir (/mnt/jffs2/satlink-campaigns on the
board), which survives reboots. Only the position and the verdicts are kept there, not the
reports: the partition is small. Scenario reports stay in the run store, on the RAM disk.
Renaming a step's id breaks a saved cursor on purpose: resuming at a step that no longer exists
fails loudly rather than resuming somewhere else.
Schedules#
schedule:
cron: "0 3 * * *" # minute hour day-of-month month day-of-week
skip_if_running: true # defaultA five-field cron expression, in the board's local time, validated at load by the same parser
that matches it. With skip_if_running, a firing while the previous run is still going is skipped
rather than queued.
Shipped campaigns#
| Campaign | Purpose |
|---|---|
release_healthcheck | Gates on clean counters and a byte-exact PL loopback, then runs every PL-loopback health-check arm: QPSK, BPSK, GMSK, convolutional, RS, RS+convolutional, scrambler off, 39 and 156 kbaud, 1024-byte frames, CCSDS TM transfer frames. Scheduled daily at 03:00 |
device_loopback_healthcheck | QPSK through the AD9361's internal digital loopback (hc_dev_qpsk) |
rf_healthcheck | QPSK, BPSK and GMSK over the air path and an external cable loopback (hc_air_*), on a board not conditioned for anything else this boot |