Operate
Upgrades, GitOps, the licence, air-gap bundles, the metadata database after a restore, and what we need when you ask for help.
Upgrades
The control plane first, then the nodes one at a time.
sudo apt install runlot-cp=<new-version>
sudo /opt/runlot/bin/runlot-admin upgrade --nodes node-1,node-2 --version <new-version>runlot-admin upgrade installs on each node over SSH, waits for that node's own readiness check, and stops at the first failure naming the nodes that are done and the ones left. --dry-run prints the plan.
The safety is not the order; it is that a node refuses to start outside the protocol window. On start, a node compares its protocol version with the control plane's and exits naming both if the control plane is more than one version behind or ahead. Installing a node ahead of its control plane leaves that node down and says why.
The version string is not the filename. rpm forbids - in a version, so the prerelease separator becomes ~:
| We ship | apt wants | dnf wants |
|---|---|---|
runlot-cp_0.1.0-rc.110_amd64.deb | runlot-cp=0.1.0~rc.110 | |
runlot-cp-0.1.0-rc.110.x86_64.rpm | runlot-cp-0.1.0~rc.110 |
apt-cache policy runlot-node prints the string the package manager accepts.
A downgrade does not revert the schema. Migrations only add, which is what makes the previous release's code safe on the new schema, but installing an older package after a schema change is not supported. If you need to go back more than one release, ask us first.
GitOps: projects from a file
There is no Kubernetes custom resource for a project, and there will not be one: it would make the cluster a second record of what the metadata database already holds. Your pipeline calls one command instead.
apiVersion: runlot/v1
kind: ProjectSet
org: acme
projects:
- name: billing-api
secrets:
DATABASE_URL: ${BILLING_DATABASE_URL} # read from the pipeline's environment
hostnames:
- billing.acme.example
access:
policy: org # public | org | password
- name: inventoryrunlot apply -f projects.yaml --dry-run
runlot apply -f projects.yamlIt is idempotent: existing projects are left alone, declared secrets are written every time, missing hostnames are added, and the access policy is set only when it differs. Four rules:
- A secret value gets in only through a variable reference like the one in the example, and an unset variable stops the run. The file is in Git, so the value must not be in the file.
- Quote anything that is not plainly text. YAML reads
0100as 100. - Nothing is deleted without
--prune, and--prunerequires the org slug retyped as--confirm acme. Deletion takes the project's database and every backup with it. --jsonmakes the output one object for your pipeline's log.
The licence
The control plane re-reads the licence file every minute, so a renewal is a file replacement:
sudo cp license /etc/runlot/license/licenseGET /v1/license and the dashboard footer show the state, the tier, the node ceiling and the expiry. The states and what each one changes are in the overview.
Air-gap
We send one bundle per version:
runlot-<version>-deb.tar.gz
packages/runlot-cp_<v>_amd64.deb
packages/runlot-node_<v>_amd64.deb
packages/runlot-eval_<v>_amd64.deb
SHA256SUMS
SHA256SUMS.sig
license/sha256sum --check SHA256SUMS
runlot-license bundle-verify --dir .The signature is checked with a key compiled into the binary: no network and no keyring. SHA256SUMS is the build's own output, byte for byte, so the signature covers what the build produced. The Helm bundle has the same shape with the chart and docker save output in place of the packages.
The metadata database was restored
The control plane watches the identity of its PostgreSQL: the cluster id, the timeline and the write position. If the database comes back from the past, which is what a restore from backup or a failover to a lagging replica looks like, the control plane stops issuing leases to the nodes. Apps keep serving, but no deploy, no restore and no wake of a parked project goes through until an operator confirms the new state.
curl -sS -H "Authorization: Bearer $RUNLOT_OPERATOR_TOKEN" http://$RUNLOT_CP_PRIVATE_ADDR:8081/v1/eramode: lease_quarantine with a reason and a drainDeadline. Once the deadline has passed:
curl -sS -X POST -H "Authorization: Bearer $RUNLOT_OPERATOR_TOKEN" http://$RUNLOT_CP_PRIVATE_ADDR:8081/v1/era/activate
sudo systemctl restart node-agent # on every nodeActivation adopts the database as it is now and raises the era number the nodes check. A planned restore should be done the same way, with the control plane stopped during the restore and a marker file written so it starts in this state deliberately; the full procedure comes with your support contract.
A plain restart of PostgreSQL, an apt upgrade for instance, does not trigger this. The three witnesses that see a rewind all read clean after a restart, and a bare restart is logged and adopted.
Support without access
We have no access to your installation and no telemetry. What we ask for instead:
sudo runlot-admin support-bundle --out bundle.tar.gzIts contents are a list in code, serialised: versions and package digests, the licence state (not the licence), the node and placement tables, the last hour of control-plane logs, the pre-install check output, the migration ledger. Never the master key, any secret, your code, your database contents, or user emails beyond the operator list.
sudo runlot-admin preflight cp # or `node`The pre-install check is the first page of every bundle and the first thing the install runs. It lists which integrations are off, so "I thought it was on" has a line to read.
Metrics are Prometheus endpoints behind RUNLOT_METRICS_TOKEN; the chart offers a ServiceMonitor. Logs are JSON on journald or stdout for your collector. We ship no agent that sends anything anywhere.