Concepts
The vocabulary CloudRunner runs on — what each thing is and how it connects to the rest. Every concept keeps the same icon and colour wherever it is named in the dashboard.
Workblocks
The containerized processing units and the images that back them.
- Workblockname, e.g. ingressor
A named processing unit in the reconstruction pipeline — a role, not an image. The id is a descriptive snake_case string like ingressor or dense_matching. One registry image can power several workblocks; the workblock defines what step this is, not which binary runs it.
- Workblock version#id
Pins a workblock to a specific ACR image (repo:tag). A workblock accumulates versions over time; each names the arm64 job image and, optionally, an amd64 variant. Choosing a version is choosing which build of the step runs.
- Configmap#id
The CLI arguments and IO declaration for a workblock version — command args, optional env vars, and named inputs/outputs. Args reference IO by ${name.path}, validated against the declared slot names. A version can carry several configmaps for different invocations.
- Workblock instance#id
One actual execution of a workblock — which version and configmap ran, the Kubernetes job behind it, its status and result. Instances are the runtime record; the workblock and version are the declaration.
Pipelines & tasks
How steps are wired into a pipeline, and what a single run of it is.
- Pipeline preset#id
A reusable pipeline definition: the full DAG of steps, their edges, and per-node config (image pins, IO bindings, retries). A preset can be marked default; the workblock versions a default preset runs are locked from edits so a run can't shift under it.
- Pipeline stepnode id, e.g. workblock:ingressor
One node in a preset's DAG. Its type says what it does — workblock:*, io:*, or frb:* — and it carries the type's config plus its outgoing edges, an optional condition, and whether it needs a GPU. Steps are the plan; task step executions are the runs of them.
- Task#id (shown t<id>)
One run of a pipeline for one scene. On creation it freezes a snapshot of the preset's DAG, so later edits to the preset never disturb an in-flight task. Its lifecycle runs created → waiting-placement → running → completed / partially-failed / failed.
- Task step execution#id
One attempt at running a pipeline step inside a task. It records the Kubernetes resource, the workblock instance used, status and timing. Retries make new executions, so several can share a step name — the history of that node.
- Task event#id per task
A user-facing milestone of a task, authored in the preset via notify:publish-event steps (e.g. "Footage preparation", "Point cloud"). Distinct from raw step status: it is the coarse progress the operator sees, mirrored back to the initiating clr-link.
- Task ownername, e.g. prod-clr-link
Who initiated a task — usually a clr-link deployment, seeded from SYSTEM_CLR_LINKS on boot. Carries the bearer token and the FrameBase URL/credentials a task's uploads use. Ownership is optional; unattributed tasks are allowed.
Compute
The nodes work runs on, how their capacity is sliced, and the queue that places work.
- Compute nodename, e.g. spark-1
A server with GPU/CPU/RAM, one-to-one with a Kubernetes node. It is discovered, not hand-written: the controller tracks its allocatable limits and whether it is currently reachable. Operator intent (active / maintenance) is kept separate from that liveness.
- Compute lane#id
A named slice of a node's capacity with its own GPU/CPU/RAM budget. A lane answers both "how many workloads fit on this box" (one lane, one active allocation) and "how big is a workload" — the lane is the single source of a pod's resource requests. It drains before it disables.
- Compute space#id
A task's resource requirement — how much GPU, CPU and RAM it needs. Created one-to-one with the task and used only to filter which lanes are big enough; what the workload actually gets comes from the lane it lands on.
- Compute queue#id
The pending placement requests. A task enters the queue and waits until a free lane on a schedulable node is large enough for its compute space. Each item records why it is still waiting — no free lane, no schedulable node, none large enough.
- Allocation#id
The actual lock on one lane for one execution attempt. The queue picks the node and lane; the allocation holds it and is the source of the pod's resources. Multiple allocations per compute space track retries; one lane is held at a time.
IO & artifacts
The typed data that flows between steps, and how steps name their inputs and outputs.
- IO slotname, e.g. rgb_images
A named input or output of a step — one asset inside the step's output directory, addressed by subpath. Slots are how a step declares what it consumes and produces; args and bindings reference them by name rather than by hard-coded paths.
- Artifacttype + version
The data that flows between steps — reconstructions, point clouds, meshes, images. Producers and consumers are wired blind through the DAG, so an artifact's registered type is the contract that lets a later step trust what an earlier one made.
- Artifact typekind / version / variation
A registered artifact specification — its kind, a version for the data format, and an optional variation. The registry is the single source of truth for cross-step compatibility; changing an output's shape or meaning means bumping its version.
Registry & runtime
Where images live and what the controller spins up in the cluster to run a task.
- ACR imagerepo:tag
A container image in Azure Container Registry (c2gridinternal.azurecr.io). Workblock versions pin an ACR repo tag; the controller also pulls its system images — storage sniffer, frb-utils — from here. Job images are arm64; the controller itself is amd64.
- Kubernetes jobname, e.g. clr-task-<id>-...
A Kubernetes Job the controller creates to run a step — a workblock, an IO transfer, or a FrameBase operation. It mounts the task volume, requests the lane's resources, and (for GPU work) carries an nvidia-smi sidecar. Its pod logs survive until the placement is closed.
- Task volume (PVC)PVC name
The per-task persistent volume every job for that task mounts. Created once when a node is chosen and kept intact after the compute is done, so outputs and logs stay reachable. It is torn down only when the task's placement is closed.
- Storage snifferper-task deployment
A per-task file browser the controller stands up over the task volume — a small deployment, service and ingress. It lets the dashboard list, download and archive the files a run produced; archiving runs as an async tar job the dashboard polls.
- FrameBase stepnode type frb:*
A pipeline step that talks to FrameBase — pulling inputs or pushing results and structures. It runs as a Kubernetes job on the clr-frb-utils image, which calls the FrameBase API, and carries a small auto-retry budget for transient failures.
- GPU VMAzure VM name
An optional Azure-managed GPU VM in a pool the gpu-vm-operator can spawn and stop on demand. It is off unless GPU_VM_OPERATOR_CONFIG is set — a way to add compute nodes elastically rather than keeping every box powered on.