MODEL below).
Everything described here lives in deployment/slurm/, and the design rationale is ADR-0009.
Prerequisites
On the cluster- An account with access to GPU nodes, and whatever network access the site requires — a VPN, at many.
- Apptainer or Singularity on the compute nodes. Postgres and the backend run as containers; only vLLM does not. Check before anything else:
If that fails, this stack will not work here as designed — see Adapting to another cluster.
- Python 3.9+ as
python3on the compute nodes (srun --pty python3 --version). Verified on 3.12. - Roughly 15 GB of quota for the serving environment, plus whatever model weights you download if the site has no shared cache.
- An SSH host alias for the login/submit node. The tooling never prompts for a password, so key auth must already work:
Verify with
ssh mycluster hostname, then put the alias incluster.env. TheUserline is load-bearing, not cosmetic: the laptop-side tooling reads it withssh -Gand substitutes it for{user}inCLUSTER_WORKto decide where the venv, the secrets file and the results live, while the job substitutes its own$USERon the node. Put your cluster username there even if it matches your laptop login. make,bash,rsync,curlandpython3(used bymake secrets,make fetchandsubmit.sh’s local validation).- The GitHub CLI (
gh), logged in, formake digestonly: it resolves an image tag through yourghsession, not through the token stored bymake secrets. The target refuses to run whenghis missing or logged out. - A classic GitHub token with the
read:packagesscope, to pull the private backend image. Fine-grained tokens are rejected by GHCR with an opaque401.
From zero to a first result
make help lists every target.
Step 1 first, or steps 2 and 3 build in the wrong place.
cluster.env holds the submit host and the work directory, and both make env and make secrets write to paths derived from it. At the UKP cluster the shipped values are already correct and you can skip straight to step 2.On any cluster other than UKP, step 4 needs an arm of your own first. The shipped
qwen3-32b arm hardcodes PARTITION=gpu, QOS=gpu, GRES=gpu:a180:1 and a model path under UKP’s shared weights. cp arms/qwen3-32b.env arms/<name>.env, set PARTITION, QOS (empty if your site has none), GRES (your site’s GPU type name) and MODEL (a HuggingFace id such as Qwen/Qwen3-32B if SHARED_WEIGHTS is empty, or an absolute path to a flat model directory), then make gate EXP=<name>. See Adapting to another cluster.make env
Submits a CPU-partition job (CLUSTER_CPU_PARTITION / CLUSTER_CPU_QOS in cluster.env) that creates the venv named by VLLM_VENV and installs env/requirements.txt. A pip install needs no GPU, and holding an A100 to download wheels denies it to someone who does.
make env only queues the job and returns at once. Watch it with make logs EXP=haico-env and wait for the line environment ready: … before running make gate: the build takes a few minutes once it starts, plus whatever the CPU queue wait is. make secrets does not depend on it and can run meanwhile. A gate submitted too early dies at source …/bin/activate after its own queue wait.
It is idempotent — re-running when the environment is already healthy is a no-op, so it is safe as the first line of any runbook. make env FORCE=1 rebuilds from scratch; make env VENV=/path/to/other builds elsewhere, which is how to test a version bump without disturbing the environment your campaigns use.
The build verifies rather than trusting pip’s exit code: it imports vLLM, httpx and torch, and reports whether the tool-call parser registry loads (advisory — make gate is the real test). That matters because a subtly broken install otherwise surfaces much later as “the model cannot emit tool calls” at the top of a campaign.
It then writes haico-requirements.lock.txt into the venv — the full resolved closure from pip freeze. requirements.txt states intent; the lock file states what a given result is actually attributable to.
Versions are pinned deliberately. A campaign’s results belong to a model and the stack that served it, so a floating range means a rebuild months later silently serves the same weights through different sampling or a different parser set. Upgrading vLLM is a decision, not housekeeping — the tool-call parser registry changes between releases and every arm file’s
TOOL_PARSER names one of its entries, so re-run make gate after any bump.torch is deliberately absent from requirements.txt. vLLM ships its own pin and installs a matching build; adding a torch line fights that and is how you get a venv that imports but cannot see the GPU. It is also why this must stay separate from any training environment — trl/peft pull a different torch.make env once and changes nothing else: VLLM_VENV is per-user by default. Point it at a shared project directory if a group wants one environment — accepting that everyone then shares its pins, and an upgrade breaks everyone at once.
make secrets
Prompts for your GitHub username and the token, verifies the token against GHCR without downloading anything, and writes both to $WORK/haico-cluster/secrets/ghcr.env with mode 600. submit.sh checks that this file exists before queueing a batch, because the job dies without it — after the queue wait and the Postgres pull.
It deliberately does not offer to reuse your gh CLI token. That credential carries whatever scopes the account granted it — typically repo, workflow, admin:org — and storing it on multi-user cluster storage for the length of a campaign, to save one token mint, trades a broad credential for a small convenience. A dedicated classic token with only read:packages can do nothing but pull the image.
It is stored under $WORK and not $HOME for a reason: $HOME is not mounted on compute nodes, so a credential there is unreadable exactly where the job needs it.
That is the only long-lived secret. JWT_SECRET, the admin password, the HAI-Co² API token and the vLLM key are all generated inside each job and die with it. HF_TOKEN is optional — needed only for gated repositories such as Llama, and irrelevant for Qwen or Mistral, which are what the shipped arms use. Add one later with bash scripts/bootstrap_secrets.sh --hf: the script rewrites the file as a whole, but it reads the stored values first and Enter at a prompt keeps them, so you are not asked to dig out the GHCR token again; a blank HF answer keeps whatever is stored. Re-running without --hf (to rotate the GHCR token) keeps a stored HF token too.
Run your first experiment
$WORK/haico-cluster/ is this directory, rsynced by every make submit / make gate / make env, plus logs/<jobname>_<jobid>.out|.err (the job logs) and secrets/ghcr.env. $WORK/haico-runs/<RUN_ID>/ holds a run’s results and is what make fetch copies. $WORK/envs/haico-serving is the venv from make env; $WORK/hf_cache, $WORK/.apptainer and $WORK/.cache are the model, image and pip caches shared across jobs. Jobs are named haico-exp-<RUN_ID> (batch), haico-gate-<RUN_ID> (gate) and haico-env (venv build).
Two log files per job. Progress goes to .out; every FATAL: diagnostic, the driver’s progress notes and the gate’s GATE FAILED verdict go to .err. make logs tails both. make fetch does not copy the job logs — read them with make logs or ssh while you need them.
Reading the gate. The gate job serves the model, makes GATE_N (default 25) tool-calling requests at temperature 0.7 (the arm’s TEMPERATURE is not used) with a 2048-token output cap and a 180 s timeout per call, and passes when at least 95 % return a parseable tool call carrying a non-empty action_and_reasoning. GATE PASSED appears in the .out log, GATE FAILED in .err. The per-prompt report is $WORK/haico-runs/<RUN_ID>/gate-report.json ({model, base_url, n, passed, pass_rate, threshold, results: [{prompt, passed, reason}]}, with gate-vllm.log beside it); make fetch EXP=<RUN_ID> pulls both, and a second gate run on the same arm overwrites them. Read reason before trusting a failure: no tool_calls emitted, action_and_reasoning missing and action_and_reasoning empty are model results; a reason beginning ConnectError, ReadTimeout or HTTPStatusError means the call never reached the model, and the gate says so with THIS IS NOT A MODEL RESULT.
Stopping a run
make cancel EXP=<arm> acts on the jobs named after the arm’s RUN_ID. A queued batch is removed from the queue. A running batch receives TERM on its batch shell only, so the job’s own teardown runs: the driver stops launching turns, records every unfinished conversation as interrupted, writes ledger.json, captures what it can with one short attempt per panel, and the database is dumped while Postgres is still up — allow a few minutes. make cancel EXP=<arm> FORCE=1 is a plain scancel: the whole cgroup is killed at once and nothing newer than the last periodic dump (every DUMP_INTERVAL seconds) survives. Gate jobs are always cancelled outright. Do not scancel a batch by hand unless you mean FORCE=1: it kills Postgres together with the driver, so the final dump fails and the ledger is not written.
Afterwards make submit refuses the RUN_ID (it has a ledger); make resume EXP=<arm> continues it. sacct shows a cancelled run as FAILED with exit code 15:0 (Slurm’s rendering of 143, 128 + SIGTERM) on purpose, so a stopped run is never mistaken for a completed one; the results up to the signal are complete.
Changing the experiment
Everything is in one arm file,arms/<name>.env. Every key that shapes the result is required and has no default: submit.sh refuses an arm that omits one. Six keys are optional and fall back to a fixed value when absent — HAICO_READY_TIMEOUT (2400 s), PHOENIX_ENABLED (0), DUMP_INTERVAL (600 s), GATE_N (25), and REASONING_PARSER / EXTRA_VLLM_ARGS (empty). Keep them in the arm anyway so the file records what ran. The driver’s grace period after a signal (150 s) is a job-script constant, not an arm key.
The arm file is what you edit; run-config.env in the run directory is what actually ran. Every attempt appends the job id, node, GPU, image digest, every knob, and for a model given as a directory the SHA-256 of its config.json and safetensors index, so a result stays attributable after the arm file has been edited for the next run.
submit.sh validates locally before anything is queued, because a typo caught on your laptop costs a second and the same typo caught by the scheduler costs a queue wait plus a model load. It rejects: missing values, a TENSOR_PARALLEL_SIZE that disagrees with GRES, a request above the ceilings in cluster.env, an output cap that leaves no room for the prompt, a tag where a digest belongs, a prompts file that does not parse (every line must be a JSON object with id and non-empty string turns, ids unique and usable as file names), a missing secrets file on the cluster, a QOS your account does not have, and a RUN_ID that already has a ledger (make resume is the way to continue one).
Every make target wraps ./submit.sh <arm> [--gate] [--resume] [--dry-run] [--no-sync]; --no-sync submits without re-syncing the tree to the cluster, for when nothing but the arm has changed.
Arm-file reference
Identity
GPU and node selection — all passed straight through to
sbatch.
CONCURRENCY=2 is measured, not chosen. The backend serialises requests — vLLM reports Running: 1 reqs in the large majority of samples regardless — so raising concurrency does not raise throughput; it queues conversations inside the backend, where they emit no bytes and the driver eventually abandons them. On identical 8-conversation runs:Lower concurrency was both more reliable and marginally faster. Raise it only once the backend genuinely serves requests in parallel.
The batch — what the driver sends, and how it waits.
Two ceilings are constants in
driver/run_batch.py, not arm keys: a turn is abandoned as turn_deadline after 1800 s in total, and as stream_idle after 900 s of keep-alives with no data frame. Raising READ_TIMEOUT does not move them; a legitimately slow model at high CONCURRENCY can hit the first — lower concurrency, or change the constant and record it with the run.
Prompt sets
PROMPTS names a file under driver/prompts/, one conversation per line: {"id": "pilot-01-report", "turns": ["first user message", "second", …]}. Blank lines and lines starting with # are ignored, so a set can carry notes or have a conversation commented out. id becomes transcripts/<id>.json and the key make resume uses, so ids must be unique within a set and usable as file names; submit.sh checks both before queueing. Keys other than id and turns are ignored. Every turn is sent verbatim as the user message; see driver/prompts/README.md for why each turn must be answerable without a clarifying question.
Two sets ship: pilot.jsonl, 8 conversations of 4 turns covering the document, plan, preferences and objective panels, and stress.jsonl, 20 conversations of 6 turns that apply three turn policies (grow: enlarge the document every turn; plan: churn todos and preferences; revise: rewrite repeatedly) across seven topics, built to push on MAX_MODEL_LEN, tool-call volume and snapshot-chain length respectively. Ids encode set, index and policy (stress-04-grow-tutorial), so outcomes can be grouped by policy.
Adapting to another cluster
Everything specific to one site lives indeployment/slurm/cluster.env. In the common case that file is the only thing you change.
CLUSTER_WORK is the one that most often bites. On many clusters $HOME is not mounted on compute nodes, so a credential or cache under it is unreadable exactly where the job runs. It is the single source of truth for where things live: the laptop side resolves the venv, the secrets file and the results from it (with {user} taken from the SSH config’s User line), and the job resolves the same paths from it with its own $USER. $SLURM_STORAGE_HOME is used only when CLUSTER_WORK is left empty.
CLUSTER_MAX_GPUS_PER_JOB and CLUSTER_MAX_CPUS_PER_JOB are ceilings submit.sh enforces on your laptop — an arm asking for more is refused before anything is queued. They are not read from the scheduler; set them to what sacctmgr show qos format=Name,MaxTRESPerJob reports for the QOS you use. CLUSTER_CPU_PARTITION / CLUSTER_CPU_QOS are used only by make env, which deliberately builds on a CPU partition; if the site has none, point both at the GPU partition and its QOS and accept holding a GPU for a pip install.
Then, in your arm file, replace the scheduler names — these differ at every site:
Three assumptions worth re-checking on a new site, because they shaped the design rather than merely configuring it:
- Apptainer runs OCI images unprivileged. If your site has Docker on compute nodes, the job could be simpler. If it has neither, this approach does not port.
- Compute nodes accept no inbound connections. That is why everything is co-located in one job. Where inbound is allowed, a split deployment becomes possible — though a requeued job landing on a different node still breaks every client endpoint.
- Compute nodes have outbound internet. The job pulls container images at start. Without it, pre-stage the SIFs on shared storage and point the job at them.
- The scheduler exports a per-job, node-local
$TMPDIR. The batch job keeps Postgres, its socket, the container images and the working results there and relies on it being wiped with the allocation; the checkpointer writes on every turn, and a shared filesystem degrades for everyone under that pattern. A site that does not set it stops the job at once withTMPDIR is not set: …— after a greenmake gate, which never touches it. Check withsrun --pty printenv TMPDIR; if it is empty, export node-local scratch at the top ofscripts/_env.sh. - Single-node PCIe GPUs without a CUDA toolkit.
scripts/_env.shexportsNCCL_P2P_DISABLE=1andNCCL_IB_DISABLE=1(PCIe A100s without NVLink; an InfiniBand probe that hung) andVLLM_USE_FLASHINFER_SAMPLER=0(nonvccon compute nodes). They cost nothing here but are wrong for NVLink nodes or multi-node jobs. They are not arm keys —submit.shforwards only the arm’s own knobs — so change them in_env.sh, or export them in the login-node shell, which--export=ALLcarries into the job. Downloaded weights land in$WORK/hf_cache(HF_HOME), shared across jobs.
Choosing a model
Every HAI-Co² tool schema carries an injected, requiredaction_and_reasoning argument. A model that omits it does not fail cleanly: _is_blank_turn treats any message bearing a tool call as a real response, so the retry wrapper never sees the problem and the turn loops. This is #184, and it is why the server default moved off gpt-4o-mini.
Two things contain it:
agent.recursion_limit(default 50) bounds the loop. Note that LangGraph does not raise when the limit is reached:create_react_agentsubstitutes a canned reply carrying no tool calls, which ends the graph normally. The backend detects that sentinel and emits anerrorframe instead, or a runaway turn would be reported as a success.make gate EXP=<arm>tells you before a campaign whether a model can satisfy the schema at all. It serves the model, calls itGATE_Ntimes (25 by default, 15 fixed prompts cycled) with a miniature of the real tool schema, and asserts every returned tool call parses and carries a non-emptyaction_and_reasoning. It exits 0 only when at least 95 % pass — atGATE_N=25one failure still passes and two fail — so it can gate a pipeline. Where the verdict and the report land is described under Run your first experiment.
TOOL_PARSER before abandoning it; if it still fails, prefer structured decoding on a model you have already validated over swapping in a new architecture.
Valid TOOL_PARSER values (vLLM 0.23.0 registers 42). The ones you are likely to want: hermes (Qwen2.5/Qwen3 — start here), qwen3_xml, qwen3_coder, llama3_json, llama4_json, llama4_pythonic, mistral, openai (gpt-oss), deepseek_v3/v31/v32/v4, glm45, glm47, granite, granite4, pythonic, xlam, kimi_k2, minimax, seed_oss, olmo3, gemma4, internlm, jamba, phi4_mini_json, step3.
Already staged on the cluster under $SHARED_WEIGHTS (cluster.env), so no download if you point MODEL at the directory: MODEL=$SHARED_WEIGHTS/models--Qwen--Qwen3-32B (62 GB bf16), …/models--Qwen--Qwen3-32B-FP8, …/models--Qwen--Qwen3-4B-Thinking-2507, …/models--mistralai--Mistral-7B-Instruct-v0.2; check for others with ls $SHARED_WEIGHTS/. A bare HuggingFace id such as Qwen/Qwen3-32B also works but re-fetches the weights into $HF_HOME. Note that Qwen3.5-27B is Qwen3_5ForConditionalGeneration — a hybrid-linear vision model, not a drop-in swap for Qwen3-32B.
How it works
Onesbatch job on one GPU node holds everything, all on 127.0.0.1:
redir forwarder on the login node — restricted to ports 5000–5100, unreachable from the TU/HRZ VPN, and killed daily at 03:00. Co-locating removes that entire failure class. Splitting backend-onto-CPU-node buys nothing either: a gpu allocation already grants 16 CPUs and 256 GB, and the backend needs ~2 CPUs and a few GB.
Why Apptainer, and why not for vLLM. Docker exists nowhere on the cluster; Apptainer 1.4.5 is on compute nodes and runs the existing OCI images directly, giving parity with what is deployed. vLLM stays outside a container because it is already installed in $WORK/envs/haico-serving (VLLM_VENV in cluster.env) and needs the GPUs, which would mean --nv passthrough for no parity benefit. (podman on the login node is a dead end: there is no /etc/subuid entry for the account, so rootless UID mapping — and apptainer build --fakeroot — is impossible.)
Apptainer does not namespace the network, so the four processes reach each other on loopback with no compose-style network. The flip side is that loopback is shared with co-tenant jobs, which is why ports derive from $SLURM_JOB_ID and are probed immediately before binding.
How traffic reaches vLLM. Entirely through the per-request config block on POST /api/query/stream_steps/sse — no application change:
POST /api/keys requires a browser session, which a throwaway database cannot provide. The job generates a fresh haico_pat_… per run and passes it in as ADMIN_DEFAULT_API_TOKEN, which migration 0017 installs for the seeded admin. You do not set this anywhere — not in an arm file, not in .env, not on the cluster. If the pinned image predates that migration the job falls back to logging in as the admin and calling POST /api/keys, so it works against any digest. See ADR-0006.
Results
make fetch EXP=<run-id> pulls runs/<run-id>/:
make fetch prints the attrition table, n_with_capture_errors and stopped_early; check the first two before calling a run clean.
Each transcript captures the final state of every workspace panel. document ({thread_id, title, content, updated_at} or null), objective ({thread_id, text, updated_at} or null), todos, preferences and artifacts (lists of rows; artifacts carry turn_index, artifact_type and payload) are the panels as they stood when the conversation ended. messages is the chat history rebuilt from the checkpointer ({role: user|assistant|tool, content, tool_calls?, tool_name?}), which is where the agent’s replies and tool calls live. graph is the conversation DAG with one node per turn (user_message cut to 140 characters, assistant_summary to 200, reasoning_steps, artifact_count). snapshots is the per-turn index only: {id, turn_index, user_message, captured_at}, newest first, at most 50. The per-turn document, objective, todo and preference bodies, and the model_id and provider recorded for each turn, are not in any transcript; they are rows of workspace_snapshots in the database dump, so keep the dump if you need per-turn history.
A conversation can be ok and still have lost a panel. capture_errors on the transcript and on its ledger row names each panel that could not be read (messages: HTTP 500 (after 3 attempts)), and the panel itself holds the reason as a string (HTTP <code>, thread not in this job's database (HTTP 404), or an exception name such as ReadTimeout) instead of data — any string in a panel slot is that marker. n_with_capture_errors in the ledger counts such conversations; treat a non-zero count as an incomplete capture. Within the same job the driver retries with a 120/240/480 s ladder. make resume re-fetches only the missing panels, and only for threads that exist in that job’s fresh database, so a panel lost in an earlier attempt comes back as thread not in this job's database; the only copy of that workspace is that attempt’s kept haico-db.before-job-*.sql.gz.
ledger.json
n_attempted counts every row, including not_attempted, interrupted and campaign_timeout rows, so it equals the size of the prompt set. n_completed is the number of ok rows. outcomes is the attrition table: one count per outcome name. stopped_early is null or one of received SIGTERM / received SIGINT (a cancel, or the pre-wall-clock signal), campaign ceiling of <n>s reached, or provider abort: <n> conversations in a row rejected immediately ('<message>'). Each row carries turns_attempted (turns the driver sent) and turns_completed (turns that ended ok); a failed conversation has turns_attempted = turns_completed + 1, and thread_id is null when turn 0 failed before the server assigned one. The ledger row and transcripts/<id>.json describe the last attempt of a conversation; earlier tries are in superseded/.
transcripts/<id>.json
A turn record does not contain the agent’s reply, its tool calls or its steps: the reply and tool calls are in
workspace.messages, and the individual ReAct steps are recorded only when PHOENIX_ENABLED=1. Token counts are not recorded per turn anywhere — vllm.log carries only aggregate throughput lines — so a study that needs them must enable Phoenix, whose LLM spans carry the provider’s usage fields.
Reading the database dump
haico-db.sql.gz is a plain-SQL pg_dump taken by Postgres 15, so restore it with psql into a server of version 15 or newer (the local stack’s postgres:15-alpine will do):
conversation_users (one row per thread_id, with the final turn_index), workspace_snapshots (one row per (thread_id, turn_index): the document, objective, todos and preferences as the agent saw them at the start of that turn, plus model_id and provider — the per-turn provenance nothing else records), documents, objectives, todos, preferences, artifacts (final state, the same as the transcript panels), and the LangGraph tables checkpoints, checkpoint_blobs, checkpoint_writes (the full message history the messages panel is rebuilt from). api_keys holds only the hash and prefix of the run’s token.
Every attempt starts a fresh database, so haico-db.sql.gz holds only the last attempt’s threads. Conversations completed by an earlier attempt of a resumed run, and every transcript in transcripts/superseded/, reference threads that exist only in the matching haico-db.before-job-<id>-<time>.sql.gz; restore each into its own database. A conversation that was cut mid-turn by a cancel has no transcript, but its partial thread is in the dump that attempt wrote.
Reading the Phoenix traces
phoenix.db.gz is the SQLite database of the Phoenix instance that ran inside the job (arizephoenix/phoenix:13.23.0, the same image the deployed stacks run). Unzip it and serve it with that image or a newer one, pointing it at the file:
haico_slurm, one trace per turn, with the ReAct steps, tool calls and timings that neither the transcript nor the database record; sessions are keyed by thread_id, which is how to find the trace for a transcript.
Attrition is a result, not a cleaning step. Conversations fail most often in exactly the arms where the model is weakest, so they are not missing at random and summarising only the survivors overstates whatever the campaign measures.
ledger.json reports n_attempted, n_completed, and a mutually-exclusive taxonomy: tool_arg_missing, recursion_exhausted, truncated, context_overflow, turn_deadline, client_timeout, stream_idle, transport, server_error, plus the harness outcomes interrupted, campaign_timeout, driver_error and not_attempted. Only the first indicts the model’s ability to drive the agent; recursion_exhausted is a step-ceiling hit with no tool-schema error visible in the turn’s steps. client_timeout is a harness artifact — the driver gave up after READ_TIMEOUT seconds of silence — and means you should raise READ_TIMEOUT or lower CONCURRENCY, not that the model failed.- Resume granularity is the conversation, never the turn. The server increments the turn index and writes a snapshot before the agent runs, so replaying a failed turn would duplicate state. A conversation that fails at any point is abandoned whole.
make resumekeeps only conversations whose transcript endedok; every other one (any failure outcome,interrupted,campaign_timeout,not_attempted, or missing) is run again from a freshthread_id, and its earlier transcript is moved totranscripts/superseded/<id>.attempt-N.jsonso the number of tries stays on record.make submitrefuses aRUN_IDthat already has a ledger;make resumeis the way to continue one. A resume of a finished run does nothing and leaves every transcript byte-identical. - A control arm is your responsibility, and this stack cannot run it.
make submitalways serves a model with vLLM, and the driver sends"API": "openai"with anendpoint_urlon every request, so it cannot target the deployed default (gemini-2.5-flash) or any non-OpenAI provider. For a hosted comparison, bring up a local stack (deployment/local) with the provider key, mint a personal access token, and run the driver by hand with the same prompt set and sampling settings:HAICO_API_TOKEN=… python driver/run_batch.py --api http://localhost:8000 --endpoint-url https://api.openai.com/v1 --model <model> --prompts driver/prompts/<set>.jsonl --out runs/<control-id> --run-id <control-id> --temperature 0.7 --max-tokens 8192 --concurrency 2. That works for OpenAI-compatible endpoints only; a Gemini or Anthropic control arm needs a driver change to send a differentAPIvalue. Record the control model and the date, since a hosted model is not pinnable.
Exit codes
A gate job exits 1 when the pass rate is below the threshold. Codes 2, 3, 4 and 143 all mean the ledger was written and
make resume continues the run; only a job that ends with no ledger.json in $WORK/haico-runs/<RUN_ID> failed in the sense that matters.
Secrets
One long-lived secret exists, and it lives on the cluster — never in the repository, never on your laptop.make secrets writes it there. Nothing else does, and nothing writes a secret locally: the token is piped over the SSH channel to a cat on the far side, so it never appears in your shell history, in a local file, or in the remote process list.
Why that path
Not$HOME. On many clusters — including the one this was built against — $HOME is not mounted on compute nodes. A credential there is unreadable at exactly the moment the job needs it, and the failure looks like a broken token rather than a missing mount.
Not the repository. deployment/slurm/secrets/ is in .gitignore, and secrets/ is excluded from the rsync in both submit.sh and the Makefile. That exclusion is doing two jobs: it stops a local secret being uploaded, and — because rsync --delete does not remove excluded paths on the receiver — it stops every make submit from wiping the token you stored.
What is in it
Inspect what is stored, without printing any secret value:
Everything else is generated per job and dies with it
Nothing below is ever stored, by you or by the tooling. Eachsbatch mints its own from /dev/urandom, and the values vanish when the allocation ends.
This is why there is no
.env to fill in for a cluster run, and why a leaked job log costs you nothing.
How the job handles the one real secret
Three properties, each deliberate:- Sourced inside the job, never exported before
sbatch. Slurm records the submission environment, so a secret placed there is more exposed than one read at runtime. unsetimmediately after the image pull, so it is not in the environment while the batch runs.APPTAINER_DOCKER_USERNAME/PASSWORDare scoped to the private pull only. They apply to every registry, so exporting them globally breaks the Docker Hub pull ofpostgreswithinvalid username/password— which reads like a broken credential rather than a misapplied one.
A second researcher
Each account has its own.$WORK is per-user and 700, so nothing is shared and nothing is inherited — a colleague runs make secrets once with their own token. If they lack read access to the private package, make secrets refuses the token immediately rather than letting them discover it twenty minutes into a job.
To revoke: delete the file (rm $WORK/haico-cluster/secrets/ghcr.env) and revoke the token on GitHub. Runs already in flight are unaffected — they pulled their images at startup.
What differs from haico.gr
Recorded deliberately indeployment/slurm/config/backend.yaml, and worth stating in any write-up:
Troubleshooting
Every symptom below that starts withFATAL: is in the job’s .err log; make logs shows both files.