Skip to main content
Run HAI-Co² against a model you host yourself on a SLURM cluster, driven by a scripted batch client through the real HTTP API. Developed and verified on the UKP Lab cluster at TU Darmstadt, so the concrete values throughout (partition names, GPU types, storage paths) are that site’s. They are examples. Everything site-specific lives in one file — see Adapting to another cluster. This is not how haico.gr is deployed — see Deployment for that, and Local installation for a laptop stack. This page is about running hundreds of scripted conversations against open weights instead of a paid API, so a study is cheap, reproducible, and pinned to an exact backend image — and to exact model weights when you pin those too (see MODEL below). Everything described here lives in deployment/slurm/, and the design rationale is ADR-0009.

Prerequisites

On the cluster
  • An account with access to GPU nodes, and whatever network access the site requires — a VPN, at many.
  • Apptainer or Singularity on the compute nodes. Postgres and the backend run as containers; only vLLM does not. Check before anything else:
    If that fails, this stack will not work here as designed — see Adapting to another cluster.
  • Python 3.9+ as python3 on the compute nodes (srun --pty python3 --version). Verified on 3.12.
  • Roughly 15 GB of quota for the serving environment, plus whatever model weights you download if the site has no shared cache.
On your laptop
  • An SSH host alias for the login/submit node. The tooling never prompts for a password, so key auth must already work:
    Verify with ssh mycluster hostname, then put the alias in cluster.env. The User line is load-bearing, not cosmetic: the laptop-side tooling reads it with ssh -G and substitutes it for {user} in CLUSTER_WORK to decide where the venv, the secrets file and the results live, while the job substitutes its own $USER on the node. Put your cluster username there even if it matches your laptop login.
  • make, bash, rsync, curl and python3 (used by make secrets, make fetch and submit.sh’s local validation).
  • The GitHub CLI (gh), logged in, for make digest only: it resolves an image tag through your gh session, not through the token stored by make secrets. The target refuses to run when gh is missing or logged out.
  • A classic GitHub token with the read:packages scope, to pull the private backend image. Fine-grained tokens are rejected by GHCR with an opaque 401.
The backend image is private. Your GitHub account needs read access to the haico-backend package, which is granted on the repository, not by the token — a token with the right scope still returns 401 without it. make secrets verifies this up front and refuses a token that cannot pull, so you find out in seconds rather than twenty minutes into a job. If it fails, ask a maintainer for package read access.

From zero to a first result

Steps 1–3 are once per account. Everything after is per experiment. make help lists every target.
Step 1 first, or steps 2 and 3 build in the wrong place. cluster.env holds the submit host and the work directory, and both make env and make secrets write to paths derived from it. At the UKP cluster the shipped values are already correct and you can skip straight to step 2.
On any cluster other than UKP, step 4 needs an arm of your own first. The shipped qwen3-32b arm hardcodes PARTITION=gpu, QOS=gpu, GRES=gpu:a180:1 and a model path under UKP’s shared weights. cp arms/qwen3-32b.env arms/<name>.env, set PARTITION, QOS (empty if your site has none), GRES (your site’s GPU type name) and MODEL (a HuggingFace id such as Qwen/Qwen3-32B if SHARED_WEIGHTS is empty, or an absolute path to a flat model directory), then make gate EXP=<name>. See Adapting to another cluster.

make env

Submits a CPU-partition job (CLUSTER_CPU_PARTITION / CLUSTER_CPU_QOS in cluster.env) that creates the venv named by VLLM_VENV and installs env/requirements.txt. A pip install needs no GPU, and holding an A100 to download wheels denies it to someone who does. make env only queues the job and returns at once. Watch it with make logs EXP=haico-env and wait for the line environment ready: … before running make gate: the build takes a few minutes once it starts, plus whatever the CPU queue wait is. make secrets does not depend on it and can run meanwhile. A gate submitted too early dies at source …/bin/activate after its own queue wait. It is idempotent — re-running when the environment is already healthy is a no-op, so it is safe as the first line of any runbook. make env FORCE=1 rebuilds from scratch; make env VENV=/path/to/other builds elsewhere, which is how to test a version bump without disturbing the environment your campaigns use. The build verifies rather than trusting pip’s exit code: it imports vLLM, httpx and torch, and reports whether the tool-call parser registry loads (advisory — make gate is the real test). That matters because a subtly broken install otherwise surfaces much later as “the model cannot emit tool calls” at the top of a campaign. It then writes haico-requirements.lock.txt into the venv — the full resolved closure from pip freeze. requirements.txt states intent; the lock file states what a given result is actually attributable to.
Versions are pinned deliberately. A campaign’s results belong to a model and the stack that served it, so a floating range means a rebuild months later silently serves the same weights through different sampling or a different parser set. Upgrading vLLM is a decision, not housekeeping — the tool-call parser registry changes between releases and every arm file’s TOOL_PARSER names one of its entries, so re-run make gate after any bump.torch is deliberately absent from requirements.txt. vLLM ships its own pin and installs a matching build; adding a torch line fights that and is how you get a venv that imports but cannot see the GPU. It is also why this must stay separate from any training environment — trl/peft pull a different torch.
Another researcher on the same cluster runs make env once and changes nothing else: VLLM_VENV is per-user by default. Point it at a shared project directory if a group wants one environment — accepting that everyone then shares its pins, and an upgrade breaks everyone at once.

make secrets

Prompts for your GitHub username and the token, verifies the token against GHCR without downloading anything, and writes both to $WORK/haico-cluster/secrets/ghcr.env with mode 600. submit.sh checks that this file exists before queueing a batch, because the job dies without it — after the queue wait and the Postgres pull. It deliberately does not offer to reuse your gh CLI token. That credential carries whatever scopes the account granted it — typically repo, workflow, admin:org — and storing it on multi-user cluster storage for the length of a campaign, to save one token mint, trades a broad credential for a small convenience. A dedicated classic token with only read:packages can do nothing but pull the image. It is stored under $WORK and not $HOME for a reason: $HOME is not mounted on compute nodes, so a credential there is unreadable exactly where the job needs it. That is the only long-lived secret. JWT_SECRET, the admin password, the HAI-Co² API token and the vLLM key are all generated inside each job and die with it. HF_TOKEN is optional — needed only for gated repositories such as Llama, and irrelevant for Qwen or Mistral, which are what the shipped arms use. Add one later with bash scripts/bootstrap_secrets.sh --hf: the script rewrites the file as a whole, but it reads the stored values first and Enter at a prompt keeps them, so you are not asked to dig out the GHCR token again; a blank HF answer keeps whatever is stored. Re-running without --hf (to rotate the GHCR token) keeps a stored HF token too.

Run your first experiment

Run make gate first for any model or parser you have not used before. It costs minutes and catches a failure that would otherwise burn the entire allocation silently — see Choosing a model.
Where things land on the cluster. $WORK/haico-cluster/ is this directory, rsynced by every make submit / make gate / make env, plus logs/<jobname>_<jobid>.out|.err (the job logs) and secrets/ghcr.env. $WORK/haico-runs/<RUN_ID>/ holds a run’s results and is what make fetch copies. $WORK/envs/haico-serving is the venv from make env; $WORK/hf_cache, $WORK/.apptainer and $WORK/.cache are the model, image and pip caches shared across jobs. Jobs are named haico-exp-<RUN_ID> (batch), haico-gate-<RUN_ID> (gate) and haico-env (venv build). Two log files per job. Progress goes to .out; every FATAL: diagnostic, the driver’s progress notes and the gate’s GATE FAILED verdict go to .err. make logs tails both. make fetch does not copy the job logs — read them with make logs or ssh while you need them. Reading the gate. The gate job serves the model, makes GATE_N (default 25) tool-calling requests at temperature 0.7 (the arm’s TEMPERATURE is not used) with a 2048-token output cap and a 180 s timeout per call, and passes when at least 95 % return a parseable tool call carrying a non-empty action_and_reasoning. GATE PASSED appears in the .out log, GATE FAILED in .err. The per-prompt report is $WORK/haico-runs/<RUN_ID>/gate-report.json ({model, base_url, n, passed, pass_rate, threshold, results: [{prompt, passed, reason}]}, with gate-vllm.log beside it); make fetch EXP=<RUN_ID> pulls both, and a second gate run on the same arm overwrites them. Read reason before trusting a failure: no tool_calls emitted, action_and_reasoning missing and action_and_reasoning empty are model results; a reason beginning ConnectError, ReadTimeout or HTTPStatusError means the call never reached the model, and the gate says so with THIS IS NOT A MODEL RESULT.

Stopping a run

make cancel EXP=<arm> acts on the jobs named after the arm’s RUN_ID. A queued batch is removed from the queue. A running batch receives TERM on its batch shell only, so the job’s own teardown runs: the driver stops launching turns, records every unfinished conversation as interrupted, writes ledger.json, captures what it can with one short attempt per panel, and the database is dumped while Postgres is still up — allow a few minutes. make cancel EXP=<arm> FORCE=1 is a plain scancel: the whole cgroup is killed at once and nothing newer than the last periodic dump (every DUMP_INTERVAL seconds) survives. Gate jobs are always cancelled outright. Do not scancel a batch by hand unless you mean FORCE=1: it kills Postgres together with the driver, so the final dump fails and the ledger is not written. Afterwards make submit refuses the RUN_ID (it has a ledger); make resume EXP=<arm> continues it. sacct shows a cancelled run as FAILED with exit code 15:0 (Slurm’s rendering of 143, 128 + SIGTERM) on purpose, so a stopped run is never mistaken for a completed one; the results up to the signal are complete.

Changing the experiment

Everything is in one arm file, arms/<name>.env. Every key that shapes the result is required and has no default: submit.sh refuses an arm that omits one. Six keys are optional and fall back to a fixed value when absent — HAICO_READY_TIMEOUT (2400 s), PHOENIX_ENABLED (0), DUMP_INTERVAL (600 s), GATE_N (25), and REASONING_PARSER / EXTRA_VLLM_ARGS (empty). Keep them in the arm anyway so the file records what ran. The driver’s grace period after a signal (150 s) is a job-script constant, not an arm key. The arm file is what you edit; run-config.env in the run directory is what actually ran. Every attempt appends the job id, node, GPU, image digest, every knob, and for a model given as a directory the SHA-256 of its config.json and safetensors index, so a result stays attributable after the arm file has been edited for the next run.
submit.sh validates locally before anything is queued, because a typo caught on your laptop costs a second and the same typo caught by the scheduler costs a queue wait plus a model load. It rejects: missing values, a TENSOR_PARALLEL_SIZE that disagrees with GRES, a request above the ceilings in cluster.env, an output cap that leaves no room for the prompt, a tag where a digest belongs, a prompts file that does not parse (every line must be a JSON object with id and non-empty string turns, ids unique and usable as file names), a missing secrets file on the cluster, a QOS your account does not have, and a RUN_ID that already has a ledger (make resume is the way to continue one). Every make target wraps ./submit.sh <arm> [--gate] [--resume] [--dry-run] [--no-sync]; --no-sync submits without re-syncing the tree to the cluster, for when nothing but the arm has changed.

Arm-file reference

Identity GPU and node selection — all passed straight through to sbatch.
A QOS the site defines is not necessarily one your account may use. At UKP, gpu-large and gpu-small exist but are absent from a plain ukp-researcher association. Find out what is actually yours:
CONCURRENCY=2 is measured, not chosen. The backend serialises requests — vLLM reports Running: 1 reqs in the large majority of samples regardless — so raising concurrency does not raise throughput; it queues conversations inside the backend, where they emit no bytes and the driver eventually abandons them. On identical 8-conversation runs:Lower concurrency was both more reliable and marginally faster. Raise it only once the backend genuinely serves requests in parallel.
vLLM server The batch — what the driver sends, and how it waits. Two ceilings are constants in driver/run_batch.py, not arm keys: a turn is abandoned as turn_deadline after 1800 s in total, and as stream_idle after 900 s of keep-alives with no data frame. Raising READ_TIMEOUT does not move them; a legitimately slow model at high CONCURRENCY can hit the first — lower concurrency, or change the constant and record it with the run.

Prompt sets

PROMPTS names a file under driver/prompts/, one conversation per line: {"id": "pilot-01-report", "turns": ["first user message", "second", …]}. Blank lines and lines starting with # are ignored, so a set can carry notes or have a conversation commented out. id becomes transcripts/<id>.json and the key make resume uses, so ids must be unique within a set and usable as file names; submit.sh checks both before queueing. Keys other than id and turns are ignored. Every turn is sent verbatim as the user message; see driver/prompts/README.md for why each turn must be answerable without a clarifying question. Two sets ship: pilot.jsonl, 8 conversations of 4 turns covering the document, plan, preferences and objective panels, and stress.jsonl, 20 conversations of 6 turns that apply three turn policies (grow: enlarge the document every turn; plan: churn todos and preferences; revise: rewrite repeatedly) across seven topics, built to push on MAX_MODEL_LEN, tool-call volume and snapshot-chain length respectively. Ids encode set, index and policy (stress-04-grow-tutorial), so outcomes can be grouped by policy.

Adapting to another cluster

Everything specific to one site lives in deployment/slurm/cluster.env. In the common case that file is the only thing you change.
CLUSTER_WORK is the one that most often bites. On many clusters $HOME is not mounted on compute nodes, so a credential or cache under it is unreadable exactly where the job runs. It is the single source of truth for where things live: the laptop side resolves the venv, the secrets file and the results from it (with {user} taken from the SSH config’s User line), and the job resolves the same paths from it with its own $USER. $SLURM_STORAGE_HOME is used only when CLUSTER_WORK is left empty. CLUSTER_MAX_GPUS_PER_JOB and CLUSTER_MAX_CPUS_PER_JOB are ceilings submit.sh enforces on your laptop — an arm asking for more is refused before anything is queued. They are not read from the scheduler; set them to what sacctmgr show qos format=Name,MaxTRESPerJob reports for the QOS you use. CLUSTER_CPU_PARTITION / CLUSTER_CPU_QOS are used only by make env, which deliberately builds on a CPU partition; if the site has none, point both at the GPU partition and its QOS and accept holding a GPU for a pip install. Then, in your arm file, replace the scheduler names — these differ at every site: Three assumptions worth re-checking on a new site, because they shaped the design rather than merely configuring it:
  1. Apptainer runs OCI images unprivileged. If your site has Docker on compute nodes, the job could be simpler. If it has neither, this approach does not port.
  2. Compute nodes accept no inbound connections. That is why everything is co-located in one job. Where inbound is allowed, a split deployment becomes possible — though a requeued job landing on a different node still breaks every client endpoint.
  3. Compute nodes have outbound internet. The job pulls container images at start. Without it, pre-stage the SIFs on shared storage and point the job at them.
  4. The scheduler exports a per-job, node-local $TMPDIR. The batch job keeps Postgres, its socket, the container images and the working results there and relies on it being wiped with the allocation; the checkpointer writes on every turn, and a shared filesystem degrades for everyone under that pattern. A site that does not set it stops the job at once with TMPDIR is not set: … — after a green make gate, which never touches it. Check with srun --pty printenv TMPDIR; if it is empty, export node-local scratch at the top of scripts/_env.sh.
  5. Single-node PCIe GPUs without a CUDA toolkit. scripts/_env.sh exports NCCL_P2P_DISABLE=1 and NCCL_IB_DISABLE=1 (PCIe A100s without NVLink; an InfiniBand probe that hung) and VLLM_USE_FLASHINFER_SAMPLER=0 (no nvcc on compute nodes). They cost nothing here but are wrong for NVLink nodes or multi-node jobs. They are not arm keys — submit.sh forwards only the arm’s own knobs — so change them in _env.sh, or export them in the login-node shell, which --export=ALL carries into the job. Downloaded weights land in $WORK/hf_cache (HF_HOME), shared across jobs.
Nothing else is site-specific: the driver, the capability gate, the arm-file format and the ledger are all plain HTTP and JSON.

Choosing a model

Every HAI-Co² tool schema carries an injected, required action_and_reasoning argument. A model that omits it does not fail cleanly: _is_blank_turn treats any message bearing a tool call as a real response, so the retry wrapper never sees the problem and the turn loops. This is #184, and it is why the server default moved off gpt-4o-mini. Two things contain it:
  1. agent.recursion_limit (default 50) bounds the loop. Note that LangGraph does not raise when the limit is reached: create_react_agent substitutes a canned reply carrying no tool calls, which ends the graph normally. The backend detects that sentinel and emits an error frame instead, or a runaway turn would be reported as a success.
  2. make gate EXP=<arm> tells you before a campaign whether a model can satisfy the schema at all. It serves the model, calls it GATE_N times (25 by default, 15 fixed prompts cycled) with a miniature of the real tool schema, and asserts every returned tool call parses and carries a non-empty action_and_reasoning. It exits 0 only when at least 95 % pass — at GATE_N=25 one failure still passes and two fail — so it can gate a pipeline. Where the verdict and the report land is described under Run your first experiment.
If a model fails, try a different TOOL_PARSER before abandoning it; if it still fails, prefer structured decoding on a model you have already validated over swapping in a new architecture.
--enable-auto-tool-choice requires --tool-call-parser, and without both vLLM never emits tool_calls at all. If the gate reports no tool_calls emitted for every prompt, suspect the flags before the model. The job always passes both.
Valid TOOL_PARSER values (vLLM 0.23.0 registers 42). The ones you are likely to want: hermes (Qwen2.5/Qwen3 — start here), qwen3_xml, qwen3_coder, llama3_json, llama4_json, llama4_pythonic, mistral, openai (gpt-oss), deepseek_v3/v31/v32/v4, glm45, glm47, granite, granite4, pythonic, xlam, kimi_k2, minimax, seed_oss, olmo3, gemma4, internlm, jamba, phi4_mini_json, step3. Already staged on the cluster under $SHARED_WEIGHTS (cluster.env), so no download if you point MODEL at the directory: MODEL=$SHARED_WEIGHTS/models--Qwen--Qwen3-32B (62 GB bf16), …/models--Qwen--Qwen3-32B-FP8, …/models--Qwen--Qwen3-4B-Thinking-2507, …/models--mistralai--Mistral-7B-Instruct-v0.2; check for others with ls $SHARED_WEIGHTS/. A bare HuggingFace id such as Qwen/Qwen3-32B also works but re-fetches the weights into $HF_HOME. Note that Qwen3.5-27B is Qwen3_5ForConditionalGeneration — a hybrid-linear vision model, not a drop-in swap for Qwen3-32B.

How it works

One sbatch job on one GPU node holds everything, all on 127.0.0.1:
Why one node. Compute nodes accept no inbound connections, so a backend anywhere else would need a redir forwarder on the login node — restricted to ports 5000–5100, unreachable from the TU/HRZ VPN, and killed daily at 03:00. Co-locating removes that entire failure class. Splitting backend-onto-CPU-node buys nothing either: a gpu allocation already grants 16 CPUs and 256 GB, and the backend needs ~2 CPUs and a few GB. Why Apptainer, and why not for vLLM. Docker exists nowhere on the cluster; Apptainer 1.4.5 is on compute nodes and runs the existing OCI images directly, giving parity with what is deployed. vLLM stays outside a container because it is already installed in $WORK/envs/haico-serving (VLLM_VENV in cluster.env) and needs the GPUs, which would mean --nv passthrough for no parity benefit. (podman on the login node is a dead end: there is no /etc/subuid entry for the account, so rootless UID mapping — and apptainer build --fakeroot — is impossible.) Apptainer does not namespace the network, so the four processes reach each other on loopback with no compose-style network. The flip side is that loopback is shared with co-tenant jobs, which is why ports derive from $SLURM_JOB_ID and are probed immediately before binding. How traffic reaches vLLM. Entirely through the per-request config block on POST /api/query/stream_steps/sse — no application change:
The endpoint hostname must contain vllm. Credential resolution returns VLLM_API_KEY only when the endpoint URL contains that substring, and otherwise falls back to OPENAI_API_KEY. The job bind-mounts an /etc/hosts mapping 127.0.0.1 vllm localhost precisely so this holds.Never set OPENAI_ENDPOINT_URL. It is a global fallback applied to every OpenAI-family model, so setting it as a deployment secret would silently re-route production traffic with no failing health check.
How the API token is obtained. POST /api/keys requires a browser session, which a throwaway database cannot provide. The job generates a fresh haico_pat_… per run and passes it in as ADMIN_DEFAULT_API_TOKEN, which migration 0017 installs for the seeded admin. You do not set this anywhere — not in an arm file, not in .env, not on the cluster. If the pinned image predates that migration the job falls back to logging in as the admin and calling POST /api/keys, so it works against any digest. See ADR-0006.

Results

make fetch EXP=<run-id> pulls runs/<run-id>/: make fetch prints the attrition table, n_with_capture_errors and stopped_early; check the first two before calling a run clean. Each transcript captures the final state of every workspace panel. document ({thread_id, title, content, updated_at} or null), objective ({thread_id, text, updated_at} or null), todos, preferences and artifacts (lists of rows; artifacts carry turn_index, artifact_type and payload) are the panels as they stood when the conversation ended. messages is the chat history rebuilt from the checkpointer ({role: user|assistant|tool, content, tool_calls?, tool_name?}), which is where the agent’s replies and tool calls live. graph is the conversation DAG with one node per turn (user_message cut to 140 characters, assistant_summary to 200, reasoning_steps, artifact_count). snapshots is the per-turn index only: {id, turn_index, user_message, captured_at}, newest first, at most 50. The per-turn document, objective, todo and preference bodies, and the model_id and provider recorded for each turn, are not in any transcript; they are rows of workspace_snapshots in the database dump, so keep the dump if you need per-turn history. A conversation can be ok and still have lost a panel. capture_errors on the transcript and on its ledger row names each panel that could not be read (messages: HTTP 500 (after 3 attempts)), and the panel itself holds the reason as a string (HTTP <code>, thread not in this job's database (HTTP 404), or an exception name such as ReadTimeout) instead of data — any string in a panel slot is that marker. n_with_capture_errors in the ledger counts such conversations; treat a non-zero count as an incomplete capture. Within the same job the driver retries with a 120/240/480 s ladder. make resume re-fetches only the missing panels, and only for threads that exist in that job’s fresh database, so a panel lost in an earlier attempt comes back as thread not in this job's database; the only copy of that workspace is that attempt’s kept haico-db.before-job-*.sql.gz.

ledger.json

n_attempted counts every row, including not_attempted, interrupted and campaign_timeout rows, so it equals the size of the prompt set. n_completed is the number of ok rows. outcomes is the attrition table: one count per outcome name. stopped_early is null or one of received SIGTERM / received SIGINT (a cancel, or the pre-wall-clock signal), campaign ceiling of <n>s reached, or provider abort: <n> conversations in a row rejected immediately ('<message>'). Each row carries turns_attempted (turns the driver sent) and turns_completed (turns that ended ok); a failed conversation has turns_attempted = turns_completed + 1, and thread_id is null when turn 0 failed before the server assigned one. The ledger row and transcripts/<id>.json describe the last attempt of a conversation; earlier tries are in superseded/.

transcripts/<id>.json

A turn record does not contain the agent’s reply, its tool calls or its steps: the reply and tool calls are in workspace.messages, and the individual ReAct steps are recorded only when PHOENIX_ENABLED=1. Token counts are not recorded per turn anywhere — vllm.log carries only aggregate throughput lines — so a study that needs them must enable Phoenix, whose LLM spans carry the provider’s usage fields.

Reading the database dump

haico-db.sql.gz is a plain-SQL pg_dump taken by Postgres 15, so restore it with psql into a server of version 15 or newer (the local stack’s postgres:15-alpine will do):
Tables worth knowing: conversation_users (one row per thread_id, with the final turn_index), workspace_snapshots (one row per (thread_id, turn_index): the document, objective, todos and preferences as the agent saw them at the start of that turn, plus model_id and provider — the per-turn provenance nothing else records), documents, objectives, todos, preferences, artifacts (final state, the same as the transcript panels), and the LangGraph tables checkpoints, checkpoint_blobs, checkpoint_writes (the full message history the messages panel is rebuilt from). api_keys holds only the hash and prefix of the run’s token. Every attempt starts a fresh database, so haico-db.sql.gz holds only the last attempt’s threads. Conversations completed by an earlier attempt of a resumed run, and every transcript in transcripts/superseded/, reference threads that exist only in the matching haico-db.before-job-<id>-<time>.sql.gz; restore each into its own database. A conversation that was cut mid-turn by a cancel has no transcript, but its partial thread is in the dump that attempt wrote.

Reading the Phoenix traces

phoenix.db.gz is the SQLite database of the Phoenix instance that ran inside the job (arizephoenix/phoenix:13.23.0, the same image the deployed stacks run). Unzip it and serve it with that image or a newer one, pointing it at the file:
The traces are in project haico_slurm, one trace per turn, with the ReAct steps, tool calls and timings that neither the transcript nor the database record; sessions are keyed by thread_id, which is how to find the trace for a transcript.
Attrition is a result, not a cleaning step. Conversations fail most often in exactly the arms where the model is weakest, so they are not missing at random and summarising only the survivors overstates whatever the campaign measures. ledger.json reports n_attempted, n_completed, and a mutually-exclusive taxonomy: tool_arg_missing, recursion_exhausted, truncated, context_overflow, turn_deadline, client_timeout, stream_idle, transport, server_error, plus the harness outcomes interrupted, campaign_timeout, driver_error and not_attempted. Only the first indicts the model’s ability to drive the agent; recursion_exhausted is a step-ceiling hit with no tool-schema error visible in the turn’s steps. client_timeout is a harness artifact — the driver gave up after READ_TIMEOUT seconds of silence — and means you should raise READ_TIMEOUT or lower CONCURRENCY, not that the model failed.
Two design points worth knowing when reading results:
  • Resume granularity is the conversation, never the turn. The server increments the turn index and writes a snapshot before the agent runs, so replaying a failed turn would duplicate state. A conversation that fails at any point is abandoned whole. make resume keeps only conversations whose transcript ended ok; every other one (any failure outcome, interrupted, campaign_timeout, not_attempted, or missing) is run again from a fresh thread_id, and its earlier transcript is moved to transcripts/superseded/<id>.attempt-N.json so the number of tries stays on record. make submit refuses a RUN_ID that already has a ledger; make resume is the way to continue one. A resume of a finished run does nothing and leaves every transcript byte-identical.
  • A control arm is your responsibility, and this stack cannot run it. make submit always serves a model with vLLM, and the driver sends "API": "openai" with an endpoint_url on every request, so it cannot target the deployed default (gemini-2.5-flash) or any non-OpenAI provider. For a hosted comparison, bring up a local stack (deployment/local) with the provider key, mint a personal access token, and run the driver by hand with the same prompt set and sampling settings: HAICO_API_TOKEN=… python driver/run_batch.py --api http://localhost:8000 --endpoint-url https://api.openai.com/v1 --model <model> --prompts driver/prompts/<set>.jsonl --out runs/<control-id> --run-id <control-id> --temperature 0.7 --max-tokens 8192 --concurrency 2. That works for OpenAI-compatible endpoints only; a Gemini or Anthropic control arm needs a driver change to send a different API value. Record the control model and the date, since a hosted model is not pinnable.

Exit codes

A gate job exits 1 when the pass rate is below the threshold. Codes 2, 3, 4 and 143 all mean the ledger was written and make resume continues the run; only a job that ends with no ledger.json in $WORK/haico-runs/<RUN_ID> failed in the sense that matters.

Secrets

One long-lived secret exists, and it lives on the cluster — never in the repository, never on your laptop.
make secrets writes it there. Nothing else does, and nothing writes a secret locally: the token is piped over the SSH channel to a cat on the far side, so it never appears in your shell history, in a local file, or in the remote process list.

Why that path

Not $HOME. On many clusters — including the one this was built against — $HOME is not mounted on compute nodes. A credential there is unreadable at exactly the moment the job needs it, and the failure looks like a broken token rather than a missing mount. Not the repository. deployment/slurm/secrets/ is in .gitignore, and secrets/ is excluded from the rsync in both submit.sh and the Makefile. That exclusion is doing two jobs: it stops a local secret being uploaded, and — because rsync --delete does not remove excluded paths on the receiver — it stops every make submit from wiping the token you stored.

What is in it

Inspect what is stored, without printing any secret value:

Everything else is generated per job and dies with it

Nothing below is ever stored, by you or by the tooling. Each sbatch mints its own from /dev/urandom, and the values vanish when the allocation ends. This is why there is no .env to fill in for a cluster run, and why a leaked job log costs you nothing.

How the job handles the one real secret

Three properties, each deliberate:
  • Sourced inside the job, never exported before sbatch. Slurm records the submission environment, so a secret placed there is more exposed than one read at runtime.
  • unset immediately after the image pull, so it is not in the environment while the batch runs.
  • APPTAINER_DOCKER_USERNAME/PASSWORD are scoped to the private pull only. They apply to every registry, so exporting them globally breaks the Docker Hub pull of postgres with invalid username/password — which reads like a broken credential rather than a misapplied one.
Use a dedicated classic token with only read:packages. make secrets warns if you paste a gho_ token — that is the gh CLI credential, which typically carries repo, workflow and admin:org. It would work, and that is the problem: a token that can delete repositories should not sit on multi-user storage for the length of a campaign, whatever the file mode.

A second researcher

Each account has its own. $WORK is per-user and 700, so nothing is shared and nothing is inherited — a colleague runs make secrets once with their own token. If they lack read access to the private package, make secrets refuses the token immediately rather than letting them discover it twenty minutes into a job. To revoke: delete the file (rm $WORK/haico-cluster/secrets/ghcr.env) and revoke the token on GitHub. Runs already in flight are unaffected — they pulled their images at startup.

What differs from haico.gr

Recorded deliberately in deployment/slurm/config/backend.yaml, and worth stating in any write-up:

Troubleshooting

Every symptom below that starts with FATAL: is in the job’s .err log; make logs shows both files.