> ## Documentation Index
> Fetch the complete documentation index at: https://dev.haico.gr/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# SLURM experiments

> Run mass HAI-Co² experiments against self-hosted open-weight models on any SLURM cluster: one self-contained job, Apptainer instead of Docker, and one file per experiment arm.

Run HAI-Co² against a **model you host yourself** on a **SLURM cluster**, driven by a scripted batch client through the real HTTP API.

Developed and verified on the UKP Lab cluster at TU Darmstadt, so the concrete values throughout (partition names, GPU types, storage paths) are that site's. They are examples. Everything site-specific lives in one file — see [Adapting to another cluster](#adapting-to-another-cluster).

This is not how haico.gr is deployed — see [Deployment](/docs/deployment) for that, and [Local installation](/docs/local-installation) for a laptop stack. This page is about running hundreds of scripted conversations against open weights instead of a paid API, so a study is cheap, reproducible, and pinned to an exact backend image — and to exact model weights when you pin those too (see `MODEL` below).

Everything described here lives in [`deployment/slurm/`](https://github.com/petrosrapto/HAICO/blob/main/deployment/slurm), and the design rationale is [ADR-0009](/docs/adr/0009-cluster-hosted-llm-experiments).

## Prerequisites

**On the cluster**

* An account with access to GPU nodes, and whatever network access the site requires — a VPN, at many.
* **Apptainer or Singularity on the compute nodes.** Postgres and the backend run as containers; only vLLM does not. Check before anything else:
  ```bash theme={null}
  srun --partition=<your-gpu-partition> --pty apptainer --version
  ```
  If that fails, this stack will not work here as designed — see [Adapting to another cluster](#adapting-to-another-cluster).
* **Python 3.9+** as `python3` on the compute nodes (`srun --pty python3 --version`). Verified on 3.12.
* Roughly **15 GB of quota** for the serving environment, plus whatever model weights you download if the site has no shared cache.

**On your laptop**

* An SSH host alias for the login/submit node. The tooling never prompts for a password, so key auth must already work:
  ```
  # ~/.ssh/config
  Host mycluster
      HostName login.cluster.example.edu
      User your-username
      IdentityFile ~/.ssh/your-key
      ServerAliveInterval 60
  ```
  Verify with `ssh mycluster hostname`, then put the alias in `cluster.env`. The `User` line is load-bearing, not cosmetic: the laptop-side tooling reads it with `ssh -G` and substitutes it for `{user}` in `CLUSTER_WORK` to decide where the venv, the secrets file and the results live, while the job substitutes its own `$USER` on the node. Put your cluster username there even if it matches your laptop login.
* `make`, `bash`, `rsync`, `curl` and `python3` (used by `make secrets`, `make fetch` and `submit.sh`'s local validation).
* The GitHub CLI (`gh`), logged in, for `make digest` only: it resolves an image tag through your `gh` session, not through the token stored by `make secrets`. The target refuses to run when `gh` is missing or logged out.
* A **classic** GitHub token with the `read:packages` scope, to pull the private backend image. Fine-grained tokens are rejected by GHCR with an opaque `401`.

<Warning>
  **The backend image is private.** Your GitHub account needs read access to the `haico-backend` package, which is granted on the repository, not by the token — a token with the right scope still returns `401` without it. `make secrets` verifies this up front and refuses a token that cannot pull, so you find out in seconds rather than twenty minutes into a job. If it fails, ask a maintainer for package read access.
</Warning>

## From zero to a first result

```bash theme={null}
git clone https://github.com/petrosrapto/HAICO.git
cd HAICO/deployment/slurm

$EDITOR cluster.env              # 1. point it at your cluster  (skip at UKP)
make env                         # 2. queue the venv build job; wait for it  (once)
make secrets                     # 3. store the GHCR pull token   (once)

make gate   EXP=qwen3-32b        # 4. can this model drive the agent at all?
make submit EXP=qwen3-32b        # 5. run the batch
make status                      # 6. watch it
make fetch  EXP=qwen3-32b-pilot  # 7. pull results into ./runs/
```

Steps 1–3 are once per account. Everything after is per experiment. `make help` lists every target.

<Note>
  **Step 1 first, or steps 2 and 3 build in the wrong place.** `cluster.env` holds the submit host and the work directory, and both `make env` and `make secrets` write to paths derived from it. At the UKP cluster the shipped values are already correct and you can skip straight to step 2.
</Note>

<Note>
  **On any cluster other than UKP, step 4 needs an arm of your own first.** The shipped `qwen3-32b` arm hardcodes `PARTITION=gpu`, `QOS=gpu`, `GRES=gpu:a180:1` and a model path under UKP's shared weights. `cp arms/qwen3-32b.env arms/<name>.env`, set `PARTITION`, `QOS` (empty if your site has none), `GRES` (your site's GPU type name) and `MODEL` (a HuggingFace id such as `Qwen/Qwen3-32B` if `SHARED_WEIGHTS` is empty, or an absolute path to a flat model directory), then `make gate EXP=<name>`. See [Adapting to another cluster](#adapting-to-another-cluster).
</Note>

### `make env`

Submits a **CPU-partition** job (`CLUSTER_CPU_PARTITION` / `CLUSTER_CPU_QOS` in `cluster.env`) that creates the venv named by `VLLM_VENV` and installs [`env/requirements.txt`](https://github.com/petrosrapto/HAICO/blob/main/deployment/slurm/env/requirements.txt). A `pip install` needs no GPU, and holding an A100 to download wheels denies it to someone who does.

`make env` only queues the job and returns at once. Watch it with `make logs EXP=haico-env` and wait for the line `environment ready: …` before running `make gate`: the build takes a few minutes once it starts, plus whatever the CPU queue wait is. `make secrets` does not depend on it and can run meanwhile. A gate submitted too early dies at `source …/bin/activate` after its own queue wait.

It is **idempotent** — re-running when the environment is already healthy is a no-op, so it is safe as the first line of any runbook. `make env FORCE=1` rebuilds from scratch; `make env VENV=/path/to/other` builds elsewhere, which is how to test a version bump without disturbing the environment your campaigns use.

The build **verifies rather than trusting pip's exit code**: it imports vLLM, httpx and torch, and reports whether the tool-call parser registry loads (advisory — `make gate` is the real test). That matters because a subtly broken install otherwise surfaces much later as *"the model cannot emit tool calls"* at the top of a campaign.

It then writes `haico-requirements.lock.txt` into the venv — the full resolved closure from `pip freeze`. `requirements.txt` states intent; the lock file states what a given result is actually attributable to.

<Note>
  **Versions are pinned deliberately.** A campaign's results belong to a model *and* the stack that served it, so a floating range means a rebuild months later silently serves the same weights through different sampling or a different parser set. Upgrading vLLM is a decision, not housekeeping — the tool-call parser registry changes between releases and every arm file's `TOOL_PARSER` names one of its entries, so **re-run `make gate` after any bump**.

  **torch is deliberately absent from `requirements.txt`.** vLLM ships its own pin and installs a matching build; adding a torch line fights that and is how you get a venv that imports but cannot see the GPU. It is also why this must stay separate from any training environment — `trl`/`peft` pull a different torch.
</Note>

Another researcher on the same cluster runs `make env` once and changes nothing else: `VLLM_VENV` is per-user by default. Point it at a shared project directory if a group wants one environment — accepting that everyone then shares its pins, and an upgrade breaks everyone at once.

### `make secrets`

Prompts for your GitHub username and the token, verifies the token against GHCR without downloading anything, and writes both to `$WORK/haico-cluster/secrets/ghcr.env` with mode `600`. `submit.sh` checks that this file exists before queueing a batch, because the job dies without it — after the queue wait and the Postgres pull.

It deliberately does **not** offer to reuse your `gh` CLI token. That credential carries whatever scopes the account granted it — typically `repo`, `workflow`, `admin:org` — and storing it on multi-user cluster storage for the length of a campaign, to save one token mint, trades a broad credential for a small convenience. A dedicated classic token with only `read:packages` can do nothing but pull the image.

It is stored under `$WORK` and not `$HOME` for a reason: **`$HOME` is not mounted on compute nodes**, so a credential there is unreadable exactly where the job needs it.

That is the only long-lived secret. `JWT_SECRET`, the admin password, the HAI-Co² API token and the vLLM key are all generated inside each job and die with it. `HF_TOKEN` is **optional** — needed only for *gated* repositories such as Llama, and irrelevant for Qwen or Mistral, which are what the shipped arms use. Add one later with `bash scripts/bootstrap_secrets.sh --hf`: the script rewrites the file as a whole, but it reads the stored values first and Enter at a prompt keeps them, so you are not asked to dig out the GHCR token again; a blank HF answer keeps whatever is stored. Re-running without `--hf` (to rotate the GHCR token) keeps a stored HF token too.

## Run your first experiment

```bash theme={null}
make gate   EXP=qwen3-32b        # can this model drive the agent at all?
make submit EXP=qwen3-32b        # run the batch
make status                      # squeue --me
make logs   EXP=qwen3-32b        # tail the newest .out and .err log for that arm's RUN_ID
make cancel EXP=qwen3-32b        # stop it: graceful for a running batch, removed if still queued
make resume EXP=qwen3-32b        # continue a RUN_ID: ok conversations kept, the rest re-run
make fetch  EXP=qwen3-32b-pilot  # pull results into ./runs/  (takes the RUN_ID, not the arm)
```

<Warning>
  Run `make gate` first for any model or parser you have not used before. It costs minutes and catches a failure that would otherwise burn the entire allocation silently — see [Choosing a model](#choosing-a-model).
</Warning>

**Where things land on the cluster.** `$WORK/haico-cluster/` is this directory, rsynced by every `make submit` / `make gate` / `make env`, plus `logs/<jobname>_<jobid>.out|.err` (the job logs) and `secrets/ghcr.env`. `$WORK/haico-runs/<RUN_ID>/` holds a run's results and is what `make fetch` copies. `$WORK/envs/haico-serving` is the venv from `make env`; `$WORK/hf_cache`, `$WORK/.apptainer` and `$WORK/.cache` are the model, image and pip caches shared across jobs. Jobs are named `haico-exp-<RUN_ID>` (batch), `haico-gate-<RUN_ID>` (gate) and `haico-env` (venv build).

**Two log files per job.** Progress goes to `.out`; every `FATAL:` diagnostic, the driver's progress notes and the gate's `GATE FAILED` verdict go to `.err`. `make logs` tails both. `make fetch` does not copy the job logs — read them with `make logs` or `ssh` while you need them.

**Reading the gate.** The gate job serves the model, makes `GATE_N` (default 25) tool-calling requests at temperature 0.7 (the arm's `TEMPERATURE` is not used) with a 2048-token output cap and a 180 s timeout per call, and passes when at least 95 % return a parseable tool call carrying a non-empty `action_and_reasoning`. `GATE PASSED` appears in the `.out` log, `GATE FAILED` in `.err`. The per-prompt report is `$WORK/haico-runs/<RUN_ID>/gate-report.json` (`{model, base_url, n, passed, pass_rate, threshold, results: [{prompt, passed, reason}]}`, with `gate-vllm.log` beside it); `make fetch EXP=<RUN_ID>` pulls both, and a second gate run on the same arm overwrites them. Read `reason` before trusting a failure: `no tool_calls emitted`, `action_and_reasoning missing` and `action_and_reasoning empty` are model results; a reason beginning `ConnectError`, `ReadTimeout` or `HTTPStatusError` means the call never reached the model, and the gate says so with `THIS IS NOT A MODEL RESULT`.

### Stopping a run

`make cancel EXP=<arm>` acts on the jobs named after the arm's `RUN_ID`. A queued batch is removed from the queue. A running batch receives `TERM` on its batch shell only, so the job's own teardown runs: the driver stops launching turns, records every unfinished conversation as `interrupted`, writes `ledger.json`, captures what it can with one short attempt per panel, and the database is dumped while Postgres is still up — allow a few minutes. `make cancel EXP=<arm> FORCE=1` is a plain `scancel`: the whole cgroup is killed at once and nothing newer than the last periodic dump (every `DUMP_INTERVAL` seconds) survives. Gate jobs are always cancelled outright. Do not `scancel` a batch by hand unless you mean `FORCE=1`: it kills Postgres together with the driver, so the final dump fails and the ledger is not written.

Afterwards `make submit` refuses the `RUN_ID` (it has a ledger); `make resume EXP=<arm>` continues it. `sacct` shows a cancelled run as `FAILED` with exit code `15:0` (Slurm's rendering of 143, 128 + SIGTERM) on purpose, so a stopped run is never mistaken for a completed one; the results up to the signal are complete.

## Changing the experiment

Everything is in one arm file, `arms/<name>.env`. Every key that shapes the result is required and has no default: `submit.sh` refuses an arm that omits one. Six keys are optional and fall back to a fixed value when absent — `HAICO_READY_TIMEOUT` (2400 s), `PHOENIX_ENABLED` (0), `DUMP_INTERVAL` (600 s), `GATE_N` (25), and `REASONING_PARSER` / `EXTRA_VLLM_ARGS` (empty). Keep them in the arm anyway so the file records what ran. The driver's grace period after a signal (150 s) is a job-script constant, not an arm key.

The arm file is what you edit; `run-config.env` in the run directory is what actually ran. Every attempt appends the job id, node, GPU, image digest, every knob, and for a model given as a directory the SHA-256 of its `config.json` and safetensors index, so a result stays attributable after the arm file has been edited for the next run.

```bash theme={null}
cp arms/_template.env arms/my-run.env
$EDITOR arms/my-run.env
make dry-run EXP=my-run     # validate and print the sbatch line; submit nothing
make submit  EXP=my-run
```

`submit.sh` validates locally before anything is queued, because a typo caught on your laptop costs a second and the same typo caught by the scheduler costs a queue wait plus a model load. It rejects: missing values, a `TENSOR_PARALLEL_SIZE` that disagrees with `GRES`, a request above the ceilings in `cluster.env`, an output cap that leaves no room for the prompt, a tag where a digest belongs, a prompts file that does not parse (every line must be a JSON object with `id` and non-empty string `turns`, ids unique and usable as file names), a missing secrets file on the cluster, a QOS your account does not have, and a `RUN_ID` that already has a ledger (`make resume` is the way to continue one).

Every `make` target wraps `./submit.sh <arm> [--gate] [--resume] [--dry-run] [--no-sync]`; `--no-sync` submits without re-syncing the tree to the cluster, for when nothing but the arm has changed.

### Arm-file reference

**Identity**

| Key | Notes |
| - | - |
| `RUN_ID` | Names the run directory and ledger row. Must be unique; results are evidence, so a re-run cannot overwrite them. |
| `MODEL` | HuggingFace id, or an absolute cluster path to a flat model directory (`$SHARED_WEIGHTS/models--Qwen--Qwen3-32B` at UKP; a bare id re-fetches tens of GB into `$HF_HOME`). vLLM serves bf16/fp16 only — merge a QLoRA adapter first, a bnb-4bit checkpoint will not serve. The image is pinned by digest; the weights are pinned only if you pin them: with an id, vLLM serves the repository's current `main`, so add `EXTRA_VLLM_ARGS=--revision <commit sha>`; with a directory, `run-config.env` records the SHA-256 of its `config.json` and safetensors index. |
| `BACKEND_DIGEST` | Pin the backend image by digest, never a tag: a tag can be repointed, silently changing what a "reproduced" run ran. Resolve one with `make digest TAG=e1023eb` (needs a logged-in `gh`; `IMAGE=<owner>/<name>` for a fork). |

**GPU and node selection** — all passed straight through to `sbatch`.

| Key | Notes |
| - | - |
| `PARTITION` | **Site-specific.** `sinfo -o "%P %a %l %D %G"`. At UKP: `gpu` (`PreemptMode=OFF`) or `yolo` (every node, but preemptible with `GraceTime=0` — killed mid-token, no warning, requeued from the start). Leave empty to omit the flag. |
| `QOS` | **Site- *and* account-specific**, and optional — many sites need none. Leave empty to omit it. `submit.sh` checks it against your account before queueing, because Slurm otherwise rejects a bad one with a terse `Invalid qos specification`. |
| `GRES` | `<type>:<count>`. `a180` is A100-80GB (penelope, rubeus) and the workhorse; also `h200nvl` 140 GB, `b6000` 96 GB, `a6000`/`l40` 48 GB, `a100` 40 GB. |
| `CONSTRAINT`, `NODELIST`, `EXCLUDE` | Optional. Feature constraint, host pinning, host avoidance. |
| `CPUS`, `MEM`, `TIME` | `CPUS` caps at `CLUSTER_MAX_CPUS_PER_JOB` (16 at UKP). Keep `MEM` honest: **memory, not GPUs, is what limits how many array tasks co-schedule** against the 512 GB/user cap. `TIME` is the whole allocation: budget it as model load (bounded by `HAICO_READY_TIMEOUT`, \~15–25 min for a 32B on a cold JIT cache) plus the campaign plus ten minutes. The driver stops launching turns ten minutes before the end and cancels every conversation still in flight (each becomes a `campaign_timeout` row, re-run by `make resume`), so the capture pass and the final `pg_dump` run with Postgres alive; Slurm's own `TERM` follows five minutes before the end, and the driver gets 150 s to write the ledger. If fewer than twenty minutes remain when the batch starts, no ceiling is armed and only the `TERM` path protects the ledger. |

<Warning>
  **A QOS the *site* defines is not necessarily one your *account* may use.** At UKP, `gpu-large` and `gpu-small` exist but are absent from a plain `ukp-researcher` association. Find out what is actually yours:

  ```bash theme={null}
  ssh <cluster> 'sacctmgr -nP show assoc user=$USER format=Account,Partition,QOS'
  ```
</Warning>

<Note>
  **`CONCURRENCY=2` is measured, not chosen.** The backend serialises requests — vLLM reports `Running: 1 reqs` in the large majority of samples regardless — so raising concurrency does not raise throughput; it queues conversations inside the backend, where they emit no bytes and the driver eventually abandons them. On identical 8-conversation runs:

  | `CONCURRENCY` | completed | failures | wall clock | slowest conversation |
  | - | - | - | - | - |
  | 4 | 6/8 | 2 × `client_timeout` | 53m32s | 1047s |
  | 2 | **8/8** | **0** | **50m27s** | **609s** |

  Lower concurrency was both more reliable and marginally faster. Raise it only once the backend genuinely serves requests in parallel.
</Note>

**vLLM server**

| Key | Notes |
| - | - |
| `TENSOR_PARALLEL_SIZE` | Must equal the GPU count in `GRES`; `submit.sh` enforces it. |
| `MAX_MODEL_LEN` | Total context. The agent's prompt carries the whole working document every turn, so this must comfortably exceed prompt + `MAX_TOKENS`. |
| `MAX_TOKENS` | Output cap per turn. Budget roughly a quarter of the window on a 32k model. The deployed environments use 32768 because their hosted models have far larger contexts. |
| `TOOL_PARSER` | How vLLM turns the model's output into OpenAI `tool_calls`. See below. |
| `REASONING_PARSER` | For models that emit `<think>` blocks. Empty to disable. |
| `GPU_MEMORY_UTILIZATION`, `MAX_NUM_SEQS`, `DTYPE`, `EXTRA_VLLM_ARGS` | Passed through; `EXTRA_VLLM_ARGS` is appended verbatim. |
| `HAICO_READY_TIMEOUT` | Seconds to wait for the server (default 2400). Dominated by `torch.compile` and CUDA-graph capture, not weight loading; the JIT caches are per job, so every run compiles cold. |

**The batch** — what the driver sends, and how it waits.

| Key | Notes |
| - | - |
| `PROMPTS` | File under `driver/prompts/`; see [Prompt sets](#prompt-sets). Required; validated locally. |
| `TEMPERATURE` | Sent in every request's `config.args`, with `MAX_TOKENS`. Part of the record of a run: two arms with different values are different experiments. Not used by `make gate`, which samples at 0.7. Required. |
| `CONCURRENCY` | Conversations in flight at once; `2` is the measured optimum on this stack (see the Note above). Each open turn also holds one unpooled Postgres connection. Required. |
| `READ_TIMEOUT` | Seconds of socket silence before the driver abandons a turn as `client_timeout`; also passed to the backend as the provider call's timeout. Heartbeats only start once the agent streams, so the window before the first frame is what this bounds. Required. |
| `DUMP_INTERVAL` | Seconds between periodic mirrors of the transcripts plus a `pg_dump` to `$WORK/haico-runs/<RUN_ID>` (default 600). These are what survive a preemption or a `FORCE=1` cancel; lower it for short validation runs. |
| `GATE_N` | Tool-calling probes `make gate` sends before scoring (default 25); the pass threshold is 95 %. Ignored by `make submit`. |
| `PHOENIX_ENABLED` | `1` runs Phoenix (`arizephoenix/phoenix:13.23.0`, the deployed pin) alongside the stack and traces every turn into it over OTLP/HTTP on the UI port (`/v1/traces`); its database is copied out at teardown. |

Two ceilings are constants in `driver/run_batch.py`, not arm keys: a turn is abandoned as `turn_deadline` after 1800 s in total, and as `stream_idle` after 900 s of keep-alives with no data frame. Raising `READ_TIMEOUT` does not move them; a legitimately slow model at high `CONCURRENCY` can hit the first — lower concurrency, or change the constant and record it with the run.

### Prompt sets

`PROMPTS` names a file under `driver/prompts/`, one conversation per line: `{"id": "pilot-01-report", "turns": ["first user message", "second", …]}`. Blank lines and lines starting with `#` are ignored, so a set can carry notes or have a conversation commented out. `id` becomes `transcripts/<id>.json` and the key `make resume` uses, so ids must be unique within a set and usable as file names; `submit.sh` checks both before queueing. Keys other than `id` and `turns` are ignored. Every turn is sent verbatim as the user message; see `driver/prompts/README.md` for why each turn must be answerable without a clarifying question.

Two sets ship: `pilot.jsonl`, 8 conversations of 4 turns covering the document, plan, preferences and objective panels, and `stress.jsonl`, 20 conversations of 6 turns that apply three turn policies (`grow`: enlarge the document every turn; `plan`: churn todos and preferences; `revise`: rewrite repeatedly) across seven topics, built to push on `MAX_MODEL_LEN`, tool-call volume and snapshot-chain length respectively. Ids encode set, index and policy (`stress-04-grow-tutorial`), so outcomes can be grouped by policy.

## Adapting to another cluster

Everything specific to one site lives in **`deployment/slurm/cluster.env`**. In the common case that file is the only thing you change.

```bash theme={null}
CLUSTER_HOST=ukp                                  # ssh alias for the submit node
CLUSTER_WORK=/storage/ukp/work/{user}             # readable from COMPUTE nodes
VLLM_VENV={work}/envs/haico-serving               # venv containing vLLM
SHARED_WEIGHTS=/storage/ukp/shared/shared_model_weights   # optional model cache
CLUSTER_MAX_GPUS_PER_JOB=4                        # enforced by submit.sh, see below
CLUSTER_MAX_CPUS_PER_JOB=16
CLUSTER_CPU_PARTITION=cpu                         # where `make env` builds the venv
CLUSTER_CPU_QOS=cpu                               # its QOS; always passed, so name one your account has
```

`CLUSTER_WORK` is the one that most often bites. On many clusters **`$HOME` is not mounted on compute nodes**, so a credential or cache under it is unreadable exactly where the job runs. It is the single source of truth for where things live: the laptop side resolves the venv, the secrets file and the results from it (with `{user}` taken from the SSH config's `User` line), and the job resolves the same paths from it with its own `$USER`. `$SLURM_STORAGE_HOME` is used only when `CLUSTER_WORK` is left empty.

`CLUSTER_MAX_GPUS_PER_JOB` and `CLUSTER_MAX_CPUS_PER_JOB` are ceilings `submit.sh` enforces on your laptop — an arm asking for more is refused before anything is queued. They are not read from the scheduler; set them to what `sacctmgr show qos format=Name,MaxTRESPerJob` reports for the QOS you use. `CLUSTER_CPU_PARTITION` / `CLUSTER_CPU_QOS` are used only by `make env`, which deliberately builds on a CPU partition; if the site has none, point both at the GPU partition and its QOS and accept holding a GPU for a `pip install`.

Then, in your arm file, replace the scheduler names — these differ at every site:

| Key | How to find your site's value |
| - | - |
| `PARTITION` | `sinfo -o "%P %a %l %D %G"` |
| `QOS` | `sacctmgr show qos format=Name,MaxTRESPerJob,MaxWall` |
| `GRES` | `sinfo -o "%n %G"` — GPU type names are site-defined (`a180`, `a100`, `h100`, …) |
| `CPUS`, `MEM`, `TIME` | whatever the QOS above permits |

**Three assumptions worth re-checking on a new site**, because they shaped the design rather than merely configuring it:

1. **Apptainer runs OCI images unprivileged.** If your site has Docker on compute nodes, the job could be simpler. If it has neither, this approach does not port.
2. **Compute nodes accept no inbound connections.** That is why everything is co-located in one job. Where inbound *is* allowed, a split deployment becomes possible — though a requeued job landing on a different node still breaks every client endpoint.
3. **Compute nodes have outbound internet.** The job pulls container images at start. Without it, pre-stage the SIFs on shared storage and point the job at them.
4. **The scheduler exports a per-job, node-local `$TMPDIR`.** The batch job keeps Postgres, its socket, the container images and the working results there and relies on it being wiped with the allocation; the checkpointer writes on every turn, and a shared filesystem degrades for everyone under that pattern. A site that does not set it stops the job at once with `TMPDIR is not set: …` — after a green `make gate`, which never touches it. Check with `srun --pty printenv TMPDIR`; if it is empty, export node-local scratch at the top of `scripts/_env.sh`.
5. **Single-node PCIe GPUs without a CUDA toolkit.** `scripts/_env.sh` exports `NCCL_P2P_DISABLE=1` and `NCCL_IB_DISABLE=1` (PCIe A100s without NVLink; an InfiniBand probe that hung) and `VLLM_USE_FLASHINFER_SAMPLER=0` (no `nvcc` on compute nodes). They cost nothing here but are wrong for NVLink nodes or multi-node jobs. They are not arm keys — `submit.sh` forwards only the arm's own knobs — so change them in `_env.sh`, or export them in the login-node shell, which `--export=ALL` carries into the job. Downloaded weights land in `$WORK/hf_cache` (`HF_HOME`), shared across jobs.

Nothing else is site-specific: the driver, the capability gate, the arm-file format and the ledger are all plain HTTP and JSON.

## Choosing a model

Every HAI-Co² tool schema carries an injected, **required** `action_and_reasoning` argument. A model that omits it does not fail cleanly: `_is_blank_turn` treats any message bearing a tool call as a real response, so the retry wrapper never sees the problem and the turn loops. This is [#184](https://github.com/petrosrapto/HAICO/issues/184), and it is why the server default moved off `gpt-4o-mini`.

Two things contain it:

1. **`agent.recursion_limit`** (default 50) bounds the loop. Note that LangGraph does *not* raise when the limit is reached: `create_react_agent` substitutes a canned reply carrying no tool calls, which ends the graph **normally**. The backend detects that sentinel and emits an `error` frame instead, or a runaway turn would be reported as a success.
2. **`make gate EXP=<arm>`** tells you *before* a campaign whether a model can satisfy the schema at all. It serves the model, calls it `GATE_N` times (25 by default, 15 fixed prompts cycled) with a miniature of the real tool schema, and asserts every returned tool call parses and carries a non-empty `action_and_reasoning`. It exits 0 only when at least 95 % pass — at `GATE_N=25` one failure still passes and two fail — so it can gate a pipeline. Where the verdict and the report land is described under [Run your first experiment](#run-your-first-experiment).

If a model fails, try a different `TOOL_PARSER` before abandoning it; if it still fails, prefer structured decoding on a model you have already validated over swapping in a new architecture.

<Warning>
  `--enable-auto-tool-choice` **requires** `--tool-call-parser`, and without both vLLM never emits `tool_calls` at all. If the gate reports `no tool_calls emitted` for every prompt, suspect the flags before the model. The job always passes both.
</Warning>

**Valid `TOOL_PARSER` values** (vLLM 0.23.0 registers 42). The ones you are likely to want: `hermes` (Qwen2.5/Qwen3 — start here), `qwen3_xml`, `qwen3_coder`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `mistral`, `openai` (gpt-oss), `deepseek_v3`/`v31`/`v32`/`v4`, `glm45`, `glm47`, `granite`, `granite4`, `pythonic`, `xlam`, `kimi_k2`, `minimax`, `seed_oss`, `olmo3`, `gemma4`, `internlm`, `jamba`, `phi4_mini_json`, `step3`.

**Already staged on the cluster** under `$SHARED_WEIGHTS` (`cluster.env`), so no download *if you point `MODEL` at the directory*: `MODEL=$SHARED_WEIGHTS/models--Qwen--Qwen3-32B` (62 GB bf16), `…/models--Qwen--Qwen3-32B-FP8`, `…/models--Qwen--Qwen3-4B-Thinking-2507`, `…/models--mistralai--Mistral-7B-Instruct-v0.2`; check for others with `ls $SHARED_WEIGHTS/`. A bare HuggingFace id such as `Qwen/Qwen3-32B` also works but re-fetches the weights into `$HF_HOME`. Note that `Qwen3.5-27B` is `Qwen3_5ForConditionalGeneration` — a hybrid-linear *vision* model, not a drop-in swap for Qwen3-32B.

## How it works

One `sbatch` job on one GPU node holds everything, all on `127.0.0.1`:

```
GPU NODE — ONE JOB ──────────────────────────────────────────────────────┐
 [1] vllm serve   $WORK/envs/haico-serving         127.0.0.1:$VPORT      │
 [2] postgres 15  apptainer                        127.0.0.1:$PGPORT     │
 [3] uvicorn      backend SIF, 1 worker            127.0.0.1:$APIPORT    │
 [4] driver       run_batch.py, Bearer haico_pat_…                       │
                                                                         │
 /etc/hosts bind-mount: 127.0.0.1 vllm localhost                         │
 results → $TMPDIR → rsync + pg_dump every DUMP_INTERVAL → $WORK/haico-runs/$RUN_ID │
─────────────────────────────────────────────────────────────────────────┘
```

**Why one node.** Compute nodes accept no inbound connections, so a backend anywhere else would need a `redir` forwarder on the login node — restricted to ports 5000–5100, unreachable from the TU/HRZ VPN, and **killed daily at 03:00**. Co-locating removes that entire failure class. Splitting backend-onto-CPU-node buys nothing either: a `gpu` allocation already grants 16 CPUs and 256 GB, and the backend needs \~2 CPUs and a few GB.

**Why Apptainer, and why not for vLLM.** Docker exists nowhere on the cluster; Apptainer 1.4.5 is on compute nodes and runs the existing OCI images directly, giving parity with what is deployed. vLLM stays outside a container because it is already installed in `$WORK/envs/haico-serving` (`VLLM_VENV` in `cluster.env`) and needs the GPUs, which would mean `--nv` passthrough for no parity benefit. (`podman` on the login node is a dead end: there is no `/etc/subuid` entry for the account, so rootless UID mapping — and `apptainer build --fakeroot` — is impossible.)

Apptainer does **not** namespace the network, so the four processes reach each other on loopback with no compose-style network. The flip side is that loopback is shared with co-tenant jobs, which is why ports derive from `$SLURM_JOB_ID` and are probed immediately before binding.

**How traffic reaches vLLM.** Entirely through the per-request `config` block on `POST /api/query/stream_steps/sse` — no application change:

```json theme={null}
{
  "query": "…",
  "thread_id": null,
  "config": {
    "API": "openai",
    "model_id": "Qwen/Qwen3-32B",
    "endpoint_url": "http://vllm:24002/v1",
    "args": { "temperature": 0.7, "max_tokens": 8192 }
  }
}
```

<Warning>
  The endpoint hostname **must contain `vllm`**. Credential resolution returns `VLLM_API_KEY` only when the endpoint URL contains that substring, and otherwise falls back to `OPENAI_API_KEY`. The job bind-mounts an `/etc/hosts` mapping `127.0.0.1 vllm localhost` precisely so this holds.

  Never set `OPENAI_ENDPOINT_URL`. It is a *global* fallback applied to every OpenAI-family model, so setting it as a deployment secret would silently re-route production traffic with no failing health check.
</Warning>

**How the API token is obtained.** `POST /api/keys` requires a browser session, which a throwaway database cannot provide. The job generates a fresh `haico_pat_…` per run and passes it in as `ADMIN_DEFAULT_API_TOKEN`, which migration `0017` installs for the seeded admin. **You do not set this anywhere** — not in an arm file, not in `.env`, not on the cluster. If the pinned image predates that migration the job falls back to logging in as the admin and calling `POST /api/keys`, so it works against any digest. See [ADR-0006](/docs/adr/0006-programmatic-api-access).

## Results

`make fetch EXP=<run-id>` pulls `runs/<run-id>/`:

| File | What |
| - | - |
| `ledger.json` | one row per conversation, the attrition table, and `stopped_early` naming why a run ended before its prompt set (signal, wall-clock ceiling, provider abort), or `null` — schema below |
| `transcripts/<id>.json` | per-turn outcomes and the captured workspace — schema below |
| `transcripts/superseded/<id>.attempt-N.json` | earlier, non-`ok` attempts of conversations that `make resume` re-ran; never delete these to "clean" a run |
| `run-config.env` | what each attempt actually ran with: job id, node, GPU, image digest, every knob, model hashes — appended per attempt |
| `haico-db.sql.gz` | `pg_dump` of the **last attempt's** database, refreshed every `DUMP_INTERVAL` seconds and at teardown |
| `haico-db.before-job-<id>-<time>.sql.gz` | the dump that was in the run directory when job `<id>` started at `<time>`, i.e. the previous attempt's last dump. Named after the job that displaced it, not the one that wrote it; kept so no attempt overwrites another's |
| `gate-report.json`, `gate-vllm.log` | written by `make gate` for the same `RUN_ID`: per-prompt pass/fail with reasons, pass rate and threshold, plus the gate's vLLM log; a second gate run overwrites them |
| `phoenix.db.gz`, `phoenix.log` | Phoenix trace database and service log, when `PHOENIX_ENABLED=1` |
| `vllm.log`, `backend.log`, `pg.log` | service logs |

`make fetch` prints the attrition table, `n_with_capture_errors` and `stopped_early`; check the first two before calling a run clean.

Each transcript captures the final state of every workspace panel. `document` (`{thread_id, title, content, updated_at}` or `null`), `objective` (`{thread_id, text, updated_at}` or `null`), `todos`, `preferences` and `artifacts` (lists of rows; artifacts carry `turn_index`, `artifact_type` and `payload`) are the panels as they stood when the conversation ended. `messages` is the chat history rebuilt from the checkpointer (`{role: user|assistant|tool, content, tool_calls?, tool_name?}`), which is where the agent's replies and tool calls live. `graph` is the conversation DAG with one node per turn (`user_message` cut to 140 characters, `assistant_summary` to 200, `reasoning_steps`, `artifact_count`). `snapshots` is the per-turn index only: `{id, turn_index, user_message, captured_at}`, newest first, at most 50. The per-turn document, objective, todo and preference bodies, and the `model_id` and `provider` recorded for each turn, are **not** in any transcript; they are rows of `workspace_snapshots` in the database dump, so keep the dump if you need per-turn history.

**A conversation can be `ok` and still have lost a panel.** `capture_errors` on the transcript and on its ledger row names each panel that could not be read (`messages: HTTP 500 (after 3 attempts)`), and the panel itself holds the reason as a string (`HTTP <code>`, `thread not in this job's database (HTTP 404)`, or an exception name such as `ReadTimeout`) instead of data — any string in a panel slot is that marker. `n_with_capture_errors` in the ledger counts such conversations; treat a non-zero count as an incomplete capture. Within the same job the driver retries with a 120/240/480 s ladder. `make resume` re-fetches only the missing panels, and only for threads that exist in that job's fresh database, so a panel lost in an earlier attempt comes back as `thread not in this job's database`; the only copy of that workspace is that attempt's kept `haico-db.before-job-*.sql.gz`.

### `ledger.json`

```json theme={null}
{
  "run_id": "qwen3-32b-pilot",
  "n_attempted": 8,
  "n_completed": 7,
  "outcomes": {"client_timeout": 1, "ok": 7},
  "n_with_capture_errors": 0,
  "stopped_early": null,
  "conversations": [
    {"id": "pilot-01-report", "thread_id": "th_…", "outcome": "ok",
     "turns_completed": 4, "turns_attempted": 4, "elapsed_s": 412.3,
     "capture_errors": []}
  ]
}
```

`n_attempted` counts every row, including `not_attempted`, `interrupted` and `campaign_timeout` rows, so it equals the size of the prompt set. `n_completed` is the number of `ok` rows. `outcomes` is the attrition table: one count per outcome name. `stopped_early` is `null` or one of `received SIGTERM` / `received SIGINT` (a cancel, or the pre-wall-clock signal), `campaign ceiling of <n>s reached`, or `provider abort: <n> conversations in a row rejected immediately ('<message>')`. Each row carries `turns_attempted` (turns the driver sent) and `turns_completed` (turns that ended `ok`); a failed conversation has `turns_attempted = turns_completed + 1`, and `thread_id` is `null` when turn 0 failed before the server assigned one. The ledger row and `transcripts/<id>.json` describe the **last** attempt of a conversation; earlier tries are in `superseded/`.

### `transcripts/<id>.json`

| Field | Meaning |
| - | - |
| `id`, `run_id`, `model` | the spec id, the `RUN_ID`, and the `MODEL` string passed to vLLM |
| `thread_id` | the server-side conversation; `null` if turn 0 failed before one was assigned |
| `outcome` | `ok`, or the outcome of the first turn that failed (later turns are never sent) |
| `turns[]` | one entry per turn sent: `index`, `query`, `outcome`, `elapsed_s`, `step_count` (SSE `step` frames), `artifact_count`, `error` (first 500 characters of the server's message, or `null`) |
| `workspace` | the eight panels described above, captured after every conversation finished |
| `capture_errors[]` | one string per panel that could not be read |
| `started_at`, `elapsed_s` | wall-clock start (epoch seconds) and duration of the turns, excluding capture |

A turn record does not contain the agent's reply, its tool calls or its steps: the reply and tool calls are in `workspace.messages`, and the individual ReAct steps are recorded only when `PHOENIX_ENABLED=1`. Token counts are not recorded per turn anywhere — `vllm.log` carries only aggregate throughput lines — so a study that needs them must enable Phoenix, whose LLM spans carry the provider's usage fields.

### Reading the database dump

`haico-db.sql.gz` is a plain-SQL `pg_dump` taken by Postgres 15, so restore it with `psql` into a server of version 15 or newer (the local stack's `postgres:15-alpine` will do):

```bash theme={null}
docker run -d --name haico-run -e POSTGRES_USER=haico -e POSTGRES_PASSWORD=x -e POSTGRES_DB=haico postgres:15-alpine
gunzip -c runs/<run-id>/haico-db.sql.gz | docker exec -i haico-run psql -q -U haico -d haico
```

Tables worth knowing: `conversation_users` (one row per `thread_id`, with the final `turn_index`), `workspace_snapshots` (one row per `(thread_id, turn_index)`: the document, objective, todos and preferences as the agent saw them at the start of that turn, plus `model_id` and `provider` — the per-turn provenance nothing else records), `documents`, `objectives`, `todos`, `preferences`, `artifacts` (final state, the same as the transcript panels), and the LangGraph tables `checkpoints`, `checkpoint_blobs`, `checkpoint_writes` (the full message history the `messages` panel is rebuilt from). `api_keys` holds only the hash and prefix of the run's token.

Every attempt starts a fresh database, so `haico-db.sql.gz` holds only the last attempt's threads. Conversations completed by an earlier attempt of a resumed run, and every transcript in `transcripts/superseded/`, reference threads that exist only in the matching `haico-db.before-job-<id>-<time>.sql.gz`; restore each into its own database. A conversation that was cut mid-turn by a cancel has no transcript, but its partial thread is in the dump that attempt wrote.

### Reading the Phoenix traces

`phoenix.db.gz` is the SQLite database of the Phoenix instance that ran inside the job (`arizephoenix/phoenix:13.23.0`, the same image the deployed stacks run). Unzip it and serve it with that image or a newer one, pointing it at the file:

```bash theme={null}
gunzip -k runs/<run-id>/phoenix.db.gz
docker run --rm -p 6006:6006 -e PHOENIX_SQL_DATABASE_URL=sqlite:////data/phoenix.db \
  -v "$PWD/runs/<run-id>:/data" arizephoenix/phoenix:13.23.0
```

The traces are in project `haico_slurm`, one trace per turn, with the ReAct steps, tool calls and timings that neither the transcript nor the database record; sessions are keyed by `thread_id`, which is how to find the trace for a transcript.

<Note>
  **Attrition is a result, not a cleaning step.** Conversations fail most often in exactly the arms where the model is weakest, so they are not missing at random and summarising only the survivors overstates whatever the campaign measures. `ledger.json` reports `n_attempted`, `n_completed`, and a mutually-exclusive taxonomy: `tool_arg_missing`, `recursion_exhausted`, `truncated`, `context_overflow`, `turn_deadline`, `client_timeout`, `stream_idle`, `transport`, `server_error`, plus the harness outcomes `interrupted`, `campaign_timeout`, `driver_error` and `not_attempted`. Only the first indicts the model's ability to drive the agent; `recursion_exhausted` is a step-ceiling hit with no tool-schema error visible in the turn's steps. `client_timeout` is a **harness artifact** — the driver gave up after `READ_TIMEOUT` seconds of silence — and means you should raise `READ_TIMEOUT` or lower `CONCURRENCY`, not that the model failed.
</Note>

Two design points worth knowing when reading results:

* **Resume granularity is the conversation, never the turn.** The server increments the turn index and writes a snapshot *before* the agent runs, so replaying a failed turn would duplicate state. A conversation that fails at any point is abandoned whole. `make resume` keeps only conversations whose transcript ended `ok`; every other one (any failure outcome, `interrupted`, `campaign_timeout`, `not_attempted`, or missing) is run again from a fresh `thread_id`, and its earlier transcript is moved to `transcripts/superseded/<id>.attempt-N.json` so the number of tries stays on record. `make submit` refuses a `RUN_ID` that already has a ledger; `make resume` is the way to continue one. A resume of a finished run does nothing and leaves every transcript byte-identical.
* **A control arm is your responsibility, and this stack cannot run it.** `make submit` always serves a model with vLLM, and the driver sends `"API": "openai"` with an `endpoint_url` on every request, so it cannot target the deployed default (`gemini-2.5-flash`) or any non-OpenAI provider. For a hosted comparison, bring up a local stack (`deployment/local`) with the provider key, mint a personal access token, and run the driver by hand with the same prompt set and sampling settings: `HAICO_API_TOKEN=… python driver/run_batch.py --api http://localhost:8000 --endpoint-url https://api.openai.com/v1 --model <model> --prompts driver/prompts/<set>.jsonl --out runs/<control-id> --run-id <control-id> --temperature 0.7 --max-tokens 8192 --concurrency 2`. That works for OpenAI-compatible endpoints only; a Gemini or Anthropic control arm needs a driver change to send a different `API` value. Record the control model and the date, since a hosted model is not pinnable.

### Exit codes

| Exit | Meaning |
| - | - |
| 0 | every conversation ran; ledger written |
| 2 | driver stopped by TERM/INT (`make cancel`, or Slurm's signal five minutes before the wall clock); ledger written, unfinished rows `interrupted` |
| 3 | provider abort after three immediate rejections; ledger written, remaining rows `not_attempted` |
| 4 | campaign ceiling (ten minutes before the wall clock) reached; ledger written, cut rows `campaign_timeout` |
| 1 | the driver rejected its arguments (not reachable through `submit.sh`, which validates the prompts file first) |
| 12 | a port in the job's window was in use before anything started |
| 13 | vLLM did not accept requests within `HAICO_READY_TIMEOUT`, or exited during start-up |
| 14 | the backend never became healthy |
| 15 | no usable API token (`could not log in as admin`, `POST /api/keys did not return a usable token`, or `the API token does not authenticate` — all three point at `backend.log`) |
| 16 | Phoenix did not become healthy |
| 143 / 130 | the job itself was signalled (TERM / INT) and ran its teardown; `sacct` shows `FAILED` with `15:0` / `2:0` on purpose, and the results are complete up to the signal |

A gate job exits 1 when the pass rate is below the threshold. Codes 2, 3, 4 and 143 all mean the ledger was written and `make resume` continues the run; only a job that ends with no `ledger.json` in `$WORK/haico-runs/<RUN_ID>` failed in the sense that matters.

## Secrets

**One long-lived secret exists, and it lives on the cluster — never in the repository, never on your laptop.**

```
$WORK/haico-cluster/secrets/ghcr.env      mode 600, in a 700 directory
```

`make secrets` writes it there. Nothing else does, and nothing writes a secret locally: the token is piped over the SSH channel to a `cat` on the far side, so it never appears in your shell history, in a local file, or in the remote process list.

### Why that path

**Not `$HOME`.** On many clusters — including the one this was built against — `$HOME` is not mounted on compute nodes. A credential there is unreadable at exactly the moment the job needs it, and the failure looks like a broken token rather than a missing mount.

**Not the repository.** `deployment/slurm/secrets/` is in `.gitignore`, and `secrets/` is excluded from the `rsync` in both `submit.sh` and the `Makefile`. That exclusion is doing two jobs: it stops a local secret being uploaded, and — because `rsync --delete` does not remove excluded paths on the receiver — it stops every `make submit` from wiping the token you stored.

### What is in it

| Variable | Needed for |
| - | - |
| `GHCR_USER`, `GHCR_PAT` | pulling the private backend image. **Required.** |
| `HF_TOKEN` | gated models only (Llama). Qwen and Mistral need none. Add with `bash scripts/bootstrap_secrets.sh --hf`. |

Inspect what is stored, without printing any secret value:

```bash theme={null}
bash scripts/bootstrap_secrets.sh --show
```

### Everything else is generated per job and dies with it

Nothing below is ever stored, by you or by the tooling. Each `sbatch` mints its own from `/dev/urandom`, and the values vanish when the allocation ends.

| Generated per job | Purpose |
| - | - |
| `JWT_SECRET` | signs the backend's session tokens |
| `ADMIN_DEFAULT_PASSWORD` | seeds the admin account in the throwaway database |
| `ADMIN_DEFAULT_API_TOKEN` | the `haico_pat_…` the driver authenticates with |
| vLLM API key | authenticates the backend to the model server |
| Postgres password | `scram-sha-256`, because loopback is shared with co-tenants |

This is why there is no `.env` to fill in for a cluster run, and why a leaked job log costs you nothing.

### How the job handles the one real secret

Three properties, each deliberate:

* **Sourced inside the job**, never exported before `sbatch`. Slurm records the submission environment, so a secret placed there is more exposed than one read at runtime.
* **`unset` immediately after the image pull**, so it is not in the environment while the batch runs.
* **`APPTAINER_DOCKER_USERNAME`/`PASSWORD` are scoped to the private pull only.** They apply to *every* registry, so exporting them globally breaks the Docker Hub pull of `postgres` with `invalid username/password` — which reads like a broken credential rather than a misapplied one.

<Warning>
  Use a **dedicated classic token with only `read:packages`**. `make secrets` warns if you paste a `gho_` token — that is the `gh` CLI credential, which typically carries `repo`, `workflow` and `admin:org`. It would work, and that is the problem: a token that can delete repositories should not sit on multi-user storage for the length of a campaign, whatever the file mode.
</Warning>

### A second researcher

Each account has its own. `$WORK` is per-user and `700`, so nothing is shared and nothing is inherited — a colleague runs `make secrets` once with their own token. If they lack read access to the private package, `make secrets` refuses the token immediately rather than letting them discover it twenty minutes into a job.

To revoke: delete the file (`rm $WORK/haico-cluster/secrets/ghcr.env`) and revoke the token on GitHub. Runs already in flight are unaffected — they pulled their images at startup.

## What differs from haico.gr

Recorded deliberately in `deployment/slurm/config/backend.yaml`, and worth stating in any write-up:

| Setting | Cluster | Deployed | Why |
| - | - | - | - |
| `observability.provider` | `none` by default | `phoenix` | Set `PHOENIX_ENABLED=1` in the arm to run Phoenix in the job. Off suits a throughput run — per-turn provider and `model_id` are already in `workspace_snapshots`. On shows what happened *inside* a turn. The trace DB is copied to the run directory at teardown. |
| `study.enabled` | `false` | `true` | The in-app study is for human participants. |
| `api_keys.rate_limit_per_minute` | `0` | `120` | The limiter is in-process and would throttle the driver, not an abuser. |
| `MAX_TOKENS` (per request) | `8192` in the shipped arms | `32768` | Sent in every turn's per-request `config.args`, which overrides `llm.max_tokens`; `config/backend.yaml` keeps 32768 as a fallback only. A self-hosted 32k model must leave room for a prompt carrying the whole document. |
| uvicorn workers | 1 | 1 | The rate limiter is in-process state; more than one worker multiplies the ceiling. |

## Troubleshooting

Every symptom below that starts with `FATAL:` is in the job's `.err` log; `make logs` shows both files.

| Symptom | Cause and fix |
| - | - |
| `unable to retrieve auth token: invalid username/password` on the **postgres** pull | `APPTAINER_DOCKER_*` applies to *every* registry, so GHCR credentials leaked into the Docker Hub pull. The job scopes them to the private pull only; if you script your own pull, do the same. |
| Gate reports `no tool_calls emitted` for every prompt | vLLM started without both `--enable-auto-tool-choice` and `--tool-call-parser`. |
| Gate reports `action_and_reasoning missing` | The model cannot satisfy the tool schema. Try another `TOOL_PARSER`, then another model. |
| Turns end in `tool_arg_missing` | Same cause, now bounded by `agent.recursion_limit` rather than looping. |
| Turns end in `recursion_exhausted` | The agent hit `agent.recursion_limit` without a visible tool-schema error; inspect the transcript's steps before blaming the parser. |
| Gate reports `arguments are not valid JSON` or `arguments are not an object` | The model emits tool calls in a format the chosen `TOOL_PARSER` does not decode. This is the parser, not the schema: try the family-specific parser (`qwen3_xml`, `llama3_json`, `mistral`, …) before changing the model. |
| Gate prints `N/25 calls never reached the model … THIS IS NOT A MODEL RESULT` | vLLM passed readiness and then died or stalled. Read `gate-vllm.log` in the run directory; do not record this as a capability result. |
| Three `server_error` rows in under a second each, every later row `not_attempted`, `stopped_early` begins `provider abort:`, job exit 3 | The first turn of three conversations in a row was rejected immediately with an authentication/permission-class error, so the driver stopped instead of burning the prompt set; the ledger is complete. The quoted message (also the first turn's `error` in each transcript) names the cause: in this stack it means the backend or the loopback vLLM rejected the request — check `backend.log`. For a hand-run hosted control arm it is that provider's credit or key. Fix it, then `make resume`. |
| Rows are `interrupted` or `campaign_timeout` | Cancelled, or the campaign ceiling fired ten minutes before `TIME` (expected, so the ledger and final dump are written with Postgres alive). Those conversations have no transcript; finished ones and the dumps are in `$WORK/haico-runs/<RUN_ID>`, and `make resume` re-runs the rest. |
| `sacct` shows the job `FAILED` with exit code 2, 3, 4 or `15:0` | The driver's own codes, or a graceful stop (143 rendered as `15:0`); see [Exit codes](#exit-codes). The ledger was written; `make fetch` and `make resume` work the same as after a clean run. |
| Turns end in `truncated` or `context_overflow` | `MAX_TOKENS` too close to `MAX_MODEL_LEN`; the prompt carries the whole document. |
| Turns end in `client_timeout`, often with `steps=0` and no `thread_id` | The driver saw total silence before the first frame. Heartbeats do not cover that window, so a bigger `READ_TIMEOUT` may not help. Measured cause: the backend serialises requests — at `CONCURRENCY=4`, vLLM reported `Running: 1 reqs` in 122 of 150 samples while three conversations queued invisibly. **Lower `CONCURRENCY` first**; it buys little throughput and costs conversations. |
| Turns end in `turn_deadline` | One turn ran longer than the 1800 s hard ceiling (`TURN_DEADLINE_SECONDS` in `driver/run_batch.py`). Not an arm key: lower `CONCURRENCY` first; raise the constant only if a slow model writing long documents legitimately needs it, and record the change with the run. |
| Turns end in `stream_idle` | The backend kept sending keep-alives for 900 s (`STREAM_IDLE_SECONDS`) with no data frame: the agent's LLM call is wedged while the connection stays alive. Server side, unlike `client_timeout` — look at `vllm.log` around the turn's timestamp. |
| Turns end in `transport`, or rows are `driver_error` | `transport`: the connection to the backend dropped mid-stream; check whether `backend.log` ends with a crash. `driver_error`: an exception inside the driver, printed as `[FAIL] <id> driver_error <Exception>` in the `.err` log; the ledger keeps the row and `make resume` retries it. |
| Ledger shows `n_with_capture_errors > 0`, or a transcript panel is a string such as `HTTP 500`, `ReadTimeout` or `thread not in this job's database` | The turns succeeded but the panel could not be read back. See the capture paragraph under [Results](#results): the same job retries, a later job cannot, and the panel is in that attempt's kept dump. |
| `FATAL: port … already in use on <node> by another tenant` (exit 12) | The pre-flight check found a socket in the job's five-port window before anything was started; any TCP socket counts, including a co-tenant's outbound connection. Resubmit: the window derives from the job id. |
| `FATAL: vLLM exited during startup` or `FATAL: vLLM did not accept requests within 2400s` (exit 13) | The last 40 lines of `vllm.log` are printed under the message; read them first. Still logging compile / CUDA-graph capture: raise `HAICO_READY_TIMEOUT`. A KV-cache or out-of-memory error: lower `MAX_MODEL_LEN` or `MAX_NUM_SEQS`, raise `GPU_MEMORY_UTILIZATION`, or ask for more GPU memory. `port N was taken by a co-tenant job between the check and the bind`: resubmit. `401 Client Error` / `GatedRepoError`: the model is gated and the job had no `HF_TOKEN` — accept the licence on huggingface.co, then `bash scripts/bootstrap_secrets.sh --hf`. At `TENSOR_PARALLEL_SIZE>1`, rank 0 logs `vLLM is using nccl` and rank 1 never does: an NCCL hang; prefer `TENSOR_PARALLEL_SIZE=1` whenever the model fits one card. `Could not find nvcc`: something in `EXTRA_VLLM_ARGS` reaches a JIT path the compute nodes cannot serve. |
| `FATAL: backend not healthy` (exit 14) | The backend did not answer `/health` within 240 s; the last 40 lines of `backend.log` follow. `PostgreSQL not reachable`: the entrypoint waits only 30 s for Postgres, so look at `pg.log`. An Alembic traceback: the pinned digest's migration chain failed against a fresh database — try a newer digest. |
| `FATAL: Phoenix did not become healthy` (exit 16) | Only with `PHOENIX_ENABLED=1`: the image pull failed or the UI port did not answer `/healthz` in 120 s; the last 30 lines of `phoenix.log` follow. Tracing is optional: `PHOENIX_ENABLED=0` and resubmit. |
| `FATAL: could not log in as admin to mint an API key` (exit 15) | The image predates migration `0017` *and* the admin seed failed. The job falls back to logging in as `admin` and calling `POST /api/keys`, so it works against any digest; check `grep -i admin backend.log`. The other two exit-15 messages (`POST /api/keys did not return a usable token`, `the API token does not authenticate`) point at `backend.log` too. |
| Boxed `WARNING: this backend image predates the step-limit guard` after the image pull | The pinned `BACKEND_DIGEST` was built before the backend detected a step-limit hit, so a turn that exhausts `agent.recursion_limit` is recorded as `ok` and `n_completed` overstates the model. The job continues on purpose (reproducing an old result needs the old image); for a campaign that measures tool-schema adherence, re-pin a newer digest. |
| Job log shows `[mirror] … db dump FAILED` or `(pg_dump failed — any previous dump left intact)` | `pg_dump` could not connect. At teardown this means Postgres already received a cgroup-wide TERM — a plain `scancel`, `make cancel FORCE=1`, or the wall clock arriving before the driver stopped — and entered smart shutdown. Cancel with `make cancel` (no `FORCE`). The last periodic dump is what you have otherwise; lower `DUMP_INTERVAL` for short or risky runs. Failing every interval from the start is a Postgres problem: read `pg.log`. |
| Batch job stops in its first second with `TMPDIR is not set` (the gate for the same arm succeeded) | This site's scheduler does not export a per-job `$TMPDIR`; see assumption 4 under [Adapting to another cluster](#adapting-to-another-cluster). |
| `apptainer pull` of `ghcr.io/petrosrapto/haico-backend@sha256:…` fails with a manifest / not-found error | The digest is not one the package has: `submit.sh` only checks the `sha256:` prefix. Resolve it again with `make digest TAG=<tag>` and paste the output verbatim. An `unauthorized` error instead means the stored token lost package access; re-run `make secrets`. |
| Job requeued from the start with no warning | `PARTITION=yolo` is preemptible with `GraceTime=0`. Use `gpu` for long batches. |
| `sbatch` appears to succeed but nothing queues | SSH connection multiplexing has been observed to swallow submissions on long-lived sessions. The tooling passes `-o ControlPath=none` throughout; re-check with `make status`. |
