0009. Run mass experiments against cluster-hosted models in a single SLURM job
Context and problem statement
HAI-Co² calls third-party LLM APIs. That is right for the deployed product, and wrong for research at volume: a campaign of hundreds of scripted conversations is expensive, the model can be deprecated underneath a study, and the exact weights behind a result are neither pinnable nor inspectable. A university GPU cluster removes all three problems, and ADR-0006 already provides the credential a scripted client needs. The question was where each piece runs. The decision was made against a concrete cluster (UKP Lab, TU Darmstadt) whose properties were verified live rather than assumed. They are typical of academic SLURM sites, and the ones that shaped the design are called out below so another site can check them:- Compute nodes accept no inbound connections and are not SSH-able. They do have full outbound internet with no proxy.
- Docker exists nowhere. Apptainer 1.4.5 is available on compute nodes only;
podmanexists on the login node but there is no/etc/subuidentry for the account, so rootless UID mapping — and thereforeapptainer build --fakeroot— is impossible. - The only sanctioned way to expose a job’s port outward is a
redirforwarder on the login node, restricted to external ports 5000–5100, reachable only from the UKP network (not the TU/HRZ VPN), and killed daily at 03:00. - QOS
gpugrants 4 GPUs and 16 CPUs per job, 5 GPUs per user, and a 3-day wall. Multi-node jobs are permitted (MaxNodes=UNLIMITED). - Postgres is not optional: the app schema and the LangGraph checkpointer both live there, and
0002_artifacts_and_snapshotsuses rawpostgresql.JSONB()with no variant, so the migration chain cannot compile on SQLite.
action_and_reasoning argument, and _is_blank_turn treats any message bearing a tool call as a real response — so the retry wrapper cannot see a tool call whose arguments fail validation. A weaker self-hosted model does not fail; it loops (#184).
Decision drivers
- The deployed environments must not change. haico.gr and dev.haico.gr keep using hosted APIs exactly as before.
- Fidelity. An experiment must measure HAI-Co² as it actually runs — real turn indexing, snapshots, artifacts and branching — not a reimplementation.
- Unattended. A campaign should survive overnight without a laptop, a VPN session, or a forwarder that dies at 03:00.
- Reproducible. A result must name the exact model and the exact backend build that produced it.
- Minimal blast radius in the repo. Prefer configuration and new standalone files to changes in the running application.
Considered options
- One self-contained SLURM job. vLLM, Postgres, the backend and the driver on one GPU node, all on
127.0.0.1. - Split services. A long-lived vLLM job on a GPU node, the backend elsewhere (a CPU-partition job, the login node, or the dev VM), connected through
rediror an SSH tunnel. - No backend at all. Drive the LangGraph agent in-process, skipping HTTP and Postgres.
Decision
Option 1. Onesbatch job, one GPU node, everything on loopback.
Apptainer runs Postgres and the backend from the existing OCI images, giving parity with what is deployed. vLLM is deliberately not containerised — it runs from a venv the stack builds itself (make env installs the pinned env/requirements.txt into VLLM_VENV from cluster.env, $WORK/envs/haico-serving by default) and needs the GPUs, which under Apptainer would mean --nv passthrough for no parity benefit.
The backend is pointed at vLLM entirely through the per-request config block on POST /api/query/stream_steps/sse, which needs no application change: API: "openai" plus an endpoint_url resolves to ChatOpenAI(base_url=…).
Two supporting changes were required:
agent.recursion_limit(default 50), passed toagent.astream, plus explicit detection of the limit being reached. This is a product fix, not an experiment workaround: without it the loop above is bounded only by the framework default, 10007 super-steps in LangGraph 1.x. The detection is not optional — LangGraph does not reliably raise on exhaustion: depending on where the limit falls in the agent/tools cycle it may substitute a canned reply with no tool calls and end the graph normally, so an unguarded limit converts a visible hang into an invisible success.- Migration
0017_seed_admin_api_key, which installs a personal access token for the admin user fromADMIN_DEFAULT_API_TOKEN, opt-in and silently skipped when unset — mirroring how0001seeds the admin account itself. A throwaway database has no browser and no mailbox, butPOST /api/keysrequires a session, so without this a batch run could not obtain the credential ADR-0006 defines.
Consequences
Portable. Everything site-specific — submit host, work directory, vLLM venv, shared model cache, per-job ceilings, the CPU partition for the venv build — lives indeployment/slurm/cluster.env, and the scheduler names live in the arm file. CLUSTER_WORK is the single source of truth for paths on both sides: the laptop tooling and the job resolve the venv, the secrets file and the results from it. The properties the design leans on (Apptainer available, no inbound to compute nodes, outbound internet from them, a per-job node-local $TMPDIR, single-node PCIe GPUs) are documented in the guide as things to re-check rather than assumed universal.
Good. No inbound ports, no tunnels, nothing to keep alive between runs, and nothing that dies at 03:00. Apptainer does not namespace the network, so the four processes reach each other on 127.0.0.1 with no compose-style network. A run pins the backend image by digest and records what it ran with (run-config.env: every knob, the node, the GPU, and for a model directory the hashes of its config and shard index), so it is reproducible; a model given as a HuggingFace id is pinned only when the arm passes --revision. The deployed environments are untouched, and backend/ gains only the recursion bound and the opt-in migration.
Bad. Every run reloads the model (minutes for a 32B at TP=2), which is wasted work for very short campaigns. The database dies with the job, so durability depends on the periodic mirror (transcripts by rsync plus a pg_dump, every DUMP_INTERVAL seconds) and on Slurm signalling the batch shell five minutes before the wall clock, which lets the driver write its ledger and the teardown dump run while Postgres is still up; a resumed run keeps every earlier attempt’s dump. Ports are shared with co-tenant jobs on the same node, so they are derived from $SLURM_JOB_ID and probed before binding.
Divergences from the deployed configuration, each deliberate and recorded in deployment/slurm/config/backend.yaml: tracing defaults off (observability.provider: none), the study is off, the per-key rate limiter is disabled (it is in-process and would throttle the driver rather than any abuser), and the effective per-turn max_tokens is lower than production’s — set per request from the arm’s MAX_TOKENS, while backend.yaml keeps the deployed 32768 as its fallback — because a self-hosted 32k model must leave room for a prompt that carries the whole working document.
Tracing is a knob rather than a removal: PHOENIX_ENABLED=1 runs Phoenix in the job and points the backend’s tracer at it. Off is right for a throughput run, because the provenance an experiment needs — per-turn provider and model_id — is already written to workspace_snapshots, and spans would only add load. On is right when the question is about what happened inside a turn: the ReAct steps, tool calls and timings, which Postgres does not record. Its SQLite database is copied out at teardown, since node-local scratch is wiped with the allocation.
Prohibited. OPENAI_ENDPOINT_URL must never be set, in any environment. It is a global fallback applied to every OpenAI-family model, so setting it as a deployment secret would silently re-route production traffic with no failing health check. Per-request endpoint_url is the only supported mechanism. Relatedly, the endpoint hostname must contain vllm — that substring is what makes credential resolution pick VLLM_API_KEY instead of a real OPENAI_API_KEY.
Rejected options
Split services (2). The backend cannot reach a compute node from outside the cluster, so this requires aredir forwarder — reaped nightly, restricted to a firewall range shared with the whole lab, and unreachable from the TU/HRZ VPN. It also introduces a node-discovery problem: a requeued vLLM job lands on a different host and every client’s endpoint silently breaks. The CPU-node variant fails for a different reason — a gpu allocation already grants 16 CPUs and 256 GB, far more than the backend’s ~2 CPUs and few GB, while krusty is the lab’s only CPU node, so splitting adds a queue dependency and buys nothing. Driving cluster vLLM from dev.haico.gr was rejected outright: synthetic conversations would land in the same tables as real study participants.
Building the image on the cluster instead of pulling it (considered and rejected). A parallel effort took a different route to the same runtime: build with podman on the login node, export an OCI archive to BeeGFS, and apptainer build from that tarball on the compute node. It is attractive because it needs no registry credential on the cluster at all, removing the one long-lived secret in this pipeline, and it does not assume compute nodes have outbound internet (they do here, verified, but that is a property of this cluster rather than of clusters).
Rejected because it trades away the property an experiment needs most. A source build produces whatever the working tree happened to contain; pulling haico-backend@sha256:… names one immutable artifact, so a result can state exactly which backend produced it and a later run can reproduce it. Pinning an older digest to re-examine an older result is a first-class use case, and a build-from-source path cannot offer it. The secret is contained instead: one classic token with read:packages, stored chmod 600 under $WORK, sourced inside the job rather than through the submission environment, and unset immediately after the pull.
No backend (3). Removes exactly the system under study. Turn indexing, per-turn snapshots, artifact extraction and branching are all server-side, and ADR-0006’s guarantee is precisely that an API-only client produces state identical to a browser user. It would also not remove Postgres, since the agent re-reads the workspace every turn.
Related
- ADR-0006 — the credential this uses, amended here by the seeding migration.
- ADR-0003 — why the checkpointer, and therefore Postgres, is load-bearing.
- SLURM experiments — the operational guide.
- #184 — the silent tool-argument loop that motivates the recursion bound and the capability gate.