Skip to content

Deploy

selfjev serve runs the API on one NVIDIA GPU with at least 16 GB (selfjev-4b is Qwen3.5-4B in bf16 plus a LoRA). Three ways to run it, from the most hands-on to the least.

On a GPU machine you have

git clone https://github.com/Jwuthri/SelfJev.git && cd SelfJev
git lfs pull --include "weights/selfjev_4b/*"
uv sync --extra serve --extra gpu
SELFJEV_API_KEYS=my-key uv run selfjev serve --host 0.0.0.0 --port 8000

The base model (about 9 GB) downloads from Hugging Face on first start, at a pinned revision. Without SELFJEV_API_KEYS (comma-separated) the server accepts every request, which suits a laptop or a private network. The my-key shown above is a placeholder: choose a long random secret and pass the same value as api_key in the SDK or SELFJEV_API_KEY in your client environment. Hugging Face tokens and Jev API keys are unrelated to this secret.

Useful flags: --max-length (state plus longest question, default 32,768 tokens), --max-batch-tokens (tokens packed per forward pass), --calibration (temperatures from selfjev calibrate), --engine vllm --model-dir <merged> (vLLM on weights merged by selfjev merge: fast for one question per request, slower for many; the default tree engine reads the text once for all questions, see speed).

Download the merged model

SelfJev-4B merged includes the complete base model, merged fine-tuning weights, tokenizer and configuration (about 9.32 GB of weights). Follow its model card to download and serve it using selfjev serve --engine vllm --model-dir <download-directory> in a compatible vLLM environment. No separate base or adapter download is needed. The default TreeServer uses the adapter installation above.

Docker

git lfs pull --include "weights/selfjev_4b/*"
docker build -f deploy/Dockerfile -t selfjev .
docker run --gpus all -p 8000:8000 -e SELFJEV_API_KEYS=my-key selfjev

The image bakes the base model in (--build-arg BAKE_MODEL=0 to download it at start instead) and has a health check on /health. docker compose -f deploy/docker-compose.yml up -d does the same with SELFJEV_API_KEYS from the environment or an .env file.

AWS, one command

Needs AWS credentials that boto3 finds (for an SSO profile: AWS_PROFILE=<profile>) and pip install "selfjev[deploy]" (or uv sync --extra deploy).

selfjev deploy aws machines                                   # presets and prices
selfjev deploy aws up --name prod --instance g6.xlarge        # prints the endpoint and a new API key
selfjev deploy aws status --name prod                         # health, uptime, cost so far
selfjev deploy aws list
selfjev deploy aws down --name prod                           # terminates the box, deletes its security group

up starts NVIDIA's Deep Learning Base AMI and, on first boot, installs this repository at --ref (default master) with only the selfjev-4b weights, pre-downloads the base model and runs selfjev serve as a systemd service on port 8000 behind the key. Setup installs packages and fetches about 9 GB of model: 5.4 min from launch to a healthy server on a g6e.xlarge in the end-to-end test below; up waits for /health (up to 30 minutes) unless --no-wait. The server compiles its GPU kernels before /health answers (the first request otherwise took 39 s). The deployment record, key included, is kept in ~/.selfjev/deployments/<name>.json; every resource is tagged Project=selfjev.

preset GPU $/h (us-east-2, on demand) for
g6.xlarge (default) L4, 24 GB 0.805 the cheapest 24 GB GPU; serving on it is not measured yet
g5.xlarge A10G, 24 GB 1.006 when L4 capacity is short
g6e.xlarge L40S, 48 GB 1.861 long texts, more traffic, fine-tuning on the box
g6e.2xlarge L40S, 48 GB 2.242 as g6e.xlarge with more CPU
p5.4xlarge H100, 80 GB 6.88 lowest latency; often out of capacity

Other options: --region (default us-east-2), --allow-cidr (who may reach the port, default everyone; the key still applies), --api-key (bring your own), --max-hours (the box terminates itself after that long: a cost cap for trials), --ssh (a key pair and port 22 from your IP, to read /var/log/selfjev-setup.log).

There is no pause: down terminates the box and billing stops; up builds a fresh one.

Fine-tuning on the server

selfjev serve --fine-tuning (or selfjev deploy aws up --fine-tuning) adds the fine-tuning routes: upload a JSONL of requests with their expected answers, start a supervised or RLCD job, and the fine-tuned model is served next to selfjev-4b as soon as the job succeeds. Jobs train on the same GPU as serving, one at a time. Our training runs used 48 GB L40S cards (g6e.2xlarge; g6e.xlarge has the same GPU); a 24 GB card next to serving is untested and likely too small. Job state lives in --home (SELFJEV_HOME, default ~/.selfjev/server). In Docker, mount a volume there and add the flag: docker run --gpus all -p 8000:8000 -v selfjev:/root/.selfjev selfjev uv run --no-sync selfjev serve --host 0.0.0.0 --fine-tuning.

To train elsewhere (a bigger box, a notebook), run selfjev finetune or selfjev rlcd there (fine-tune and RLCD) and serve the resulting adapter with selfjev serve --adapter <run>/adapter.

Tested end to end

scripts/aws/e2e.py deploys a commit this way (--fine-tuning), then checks through the SDK: health, auth and error codes, every question type on a ticket with obvious answers, Jev's paths and model names, 16 concurrent requests, a supervised and an RLCD fine-tuning job over HTTP and their models; it records every HTTP call and tears the box down. Last run: 2026-09-28, 14 of 14 checks passed on an L40S in 26 minutes (≈ $0.81). A 75-second video replays that run from its record (scripts/docs/e2e_video.py: every answer and number on screen comes from the transcript).

Operating it

  • GET /health: 200 with the queue depth once the model is loaded (open, for load balancers).
  • GET /metrics: Prometheus text: requests by route and status, latency, questions and tokens, queue depth (open).
  • Every response carries x-request-id; errors are JSON ({"error": {"type", "message", "param"}}). An unhandled error is a 500 whose message names the request id; the server logs its traceback under that id.
  • Concurrent requests are batched into shared forward passes (up to 32 requests, 5 ms wait). When 256 requests are waiting the server answers 529 with Retry-After; the SDK retries 429, 529 and 5xx with backoff.