Skip to content

Reproduce

Install

uv sync                  # Python 3.12, torch, transformers, peft, fastapi, pytest (pinned in uv.lock)
uv sync --group data     # + datasets/pyarrow, only needed to rebuild data/hf.jsonl
uv sync --extra gpu      # Linux CUDA boxes: fast Gated DeltaNet kernels (flash-linear-attention)

The base model (Qwen3.5-4B, ≈ 9 GB) is pinned to a revision and downloads on first use. The adapter, selfjev-4b, is in weights/ (Git LFS, with its model.json: git lfs pull --include "weights/selfjev_4b/*"). New runs write to runs/, which is not in git; the training logs of past runs are in reports/train_meta/.

Use

The default model needs a CUDA GPU. Serve it and call it with the SDK (API, deploy), or from Python directly:

import json

from selfjev.core.answers import classify
from selfjev.core.options import with_options
from selfjev.engine.tree import TreeServer

scorer = TreeServer("weights/selfjev_4b")  # loads Qwen3.5-4B + the LoRA (merged) on the GPU
request = json.load(open("examples/request.json"))  # the internal schema: {"state": ..., "questions": [...]}
request["questions"] = [with_options(q, q["id"]) for q in request["questions"]]  # the option lists, as the server adds them
result = classify(scorer, request)  # {"questions": [one decision each], "meta": {...}}
uv run selfjev classify examples/request.json   # from the command line; scores the request as given (no option lists)
uv run selfjev --help                           # serve / classify / eval / calibrate / compare / bench / finetune / rlcd / merge / deploy

Serve the best model (CUDA GPU)

# selfjev-4b (weights/selfjev_4b, eval2 95.8): the shared-prefix tree, exact, any number of questions
uv run selfjev serve --host 0.0.0.0 --port 8000
# or through vLLM (separate venv with vllm): exact, fast for one question, slow for many (speed.md)
uv run selfjev merge --adapter weights/selfjev_4b --out runs/selfjev_4b/merged
~/vllm-env/bin/python -m selfjev.cli serve --engine vllm --model-dir runs/selfjev_4b/merged

Both answer Jev's POST /v1/systemone (also at /api/alpha/decisions and /v1/decisions, API) and the internal POST /classify. The vLLM venv: uv venv ~/vllm-env --python 3.12 && uv pip install --python ~/vllm-env/bin/python vllm (0.30.0 measured), run with PYTHONPATH=src.

Fine-tune on your data, then RLCD (CUDA GPU)

uv run selfjev finetune --data my_train.jsonl --out runs/mine --init weights/selfjev_4b
uv run selfjev rlcd --data my_train.jsonl --out runs/mine_rlcd --init runs/mine/adapter

Data format, outputs and what RLCD optimizes: fine-tune and RLCD.

Tests

uv run pytest -q                              # everything, CPU, ~20 s (tiny random models; nothing is downloaded)
uv run pytest tests/engine                    # the Qwen3.5 tree: scores and gradients equal full sequences, the tree server
uv run pytest tests/training                  # RLCD learns calibrated probabilities; finetune + rlcd end to end

Data

New training data is a batch (scripts/data/grow_batch.sh, procedure in data/README.md); the sources are rebuilt with:

uv run python scripts/data/build_hf.py                                     # data/hf.jsonl from pinned HF revisions
uv run python scripts/data/build_data.py                                   # eval.jsonl + synthetic.jsonl, hash splits, overlap check
zsh -ic 'uv run python scripts/data/gen_hardcases.py --model openai/gpt-6-luna --budget 2 --max-questions 3000'  # resumes
zsh -ic 'uv run python scripts/data/judge_hardcases.py --jev'              # blind judge; never run two at once
uv run python scripts/data/build_hardcases.py                              # keep author = judge, leakage guard
uv run python scripts/data/build_all.py                                    # data/all.jsonl.gz + the catalog

Train and evaluate (on a GPU box, never on the laptop)

# a tagged box with an SSH-only security group and a shutdown cap (clean-up commands in the script's header)
scripts/aws/aws_launch.sh selfjev-mine 10 "us-east-2:g6e.2xlarge us-east-1:g6e.2xlarge us-west-2:g6e.2xlarge"
# the default model, selfjev-4b: every non-test question of data/all.jsonl.gz, 0.5 × label + 0.5 × Jev
uv run python scripts/train/jev_soft_targets.py  # local and free: runs/jev_all/{train,val}.jsonl.gz
nohup bash scripts/train/selfjev_4b.sh mine > train.log 2>&1 < /dev/null &   # on the box (what to sync: the script's header):
    # a GPU preflight against selfjev-4b's eval2 report, selfjev finetune (new LoRA r64, lr 2e-4, texts up to 16K), then
    # eval2, the dev benchmark and eval_llm for the best and the last checkpoint (≈ 9–10 h, one L40S)
# score any Qwen3.5 adapter: eval2, the dev benchmark (hf + eval test rows), eval_llm
uv run selfjev eval --adapter runs/mine/adapter --data data/ova/eval2.jsonl --out reports/mine/eval2
uv run selfjev eval --adapter runs/mine/adapter --data data/ova/hf.jsonl data/ova/eval.jsonl --split test --out reports/mine/test
uv run selfjev eval --adapter runs/mine/adapter --data data/ova/eval_llm.jsonl --out reports/mine/eval_llm
uv run python scripts/eval/eval2_summary.py      # regenerate reports/eval2/summary.md
uv run python scripts/eval/ledger.py             # regenerate the ledger table in docs/experiments.md
uv run python scripts/eval/calibration_table.py qwen35_4b_tree_scratch_jevall_ qwen35_4b_tree jev   # Brier, ECE, confident mistakes, McNemar

Older recipes

The previous default qwen35_4b_tree (adapter in weights/qwen35_4b_tree), the Qwen3 tree models (tree_4b_combo and earlier), the stock reranker pipeline, the forked-cache Qwen3.5 engine, the custom, jina and T5Gemma models and their training scripts and configs are at the tag archive/pre-cleanup-2026-09-27, where the package is src/personal_jev/, the CLI uv run pjev ... and the scripts sit directly in scripts/. Their round-2b and data/ova/ training copies rebuild byte for byte, from a checkout of the tag, with:

uv run python scripts/rebalance_nota.py && uv run python scripts/options_in_question.py data/synthetic.jsonl data/hardcases.jsonl data/hardcases_nb.jsonl data/hardcases_r3.jsonl

Jev and GPT-6 Astra on the same questions (responses cached under reports/external/cache/, so Jev reruns cost nothing; Astra's request asks for 8,192 output tokens since 2026-09-27, 6,000 before, so an Astra rerun misses the old cache and pays again; the run stops at --budget USD):

zsh -ic 'uv run python scripts/eval/compare_external.py --per-hf-family 300 --tag full --only jev --budget 5'
zsh -ic 'uv run python scripts/eval/compare_external.py --data data/eval2.jsonl --ours tree_4b/eval2 --only jev --tag eval2'
zsh -ic 'uv run python scripts/eval/audit_failures.py relabel --limit 40'   # PAID (user OK): blind relabel of failed test questions
uv run python scripts/eval/audit_failures.py report                         # free: reports/audit_2026-09-26/AUDIT.md

selfjev calibrate fits temperatures only on a report whose every prediction is from the calibration split and selects thresholds only on validation; anything else raises LeakageError. A calibration file is bound to the model revision, adapter sha256 and prompt sha that produced it.

This site

The site is built with Zensical from docs/ and zensical.toml, and deployed to GitHub Pages by .github/workflows/docs.yml on every push to master that touches the docs.

uv run --no-project --with zensical==0.0.65 python -m zensical serve      # http://127.0.0.1:8000, live reload
uv run --no-project --with zensical==0.0.65 python -m zensical build      # static site in site/

Use python -m zensical from the repo root: scripts/docs/docs_links.py, which points links to repo files outside docs/ at GitHub, must be importable. The leaderboard embeds the generated tables (reports/eval2/summary.md and the ledger section of docs/experiments.md), so re-running eval2_summary.py and ledger.py updates the site.

Layout

src/selfjev/       client.py + types.py (the SDK)   server/ (the HTTP API, fine-tuning jobs)   engine/ (qwen35: model,
                   prompt, readout; tree: shared-prefix tree, TreeServer; vllm)   core/ (request schema, option lists,
                   typed answers)   training/ (finetune, RLCD, losses, tree batching)   evaluation/ (eval, calibration,
                   benchmark, stats)   data/ (loading, validation, catalog, paid API clients)   deploy/ (AWS)   cli.py
scripts/           data/ (builders, generators, judges, batches), eval/ (summaries, Jev comparison, calibration),
                   train/ (the selfjev-4b recipe), aws/ (aws_launch.sh), docs/ (the site's link extension)
data/              hf / eval / synthetic / hardcases* / batches / eval2 / eval_llm (+ briefs and reviews), ova/ (the test
                   sets with option lists), all.jsonl.gz (built locally, not in git); data/README.md
deploy/            Dockerfile, docker-compose.yml
weights/           selfjev_4b (Git LFS) with model.json: base model, revision, recipe, scores
reports/           every eval report, benchmark, review and summary
docs/              this site: API, deploy, findings, write-ups, ledger (experiments.md) and journal
tests/             mirrors src/selfjev: core, data, engine, training, evaluation, server, client, deploy (CPU only)