Fine-tune and RLCD¶
Two commands train the best recipe (Qwen3.5-4B, shared-prefix tree, LoRA) on your own data: selfjev finetune for the
supervised step and selfjev rlcd for calibration training, which Jev calls RLCD. Both run on one CUDA GPU (an
AWS box, never the laptop) and write a run directory you can serve directly.
# 1. supervised fine-tune, from scratch or on top of the best model
uv run selfjev finetune --data my_train.jsonl --out runs/mine --init weights/selfjev_4b
# 2. RLCD on top of it
uv run selfjev rlcd --data my_train.jsonl --out runs/mine_rlcd --init runs/mine/adapter
# 3. serve the result (same option-list transform as in training)
uv run selfjev serve --adapter runs/mine_rlcd/adapter
The tree server scores exactly like standalone sequences (CPU test on a tiny model) but has not been timed on a GPU yet.
The vLLM path (selfjev merge, then selfjev serve --engine vllm) is the one measured on the real model (with
qwen35_4b_tree): same accuracy, fast for one question, slow for many (speed). Over HTTP,
selfjev serve --fine-tuning runs the same two commands as jobs and serves the result next to selfjev-4b
(API).
Code: src/selfjev/training/ (finetune.py the loop, rlcd.py the objective). Tests (CPU, a tiny random model):
tests/training/test_finetune.py.
Part 1: fine-tune vs RLCD in plain words¶
Picture a kid learning to guess which of two boxes holds the candy, and to say how sure they are.
- Fine-tuning is the answer key. After each guess the teacher opens the right box, and the kid adjusts to point at it more firmly. A confident wrong guess gets the biggest correction of all, even when the answer key itself has a mistake. On messy lessons the kid ends up very sure, even about things nobody can be sure of.
- RLCD is a points game. The kid says "I'm 70% sure it's this box". The box is opened and the kid gets points: sure and right earns a lot, sure and wrong loses a lot, unsure earns or loses a little. The kid tries a few different "how sure" answers, sees which ones earned more points, and leans toward those. The points are designed so that over many rounds the best strategy is to say exactly how sure you really are, and you can choose how much a confident miss costs. The points are computed from the same answer key: it is a different way of grading, not a different kind of learning.
- The points are the reward. A reward is a number that grades one attempt: higher is better.
What changes inside the model is the same in both: the same small set of adjustable weights (the LoRA adapter, 1.4% of the model) gets nudged a little after every batch. Fine-tuning nudges it toward "this is the answer"; RLCD nudges it toward "be this sure". RLCD does not train a different part of the model, and it adds no new layer.
Part 2: step by step, the data and one update of each¶
The data¶
| fine-tune | RLCD | |
|---|---|---|
| each row | text, question, options, target | the same row |
| also needs | nothing | the starting model's own probabilities for the row (computed once before training: the anchor) and a reward, the rule that grades an attempt |
| where the grade comes from | here from the target (log, Brier and spherical score); it could be a cost table, a judge's verdict or what happened next, with no target at all |
In our runs both commands read the same file, in order: selfjev finetune first, then selfjev rlcd --init on its result.
One update¶
Both start identically: the model reads the text through the tree and gives each option a score, here billing 0.49 and tech 0.00, which a softmax turns into 62% billing, 38% tech. Then:
Fine-tune
- Loss = −log p(target) = −log 0.62 = 0.48.
- The push on each score is p − target: billing 0.62 − 1 = −0.38 (raise it), tech 0.38 − 0 = +0.38 (lower it). The push is exactly how wrong the model was.
- Backpropagation carries the push through all 32 layers into the LoRA matrices, and the optimizer moves them a small step (learning rate 2e-4). The next text like this one gets a higher p(billing).
RLCD
- Try 8 answers near the model's own: add a little random noise (sd 0.3) to the scores. Five of them, as p(billing): 0.48, 0.55, 0.62, 0.70, 0.78.
- Grade each try with the reward (log + Brier + spherical score against billing): −0.60, −0.23, +0.09, +0.38, +0.62. The average is +0.05.
- Advantage = grade − average: −0.65, −0.28, +0.03, +0.33, +0.57. Tries above average should become more likely and tries below less likely; here that means raise billing and lower tech.
- Add the anchor: a small pull (KL penalty, weight 0.05) back toward the starting model's probabilities, so a few batches cannot drag it far.
- Backpropagation into the same LoRA matrices, with smaller steps (learning rate 5e-5).
Despite the name, none of this is reinforcement learning. RL is an agent acting in an environment: its actions change what happens next, rewards can come late, and it must explore to learn what each action does. Here there is no environment and no sequence of actions, and the grade of every possible answer is known from the label. The sampled tries only borrow an RL tool, a policy-gradient estimator: averaged over many tries it follows the slope of the expected score, a slope that could be computed exactly. "RLCD" is the name Jev (and Laya) use; what it does here is supervised fine-tuning with a calibration score as the loss.
How the reward changes the model¶
In fine-tuning the push on a score is fixed by one formula (p − target). In RLCD the push comes from the grades, so the reward decides how hard each example pulls. How hard each score pulls the correct option's score up, with two options, in three situations:
| situation | p(target) | log score (= fine-tuning) | Brier | spherical |
|---|---|---|---|---|
| right but unsure | 0.62 | 0.38 | 0.36 | 0.23 |
| right and sure | 0.95 | 0.05 | 0.01 | 0.003 |
| sure and wrong | 0.05 | 0.95 | 0.18 | 0.05 |
- The log score pushes hardest on confident mistakes. When the label is right, that fixes a real error; when the label is noise or the text is genuinely ambiguous, the model bends to fit it and learns to be sure where it should not be.
- Brier and spherical pull much less on those cases, so the model can stay unsure there. That is where hedging comes from. On examples where the model is unsure but right, all three pull about the same.
- All three are proper scoring rules: over many similar texts that are billing 70% of the time, each one is best at saying 70%. They differ in how much a single example can pull.
- The default reward mixes all three (
--reward log=1,brier=1,spherical=1, so confident mistakes still get pushed about as hard as in fine-tuning).--reward brier=1,spherical=1drops the log score for the most hedging.
Data¶
One question per line, the same format as every eval file in data/:
{"state": "Ticket 4411: the invoice was charged twice ...",
"question": {"type": "multiclass", "instruction": "Which team should handle this?",
"candidates": [{"id": "billing", "description": "billing and refunds"},
{"id": "tech", "description": "outages and bugs"}]},
"target": "billing"}
typeisbinary(targettrue/false, no candidates),multiclass(target: one candidate id) ormultilabel(target: a list of candidate ids, possibly empty).idandfamilyare optional. Several questions about the samestateshare one tree: the text is encoded once.soft(optional): a teacher's probabilities, P(yes) for binary or{candidate id: p}otherwise. Training then targets (1 −--soft-weight) × label +--soft-weight×soft(default 0.5, so the label stays the answer); validation stays on the labels.scripts/train/jev_soft_targets.pywrites Jev's, fromdata/all.jsonl.gz.--valgives a validation file; without it, 5% of--data(at most 1,000 questions) is held out.- Every option is listed in the question text (
options.py), as for the best model;--no-options-in-questionturns that off. Serve with the same setting. - Questions longer than
--max-length(default 16,384 tokens: root + question + longest candidate) are dropped and counted intrain_meta.json, never truncated.
Output¶
<out>/adapter (best validation checkpoint), <out>/adapter_last, <out>/train_meta.json (base model and revision,
data sha256, settings, every validation). Validation reports accuracy, cross-entropy, Brier score, expected calibration
error (10 bins) and the number of confidently wrong decisions (confidence ≥ 0.9). finetune keeps the checkpoint with
the lowest validation cross-entropy, rlcd the one with the lowest Brier score. To keep a result with the repo, copy
the adapter into weights/.
finetune¶
Cross-entropy on the targets (softmax over a multiclass question's candidates, a sigmoid per yes/no decision): the
recipe that produced qwen35_4b_tree (the previous default, archived) and, from scratch on all the non-test data with
texts up to 16K tokens and Jev's probabilities as soft targets, the default weights/selfjev_4b: LoRA r64 on the
attention and DeltaNet projections, lr 2e-4 with 5% warm-up and linear decay, batches of whole states packed to 16,384
tokens × 2 accumulation steps (the defaults; qwen35_4b_tree used 8,192 × 4), per-layer activation checkpointing.
Start from --init weights/selfjev_4b to adapt the default model to a new domain, or without --init to train a
fresh adapter (--lora-r, default 64). Every run uses Qwen3.5-4B; the Qwen3.5-2B option for quick runs is at the tag
archive/pre-cleanup-2026-09-27.
rlcd settings¶
--reward: weights oflog,brier,spherical(strictly proper scoring rules) andaccuracy(the argmax decision is right; not proper on its own, keep a proper term next to it).confident_missin--rewardis a cost, not a score: −1 for every decision made with confidence ≥ 0.9 that the target says is wrong (expected under a soft target).confident_miss=5makes a confident mistake cost 5×; with a 95% target the best report drops from 0.95 to just under 0.9 (test). It is not differentiable: the sampled tries are what let RLCD optimize it, the one reward here a fine-tune cannot.--samples(8) tries per question, drawn from a Gaussian around the model's scores with sd--sigma(0.3).--beta(0.05): the KL penalty to the starting model's probabilities, computed once before training.--lrdefaults to 5e-5 (fine-tune: 2e-4); the best checkpoint is the one with the lowest validation Brier score.
tests/training/test_finetune.py checks the property that matters: trained only with this objective, on a question whose
answer is "yes" 70% of the time, the model learns to say 70% (binary), and 60% for a three-way choice with a
60/30/10 split.
Jev says it is trained with "RLCD" and gives no details; Laya, an open Jev-like engine, uses the name for policy
gradient on proper-scoring-rule rewards (landscape), which is what selfjev rlcd implements.
First test on the real model (2026-09-25): no gain¶
RLCD from qwen35_4b_tree (then the default) on 4,412 labeled questions it never trained on, against a plain fine-tune on the
same questions (JOURNAL, reports/rlcd_2026-09-25/):
| eval2 | accuracy | ECE | confidently wrong (≥ 0.9) |
|---|---|---|---|
| start | 95.58 | 0.005 | 30 |
| + RLCD | 95.63 | 0.010 | 34 |
| + fine-tune (control) | 95.43 | 0.005 | 31 |
| Jev | 97.24 | 0.041 | 7 |
- On questions the model had already trained on, RLCD made it overconfident within 100 steps: keep RLCD data fresh.
- With one hard label per question, a proper-score reward is the same signal as cross-entropy, only noisier: it cannot say "be less sure here". The model is already calibrated (ECE 0.005, Jev 0.041); Jev's edge is fewer mistakes and hedged ones.
- The sampled reports also bias the optimum toward overconfidence as
--sigmagrows (70% base rate: 0.695 at 0.3, 0.754 at 1.0); keep--sigmasmall. - For RLCD to matter, the reward has to carry more than the label: soft targets (judges' agreement, a teacher's probabilities) or a cost that punishes confident mistakes more than it rewards confident right answers.
Second test (2026-09-26): Jev's probabilities as soft targets, all the data¶
From qwen35_4b_tree, one epoch over 69,528 non-test questions of data/all.jsonl.gz, target 0.5 × label +
0.5 × Jev (Jev is ≥ 0.9 sure of the label on 77% of them, hedges on 15%, disagrees on 8%). Final checkpoints
(JOURNAL, reports/rlcd_jev_2026-09-26/):
| eval2 (1,991 q) | accuracy (paired vs start) | Brier | ECE | wrong decisions | ≥ 0.9 sure | mean confidence when wrong |
|---|---|---|---|---|---|---|
qwen35_4b_tree (start) |
95.58 | 0.0480 | 0.0051 | 99 | 30 | 0.755 |
| A: RLCD on Jev targets | 95.43 (15 / 18, p = 0.73) | 0.0446 | 0.0119 | 100 | 22 | 0.724 |
| B: fine-tune on Jev targets | 95.73 (19 / 16, p = 0.74) | 0.0438 | 0.0245 | 94 | 14 | 0.708 |
| Jev | 97.24 | 0.0335 | 0.0406 | 57 | 7 | 0.670 |
| dev benchmark (3,471 q) | accuracy (paired vs start) | Brier | ECE | wrong decisions | ≥ 0.9 sure | mean confidence when wrong |
|---|---|---|---|---|---|---|
qwen35_4b_tree (start) |
84.44 | 0.1787 | 0.0056 | 582 | 83 | 0.686 |
| A: RLCD on Jev targets | 84.50 (76 / 74, p = 0.93) | 0.1816 | 0.0113 | 593 | 88 | 0.701 |
| B: fine-tune on Jev targets | 84.41 (87 / 88, p = 1) | 0.1820 | 0.0127 | 603 | 75 | 0.693 |
| Jev | 82.71 | 0.2105 | 0.0444 | 708 | 253 | 0.791 |
- Soft targets do what hard labels could not: confident mistakes on eval2 halve (30 → 14, Jev 7) and Brier improves 9%, at the same accuracy. The dev benchmark is flat.
- RLCD on the same targets moves less (the KL anchor) and is no better than the fine-tune: A vs B 4 / 10 on eval2 (p = 0.18). All its rewards peak at the same target as cross-entropy, so the sampling only adds noise.
- For these runs from a trained adapter, the best-by-validation checkpoint was step 0 (hard-label loss and Brier
penalize hedging): use
adapter_last. From scratch,selfjev-4b's best was step 1,800 of 1,803.
C: RLCD with a confident-mistake cost, on top of the fine-tune¶
--reward log=1,brier=1,spherical=1,confident_miss=5 from B's result, same data and settings:
| eval2 (1,991 q) | accuracy (paired vs start) | Brier | ECE | wrong decisions | ≥ 0.9 sure | mean confidence when wrong |
|---|---|---|---|---|---|---|
qwen35_4b_tree (start) |
95.58 | 0.0480 | 0.0051 | 99 | 30 | 0.755 |
| B: fine-tune on Jev targets | 95.73 (19 / 16, p = 0.74) | 0.0438 | 0.0245 | 94 | 14 | 0.708 |
| C: B + RLCD with a 5× confident-mistake cost | 95.68 (24 / 22, p = 0.88; vs B 9 / 10, p = 1) | 0.0518 | 0.0431 | 96 | 8 | 0.690 |
| Jev | 97.24 | 0.0335 | 0.0406 | 57 | 7 | 0.670 |
| dev benchmark (3,471 q) | accuracy | Brier | ECE | wrong decisions | ≥ 0.9 sure | mean confidence when wrong |
|---|---|---|---|---|---|---|
qwen35_4b_tree (start) |
84.44 | 0.1787 | 0.0056 | 582 | 83 | 0.686 |
| B | 84.41 | 0.1820 | 0.0127 | 603 | 75 | 0.693 |
| C | 84.41 (vs B 17 / 17) | 0.1835 | 0.0298 | 597 | 48 | 0.672 |
| Jev | 82.71 | 0.2105 | 0.0444 | 708 | 253 | 0.791 |
- The cost does what fine-tuning cannot: confident mistakes 14 → 8 on eval2 (Jev 7), 75 → 48 on the dev benchmark, ECE and confidence when wrong at Jev's level, same accuracy.
- The price is sharpness: Brier 0.0438 → 0.0518, because the model also backs off on answers it gets right. Use B for the best probabilities, C for Jev-like caution; a smaller weight trades between the two.
- But compared at equal coverage, C is no better than B: it gets its few confident mistakes by being less sure overall (≥ 0.9 sure on 77% of decisions, B 88.8%), the same trade a higher threshold on B gives for free:
| mistakes among the k most confident decisions (eval2, 3,491) | 70% | 77% | 85% | 88% | 92% |
|---|---|---|---|---|---|
qwen35_4b_tree |
3 | 8 | 14 | 16 | 29 |
| B | 4 | 5 | 9 | 13 | 20 |
| C | 4 | 8 | 10 | 12 | 24 |
| Jev | 3 | 3 | 4 | 7 | 8 |
- So the cost changed the model's confidence scale, not its ranking. Jev's real edge is ranking: it knows which answers are risky.