SelfJev¶
Can an open model do what Jev does? Jev turns a text and a list of typed questions (yes/no, pick one, pick many) into calibrated decisions in about 150 ms, with no text generation. SelfJev rebuilds that interface on an open model, Qwen3.5-4B with a small LoRA adapter (Qwen3 models before it), and measures every step against Jev on the same questions.
New to the shorthand (eval2, round 2b, r64, stock pairs…)? Hover any dotted-underlined term, or read the glossary.
The best model, selfjev-4b (weights/selfjev_4b), is Qwen3.5-4B with a rank-64 LoRA adapter
trained with our shared-prefix tree on all 80K non-test questions (public datasets plus LLM-written cases that a blind
judge confirmed, texts up to 16K tokens), with half-weight Jev probabilities as soft targets and every option listed in
the question. It scores 95.8% on eval2 (Jev 97.2%) and 93.1% on eval_llm (Jev 92.5%), with 11 confident mistakes on
eval2 where its predecessor qwen35_4b_tree (95.6%, archived) made 30. The fastest one measured
was the Qwen3 tree tree_4b_combo (94.5%, archived), which on vLLM was cheaper per request than Jev on a busy GPU;
selfjev-4b's own tree server is not timed on a GPU yet.
The story in twelve steps¶
| when | step | result |
|---|---|---|
| 09-22 | Qwen3-Reranker-0.6B + LoRA on the laptop | 61.0 → 73.5% on the dev benchmark |
| 09-23 | Jev and GPT-6 Astra on the same questions; scale to 4B and 8B on AWS | Jev 82.7, Astra 85.8; 4B 80.3 = 8B 80.7: fine-tuning matters far more than size |
| 09-23 | A custom cross-attention model that reads the text once | 38–43× faster but 39–58% accurate: yes/no never learned |
| 09-23 | The shared-prefix tree: read once, but every branch attends to the text in every layer | 81.6%, 32–37× faster than stock pairs |
| 09-23 | Learning curves, bigger adapters, an 8B, an Instruct base | everything lands at 80–82% on the dev benchmark |
| 09-23 | Round-2 data: 10K hard cases written by 5 models, kept only when a blind judge agrees | better on the traps, but CLINC over-rejection |
| 09-24 | eval2: a frozen 1,991-question target-task test set | the hidden effects appear: data +5.5, Instruct base +3.0, rank 64 +2.1 |
| 09-24 | Combine the levers: Instruct base, rank 64, round-2b data | 92.7% on eval2 |
| 09-24 | Round 3 (38.6K more verified questions) + every option listed in the question | 94.5%: each adds about a point, and they stack |
| 09-25 | Qwen3.5-4B, a hybrid model (3 recurrent Gated DeltaNet layers per attention layer), same levers | 94.5% trained on full sequences capped at 2K tokens |
| 09-25 | A tree for Qwen3.5: its recurrent layers run level by level from copied states | 95.6%, trained on texts up to 8K in 4.1 h |
| 09-26 | selfjev-4b: a new adapter on all 80K non-test questions, texts up to 16K, Jev's probabilities as half-weight soft targets |
95.8%, eval_llm 93.1% (Jev 92.5%), confident mistakes on eval2 30 → 11 |
Along the way: Qwen3.5 on vLLM keeps its accuracy but is slow with many questions (vLLM reuses its recurrent state only every 528 tokens), and RLCD (Jev's name for training on calibration scores; despite the name, no reinforcement learning is involved) gave no gain on hard labels; Jev's probabilities as soft targets are what cut the confident mistakes (fine-tune and RLCD).
Where to go¶
-
Jev's decisions API on
selfjev-4b, the Python client, and fine-tuning jobs over HTTP. -
Self-host on your GPU, in Docker, or with one command on AWS.
-
What we learned, each finding with its numbers and evidence.
-
Every model on eval2 and the dev benchmark, with all slices.
-
Tree vs pairs, end-to-end latency vs Jev, vLLM, cost per request.
-
Negative results, so nobody pays for them twice.
-
The architecture behind every best model.
-
Public sets, LLM-written hard cases, blind judging, eval2.
-
Train the best recipe on your own data, then calibration training (what Jev calls RLCD).
-
What to do next to close the last 1.4 points; speed is a deployment question since the H100 sweep (2026-09-26).
-
The live ledger and the dated journal the working sessions write.
Read the numbers with care
- eval2 labels are written by LLMs and kept only when two blind LLM judges agree: LLM-verified, not human-verified.
- The dev benchmark has been reused for many decisions, and eval2 has now informed the research direction: neither is a clean final test.
- Every result is a single run with one seed. Paired McNemar tests give the significance; wins with p ≈ 0.06–0.07 need a second seed.