Skip to content

Leaderboard

eval2: the target task

1,991 authored questions (how it was built). This is the benchmark that decides between models, with eval_llm for the LLM-evaluation use cases.

eval2 question accuracy, %. Bars start at 0; the dashed line is Jev. All ours use LoRA.

run base recipe eval2 binary multiclass multilabel EM dev benchmark
Jev (API) undisclosed undisclosed 97.2 97.8 98.1 94.2 82.7
selfjev-4b (qwen35_4b_tree_scratch_jevall_), the default Qwen3.5-4B tree, trained from scratch on 79.9K non-test questions of data/all.jsonl.gz (texts ≤ 16K, incl. the LLM-evaluation data, llm_multilabel_v1, numdate_neg_v1), target 0.5 × label + 0.5 × Jev, r64, all options in the question; eval_llm 93.1 (Jev 92.5) 95.8 96.8 97.0 91.4 83.8
qwen35_4b_tree, the previous default Qwen3.5-4B shared-prefix tree in training (src/selfjev/engine/tree.py, texts ≤ 8K: 51.8K q), text shared at inference, r64, round-2b + round-3 data, all options in the question 95.6 96.9 96.8 90.1 84.4
qwen35_4b_combo Qwen3.5-4B each option trained as a full sequence (no tree in training, texts ≤ 2K: 43.8K q), text shared at inference, r64, round-2b + round-3 data, all options in the question 94.5 96.2 95.8 88.2 84.3
tree_4b_combo Qwen3-4B-Instruct-2507 tree, r64, round-2b + round-3 data (51.8K q), hard cases not capped, all options in the question 94.5 96.0 97.5 85.9 82.7
tree_4b_instruct_r3 Qwen3-4B-Instruct-2507 tree, r64, round-2b + round-3 data (51.9K q), hard cases not capped 93.3 95.1 96.1 84.3 82.8
tree_4b_combo_r2 Qwen3-4B-Instruct-2507 tree, r64, round-2b data, hard cases not capped, all options in the question (served: merged + vLLM, 93.0) 92.9 94.4 95.8 84.3 83.5
tree_4b_instruct_r2x64 Qwen3-4B-Instruct-2507 tree, r64, round-2b data, hard cases not capped per family 92.7 94.6 95.4 83.5 82.7
tree_4b_combo_ptr Qwen3-4B-Instruct-2507 as tree_4b_combo_r2, options numbered once, leaves say "option k" 92.0 94.9 94.3 80.9 82.7
tree_4b_ova Qwen3-Reranker-4B tree, r16, round-2b data, all options in the question 91.6 93.8 94.9 80.4 82.6
Qwen3.8-27B-FP8, zero-shot (teacher) Qwen3.8-27B no training 91.4 93.4 97.8 75.9 —
curve/tree_4b_r2b_r64_mlp Qwen3-Reranker-4B tree, r64 + MLP, round-2b data 91.3 93.7 94.3 80.4 82.4
tree_4b_ova_kd Qwen3-Reranker-4B as tree_4b_ova + 27B soft targets 91.0 92.9 95.6 78.5 82.5
tree_4b_instruct_r3_step300 Qwen3-4B-Instruct-2507 round-3 run, stopped; step 300 of 915 90.8 93.1 95.1 78.0 81.5
tree_4b_r2b Qwen3-Reranker-4B tree, r16, round-2 data, none capped 90.6 92.9 94.1 78.8 81.2
tree_4b_r2 Qwen3-Reranker-4B tree, r16, round-2 data 90.4 92.8 95.8 75.4 80.6
curve/tree_4b_r2b_r64 Qwen3-Reranker-4B tree, r64, round-2b data 90.3 92.9 94.9 75.9 81.2
tree_4b_instruct Qwen3-4B-Instruct-2507 tree, r16, round-1 data 88.1 91.4 92.7 72.3 80.4
curve/tree_4b_r64 Qwen3-Reranker-4B tree, r64, round-1 data 87.2 90.5 92.7 69.9 81.5
lora_4b Qwen3-Reranker-4B stock pairs, r16, round-1 data 86.5 89.7 91.6 70.4 80.3
curve/tree_4b_mlp Qwen3-Reranker-4B tree, r16 + MLP, round-1 data 86.3 90.0 91.9 68.1 81.6
tree_4b Qwen3-Reranker-4B tree, r16, round-1 data 85.1 89.4 91.6 63.9 81.6
T5Gemma 2, decoder-only LoRA (audit) t5gemma-2-1b-1b shared encoder, round-2b data 76.8 82.6 83.6 50.8 73.8
jina_r2b jina-reranker-v3.5 (0.6B) listwise, round-2b data 73.3 79.8 83.5 40.1 76.6
T5Gemma 2 (t5_r1) t5gemma-2-1b-1b shared encoder, round-1 data 73.0 79.5 79.8 45.0 75.4
lora_pilot Qwen3-Reranker-0.6B stock pairs, r16, round-1 data 68.8 75.7 75.7 39.8 73.5
jina_zeroshot jina-reranker-v3.5 (0.6B) no training 46.5 55.8 54.5 9.4 —
  • Sources: reports/eval2/summary.md (generated by scripts/eval/eval2_summary.py) and the 27B teacher's predictions in reports/teacher/qwen38_27b_eval2.jsonl.
  • GPT-6 Astra is not scored on eval2: it was one of the two judges that decided which questions were kept.
  • "Round-1 data" is the 10,112-question mix (public sets + synthetic); round 2 adds about 10K verified hard cases (data).
  • Only selfjev-4b's adapter is on master (weights/selfjev_4b). qwen35_4b_tree and tree_4b_combo are at tag archive/pre-cleanup-2026-09-27, with the code of every other architecture (Qwen3 trees, stock pairs, jina, T5Gemma, the 27B teacher). All these reports were scored before the 2026-09-27 cleanup, the Qwen3.5 rows by the forked-cache engine that TreeServer replaced.
All eval2 slices: type, tier, author, text length, trap, paired tests (generated)

eval2 results (frozen target-task test set)

Generated by uv run python scripts/eval/eval2_summary.py; do not hand-edit. 1991 questions, data/eval2.jsonl (see data/eval2/REVIEW.md). Never train, select prompts or fit calibration on it. GPT-6 Astra is one of the two judges that decided which questions were kept, so it is not a fair competitor here and isn't scored.

overall

slice n jev qwen35_4b_tree_scratch_jevall_ qwen35_4b_tree_sft_jevall__last qwen35_4b_tree_cost_jevall__last selfjev_4b_treeserver qwen35_4b_tree_rlcd_fresh_ qwen35_4b_tree_scratch_jevall__last qwen35_4b_tree qwen35_4b_tree_rlcd_jevall__last qwen35_4b_tree_sft_fresh_ selfjev_4b_v2 selfjev_4b_repro qwen35_4b_combo tree_4b_combo qwen35_4b_r2x64 tree_4b_instruct_r3 tree_4b_combo_r2 eikos_4b tree_4b_instruct_r2x64 tree_4b_combo_ptr tree_4b_ova curve/tree_4b_r2b_r64_mlp tree_4b_ova_kd tree_4b_instruct_r3_step300 tree_4b_r2b tree_4b_r2 curve/tree_4b_r2b_r64 tree_4b_instruct curve/tree_4b_r64 lora_4b curve/tree_4b_mlp tree_4b qwen35_r1 t5_round2b_2026-09-24/decoder_only_audit/t5_r2b jina_r2b t5_round2b_2026-09-24/t5_r1 lora_pilot jina_zeroshot
all 1991 97.2 95.8 95.7 95.7 95.7 95.6 95.6 95.6 95.4 95.4 95.4 95.1 94.5 94.5 93.7 93.3 92.9 92.8 92.7 92.0 91.6 91.3 91.0 90.8 90.6 90.4 90.3 88.1 87.2 86.5 86.3 85.1 84.3 76.8 73.3 73.0 68.8 46.5

type

slice n jev qwen35_4b_tree_scratch_jevall_ qwen35_4b_tree_sft_jevall__last qwen35_4b_tree_cost_jevall__last selfjev_4b_treeserver qwen35_4b_tree_rlcd_fresh_ qwen35_4b_tree_scratch_jevall__last qwen35_4b_tree qwen35_4b_tree_rlcd_jevall__last qwen35_4b_tree_sft_fresh_ selfjev_4b_v2 selfjev_4b_repro qwen35_4b_combo tree_4b_combo qwen35_4b_r2x64 tree_4b_instruct_r3 tree_4b_combo_r2 eikos_4b tree_4b_instruct_r2x64 tree_4b_combo_ptr tree_4b_ova curve/tree_4b_r2b_r64_mlp tree_4b_ova_kd tree_4b_instruct_r3_step300 tree_4b_r2b tree_4b_r2 curve/tree_4b_r2b_r64 tree_4b_instruct curve/tree_4b_r64 lora_4b curve/tree_4b_mlp tree_4b qwen35_r1 t5_round2b_2026-09-24/decoder_only_audit/t5_r2b jina_r2b t5_round2b_2026-09-24/t5_r1 lora_pilot jina_zeroshot
binary 1016 97.8 96.8 96.8 96.9 96.7 96.9 96.8 96.9 96.9 96.8 96.9 96.1 96.2 96.0 95.6 95.1 94.4 94.9 94.6 94.9 93.8 93.7 92.9 93.1 92.9 92.8 92.9 91.4 90.5 89.7 90.0 89.4 87.5 82.6 79.8 79.5 75.7 55.8
multiclass 593 98.1 97.0 97.5 97.6 97.0 97.0 96.6 96.8 97.1 97.0 96.5 96.3 95.8 97.5 95.1 96.1 95.8 95.8 95.4 94.3 94.9 94.3 95.6 95.1 94.1 95.8 94.9 92.7 92.7 91.6 91.9 91.6 88.7 83.6 83.5 79.8 75.7 54.5
multilabel 382 94.2 91.4 90.3 89.5 91.1 90.1 91.1 90.1 89.0 89.5 89.5 90.6 88.2 85.9 86.6 84.3 84.3 82.5 83.5 80.9 80.4 80.4 78.5 78.0 78.8 75.4 75.9 72.3 69.9 70.4 68.1 63.9 68.8 50.8 40.1 45.0 39.8 9.4

tier

slice n jev qwen35_4b_tree_scratch_jevall_ qwen35_4b_tree_sft_jevall__last qwen35_4b_tree_cost_jevall__last selfjev_4b_treeserver qwen35_4b_tree_rlcd_fresh_ qwen35_4b_tree_scratch_jevall__last qwen35_4b_tree qwen35_4b_tree_rlcd_jevall__last qwen35_4b_tree_sft_fresh_ selfjev_4b_v2 selfjev_4b_repro qwen35_4b_combo tree_4b_combo qwen35_4b_r2x64 tree_4b_instruct_r3 tree_4b_combo_r2 eikos_4b tree_4b_instruct_r2x64 tree_4b_combo_ptr tree_4b_ova curve/tree_4b_r2b_r64_mlp tree_4b_ova_kd tree_4b_instruct_r3_step300 tree_4b_r2b tree_4b_r2 curve/tree_4b_r2b_r64 tree_4b_instruct curve/tree_4b_r64 lora_4b curve/tree_4b_mlp tree_4b qwen35_r1 t5_round2b_2026-09-24/decoder_only_audit/t5_r2b jina_r2b t5_round2b_2026-09-24/t5_r1 lora_pilot jina_zeroshot
e2_hard 673 96.7 94.5 95.5 95.2 94.4 94.7 94.4 94.7 94.9 94.1 93.9 93.8 93.6 93.8 92.3 92.3 90.6 90.8 90.5 90.3 90.8 89.6 89.7 89.0 89.9 89.3 89.5 85.6 85.3 84.2 83.8 83.2 80.4 74.7 70.7 68.2 64.6 50.4
e2_simple 664 98.5 98.5 98.5 98.5 98.5 98.6 98.5 98.6 98.5 98.5 98.6 97.9 98.0 96.8 97.7 96.8 97.3 95.9 96.8 96.5 94.9 94.7 95.3 95.9 95.2 94.6 93.8 92.8 94.4 92.8 93.4 92.3 92.3 85.1 80.6 83.4 80.1 46.1
e2_very_hard 654 96.5 94.3 93.1 93.3 94.2 93.6 94.0 93.4 92.8 93.7 93.6 93.6 91.9 92.8 91.1 90.8 90.7 91.6 90.8 89.1 89.0 89.6 87.8 87.5 86.5 87.2 87.5 86.1 81.8 82.6 81.8 79.8 80.1 70.5 68.5 67.3 61.6 43.0

author

slice n jev qwen35_4b_tree_scratch_jevall_ qwen35_4b_tree_sft_jevall__last qwen35_4b_tree_cost_jevall__last selfjev_4b_treeserver qwen35_4b_tree_rlcd_fresh_ qwen35_4b_tree_scratch_jevall__last qwen35_4b_tree qwen35_4b_tree_rlcd_jevall__last qwen35_4b_tree_sft_fresh_ selfjev_4b_v2 selfjev_4b_repro qwen35_4b_combo tree_4b_combo qwen35_4b_r2x64 tree_4b_instruct_r3 tree_4b_combo_r2 eikos_4b tree_4b_instruct_r2x64 tree_4b_combo_ptr tree_4b_ova curve/tree_4b_r2b_r64_mlp tree_4b_ova_kd tree_4b_instruct_r3_step300 tree_4b_r2b tree_4b_r2 curve/tree_4b_r2b_r64 tree_4b_instruct curve/tree_4b_r64 lora_4b curve/tree_4b_mlp tree_4b qwen35_r1 t5_round2b_2026-09-24/decoder_only_audit/t5_r2b jina_r2b t5_round2b_2026-09-24/t5_r1 lora_pilot jina_zeroshot
claude-opus-5.5 560 96.1 94.5 93.6 93.2 94.5 94.3 94.1 94.5 93.2 94.6 93.2 93.0 92.7 93.4 91.4 91.4 89.8 91.1 90.9 88.9 89.1 88.8 88.0 88.0 87.7 87.5 87.0 85.4 83.4 83.2 83.6 82.5 80.5 72.3 68.2 69.3 66.4 40.0
glm-5.3 782 97.3 96.0 96.0 96.2 95.9 95.7 95.9 95.5 95.7 95.1 95.5 95.0 95.3 93.1 93.7 93.1 92.7 92.6 91.3 92.3 90.9 91.6 91.0 90.8 90.3 89.8 90.2 87.5 86.7 87.3 85.8 84.0 84.8 76.6 73.9 71.2 68.3 47.7
kimi-k3 649 98.2 96.6 97.2 97.2 96.5 96.8 96.6 96.6 97.1 96.5 97.1 96.9 95.2 97.1 95.7 95.2 95.7 94.5 96.0 94.3 94.5 93.2 93.4 93.2 93.4 93.5 93.2 91.4 91.1 88.4 89.4 88.8 86.9 80.9 76.9 78.3 71.5 50.7

text length (tokens)

slice n jev qwen35_4b_tree_scratch_jevall_ qwen35_4b_tree_sft_jevall__last qwen35_4b_tree_cost_jevall__last selfjev_4b_treeserver qwen35_4b_tree_rlcd_fresh_ qwen35_4b_tree_scratch_jevall__last qwen35_4b_tree qwen35_4b_tree_rlcd_jevall__last qwen35_4b_tree_sft_fresh_ selfjev_4b_v2 selfjev_4b_repro qwen35_4b_combo tree_4b_combo qwen35_4b_r2x64 tree_4b_instruct_r3 tree_4b_combo_r2 eikos_4b tree_4b_instruct_r2x64 tree_4b_combo_ptr tree_4b_ova curve/tree_4b_r2b_r64_mlp tree_4b_ova_kd tree_4b_instruct_r3_step300 tree_4b_r2b tree_4b_r2 curve/tree_4b_r2b_r64 tree_4b_instruct curve/tree_4b_r64 lora_4b curve/tree_4b_mlp tree_4b qwen35_r1 t5_round2b_2026-09-24/decoder_only_audit/t5_r2b jina_r2b t5_round2b_2026-09-24/t5_r1 lora_pilot jina_zeroshot
≤64 439 97.3 95.4 96.1 95.7 95.4 96.4 95.2 96.1 95.9 95.7 95.7 95.2 95.9 94.3 92.7 91.6 91.8 90.2 90.0 91.3 90.9 90.0 89.5 90.0 89.1 89.1 88.2 85.0 85.4 85.0 85.4 85.2 80.9 75.4 74.9 71.3 67.2 49.0
65–256 508 96.7 94.5 93.3 92.9 94.3 93.9 94.3 94.3 93.1 93.9 93.1 92.9 92.1 93.3 91.3 92.1 91.9 90.9 92.3 89.8 91.3 90.0 90.4 90.2 88.6 89.0 88.8 87.8 87.4 86.2 85.8 83.9 83.1 74.4 77.4 71.5 68.5 50.6
257–1K 469 96.6 95.3 95.3 95.7 95.1 95.5 95.1 95.1 95.1 95.5 95.1 94.7 93.4 93.4 93.0 93.0 91.0 94.5 92.1 91.9 90.0 90.6 90.0 90.4 90.8 90.0 90.4 88.1 86.6 84.9 86.6 84.6 84.0 75.9 74.2 72.9 71.6 46.9
1K–4K 409 98.0 97.6 97.3 97.8 97.6 96.6 97.6 96.6 97.1 96.3 97.3 97.3 96.6 96.6 97.3 96.6 96.3 94.4 96.1 95.4 93.2 93.9 94.6 93.4 93.9 93.2 92.7 90.5 89.7 90.5 88.5 87.8 89.7 81.7 68.7 76.5 69.2 41.1
>4K 166 98.8 97.6 99.4 98.8 97.6 97.0 97.6 97.0 98.2 97.0 97.6 97.0 96.4 96.4 97.0 94.6 95.2 96.4 94.6 92.8 94.6 94.6 90.4 89.8 91.6 92.2 94.0 92.2 86.7 86.7 84.3 83.7 84.3 78.3 65.1 73.5 65.1 39.8

trap

slice n jev qwen35_4b_tree_scratch_jevall_ qwen35_4b_tree_sft_jevall__last qwen35_4b_tree_cost_jevall__last selfjev_4b_treeserver qwen35_4b_tree_rlcd_fresh_ qwen35_4b_tree_scratch_jevall__last qwen35_4b_tree qwen35_4b_tree_rlcd_jevall__last qwen35_4b_tree_sft_fresh_ selfjev_4b_v2 selfjev_4b_repro qwen35_4b_combo tree_4b_combo qwen35_4b_r2x64 tree_4b_instruct_r3 tree_4b_combo_r2 eikos_4b tree_4b_instruct_r2x64 tree_4b_combo_ptr tree_4b_ova curve/tree_4b_r2b_r64_mlp tree_4b_ova_kd tree_4b_instruct_r3_step300 tree_4b_r2b tree_4b_r2 curve/tree_4b_r2b_r64 tree_4b_instruct curve/tree_4b_r64 lora_4b curve/tree_4b_mlp tree_4b qwen35_r1 t5_round2b_2026-09-24/decoder_only_audit/t5_r2b jina_r2b t5_round2b_2026-09-24/t5_r1 lora_pilot jina_zeroshot
distractor 474 97.0 93.9 94.7 94.7 93.7 93.2 93.7 92.8 93.9 93.2 93.9 92.6 92.0 92.6 91.8 91.1 90.7 92.4 91.1 90.1 89.7 89.5 88.2 88.8 88.2 88.2 88.6 87.1 84.4 81.9 82.7 80.4 80.6 72.8 65.0 70.7 62.9 42.8
multi_positive 322 96.0 91.0 90.7 90.1 91.0 91.3 91.0 91.3 89.8 91.0 91.9 91.9 88.5 87.0 88.2 84.8 85.4 86.6 86.3 82.0 81.1 81.7 78.3 81.1 79.5 75.2 76.7 75.8 73.3 72.4 71.4 66.8 71.4 57.5 43.5 51.9 44.1 15.2
negation 257 98.8 98.1 97.3 96.1 98.1 96.9 98.1 97.3 96.5 96.1 95.7 98.1 97.3 97.7 96.9 96.5 96.5 96.1 94.9 92.2 94.6 93.8 93.0 91.8 92.2 91.8 93.0 89.9 87.2 88.3 85.2 83.3 83.7 74.3 75.1 71.2 67.7 52.9
numeric_reasoning 224 92.0 87.5 87.1 87.1 87.5 88.4 87.1 88.4 86.6 88.8 86.6 87.9 86.2 87.1 80.4 80.4 82.1 83.9 82.1 80.4 83.0 82.1 80.4 78.6 78.6 79.0 77.2 79.5 71.9 75.0 72.3 67.0 70.1 62.1 59.8 58.5 54.9 42.0
paraphrase 218 95.4 93.1 91.7 91.7 92.2 91.7 92.2 92.2 92.2 91.3 92.2 91.7 93.1 91.3 89.4 89.9 86.2 90.4 88.1 87.6 89.0 88.1 86.2 86.7 83.9 85.8 86.2 83.5 83.0 83.0 83.9 83.5 80.7 71.1 74.8 66.5 67.4 49.1
exception 213 95.3 96.2 93.9 93.9 96.2 94.8 96.2 93.9 93.9 95.3 93.0 93.0 89.7 91.1 92.0 90.6 89.7 91.5 92.0 88.3 89.7 89.2 88.7 86.4 87.3 85.9 88.3 85.9 80.3 80.8 81.2 78.4 83.1 71.8 72.3 68.5 66.2 50.2
temporal_reasoning 203 89.2 88.2 84.2 85.7 88.2 85.2 87.2 85.7 85.7 84.2 85.2 83.7 80.3 85.2 82.8 81.8 76.4 82.3 82.8 81.3 79.3 80.8 80.3 74.9 77.8 79.8 78.8 74.9 71.9 70.0 71.4 69.5 69.0 62.1 59.6 59.6 53.7 44.3
lexical_overlap 201 98.5 97.0 97.5 97.5 96.5 97.0 96.5 96.5 97.0 96.5 98.0 98.5 96.5 97.0 96.5 97.0 96.0 94.5 94.5 94.5 94.5 95.5 95.0 91.5 93.0 93.5 95.5 91.5 90.5 89.1 88.6 88.6 88.6 80.6 77.6 73.1 70.1 49.3
long_state 191 99.0 98.4 97.4 98.4 98.4 95.8 98.4 95.8 97.4 95.8 98.4 97.4 95.3 94.8 95.3 94.8 94.8 96.3 92.7 92.7 91.1 92.1 89.0 89.5 89.0 89.5 90.6 90.6 84.3 87.4 82.7 82.7 81.7 77.5 59.7 74.3 59.2 36.6
role_reversal 187 96.8 94.1 95.7 96.3 93.6 94.1 93.6 94.1 94.7 94.7 95.2 95.2 92.5 94.7 94.7 95.2 93.0 94.7 91.4 89.3 88.8 93.0 89.8 90.9 93.0 90.9 92.5 88.8 85.0 82.4 85.0 86.6 82.9 74.3 69.5 67.9 59.4 43.9
contradiction 172 98.8 97.7 96.5 96.5 97.1 95.9 97.1 95.9 96.5 95.9 94.8 95.9 95.9 95.9 94.2 94.8 95.3 95.9 95.3 94.2 94.8 93.6 93.0 91.3 92.4 93.6 95.3 91.9 85.5 89.5 86.0 86.0 86.6 76.7 73.3 71.5 64.0 52.3
injection 151 97.4 95.4 92.7 92.7 95.4 94.7 95.4 94.7 93.4 95.4 96.7 93.4 94.7 94.0 89.4 92.1 87.4 92.1 88.7 86.8 87.4 86.1 83.4 88.7 82.1 85.4 84.1 79.5 74.2 76.8 72.8 71.5 73.5 70.2 57.6 62.3 58.9 41.7
multi_turn 149 96.0 92.6 92.6 92.6 93.3 94.0 92.6 94.0 91.9 94.0 93.3 93.3 90.6 94.0 93.3 90.6 93.3 91.3 92.6 91.9 93.3 90.6 92.6 85.2 91.3 92.6 91.9 87.2 84.6 87.9 83.9 83.9 84.6 77.9 73.8 69.1 67.8 45.6
sarcasm 149 97.3 96.0 96.0 96.0 95.3 94.6 96.0 95.3 95.3 94.6 94.6 96.0 94.0 91.3 92.6 91.9 91.3 85.9 86.6 91.9 85.2 87.2 83.2 86.6 82.6 81.9 83.9 76.5 75.2 82.6 74.5 71.8 71.8 69.1 67.1 59.7 55.0 45.6
missing_evidence 141 97.9 96.5 98.6 97.9 96.5 97.9 96.5 97.9 97.9 97.9 96.5 98.6 100.0 97.9 95.0 97.9 97.2 95.0 95.0 95.7 93.6 94.3 94.3 96.5 94.3 96.5 95.0 89.4 92.9 88.7 90.8 92.9 86.5 82.3 80.1 77.3 70.9 57.4
hypothetical 128 98.4 96.1 96.1 97.7 96.1 94.5 96.1 94.5 96.9 94.5 95.3 96.1 93.0 93.8 96.9 93.8 92.2 93.0 93.0 95.3 94.5 91.4 95.3 90.6 92.2 92.2 93.8 89.1 89.1 88.3 88.3 92.2 84.4 81.2 75.0 74.2 70.3 51.6
double_negation 126 99.2 96.8 95.2 94.4 96.8 94.4 96.8 94.4 94.4 93.7 96.8 96.0 93.7 96.0 91.3 94.4 92.9 92.9 90.5 88.1 88.1 91.3 88.9 92.9 87.3 87.3 88.9 88.1 87.3 85.7 87.3 85.7 75.4 65.9 68.3 59.5 54.0 41.3
nota 118 95.8 95.8 95.8 94.1 95.8 94.1 94.9 93.2 94.9 94.1 92.4 93.2 92.4 94.1 91.5 92.4 89.8 88.1 89.8 88.1 86.4 88.1 86.4 89.8 89.0 91.5 89.0 89.0 82.2 80.5 82.2 83.1 81.4 72.0 63.6 66.9 58.5 30.5
evidence_middle 114 99.1 99.1 99.1 100.0 99.1 98.2 99.1 98.2 99.1 98.2 100.0 97.4 99.1 97.4 97.4 96.5 96.5 98.2 97.4 94.7 93.9 94.7 93.0 91.2 94.7 93.9 92.1 93.9 90.4 91.2 89.5 87.7 89.5 82.5 66.7 77.2 69.3 43.0
zero_positive 73 95.9 94.5 93.2 93.2 94.5 93.2 94.5 93.2 91.8 93.2 93.2 90.4 91.8 95.9 93.2 87.7 91.8 90.4 87.7 89.0 90.4 89.0 93.2 82.2 86.3 87.7 89.0 87.7 79.5 89.0 83.6 80.8 78.1 79.5 75.3 68.5 65.8 75.3
evidence_end 57 100.0 98.2 96.5 96.5 98.2 96.5 98.2 96.5 96.5 96.5 96.5 98.2 94.7 94.7 98.2 93.0 94.7 98.2 91.2 93.0 94.7 94.7 93.0 94.7 89.5 93.0 94.7 86.0 84.2 91.2 84.2 82.5 86.0 80.7 59.6 77.2 56.1 31.6
evidence_start 17 94.1 100.0 94.1 100.0 100.0 94.1 100.0 94.1 94.1 94.1 100.0 100.0 94.1 94.1 88.2 100.0 100.0 94.1 100.0 100.0 82.4 94.1 70.6 88.2 82.4 88.2 88.2 94.1 76.5 88.2 70.6 76.5 64.7 70.6 47.1 70.6 52.9 47.1
zero_positive_distractor 1 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 0.0 100.0 100.0 100.0 0.0

paired exact McNemar (questions only the row / only the other gets right)

run vs selfjev-4b (qwen35_4b_tree_scratch_jevall_) vs Jev
jev 52 / 23 (p = 0.0011) —
qwen35_4b_tree_scratch_jevall_ — 23 / 52 (p = 0.0011)
qwen35_4b_tree_sft_jevall__last 27 / 28 (p = 1) 27 / 57 (p = 0.0014)
qwen35_4b_tree_cost_jevall__last 29 / 31 (p = 0.9) 26 / 57 (p = 0.00088)
selfjev_4b_treeserver 1 / 3 (p = 0.62) 24 / 55 (p = 0.00064)
qwen35_4b_tree_rlcd_fresh_ 30 / 33 (p = 0.8) 30 / 62 (p = 0.0011)
qwen35_4b_tree_scratch_jevall__last 0 / 3 (p = 0.25) 23 / 55 (p = 0.00038)
qwen35_4b_tree 29 / 33 (p = 0.7) 31 / 64 (p = 0.00092)
qwen35_4b_tree_rlcd_jevall__last 22 / 29 (p = 0.4) 26 / 62 (p = 0.00016)
qwen35_4b_tree_sft_fresh_ 29 / 36 (p = 0.46) 30 / 66 (p = 0.00031)
selfjev_4b_v2 33 / 41 (p = 0.42) 24 / 61 (p = 7.4e-05)
selfjev_4b_repro 28 / 42 (p = 0.12) 22 / 65 (p = 4.3e-06)
qwen35_4b_combo 32 / 57 (p = 0.011) 23 / 77 (p = 5.5e-08)
tree_4b_combo 38 / 64 (p = 0.013) 27 / 82 (p = 1.3e-07)
qwen35_4b_r2x64 24 / 65 (p = 1.6e-05) 21 / 91 (p = 1.4e-11)
tree_4b_instruct_r3 31 / 80 (p = 3.7e-06) 21 / 99 (p = 2.7e-13)
tree_4b_combo_r2 31 / 89 (p = 1.1e-07) 18 / 105 (p = 4e-16)
eikos_4b 25 / 85 (p = 7.8e-09) 20 / 109 (p = 5.1e-16)
tree_4b_instruct_r2x64 31 / 92 (p = 3.3e-08) 25 / 115 (p = 5.4e-15)
tree_4b_combo_ptr 33 / 108 (p = 1.7e-10) 23 / 127 (p = 1.3e-18)
tree_4b_ova 34 / 118 (p = 4.6e-12) 24 / 137 (p = 2e-20)
curve/tree_4b_r2b_r64_mlp 30 / 119 (p = 9.6e-14) 26 / 144 (p = 5.3e-21)
tree_4b_ova_kd 40 / 136 (p = 1.9e-13) 29 / 154 (p = 8.9e-22)
tree_4b_instruct_r3_step300 21 / 120 (p = 4.8e-18) 11 / 139 (p = 2.3e-29)
tree_4b_r2b 30 / 134 (p = 6.8e-17) 24 / 157 (p = 3.8e-25)
tree_4b_r2 34 / 142 (p = 6.8e-17) 26 / 163 (p = 1.9e-25)
curve/tree_4b_r2b_r64 29 / 139 (p = 2e-18) 22 / 161 (p = 2.7e-27)
tree_4b_instruct 25 / 177 (p = 2.1e-29) 15 / 196 (p = 2.2e-41)
curve/tree_4b_r64 27 / 198 (p = 2.5e-33) 14 / 214 (p = 3.9e-47)
lora_4b 32 / 216 (p = 1e-34) 20 / 233 (p = 3.3e-47)
curve/tree_4b_mlp 29 / 217 (p = 9e-37) 17 / 234 (p = 6e-50)
tree_4b 30 / 242 (p = 2.3e-42) 20 / 261 (p = 1.1e-54)
qwen35_r1 26 / 255 (p = 2e-48) 16 / 274 (p = 8.4e-62)
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b 23 / 401 (p = 2.8e-90) 17 / 424 (p = 6.8e-103)
jina_r2b 32 / 480 (p = 1.1e-103) 22 / 499 (p = 1e-118)
t5_round2b_2026-09-24/t5_r1 20 / 474 (p = 8.5e-114) 12 / 495 (p = 2.6e-129)
lora_pilot 19 / 556 (p = 2.8e-138) 11 / 577 (p = 1.3e-154)
jina_zeroshot 27 / 1008 (p = 9.2e-259) 22 / 1032 (p = 2.4e-272)

eval_llm: LLM-evaluation use cases

946 frozen questions about LLM prompts, reasoning traces and outputs (score, judge, verify, guardrail, jailbreak), written by eval2's authors and kept only when two blind judges agree (how it was built). Scored so far:

run eval_llm binary multiclass multilabel EM
Jev (API) 92.5 95.1 95.3 81.3
selfjev-4b 93.1 94.3 95.3 86.8
qwen35_4b_tree_sft_jevall__last: qwen35_4b_tree + one epoch on Jev's soft targets (incl. the LLM-evaluation data) 90.1 92.8 93.1 78.0
qwen35_4b_tree: never trained on LLM-evaluation data 82.1 83.8 88.4 68.1

selfjev-4b vs Jev: 33 / 27, p = 0.52. Sources: reports/*/eval_llm/report.json; Jev from data/eval_llm/review/jev_answers.jsonl (llm_eval_data.md).

Dev benchmark: the original test split

3,471 questions: 3,300 from public datasets (1,800 of them from families never trained on) and 171 authored. Frontier models top out at 86% here because of public-label noise, so it separates small models from large ones but not good models from each other. Selected rows:

model question acc % binary AUROC multiclass macro-F1 % multilabel EM % binary ECE
GPT-6 Astra (reasoning low) 85.8 0.971 95.3 55.8 0.050
Jev 82.7 0.981 91.6 40.1 0.045
selfjev-4b: Qwen3.5-4B tree, from scratch on 80K questions with Jev's soft targets 83.8 0.974 92.7 52.3 0.048
Qwen3.5-4B trained with the tree, r64, round-2b + round-3 data, all options in question (qwen35_4b_tree) 84.4 0.972 92.6 59.3 0.038
Qwen3.5-4B, r64, round-2b + round-3 data, all options in question 84.3 0.963 92.2 62.2 0.047
tree Instruct-4B, r64, round-2b + round-3 data, all options in question 82.7 0.959 89.3 57.6 0.064
tree Instruct-4B, r64, round-2b data, all options in question 83.5 0.966 89.3 59.0 0.038
tree Instruct-4B, r64, round-2b data 82.7 0.965 88.0 52.9 0.045
tree Reranker-4B, round-1 data 81.6 0.953 89.7 51.7 0.077
stock 8B + LoRA 80.7 0.945 88.7 48.0 0.051
stock 4B + LoRA 80.3 0.945 86.2 47.7 0.056
stock 0.6B + LoRA 73.5 0.863 82.3 31.1 0.066
stock 8B, untrained 66.2 0.658 86.5 10.2 0.364
stock 4B, untrained 62.8 0.604 83.4 2.0 0.357
stock 0.6B, untrained 61.0 0.605 79.6 1.2 0.384

Per-family tables for these models: stock reranker and tree scorer. Jev and GPT-6 Astra answers are cached in reports/external/cache/, keyed by the request body: Jev reruns cost nothing, but GPT-6 Astra's requests changed on 2026-09-27 (max_tokens 6000 → 8192), so its reruns miss the cache and pay again. selfjev-4b's 52.3 multilabel exact match is the emotion-set effect of Jev's targets (finding 1).

Every run (generated)

Written by uv run python scripts/eval/ledger.py into the experiment ledger; sorted by eval2, then by the dev benchmark.

run base architecture trained on eval2 acc % (target task, 1,991 q) dev benchmark acc % (old test, 3,471 q) binary acc % binary AUROC multiclass acc % multilabel EM % authored eval_* acc % (n=171) report weights
external ~typesafe/jev-latest external API — 97.2 82.7 93.5 0.981 84.6 40.1 94.7 report —
qwen35_4b_tree_scratch_jevall_ Qwen3.5-4B shared document / forked native cache branches 79,943 q / 1803 steps 95.8 83.8 91.5 0.974 85.3 52.3 91.2 report weights/selfjev_4b
qwen35_4b_tree_sft_jevall__last Qwen3.5-4B shared document / forked native cache branches 69,528 q / 1230 steps 95.7 84.4 90.5 0.976 85.7 59.0 92.4 report —
qwen35_4b_tree_cost_jevall__last Qwen3.5-4B shared document / forked native cache branches 69,528 q / 1230 steps 95.7 84.4 90.7 0.975 85.6 59.3 92.4 report —
selfjev_4b_treeserver Qwen3.5-4B shared-prefix tree, forward only: text once per request, each question once, then each candidate — 95.7 83.8 91.6 0.974 85.2 52.6 91.2 report —
qwen35_4b_tree_rlcd_fresh_ Qwen3.5-4B shared document / forked native cache branches 4,412 q / 97 steps 95.6 84.4 90.2 0.972 85.6 59.9 90.1 report —
qwen35_4b_tree_scratch_jevall__last Qwen3.5-4B shared document / forked native cache branches 79,943 q / 1803 steps 95.6 83.8 91.7 0.974 85.3 52.0 91.2 report —
qwen35_4b_tree Qwen3.5-4B shared document / forked native cache branches 51,774 q / 941 steps 95.6 84.4 90.5 0.972 85.7 59.3 90.1 report —
qwen35_4b_tree_sft_fresh_ Qwen3.5-4B shared document / forked native cache branches 4,412 q / 97 steps 95.4 84.5 90.2 0.971 85.8 59.9 90.1 report —
qwen35_4b_tree_rlcd_jevall__last Qwen3.5-4B shared document / forked native cache branches 69,528 q / 1230 steps 95.4 84.5 90.7 0.975 85.5 60.8 92.4 report —
selfjev_4b_v2 Qwen3.5-4B shared-prefix tree, forward only: text once per request, each question once, then each candidate 83,581 q / 1912 steps 95.4 83.5 91.2 0.973 85.2 51.7 88.9 report —
selfjev_4b_repro Qwen3.5-4B shared-prefix tree, forward only: text once per request, each question once, then each candidate 79,943 q / 1803 steps 95.1 83.8 91.8 0.973 85.2 52.3 91.8 report —
qwen35_4b_combo Qwen3.5-4B shared document / forked native cache branches 43,827 q / 2474 steps 94.5 84.3 90.0 0.963 85.3 62.2 84.2 report —
tree_4b_combo Qwen3-4B-Instruct-2507 shared-prefix tree (tree-v1) 51,826 q / 967 steps 94.5 82.7 89.1 0.959 83.8 57.6 86.0 report —
qwen35_4b_r2x64 Qwen3.5-4B shared document / forked native cache branches 17,435 q / 769 steps 93.7 83.4 90.9 0.971 84.0 58.4 87.7 report —
tree_4b_instruct_r3 Qwen3-4B-Instruct-2507 shared-prefix tree (tree-v1) 51,853 q / 910 steps 93.3 82.8 89.8 0.966 83.9 56.1 85.4 report —
tree_4b_combo_r2 Qwen3-4B-Instruct-2507 shared-prefix tree (tree-v1) 18,676 q / 239 steps 92.9 83.5 90.0 0.966 84.4 59.0 84.8 report —
tree_4b_instruct_r2x64 Qwen3-4B-Instruct-2507 shared-prefix tree (tree-v1) 18,681 q / 218 steps 92.7 82.7 90.9 0.965 83.7 52.9 81.9 report —
tree_4b_combo_ptr Qwen3-4B-Instruct-2507 shared-prefix tree (tree-v1) 18,676 q / 232 steps 92.0 82.7 89.6 0.965 83.6 57.0 81.3 report —
tree_4b_ova Qwen3-Reranker-4B shared-prefix tree (tree-v1) 16,372 q / 241 steps 91.6 82.6 89.2 0.962 83.6 57.8 84.2 report —
curve/tree_4b_r2b_r64_mlp Qwen3-Reranker-4B shared-prefix tree (tree-v1) 16,375 q / 208 steps 91.3 82.4 89.0 0.962 83.5 57.0 84.2 report —
tree_4b_ova_kd Qwen3-Reranker-4B shared-prefix tree (tree-v1) 16,372 q / 241 steps 91.0 82.5 88.1 0.959 83.8 58.4 84.8 report —
tree_4b_instruct_r3_step300 Qwen3-4B-Instruct-2507 shared-prefix tree (tree-v1) — 90.8 81.5 89.4 0.954 82.5 52.6 81.3 report —
tree_4b_r2b Qwen3-Reranker-4B shared-prefix tree (tree-v1) 16,375 q / 225 steps 90.6 81.2 88.1 0.955 82.8 51.7 83.0 report —
tree_4b_r2 Qwen3-Reranker-4B shared-prefix tree (tree-v1) 16,357 q / 224 steps 90.4 80.6 87.6 0.956 82.2 50.9 81.9 report —
curve/tree_4b_r2b_r64 Qwen3-Reranker-4B shared-prefix tree (tree-v1) 16,375 q / 208 steps 90.3 81.2 87.8 0.962 82.5 54.4 82.5 report —
tree_4b_instruct Qwen3-4B-Instruct-2507 shared-prefix tree (tree-v1) 10,112 q / 78 steps 88.1 80.4 87.8 0.958 82.3 46.8 79.5 report —
curve/tree_4b_r64 Qwen3-Reranker-4B shared-prefix tree (tree-v1) 10,080 q / 84 steps 87.2 81.5 88.5 0.956 83.0 52.3 77.8 report —
lora_4b Qwen3-Reranker-4B stock pairs (answer-v1) 10,112 q / 216 steps 86.5 80.3 87.5 0.945 82.3 47.7 70.8 report —
curve/tree_4b_mlp Qwen3-Reranker-4B shared-prefix tree (tree-v1) 10,080 q / 84 steps 86.3 81.6 87.9 0.955 83.4 51.7 78.9 report —
tree_4b Qwen3-Reranker-4B shared-prefix tree (tree-v1) 10,112 q / 84 steps 85.1 81.6 87.1 0.953 83.9 51.7 78.4 report —
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b google/t5gemma-2-1b-1b pretrained T5Gemma encoder/decoder; shared document cross-KV views 16,375 q / 228 steps 76.8 73.8 82.7 0.893 75.1 40.4 62.6 report —
jina_r2b jinaai/jina-reranker-v3.5 jina listwise (jina-v1) 16,301 q / 345 steps 73.3 76.6 81.4 0.888 79.4 45.1 60.2 report —
t5_round2b_2026-09-24/t5_r1 google/t5gemma-2-1b-1b pretrained T5Gemma encoder/decoder; shared document cross-KV views 10,080 q / 823 steps 73.0 75.4 83.0 0.887 76.6 45.6 62.0 report —
lora_pilot Qwen3-Reranker-0.6B stock pairs (task-v1) 10,112 q / 183 steps 68.8 73.5 77.8 0.863 78.3 31.1 55.0 report —
external openai/gpt-6-astra external API — — 85.8 93.8 0.971 87.0 55.8 100.0 report —
curve/boolq_300 Qwen3-Reranker-4B stock pairs (answer-v1) 10,412 q / 218 steps — 80.9 88.0 0.940 82.5 50.9 71.9 report —
curve/boolq_1000 Qwen3-Reranker-4B stock pairs (answer-v1) 11,112 q / 224 steps — 80.9 88.8 0.954 82.5 48.8 73.7 report —
curve/boolq_3000 Qwen3-Reranker-4B stock pairs (answer-v1) 13,112 q / 240 steps — 80.9 88.9 0.954 82.5 47.7 73.1 report —
lora_8b Qwen3-Reranker-8B stock pairs (task-v1) 10,112 q / 185 steps — 80.7 87.3 0.945 82.9 48.0 74.9 report —
curve/instruct_lora Qwen3-4B-Instruct-2507 stock pairs (answer-v1) 10,112 q / 216 steps — 80.6 87.4 0.944 82.6 48.3 75.4 report —
curve/boolq_100 Qwen3-Reranker-4B stock pairs (answer-v1) 10,212 q / 217 steps — 80.0 86.9 0.945 82.0 48.3 70.8 report —
curve/vol100_e2 Qwen3-Reranker-4B stock pairs (answer-v1) 10,112 q / 431 steps — 79.8 87.0 0.951 81.9 46.2 69.6 report —
curve/vol50_e2 Qwen3-Reranker-4B stock pairs (answer-v1) 5,056 q / 214 steps — 79.1 84.8 0.940 81.7 46.8 72.5 report —
curve/vol25_e2 Qwen3-Reranker-4B stock pairs (answer-v1) 2,528 q / 107 steps — 79.1 85.7 0.935 81.4 45.6 73.7 report —
curve/vol50_e1 Qwen3-Reranker-4B stock pairs (answer-v1) 5,056 q / 107 steps — 78.9 84.5 0.934 81.9 44.5 70.2 report —
curve/vol25_e1 Qwen3-Reranker-4B stock pairs (answer-v1) 2,528 q / 53 steps — 77.0 83.8 0.930 80.8 33.4 71.9 report —
curve/instruct_zero Qwen3-4B-Instruct-2507 stock pairs (answer-v1) — — 71.3 84.2 0.911 73.5 20.1 63.2 report —
baseline_8b Qwen3-Reranker-8B stock pairs (task-v1) — — 66.2 56.1 0.658 79.8 10.2 50.3 report —
baseline_4b Qwen3-Reranker-4B stock pairs (answer-v1) — — 62.8 55.0 0.604 76.2 2.0 44.4 report —
baseline Qwen3-Reranker-0.6B stock pairs (task-v1) — — 61.0 55.3 0.605 73.3 1.2 39.2 report —
custom_sim_lora Qwen3-Reranker-0.6B shared-state cross-attention (custom-v1) 10,112 q / 1176 steps — 58.2 51.7 0.514 70.2 1.7 42.1 report —
custom_sim_frozen Qwen3-Reranker-0.6B shared-state cross-attention (custom-v1) 10,112 q / 1176 steps — 48.0 51.4 0.524 53.9 1.5 38.0 report —
custom_distill_lora Qwen3-Reranker-0.6B shared-state cross-attention (custom-v1) 16,671 q / 2028 steps — 39.7 51.3 0.451 40.6 0.9 35.1 report —
custom_frozen Qwen3-Reranker-0.6B shared-state cross-attention (custom-v1) 10,112 q / 1764 steps — 39.0 51.9 0.536 39.1 1.2 36.3 report —
custom_distill_frozen Qwen3-Reranker-0.6B shared-state cross-attention (custom-v1) 16,671 q / 2028 steps — 37.7 52.8 0.507 36.7 0.6 36.8 report —