Leaderboard
eval2 : the target task
1,991 authored questions (how it was built ). This is the benchmark
that decides between models, with eval_llm for the LLM-evaluation use cases.
Jev 97.2
selfjev-4b: Qwen3.5-4B tree, trained from scratch on all 80K questions, texts to 16K, Jev soft targets
95.8
Qwen3.5-4B trained with the tree, r64, round-2b + round-3 data, all options in question
95.6
Qwen3.5-4B, r64, round-2b + round-3 data, all options in question
94.5
Tree, Qwen3-4B-Instruct, r64, round-2b + round-3 data, all options in question
94.5
Tree, Qwen3-4B-Instruct, r64, round-2b + round-3 data
93.3
Tree, Qwen3-4B-Instruct, r64, round-2b data, all options in question
92.9
Tree, Qwen3-4B-Instruct, r64, round-2b data
92.7
Tree, Reranker-4B, all options in question
91.6
Tree, Reranker-4B, round-2b data
90.6
Tree, Qwen3-4B-Instruct, round-1 data
88.1
Stock pairs, Reranker-4B, round-1 data
86.5
Tree, Reranker-4B, round-1 data
85.1
jina-reranker-v3.5 0.6B, round-2b data
73.3
T5Gemma 2 1B–1B, round-1 data
73.0
Stock pairs, Reranker-0.6B, round-1 data
68.8
eval2 question accuracy, %. Bars start at 0; the dashed line is Jev. All ours use LoRA.
run
base
recipe
eval2
binary
multiclass
multilabel EM
dev benchmark
Jev (API)
undisclosed
undisclosed
97.2
97.8
98.1
94.2
82.7
selfjev-4b (qwen35_4b_tree_scratch_jevall_), the default
Qwen3.5-4B
tree, trained from scratch on 79.9K non-test questions of data/all.jsonl.gz (texts ≤ 16K, incl. the LLM-evaluation data, llm_multilabel_v1, numdate_neg_v1), target 0.5 × label + 0.5 × Jev, r64 , all options in the question; eval_llm 93.1 (Jev 92.5)
95.8
96.8
97.0
91.4
83.8
qwen35_4b_tree , the previous default
Qwen3.5-4B
shared-prefix tree in training (src/selfjev/engine/tree.py, texts ≤ 8K: 51.8K q), text shared at inference, r64 , round-2b + round-3 data, all options in the question
95.6
96.9
96.8
90.1
84.4
qwen35_4b_combo
Qwen3.5-4B
each option trained as a full sequence (no tree in training, texts ≤ 2K: 43.8K q), text shared at inference, r64 , round-2b + round-3 data, all options in the question
94.5
96.2
95.8
88.2
84.3
tree_4b_combo
Qwen3-4B-Instruct-2507
tree, r64 , round-2b + round-3 data (51.8K q), hard cases not capped, all options in the question
94.5
96.0
97.5
85.9
82.7
tree_4b_instruct_r3
Qwen3-4B-Instruct-2507
tree, r64 , round-2b + round-3 data (51.9K q), hard cases not capped
93.3
95.1
96.1
84.3
82.8
tree_4b_combo_r2
Qwen3-4B-Instruct-2507
tree, r64 , round-2b data, hard cases not capped, all options in the question (served: merged + vLLM , 93.0)
92.9
94.4
95.8
84.3
83.5
tree_4b_instruct_r2x64
Qwen3-4B-Instruct-2507
tree, r64 , round-2b data, hard cases not capped per family
92.7
94.6
95.4
83.5
82.7
tree_4b_combo_ptr
Qwen3-4B-Instruct-2507
as tree_4b_combo_r2, options numbered once, leaves say "option k"
92.0
94.9
94.3
80.9
82.7
tree_4b_ova
Qwen3-Reranker-4B
tree, r16 , round-2b data, all options in the question
91.6
93.8
94.9
80.4
82.6
Qwen3.8-27B-FP8, zero-shot (teacher)
Qwen3.8-27B
no training
91.4
93.4
97.8
75.9
—
curve/tree_4b_r2b_r64_mlp
Qwen3-Reranker-4B
tree, r64 + MLP , round-2b data
91.3
93.7
94.3
80.4
82.4
tree_4b_ova_kd
Qwen3-Reranker-4B
as tree_4b_ova + 27B soft targets
91.0
92.9
95.6
78.5
82.5
tree_4b_instruct_r3_step300
Qwen3-4B-Instruct-2507
round-3 run, stopped; step 300 of 915
90.8
93.1
95.1
78.0
81.5
tree_4b_r2b
Qwen3-Reranker-4B
tree, r16 , round-2 data, none capped
90.6
92.9
94.1
78.8
81.2
tree_4b_r2
Qwen3-Reranker-4B
tree, r16 , round-2 data
90.4
92.8
95.8
75.4
80.6
curve/tree_4b_r2b_r64
Qwen3-Reranker-4B
tree, r64 , round-2b data
90.3
92.9
94.9
75.9
81.2
tree_4b_instruct
Qwen3-4B-Instruct-2507
tree, r16 , round-1 data
88.1
91.4
92.7
72.3
80.4
curve/tree_4b_r64
Qwen3-Reranker-4B
tree, r64 , round-1 data
87.2
90.5
92.7
69.9
81.5
lora_4b
Qwen3-Reranker-4B
stock pairs , r16 , round-1 data
86.5
89.7
91.6
70.4
80.3
curve/tree_4b_mlp
Qwen3-Reranker-4B
tree, r16 + MLP , round-1 data
86.3
90.0
91.9
68.1
81.6
tree_4b
Qwen3-Reranker-4B
tree, r16 , round-1 data
85.1
89.4
91.6
63.9
81.6
T5Gemma 2, decoder-only LoRA (audit)
t5gemma-2-1b-1b
shared encoder, round-2b data
76.8
82.6
83.6
50.8
73.8
jina_r2b
jina-reranker-v3.5 (0.6B)
listwise, round-2b data
73.3
79.8
83.5
40.1
76.6
T5Gemma 2 (t5_r1)
t5gemma-2-1b-1b
shared encoder, round-1 data
73.0
79.5
79.8
45.0
75.4
lora_pilot
Qwen3-Reranker-0.6B
stock pairs , r16 , round-1 data
68.8
75.7
75.7
39.8
73.5
jina_zeroshot
jina-reranker-v3.5 (0.6B)
no training
46.5
55.8
54.5
9.4
—
Sources: reports/eval2 /summary.md (generated by scripts/eval/eval2_summary.py) and the
27B teacher's predictions in reports/teacher/qwen38_27b_eval2.jsonl.
GPT-6 Astra is not scored on eval2 : it was one of the two judges that decided which questions were kept.
"Round-1 data" is the 10,112-question mix (public sets + synthetic); round 2 adds about 10K verified hard cases
(data ).
Only selfjev-4b's adapter is on master (weights/selfjev_4b). qwen35_4b_tree and tree_4b_combo are at tag
archive/pre-cleanup-2026-09-27, with the code of every other architecture (Qwen3 trees, stock pairs , jina,
T5Gemma, the 27B teacher). All these reports were scored before the 2026-09-27 cleanup, the Qwen3.5 rows by the
forked-cache engine that TreeServer replaced.
All eval2 slices: type, tier, author, text length, trap, paired tests (generated)
eval2 results (frozen target-task test set)
Generated by uv run python scripts/eval/eval2_summary.py; do not hand-edit. 1991 questions, data/eval2.jsonl (see data/eval2/REVIEW.md). Never train, select prompts or fit calibration on it. GPT-6 Astra is one of the two judges that decided which questions were kept, so it is not a fair competitor here and isn't scored.
overall
slice
n
jev
qwen35_4b_tree_scratch_jevall_
qwen35_4b_tree_sft_jevall__last
qwen35_4b_tree_cost_jevall__last
selfjev_4b_treeserver
qwen35_4b_tree_rlcd_fresh_
qwen35_4b_tree_scratch_jevall__last
qwen35_4b_tree
qwen35_4b_tree_rlcd_jevall__last
qwen35_4b_tree_sft_fresh_
selfjev_4b_v2
selfjev_4b_repro
qwen35_4b_combo
tree_4b_combo
qwen35_4b_r2x64
tree_4b_instruct_r3
tree_4b_combo_r2
eikos_4b
tree_4b_instruct_r2x64
tree_4b_combo_ptr
tree_4b_ova
curve/tree_4b_r2b_r64_mlp
tree_4b_ova_kd
tree_4b_instruct_r3_step300
tree_4b_r2b
tree_4b_r2
curve/tree_4b_r2b_r64
tree_4b_instruct
curve/tree_4b_r64
lora_4b
curve/tree_4b_mlp
tree_4b
qwen35_r1
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b
jina_r2b
t5_round2b_2026-09-24/t5_r1
lora_pilot
jina_zeroshot
all
1991
97.2
95.8
95.7
95.7
95.7
95.6
95.6
95.6
95.4
95.4
95.4
95.1
94.5
94.5
93.7
93.3
92.9
92.8
92.7
92.0
91.6
91.3
91.0
90.8
90.6
90.4
90.3
88.1
87.2
86.5
86.3
85.1
84.3
76.8
73.3
73.0
68.8
46.5
type
slice
n
jev
qwen35_4b_tree_scratch_jevall_
qwen35_4b_tree_sft_jevall__last
qwen35_4b_tree_cost_jevall__last
selfjev_4b_treeserver
qwen35_4b_tree_rlcd_fresh_
qwen35_4b_tree_scratch_jevall__last
qwen35_4b_tree
qwen35_4b_tree_rlcd_jevall__last
qwen35_4b_tree_sft_fresh_
selfjev_4b_v2
selfjev_4b_repro
qwen35_4b_combo
tree_4b_combo
qwen35_4b_r2x64
tree_4b_instruct_r3
tree_4b_combo_r2
eikos_4b
tree_4b_instruct_r2x64
tree_4b_combo_ptr
tree_4b_ova
curve/tree_4b_r2b_r64_mlp
tree_4b_ova_kd
tree_4b_instruct_r3_step300
tree_4b_r2b
tree_4b_r2
curve/tree_4b_r2b_r64
tree_4b_instruct
curve/tree_4b_r64
lora_4b
curve/tree_4b_mlp
tree_4b
qwen35_r1
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b
jina_r2b
t5_round2b_2026-09-24/t5_r1
lora_pilot
jina_zeroshot
binary
1016
97.8
96.8
96.8
96.9
96.7
96.9
96.8
96.9
96.9
96.8
96.9
96.1
96.2
96.0
95.6
95.1
94.4
94.9
94.6
94.9
93.8
93.7
92.9
93.1
92.9
92.8
92.9
91.4
90.5
89.7
90.0
89.4
87.5
82.6
79.8
79.5
75.7
55.8
multiclass
593
98.1
97.0
97.5
97.6
97.0
97.0
96.6
96.8
97.1
97.0
96.5
96.3
95.8
97.5
95.1
96.1
95.8
95.8
95.4
94.3
94.9
94.3
95.6
95.1
94.1
95.8
94.9
92.7
92.7
91.6
91.9
91.6
88.7
83.6
83.5
79.8
75.7
54.5
multilabel
382
94.2
91.4
90.3
89.5
91.1
90.1
91.1
90.1
89.0
89.5
89.5
90.6
88.2
85.9
86.6
84.3
84.3
82.5
83.5
80.9
80.4
80.4
78.5
78.0
78.8
75.4
75.9
72.3
69.9
70.4
68.1
63.9
68.8
50.8
40.1
45.0
39.8
9.4
tier
slice
n
jev
qwen35_4b_tree_scratch_jevall_
qwen35_4b_tree_sft_jevall__last
qwen35_4b_tree_cost_jevall__last
selfjev_4b_treeserver
qwen35_4b_tree_rlcd_fresh_
qwen35_4b_tree_scratch_jevall__last
qwen35_4b_tree
qwen35_4b_tree_rlcd_jevall__last
qwen35_4b_tree_sft_fresh_
selfjev_4b_v2
selfjev_4b_repro
qwen35_4b_combo
tree_4b_combo
qwen35_4b_r2x64
tree_4b_instruct_r3
tree_4b_combo_r2
eikos_4b
tree_4b_instruct_r2x64
tree_4b_combo_ptr
tree_4b_ova
curve/tree_4b_r2b_r64_mlp
tree_4b_ova_kd
tree_4b_instruct_r3_step300
tree_4b_r2b
tree_4b_r2
curve/tree_4b_r2b_r64
tree_4b_instruct
curve/tree_4b_r64
lora_4b
curve/tree_4b_mlp
tree_4b
qwen35_r1
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b
jina_r2b
t5_round2b_2026-09-24/t5_r1
lora_pilot
jina_zeroshot
e2_hard
673
96.7
94.5
95.5
95.2
94.4
94.7
94.4
94.7
94.9
94.1
93.9
93.8
93.6
93.8
92.3
92.3
90.6
90.8
90.5
90.3
90.8
89.6
89.7
89.0
89.9
89.3
89.5
85.6
85.3
84.2
83.8
83.2
80.4
74.7
70.7
68.2
64.6
50.4
e2_simple
664
98.5
98.5
98.5
98.5
98.5
98.6
98.5
98.6
98.5
98.5
98.6
97.9
98.0
96.8
97.7
96.8
97.3
95.9
96.8
96.5
94.9
94.7
95.3
95.9
95.2
94.6
93.8
92.8
94.4
92.8
93.4
92.3
92.3
85.1
80.6
83.4
80.1
46.1
e2_very_hard
654
96.5
94.3
93.1
93.3
94.2
93.6
94.0
93.4
92.8
93.7
93.6
93.6
91.9
92.8
91.1
90.8
90.7
91.6
90.8
89.1
89.0
89.6
87.8
87.5
86.5
87.2
87.5
86.1
81.8
82.6
81.8
79.8
80.1
70.5
68.5
67.3
61.6
43.0
author
slice
n
jev
qwen35_4b_tree_scratch_jevall_
qwen35_4b_tree_sft_jevall__last
qwen35_4b_tree_cost_jevall__last
selfjev_4b_treeserver
qwen35_4b_tree_rlcd_fresh_
qwen35_4b_tree_scratch_jevall__last
qwen35_4b_tree
qwen35_4b_tree_rlcd_jevall__last
qwen35_4b_tree_sft_fresh_
selfjev_4b_v2
selfjev_4b_repro
qwen35_4b_combo
tree_4b_combo
qwen35_4b_r2x64
tree_4b_instruct_r3
tree_4b_combo_r2
eikos_4b
tree_4b_instruct_r2x64
tree_4b_combo_ptr
tree_4b_ova
curve/tree_4b_r2b_r64_mlp
tree_4b_ova_kd
tree_4b_instruct_r3_step300
tree_4b_r2b
tree_4b_r2
curve/tree_4b_r2b_r64
tree_4b_instruct
curve/tree_4b_r64
lora_4b
curve/tree_4b_mlp
tree_4b
qwen35_r1
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b
jina_r2b
t5_round2b_2026-09-24/t5_r1
lora_pilot
jina_zeroshot
claude-opus-5.5
560
96.1
94.5
93.6
93.2
94.5
94.3
94.1
94.5
93.2
94.6
93.2
93.0
92.7
93.4
91.4
91.4
89.8
91.1
90.9
88.9
89.1
88.8
88.0
88.0
87.7
87.5
87.0
85.4
83.4
83.2
83.6
82.5
80.5
72.3
68.2
69.3
66.4
40.0
glm-5.3
782
97.3
96.0
96.0
96.2
95.9
95.7
95.9
95.5
95.7
95.1
95.5
95.0
95.3
93.1
93.7
93.1
92.7
92.6
91.3
92.3
90.9
91.6
91.0
90.8
90.3
89.8
90.2
87.5
86.7
87.3
85.8
84.0
84.8
76.6
73.9
71.2
68.3
47.7
kimi-k3
649
98.2
96.6
97.2
97.2
96.5
96.8
96.6
96.6
97.1
96.5
97.1
96.9
95.2
97.1
95.7
95.2
95.7
94.5
96.0
94.3
94.5
93.2
93.4
93.2
93.4
93.5
93.2
91.4
91.1
88.4
89.4
88.8
86.9
80.9
76.9
78.3
71.5
50.7
text length (tokens)
slice
n
jev
qwen35_4b_tree_scratch_jevall_
qwen35_4b_tree_sft_jevall__last
qwen35_4b_tree_cost_jevall__last
selfjev_4b_treeserver
qwen35_4b_tree_rlcd_fresh_
qwen35_4b_tree_scratch_jevall__last
qwen35_4b_tree
qwen35_4b_tree_rlcd_jevall__last
qwen35_4b_tree_sft_fresh_
selfjev_4b_v2
selfjev_4b_repro
qwen35_4b_combo
tree_4b_combo
qwen35_4b_r2x64
tree_4b_instruct_r3
tree_4b_combo_r2
eikos_4b
tree_4b_instruct_r2x64
tree_4b_combo_ptr
tree_4b_ova
curve/tree_4b_r2b_r64_mlp
tree_4b_ova_kd
tree_4b_instruct_r3_step300
tree_4b_r2b
tree_4b_r2
curve/tree_4b_r2b_r64
tree_4b_instruct
curve/tree_4b_r64
lora_4b
curve/tree_4b_mlp
tree_4b
qwen35_r1
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b
jina_r2b
t5_round2b_2026-09-24/t5_r1
lora_pilot
jina_zeroshot
≤64
439
97.3
95.4
96.1
95.7
95.4
96.4
95.2
96.1
95.9
95.7
95.7
95.2
95.9
94.3
92.7
91.6
91.8
90.2
90.0
91.3
90.9
90.0
89.5
90.0
89.1
89.1
88.2
85.0
85.4
85.0
85.4
85.2
80.9
75.4
74.9
71.3
67.2
49.0
65–256
508
96.7
94.5
93.3
92.9
94.3
93.9
94.3
94.3
93.1
93.9
93.1
92.9
92.1
93.3
91.3
92.1
91.9
90.9
92.3
89.8
91.3
90.0
90.4
90.2
88.6
89.0
88.8
87.8
87.4
86.2
85.8
83.9
83.1
74.4
77.4
71.5
68.5
50.6
257–1K
469
96.6
95.3
95.3
95.7
95.1
95.5
95.1
95.1
95.1
95.5
95.1
94.7
93.4
93.4
93.0
93.0
91.0
94.5
92.1
91.9
90.0
90.6
90.0
90.4
90.8
90.0
90.4
88.1
86.6
84.9
86.6
84.6
84.0
75.9
74.2
72.9
71.6
46.9
1K–4K
409
98.0
97.6
97.3
97.8
97.6
96.6
97.6
96.6
97.1
96.3
97.3
97.3
96.6
96.6
97.3
96.6
96.3
94.4
96.1
95.4
93.2
93.9
94.6
93.4
93.9
93.2
92.7
90.5
89.7
90.5
88.5
87.8
89.7
81.7
68.7
76.5
69.2
41.1
>4K
166
98.8
97.6
99.4
98.8
97.6
97.0
97.6
97.0
98.2
97.0
97.6
97.0
96.4
96.4
97.0
94.6
95.2
96.4
94.6
92.8
94.6
94.6
90.4
89.8
91.6
92.2
94.0
92.2
86.7
86.7
84.3
83.7
84.3
78.3
65.1
73.5
65.1
39.8
trap
slice
n
jev
qwen35_4b_tree_scratch_jevall_
qwen35_4b_tree_sft_jevall__last
qwen35_4b_tree_cost_jevall__last
selfjev_4b_treeserver
qwen35_4b_tree_rlcd_fresh_
qwen35_4b_tree_scratch_jevall__last
qwen35_4b_tree
qwen35_4b_tree_rlcd_jevall__last
qwen35_4b_tree_sft_fresh_
selfjev_4b_v2
selfjev_4b_repro
qwen35_4b_combo
tree_4b_combo
qwen35_4b_r2x64
tree_4b_instruct_r3
tree_4b_combo_r2
eikos_4b
tree_4b_instruct_r2x64
tree_4b_combo_ptr
tree_4b_ova
curve/tree_4b_r2b_r64_mlp
tree_4b_ova_kd
tree_4b_instruct_r3_step300
tree_4b_r2b
tree_4b_r2
curve/tree_4b_r2b_r64
tree_4b_instruct
curve/tree_4b_r64
lora_4b
curve/tree_4b_mlp
tree_4b
qwen35_r1
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b
jina_r2b
t5_round2b_2026-09-24/t5_r1
lora_pilot
jina_zeroshot
distractor
474
97.0
93.9
94.7
94.7
93.7
93.2
93.7
92.8
93.9
93.2
93.9
92.6
92.0
92.6
91.8
91.1
90.7
92.4
91.1
90.1
89.7
89.5
88.2
88.8
88.2
88.2
88.6
87.1
84.4
81.9
82.7
80.4
80.6
72.8
65.0
70.7
62.9
42.8
multi_positive
322
96.0
91.0
90.7
90.1
91.0
91.3
91.0
91.3
89.8
91.0
91.9
91.9
88.5
87.0
88.2
84.8
85.4
86.6
86.3
82.0
81.1
81.7
78.3
81.1
79.5
75.2
76.7
75.8
73.3
72.4
71.4
66.8
71.4
57.5
43.5
51.9
44.1
15.2
negation
257
98.8
98.1
97.3
96.1
98.1
96.9
98.1
97.3
96.5
96.1
95.7
98.1
97.3
97.7
96.9
96.5
96.5
96.1
94.9
92.2
94.6
93.8
93.0
91.8
92.2
91.8
93.0
89.9
87.2
88.3
85.2
83.3
83.7
74.3
75.1
71.2
67.7
52.9
numeric_reasoning
224
92.0
87.5
87.1
87.1
87.5
88.4
87.1
88.4
86.6
88.8
86.6
87.9
86.2
87.1
80.4
80.4
82.1
83.9
82.1
80.4
83.0
82.1
80.4
78.6
78.6
79.0
77.2
79.5
71.9
75.0
72.3
67.0
70.1
62.1
59.8
58.5
54.9
42.0
paraphrase
218
95.4
93.1
91.7
91.7
92.2
91.7
92.2
92.2
92.2
91.3
92.2
91.7
93.1
91.3
89.4
89.9
86.2
90.4
88.1
87.6
89.0
88.1
86.2
86.7
83.9
85.8
86.2
83.5
83.0
83.0
83.9
83.5
80.7
71.1
74.8
66.5
67.4
49.1
exception
213
95.3
96.2
93.9
93.9
96.2
94.8
96.2
93.9
93.9
95.3
93.0
93.0
89.7
91.1
92.0
90.6
89.7
91.5
92.0
88.3
89.7
89.2
88.7
86.4
87.3
85.9
88.3
85.9
80.3
80.8
81.2
78.4
83.1
71.8
72.3
68.5
66.2
50.2
temporal_reasoning
203
89.2
88.2
84.2
85.7
88.2
85.2
87.2
85.7
85.7
84.2
85.2
83.7
80.3
85.2
82.8
81.8
76.4
82.3
82.8
81.3
79.3
80.8
80.3
74.9
77.8
79.8
78.8
74.9
71.9
70.0
71.4
69.5
69.0
62.1
59.6
59.6
53.7
44.3
lexical_overlap
201
98.5
97.0
97.5
97.5
96.5
97.0
96.5
96.5
97.0
96.5
98.0
98.5
96.5
97.0
96.5
97.0
96.0
94.5
94.5
94.5
94.5
95.5
95.0
91.5
93.0
93.5
95.5
91.5
90.5
89.1
88.6
88.6
88.6
80.6
77.6
73.1
70.1
49.3
long_state
191
99.0
98.4
97.4
98.4
98.4
95.8
98.4
95.8
97.4
95.8
98.4
97.4
95.3
94.8
95.3
94.8
94.8
96.3
92.7
92.7
91.1
92.1
89.0
89.5
89.0
89.5
90.6
90.6
84.3
87.4
82.7
82.7
81.7
77.5
59.7
74.3
59.2
36.6
role_reversal
187
96.8
94.1
95.7
96.3
93.6
94.1
93.6
94.1
94.7
94.7
95.2
95.2
92.5
94.7
94.7
95.2
93.0
94.7
91.4
89.3
88.8
93.0
89.8
90.9
93.0
90.9
92.5
88.8
85.0
82.4
85.0
86.6
82.9
74.3
69.5
67.9
59.4
43.9
contradiction
172
98.8
97.7
96.5
96.5
97.1
95.9
97.1
95.9
96.5
95.9
94.8
95.9
95.9
95.9
94.2
94.8
95.3
95.9
95.3
94.2
94.8
93.6
93.0
91.3
92.4
93.6
95.3
91.9
85.5
89.5
86.0
86.0
86.6
76.7
73.3
71.5
64.0
52.3
injection
151
97.4
95.4
92.7
92.7
95.4
94.7
95.4
94.7
93.4
95.4
96.7
93.4
94.7
94.0
89.4
92.1
87.4
92.1
88.7
86.8
87.4
86.1
83.4
88.7
82.1
85.4
84.1
79.5
74.2
76.8
72.8
71.5
73.5
70.2
57.6
62.3
58.9
41.7
multi_turn
149
96.0
92.6
92.6
92.6
93.3
94.0
92.6
94.0
91.9
94.0
93.3
93.3
90.6
94.0
93.3
90.6
93.3
91.3
92.6
91.9
93.3
90.6
92.6
85.2
91.3
92.6
91.9
87.2
84.6
87.9
83.9
83.9
84.6
77.9
73.8
69.1
67.8
45.6
sarcasm
149
97.3
96.0
96.0
96.0
95.3
94.6
96.0
95.3
95.3
94.6
94.6
96.0
94.0
91.3
92.6
91.9
91.3
85.9
86.6
91.9
85.2
87.2
83.2
86.6
82.6
81.9
83.9
76.5
75.2
82.6
74.5
71.8
71.8
69.1
67.1
59.7
55.0
45.6
missing_evidence
141
97.9
96.5
98.6
97.9
96.5
97.9
96.5
97.9
97.9
97.9
96.5
98.6
100.0
97.9
95.0
97.9
97.2
95.0
95.0
95.7
93.6
94.3
94.3
96.5
94.3
96.5
95.0
89.4
92.9
88.7
90.8
92.9
86.5
82.3
80.1
77.3
70.9
57.4
hypothetical
128
98.4
96.1
96.1
97.7
96.1
94.5
96.1
94.5
96.9
94.5
95.3
96.1
93.0
93.8
96.9
93.8
92.2
93.0
93.0
95.3
94.5
91.4
95.3
90.6
92.2
92.2
93.8
89.1
89.1
88.3
88.3
92.2
84.4
81.2
75.0
74.2
70.3
51.6
double_negation
126
99.2
96.8
95.2
94.4
96.8
94.4
96.8
94.4
94.4
93.7
96.8
96.0
93.7
96.0
91.3
94.4
92.9
92.9
90.5
88.1
88.1
91.3
88.9
92.9
87.3
87.3
88.9
88.1
87.3
85.7
87.3
85.7
75.4
65.9
68.3
59.5
54.0
41.3
nota
118
95.8
95.8
95.8
94.1
95.8
94.1
94.9
93.2
94.9
94.1
92.4
93.2
92.4
94.1
91.5
92.4
89.8
88.1
89.8
88.1
86.4
88.1
86.4
89.8
89.0
91.5
89.0
89.0
82.2
80.5
82.2
83.1
81.4
72.0
63.6
66.9
58.5
30.5
evidence_middle
114
99.1
99.1
99.1
100.0
99.1
98.2
99.1
98.2
99.1
98.2
100.0
97.4
99.1
97.4
97.4
96.5
96.5
98.2
97.4
94.7
93.9
94.7
93.0
91.2
94.7
93.9
92.1
93.9
90.4
91.2
89.5
87.7
89.5
82.5
66.7
77.2
69.3
43.0
zero_positive
73
95.9
94.5
93.2
93.2
94.5
93.2
94.5
93.2
91.8
93.2
93.2
90.4
91.8
95.9
93.2
87.7
91.8
90.4
87.7
89.0
90.4
89.0
93.2
82.2
86.3
87.7
89.0
87.7
79.5
89.0
83.6
80.8
78.1
79.5
75.3
68.5
65.8
75.3
evidence_end
57
100.0
98.2
96.5
96.5
98.2
96.5
98.2
96.5
96.5
96.5
96.5
98.2
94.7
94.7
98.2
93.0
94.7
98.2
91.2
93.0
94.7
94.7
93.0
94.7
89.5
93.0
94.7
86.0
84.2
91.2
84.2
82.5
86.0
80.7
59.6
77.2
56.1
31.6
evidence_start
17
94.1
100.0
94.1
100.0
100.0
94.1
100.0
94.1
94.1
94.1
100.0
100.0
94.1
94.1
88.2
100.0
100.0
94.1
100.0
100.0
82.4
94.1
70.6
88.2
82.4
88.2
88.2
94.1
76.5
88.2
70.6
76.5
64.7
70.6
47.1
70.6
52.9
47.1
zero_positive_distractor
1
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
0.0
100.0
100.0
100.0
0.0
paired exact McNemar (questions only the row / only the other gets right)
run
vs selfjev-4b (qwen35_4b_tree_scratch_jevall_)
vs Jev
jev
52 / 23 (p = 0.0011)
—
qwen35_4b_tree_scratch_jevall_
—
23 / 52 (p = 0.0011)
qwen35_4b_tree_sft_jevall__last
27 / 28 (p = 1)
27 / 57 (p = 0.0014)
qwen35_4b_tree_cost_jevall__last
29 / 31 (p = 0.9)
26 / 57 (p = 0.00088)
selfjev_4b_treeserver
1 / 3 (p = 0.62)
24 / 55 (p = 0.00064)
qwen35_4b_tree_rlcd_fresh_
30 / 33 (p = 0.8)
30 / 62 (p = 0.0011)
qwen35_4b_tree_scratch_jevall__last
0 / 3 (p = 0.25)
23 / 55 (p = 0.00038)
qwen35_4b_tree
29 / 33 (p = 0.7)
31 / 64 (p = 0.00092)
qwen35_4b_tree_rlcd_jevall__last
22 / 29 (p = 0.4)
26 / 62 (p = 0.00016)
qwen35_4b_tree_sft_fresh_
29 / 36 (p = 0.46)
30 / 66 (p = 0.00031)
selfjev_4b_v2
33 / 41 (p = 0.42)
24 / 61 (p = 7.4e-05)
selfjev_4b_repro
28 / 42 (p = 0.12)
22 / 65 (p = 4.3e-06)
qwen35_4b_combo
32 / 57 (p = 0.011)
23 / 77 (p = 5.5e-08)
tree_4b_combo
38 / 64 (p = 0.013)
27 / 82 (p = 1.3e-07)
qwen35_4b_r2x64
24 / 65 (p = 1.6e-05)
21 / 91 (p = 1.4e-11)
tree_4b_instruct_r3
31 / 80 (p = 3.7e-06)
21 / 99 (p = 2.7e-13)
tree_4b_combo_r2
31 / 89 (p = 1.1e-07)
18 / 105 (p = 4e-16)
eikos_4b
25 / 85 (p = 7.8e-09)
20 / 109 (p = 5.1e-16)
tree_4b_instruct_r2x64
31 / 92 (p = 3.3e-08)
25 / 115 (p = 5.4e-15)
tree_4b_combo_ptr
33 / 108 (p = 1.7e-10)
23 / 127 (p = 1.3e-18)
tree_4b_ova
34 / 118 (p = 4.6e-12)
24 / 137 (p = 2e-20)
curve/tree_4b_r2b_r64_mlp
30 / 119 (p = 9.6e-14)
26 / 144 (p = 5.3e-21)
tree_4b_ova_kd
40 / 136 (p = 1.9e-13)
29 / 154 (p = 8.9e-22)
tree_4b_instruct_r3_step300
21 / 120 (p = 4.8e-18)
11 / 139 (p = 2.3e-29)
tree_4b_r2b
30 / 134 (p = 6.8e-17)
24 / 157 (p = 3.8e-25)
tree_4b_r2
34 / 142 (p = 6.8e-17)
26 / 163 (p = 1.9e-25)
curve/tree_4b_r2b_r64
29 / 139 (p = 2e-18)
22 / 161 (p = 2.7e-27)
tree_4b_instruct
25 / 177 (p = 2.1e-29)
15 / 196 (p = 2.2e-41)
curve/tree_4b_r64
27 / 198 (p = 2.5e-33)
14 / 214 (p = 3.9e-47)
lora_4b
32 / 216 (p = 1e-34)
20 / 233 (p = 3.3e-47)
curve/tree_4b_mlp
29 / 217 (p = 9e-37)
17 / 234 (p = 6e-50)
tree_4b
30 / 242 (p = 2.3e-42)
20 / 261 (p = 1.1e-54)
qwen35_r1
26 / 255 (p = 2e-48)
16 / 274 (p = 8.4e-62)
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b
23 / 401 (p = 2.8e-90)
17 / 424 (p = 6.8e-103)
jina_r2b
32 / 480 (p = 1.1e-103)
22 / 499 (p = 1e-118)
t5_round2b_2026-09-24/t5_r1
20 / 474 (p = 8.5e-114)
12 / 495 (p = 2.6e-129)
lora_pilot
19 / 556 (p = 2.8e-138)
11 / 577 (p = 1.3e-154)
jina_zeroshot
27 / 1008 (p = 9.2e-259)
22 / 1032 (p = 2.4e-272)
eval_llm : LLM-evaluation use cases
946 frozen questions about LLM prompts, reasoning traces and outputs (score, judge, verify, guardrail, jailbreak),
written by eval2 's authors and kept only when two blind judges agree (how it was built ). Scored
so far:
run
eval_llm
binary
multiclass
multilabel EM
Jev (API)
92.5
95.1
95.3
81.3
selfjev-4b
93.1
94.3
95.3
86.8
qwen35_4b_tree_sft_jevall__last: qwen35_4b_tree + one epoch on Jev's soft targets (incl. the LLM-evaluation data)
90.1
92.8
93.1
78.0
qwen35_4b_tree: never trained on LLM-evaluation data
82.1
83.8
88.4
68.1
selfjev-4b vs Jev: 33 / 27, p = 0.52. Sources: reports/*/eval_llm/report.json; Jev from
data/eval_llm/review/jev_answers.jsonl (llm_eval_data.md ).
Dev benchmark: the original test split
3,471 questions: 3,300 from public datasets (1,800 of them from families never trained on) and 171 authored. Frontier
models top out at 86% here because of public-label noise, so it separates small models from large ones but not good
models from each other. Selected rows:
model
question acc %
binary AUROC
multiclass macro-F1 %
multilabel EM %
binary ECE
GPT-6 Astra (reasoning low)
85.8
0.971
95.3
55.8
0.050
Jev
82.7
0.981
91.6
40.1
0.045
selfjev-4b: Qwen3.5-4B tree, from scratch on 80K questions with Jev's soft targets
83.8
0.974
92.7
52.3
0.048
Qwen3.5-4B trained with the tree, r64 , round-2b + round-3 data, all options in question (qwen35_4b_tree)
84.4
0.972
92.6
59.3
0.038
Qwen3.5-4B, r64 , round-2b + round-3 data, all options in question
84.3
0.963
92.2
62.2
0.047
tree Instruct-4B, r64 , round-2b + round-3 data, all options in question
82.7
0.959
89.3
57.6
0.064
tree Instruct-4B, r64 , round-2b data, all options in question
83.5
0.966
89.3
59.0
0.038
tree Instruct-4B, r64 , round-2b data
82.7
0.965
88.0
52.9
0.045
tree Reranker-4B, round-1 data
81.6
0.953
89.7
51.7
0.077
stock 8B + LoRA
80.7
0.945
88.7
48.0
0.051
stock 4B + LoRA
80.3
0.945
86.2
47.7
0.056
stock 0.6B + LoRA
73.5
0.863
82.3
31.1
0.066
stock 8B, untrained
66.2
0.658
86.5
10.2
0.364
stock 4B, untrained
62.8
0.604
83.4
2.0
0.357
stock 0.6B, untrained
61.0
0.605
79.6
1.2
0.384
Per-family tables for these models: stock reranker and tree scorer . Jev and GPT-6
Astra answers are cached in reports/external/cache/, keyed by the request body: Jev reruns cost nothing, but GPT-6
Astra's requests changed on 2026-09-27 (max_tokens 6000 → 8192), so its reruns miss the cache and pay again.
selfjev-4b's 52.3 multilabel exact match is the emotion-set effect of Jev's targets (finding 1).
Every run (generated)
Written by uv run python scripts/eval/ledger.py into the experiment ledger ; sorted by eval2 , then by the
dev benchmark .
run
base
architecture
trained on
eval2 acc % (target task, 1,991 q)
dev benchmark acc % (old test , 3,471 q)
binary acc %
binary AUROC
multiclass acc %
multilabel EM %
authored eval_* acc % (n=171)
report
weights
external
~typesafe/jev-latest
external API
—
97.2
82.7
93.5
0.981
84.6
40.1
94.7
report
—
qwen35_4b_tree_scratch_jevall_
Qwen3.5-4B
shared document / forked native cache branches
79,943 q / 1803 steps
95.8
83.8
91.5
0.974
85.3
52.3
91.2
report
weights/selfjev_4b
qwen35_4b_tree_sft_jevall__last
Qwen3.5-4B
shared document / forked native cache branches
69,528 q / 1230 steps
95.7
84.4
90.5
0.976
85.7
59.0
92.4
report
—
qwen35_4b_tree_cost_jevall__last
Qwen3.5-4B
shared document / forked native cache branches
69,528 q / 1230 steps
95.7
84.4
90.7
0.975
85.6
59.3
92.4
report
—
selfjev_4b_treeserver
Qwen3.5-4B
shared-prefix tree, forward only: text once per request, each question once, then each candidate
—
95.7
83.8
91.6
0.974
85.2
52.6
91.2
report
—
qwen35_4b_tree_rlcd_fresh_
Qwen3.5-4B
shared document / forked native cache branches
4,412 q / 97 steps
95.6
84.4
90.2
0.972
85.6
59.9
90.1
report
—
qwen35_4b_tree_scratch_jevall__last
Qwen3.5-4B
shared document / forked native cache branches
79,943 q / 1803 steps
95.6
83.8
91.7
0.974
85.3
52.0
91.2
report
—
qwen35_4b_tree
Qwen3.5-4B
shared document / forked native cache branches
51,774 q / 941 steps
95.6
84.4
90.5
0.972
85.7
59.3
90.1
report
—
qwen35_4b_tree_sft_fresh_
Qwen3.5-4B
shared document / forked native cache branches
4,412 q / 97 steps
95.4
84.5
90.2
0.971
85.8
59.9
90.1
report
—
qwen35_4b_tree_rlcd_jevall__last
Qwen3.5-4B
shared document / forked native cache branches
69,528 q / 1230 steps
95.4
84.5
90.7
0.975
85.5
60.8
92.4
report
—
selfjev_4b_v2
Qwen3.5-4B
shared-prefix tree, forward only: text once per request, each question once, then each candidate
83,581 q / 1912 steps
95.4
83.5
91.2
0.973
85.2
51.7
88.9
report
—
selfjev_4b_repro
Qwen3.5-4B
shared-prefix tree, forward only: text once per request, each question once, then each candidate
79,943 q / 1803 steps
95.1
83.8
91.8
0.973
85.2
52.3
91.8
report
—
qwen35_4b_combo
Qwen3.5-4B
shared document / forked native cache branches
43,827 q / 2474 steps
94.5
84.3
90.0
0.963
85.3
62.2
84.2
report
—
tree_4b_combo
Qwen3-4B-Instruct-2507
shared-prefix tree (tree-v1)
51,826 q / 967 steps
94.5
82.7
89.1
0.959
83.8
57.6
86.0
report
—
qwen35_4b_r2x64
Qwen3.5-4B
shared document / forked native cache branches
17,435 q / 769 steps
93.7
83.4
90.9
0.971
84.0
58.4
87.7
report
—
tree_4b_instruct_r3
Qwen3-4B-Instruct-2507
shared-prefix tree (tree-v1)
51,853 q / 910 steps
93.3
82.8
89.8
0.966
83.9
56.1
85.4
report
—
tree_4b_combo_r2
Qwen3-4B-Instruct-2507
shared-prefix tree (tree-v1)
18,676 q / 239 steps
92.9
83.5
90.0
0.966
84.4
59.0
84.8
report
—
tree_4b_instruct_r2x64
Qwen3-4B-Instruct-2507
shared-prefix tree (tree-v1)
18,681 q / 218 steps
92.7
82.7
90.9
0.965
83.7
52.9
81.9
report
—
tree_4b_combo_ptr
Qwen3-4B-Instruct-2507
shared-prefix tree (tree-v1)
18,676 q / 232 steps
92.0
82.7
89.6
0.965
83.6
57.0
81.3
report
—
tree_4b_ova
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
16,372 q / 241 steps
91.6
82.6
89.2
0.962
83.6
57.8
84.2
report
—
curve/tree_4b_r2b_r64_mlp
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
16,375 q / 208 steps
91.3
82.4
89.0
0.962
83.5
57.0
84.2
report
—
tree_4b_ova_kd
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
16,372 q / 241 steps
91.0
82.5
88.1
0.959
83.8
58.4
84.8
report
—
tree_4b_instruct_r3_step300
Qwen3-4B-Instruct-2507
shared-prefix tree (tree-v1)
—
90.8
81.5
89.4
0.954
82.5
52.6
81.3
report
—
tree_4b_r2b
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
16,375 q / 225 steps
90.6
81.2
88.1
0.955
82.8
51.7
83.0
report
—
tree_4b_r2
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
16,357 q / 224 steps
90.4
80.6
87.6
0.956
82.2
50.9
81.9
report
—
curve/tree_4b_r2b_r64
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
16,375 q / 208 steps
90.3
81.2
87.8
0.962
82.5
54.4
82.5
report
—
tree_4b_instruct
Qwen3-4B-Instruct-2507
shared-prefix tree (tree-v1)
10,112 q / 78 steps
88.1
80.4
87.8
0.958
82.3
46.8
79.5
report
—
curve/tree_4b_r64
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
10,080 q / 84 steps
87.2
81.5
88.5
0.956
83.0
52.3
77.8
report
—
lora_4b
Qwen3-Reranker-4B
stock pairs (answer-v1)
10,112 q / 216 steps
86.5
80.3
87.5
0.945
82.3
47.7
70.8
report
—
curve/tree_4b_mlp
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
10,080 q / 84 steps
86.3
81.6
87.9
0.955
83.4
51.7
78.9
report
—
tree_4b
Qwen3-Reranker-4B
shared-prefix tree (tree-v1)
10,112 q / 84 steps
85.1
81.6
87.1
0.953
83.9
51.7
78.4
report
—
t5_round2b_2026-09-24/decoder_only_audit/t5_r2b
google/t5gemma-2-1b-1b
pretrained T5Gemma encoder/decoder; shared document cross-KV views
16,375 q / 228 steps
76.8
73.8
82.7
0.893
75.1
40.4
62.6
report
—
jina_r2b
jinaai/jina-reranker-v3.5
jina listwise (jina-v1)
16,301 q / 345 steps
73.3
76.6
81.4
0.888
79.4
45.1
60.2
report
—
t5_round2b_2026-09-24/t5_r1
google/t5gemma-2-1b-1b
pretrained T5Gemma encoder/decoder; shared document cross-KV views
10,080 q / 823 steps
73.0
75.4
83.0
0.887
76.6
45.6
62.0
report
—
lora_pilot
Qwen3-Reranker-0.6B
stock pairs (task-v1)
10,112 q / 183 steps
68.8
73.5
77.8
0.863
78.3
31.1
55.0
report
—
external
openai/gpt-6-astra
external API
—
—
85.8
93.8
0.971
87.0
55.8
100.0
report
—
curve/boolq_300
Qwen3-Reranker-4B
stock pairs (answer-v1)
10,412 q / 218 steps
—
80.9
88.0
0.940
82.5
50.9
71.9
report
—
curve/boolq_1000
Qwen3-Reranker-4B
stock pairs (answer-v1)
11,112 q / 224 steps
—
80.9
88.8
0.954
82.5
48.8
73.7
report
—
curve/boolq_3000
Qwen3-Reranker-4B
stock pairs (answer-v1)
13,112 q / 240 steps
—
80.9
88.9
0.954
82.5
47.7
73.1
report
—
lora_8b
Qwen3-Reranker-8B
stock pairs (task-v1)
10,112 q / 185 steps
—
80.7
87.3
0.945
82.9
48.0
74.9
report
—
curve/instruct_lora
Qwen3-4B-Instruct-2507
stock pairs (answer-v1)
10,112 q / 216 steps
—
80.6
87.4
0.944
82.6
48.3
75.4
report
—
curve/boolq_100
Qwen3-Reranker-4B
stock pairs (answer-v1)
10,212 q / 217 steps
—
80.0
86.9
0.945
82.0
48.3
70.8
report
—
curve/vol100_e2
Qwen3-Reranker-4B
stock pairs (answer-v1)
10,112 q / 431 steps
—
79.8
87.0
0.951
81.9
46.2
69.6
report
—
curve/vol50_e2
Qwen3-Reranker-4B
stock pairs (answer-v1)
5,056 q / 214 steps
—
79.1
84.8
0.940
81.7
46.8
72.5
report
—
curve/vol25_e2
Qwen3-Reranker-4B
stock pairs (answer-v1)
2,528 q / 107 steps
—
79.1
85.7
0.935
81.4
45.6
73.7
report
—
curve/vol50_e1
Qwen3-Reranker-4B
stock pairs (answer-v1)
5,056 q / 107 steps
—
78.9
84.5
0.934
81.9
44.5
70.2
report
—
curve/vol25_e1
Qwen3-Reranker-4B
stock pairs (answer-v1)
2,528 q / 53 steps
—
77.0
83.8
0.930
80.8
33.4
71.9
report
—
curve/instruct_zero
Qwen3-4B-Instruct-2507
stock pairs (answer-v1)
—
—
71.3
84.2
0.911
73.5
20.1
63.2
report
—
baseline_8b
Qwen3-Reranker-8B
stock pairs (task-v1)
—
—
66.2
56.1
0.658
79.8
10.2
50.3
report
—
baseline_4b
Qwen3-Reranker-4B
stock pairs (answer-v1)
—
—
62.8
55.0
0.604
76.2
2.0
44.4
report
—
baseline
Qwen3-Reranker-0.6B
stock pairs (task-v1)
—
—
61.0
55.3
0.605
73.3
1.2
39.2
report
—
custom_sim_lora
Qwen3-Reranker-0.6B
shared-state cross-attention (custom-v1)
10,112 q / 1176 steps
—
58.2
51.7
0.514
70.2
1.7
42.1
report
—
custom_sim_frozen
Qwen3-Reranker-0.6B
shared-state cross-attention (custom-v1)
10,112 q / 1176 steps
—
48.0
51.4
0.524
53.9
1.5
38.0
report
—
custom_distill_lora
Qwen3-Reranker-0.6B
shared-state cross-attention (custom-v1)
16,671 q / 2028 steps
—
39.7
51.3
0.451
40.6
0.9
35.1
report
—
custom_frozen
Qwen3-Reranker-0.6B
shared-state cross-attention (custom-v1)
10,112 q / 1764 steps
—
39.0
51.9
0.536
39.1
1.2
36.3
report
—
custom_distill_frozen
Qwen3-Reranker-0.6B
shared-state cross-attention (custom-v1)
16,671 q / 2028 steps
—
37.7
52.8
0.507
36.7
0.6
36.8
report
—