Speed and cost¶
Bottom line
- Sharing the text is the big win. The shared-prefix tree is 32–37× faster than stock pairs for 16 questions × 3 candidates on 8K–16K-token texts, at the same quality.
- Jev is flat at ~150 ms, from 8 to 4,096 tokens and 1 to 16 questions. Our Qwen3 tree 4B on one A10G matches it only for short texts with one question and is 5× slower at 4,096 tokens. On one H100 the same model is faster than Jev inside the machine at every size with one question (22 vs 110 ms server-side for short texts, 82 vs 122 ms at 4,096 tokens) and with 16 questions up to 2,048 tokens (4,096 × 16: 189 vs 134 ms); the remaining end-to-end gap is network distance. FP8 adds nothing.
- Serving: for an attention model, vLLM with the prefix cache and merged weights is the best general option. A shorter "compact" tree format did not pay off.
- Cost: a fully busy A10G is cheaper per request than Jev (up to 3×, less with many questions); an idle one is not. On an L40S the Qwen3 tree on vLLM is cheaper than Jev in every cell measured.
- The default model,
selfjev-4b(Qwen3.5), has no NVIDIA GPU latency yet. The numbers above are the Qwen3 tree's (tree_4b_combo, weights at tagarchive/pre-cleanup-2026-09-27). The same Qwen3.5 architecture on vLLM keeps its accuracy and is fast for one question (87–131 ms server side up to 2K tokens on an L40S), but slow for many: vLLM reuses a hybrid model's recurrent state only every 528 tokens, so each candidate recomputes part of the text.selfjev servetherefore serves it with its own tree (TreeServer, the default engine), which does the Qwen3 tree's work. A separate Apple M5 Pro MPS experiment is reported below; it is not comparable to the earlier NVIDIA sweep.
Tree vs stock pairs (same A10G, bf16, unmerged LoRA)¶
End-to-end p50, one request at a time. Sources: reports/bench/gpu_{tree_4b,stock_lora_4b}_bf16/ and …_long_bf16/.
| text tokens | questions × candidates | stock 4B + LoRA | tree 4B + LoRA | speed-up |
|---|---|---|---|---|
| 512 | 1 × 3 | 375 ms | 192 ms | 2.0× |
| 512 | 16 × 3 | 5,414 ms | 686 ms | 7.9× |
| 2,048 | 16 × 3 | 20,460 ms | 1,091 ms | 18.8× |
| 8,192 | 16 × 3 | 90,449 ms | 2,798 ms | 32.3× |
| 16,384 | 1 binary | 4,324 ms | 4,483 ms | 1.0× |
| 16,384 | 16 × 3 | 206,915 ms | 5,638 ms | 36.7× |
| 32,000 | 16 × 3 | — | 14,235 ms (14.7 GB peak) |
The stock model pays pairs × text tokens. The tree pays for the text once plus a short branch per question and candidate, so adding questions barely changes its time. With one binary question there is nothing to share.
End to end vs Jev (2026-09-23)¶
Same decisions-API requests from a Mac in California, one at a time, endpoints in random order. Jev through OpenRouter (edge 12 ms away); ours, tree 4B + LoRA, on one AWS A10G in us-east-1 (71 ms away). p50 over 10 timed rounds; choice questions with 3 options. Full tables with p95 and server-side times: reports/latency/summary.md.

| text tokens | Jev, 1 q | ours vLLM, 1 q | ours transformers, 1 q | Jev, 16 q | ours vLLM, 16 q |
|---|---|---|---|---|---|
| 8 | 148 ms | 120 ms | 196 ms | 159 ms | 336 ms |
| 512 | 156 ms | 197 ms | 263 ms | 160 ms | 467 ms |
| 2,048 | 144 ms | 424 ms | 598 ms | 156 ms | 767 ms |
| 4,096 | 154 ms | 755 ms | 1,062 ms | 174 ms | 1,195 ms |
How Jev stays flat. A fit on its 400 requests gives 132–137 ms fixed + 2.2–2.6 ms per 1,000 input tokens: about 400K tokens/s of marginal speed for one request, against ~6K tokens/s for our tree on the A10G. Every architecture must touch each token, so the flat curve means a very small cost per token: consistent with a ~0.5–1B-parameter model (or an MoE with ~1B active) on H100/B200-class GPUs, or a larger model split over several GPUs. Jev does read the whole text: 97.2% on round-2 questions whose evidence is at the end of a text over 4K tokens (up to 17.6K).
The same model on an H100 (2026-09-26)¶
The same sweep, same day, same client: Jev through OpenRouter (11 ms away) against tree_4b_combo merged on vLLM 0.30
on one p5.4xlarge spot (1× H100 80 GB, ≈ $2.54/h, us-east-2, 61 ms away), in bf16 and with vLLM's dynamic FP8
(--quantization fp8). The A10G and L40S columns are the earlier sweeps of the same model. p50 ms, wall at the client /
server inside the box (Jev: inside OpenRouter). Raw rows: reports/latency/requests_h100.jsonl. The model's weights,
its vLLM code and the sweep script (scripts/latency_sweep.py) are at tag archive/pre-cleanup-2026-09-27.
| text tokens | questions | Jev wall / server | A10G | L40S | H100 bf16 | H100 FP8 |
|---|---|---|---|---|---|---|
| 8 | 1 | 130 / 110 | 125 / 55 | 158 / 36 | 140 / 22 | 139 / 22 |
| 512 | 1 | 135 / 115 | 204 / 136 | 177 / 55 | 149 / 30 | 146 / 29 |
| 2,048 | 1 | 132 / 110 | 429 / 361 | 244 / 121 | 164 / 47 | 163 / 46 |
| 4,096 | 1 | 142 / 122 | 770 / 698 | 356 / 228 | 200 / 82 | 195 / 75 |
| 8 | 16 | 140 / 117 | 471 / 401 | 199 / 135 | 179 / 58 | 175 / 56 |
| 512 | 16 | 145 / 123 | 566 / 496 | 279 / 163 | 166 / 76 | 190 / 71 |
| 2,048 | 16 | 153 / 134 | 882 / 808 | 342 / 268 | 186 / 120 | 231 / 114 |
| 4,096 | 16 | 156 / 134 | 1,336 / 1,263 | 505 / 424 | 250 / 189 | 245 / 179 |
- Inside the machine the H100 is 2–5× faster than Jev's server time, at every size with one question and up to 2,048 tokens with 16; the one slower cell is 4,096 tokens × 16 questions (189 vs 134 ms).
- End to end we trail by 10–100 ms, and that is the network: 61 ms to Ohio against 11 ms to OpenRouter's edge. Served from a point as close as theirs, this model beats Jev's latency.
- FP8 changes nothing: a 4B on an H100 is bound by per-request overhead at these sizes, not by arithmetic.
- Per-token cost, one question, server-side: ours ≈ 20 ms fixed + ≈ 15 ms per 1,000 text tokens (≈ 68K tokens/s); Jev 132–137 ms fixed + 2.2–2.6 ms per 1,000 (≈ 400K tokens/s). Jev's marginal cost per token is still ≈ 6× lower, so it is a smaller model or more GPUs per request, but at request sizes up to 4K tokens the fixed costs decide.
- Accuracy is kept: eval2 through vLLM on the H100 scores 94.42 in bf16 (5 of 1,991 decisions differ from the transformers run's 94.48) and 94.48 with FP8 (25 differ); all 1,991 questions take 16–18 s (A10G: 144 s).
- Cost, GPU fully busy (
reports/latency/throughput_h100.json, at $2.63/h): 294 requests/s and $0.0025 per 1,000 at 8 tokens × 1 question (Jev $0.016); 105 requests/s and $0.007 at 512 × 1 (Jev $0.038); 5.9 requests/s and $0.124 at 4,096 × 16 (Jev $0.265). 2.1–6.5× cheaper than Jev in every cell on spot; at $6.88/h on demand still below Jev everywhere except within 10% at 4,096 × 16. - So the speed gap was hardware, not architecture. For this model speed is a deployment question (GPU class and
placement); the default
selfjev-4bstill has to be timed on its own engine (below).
Serving optimizations (2026-09-24)¶
A controlled experiment on one A10G, R1 LoRA, 2,048-token text × 16 questions × 3 candidates, new document each request, model resident, network excluded (conclusions):
| implementation | p50 |
|---|---|
| original transformers tree | 966.62 ms |
| + direct branch-mask construction | 960.63 ms |
| compact tree format, transformers | 743.50 ms |
| tree on vLLM (merged bf16, prefix cache, CUDA graphs) | 700.37 ms |
| compact tree format on vLLM | 648.49 ms |
- vLLM keeps quality: 81.53% on the dev benchmark vs 81.50% for merged transformers.
- Merging the LoRA into bf16 weights costs a little precision: unmerged 81.68% vs merged 81.50%.
- The compact format moves repeated instructions into the shared root: 38.5% fewer branch tokens (2,119 → 1,303).
It is 22.6% faster on transformers but only 7.4% on vLLM, below the predeclared 15% gate, and it adds CLINC
over-rejection (94.67 → 91.67%, all nine changes to
none). Experimental only. - Cache matters: with the document root already cached, R1 vLLM answers new questions in 409 ms; a fully repeated request takes 185 ms. Those are different workloads from a new document.
- Faster GPUs (L40S, H100, Blackwell, A100) could not be launched in three regions that day; the L40S (2026-09-25) and H100 (2026-09-26) numbers on this page came later.
Qwen3.5 on vLLM (2026-09-25, L40S)¶
qwen35_4b_tree (the same architecture as selfjev-4b) merged into Qwen3.5-4B and served by vLLM 0.30 (the code is
now src/selfjev/engine/vllm.py: selfjev merge, then selfjev serve --engine vllm), next to the Qwen3 tree
(tree_4b_combo) on the same GPU, with Jev in the same sweep from California. Accuracy through vLLM is unchanged: eval2
95.58% (transformers 95.58%), and 94.53% for the Qwen3 tree (94.48%). Server-side p50 (ms; Jev: time inside
OpenRouter):
| text tokens × questions | Jev | Qwen3 tree, vLLM | Qwen3.5, vLLM |
|---|---|---|---|
| 512 × 1 | 106 | 55 | 87 |
| 2,048 × 1 | 102 | 121 | 131 |
| 4,096 × 1 | 114 | 228 | 230 |
| 512 × 16 | 106 | 163 | 454 |
| 2,048 × 16 | 123 | 268 | 908 |
| 4,096 × 16 | 130 | 424 | 1,138 |
Both models run at about the same speed per token (17–24K tokens/s); the difference is how much of the text each has to recompute. vLLM's prefix cache shares an attention model's text at any 16-token boundary, so the Qwen3 tree computes the text once and each candidate only adds its own tokens. For Qwen3.5 the cache has to store the recurrent state itself (tens of MB), and vLLM sets its block to 528 tokens for that; every candidate prompt recomputes the text after the last block boundary plus its question:
| 16 questions | Qwen3 tree: tokens computed → time | Qwen3.5: tokens computed → time |
|---|---|---|
| 256 tokens | 3,346 → 139 ms | 20,064 → 993 ms (nothing shared below 528 tokens) |
| 2,048 tokens | 5,138 → 254 ms | 16,320 → 855 ms |
| 4,096 tokens | 7,186 → 400 ms | 17,424 → 1,025 ms |
mamba_block_size does not change this in vLLM 0.30; caching the state in bf16 halves the block to 272 tokens, with
mixed effects (256 tokens 407 ms, 1K 744 ms, 4K 689 ms). Sources: JOURNAL 2026-09-25 18:05,
reports/latency/requests_qwen35.jsonl, the probe script reports/qwen35_4b_tree/vllm/cache_probe.py (its log is not
in git).
What serves selfjev-4b now. The fix is to serve Qwen3.5 with its own tree, as in training: TreeServer
(src/selfjev/engine/tree.py) computes the text once, each question once and each candidate's own tokens, the same
work as the Qwen3 tree. Since the 2026-09-27 cleanup it is the default engine of selfjev serve, eval and bench
(vLLM stays available with --engine vllm). It matches standalone sequences in a CPU test
(tests/engine/test_tree.py) and, on GPU, the forked-cache engine's answers: the same selfjev-4b weights score the same on both engines: eval2 95.68 vs 95.78, dev benchmark 83.78 vs 83.75, eval_llm 93.13 vs 93.13 (99.8%, 99.8% and 100% of decisions identical)
(JOURNAL 2026-09-27 11:55). But its latency has not been measured on an NVIDIA GPU, including the L4
(g6.xlarge) that selfjev deploy aws picks by default: there are no benchmark numbers for it yet. The end-to-end test
(2026-09-28: an L40S, the adapter unmerged for fine-tuning, round trips including ~70 ms of network) saw 0.2–0.3 s
for one warm request with six questions, 39 s for the very first request (kernel compilation, now done at startup),
and 16 concurrent requests, batched together, returning after 18.6 s cold and 4.4–6.4 s warm while a fine-tuning job
shared the GPU: requests in one batch all wait for it. selfjev bench on a GPU box is the first step. For reference, the forked-cache engine it replaced took 156 and 252 ms for one question at 512 and
2,048 tokens, and 341 and 508 ms for 16 questions (in-process p50 on the L40S, reports/qwen35_4b_tree/bench.json).
Current model on Apple Silicon (2026-09-28)¶
The current selfjev-4b adapter merged into Qwen3.5-4B runs through TreeServer on a local Apple M5 Pro (20-core GPU,
48 GB unified memory) with PyTorch MPS in bf16. The CLI still defaults to CUDA; this is a lower-level engine experiment,
not a validated Mac server. Each cell uses one request with synthetic repeated text and three options per question. Times include
tokenization and scoring but exclude HTTP/network. Median of 10 timed calls after two warm-ups, with MPS synchronized at
both ends. Raw samples · reproduction script.
| Text tokens | 1 question | 16 questions |
|---|---|---|
| 8 | 688 ms | 6,786 ms |
| 512 | 2,018 ms | 8,599 ms |
flash-linear-attention is not installed for MPS, so the gated recurrent operation uses the correct but slower PyTorch
reference path. The NVIDIA figures above belong to an older Qwen3 model on vLLM; their difference from these Mac times
does not isolate the hardware effect. No controlled NVIDIA latency measurement exists for the current model.
Cost vs Jev¶
Jev charges $0.042 per million input tokens and bills about 372 tokens of overhead per request plus about 113 per extra 3-option question. Our cost assumes the A10G ($1.006/h) is fully busy (batched throughput):
| request | ours, vLLM | Jev |
|---|---|---|
| 8 tokens, 1 question | $0.0051 / 1K requests | $0.0160 |
| 512 tokens, 1 question | $0.0212 | $0.0379 |
| 512 tokens, 16 questions | $0.0902 | $0.1094 |
| 2,048 tokens, 16 questions | $0.1607 | $0.1762 |
| 4,096 tokens, 16 questions | $0.2675 | $0.2653 |
A g5.xlarge left on all month (≈ $734) beats Jev only above about 7 requests/s sustained, for 512-token requests.
On an L40S ($2.242/h, fully busy; reports/latency/throughput_*_l40s.json), per 1,000 requests:
| request | Qwen3 tree, vLLM | Qwen3.5, vLLM | Jev |
|---|---|---|---|
| 8 tokens, 1 question | $0.0054 | $0.0168 | $0.0160 |
| 1,024 tokens, 1 question | $0.0320 | $0.0422 | $0.0602 |
| 4,096 tokens, 1 question | $0.1239 | $0.1268 | $0.1938 |
| 256 tokens, 16 questions | $0.0787 | $0.6330 | $0.0982 |
| 4,096 tokens, 16 questions | $0.2426 | $0.6330 | $0.2653 |
Other backends, for scale¶
| backend | request | latency | source |
|---|---|---|---|
| custom cross-attention (0.6B) | 8K tokens, 16 × 3, A10G | 626 ms (stock 0.6B + LoRA: 24,024 ms) | custom model |
| jina-reranker-v3.5 (0.6B) | eval2, per question, A10G | 58 ms (tree 4B r2b: 118 ms) | jina |
| T5Gemma 2 1B–1B | 2K × 16 × 3, L40S | 203 ms (merged tree 4B: 297 ms) | challengers |
| Qwen3.5-2B, forked cache | 8K × 16, L40S | 765 ms (tree 4B: 1,050 ms) | challengers |
| stock 0.6B / 4B / 8B + LoRA | 512 × 1 × 3, L40S | 43 / 116 / 188 ms | stock reranker |