Spend¶
Real costs logged in the journal, BYOK upstream charges included. About $829 logged as of 2026-09-27: the sum of every row of the journal's spend table (values marked ≈ counted as written), of which data ≈ $563, evaluation ≈ $24 and AWS GPU ≈ $242. Not in it: a few GPU boxes that were never reconciled (the early tree and custom-model runs, the jina box ≈ $1.60, the latency/compact and T5Gemma L40S boxes).
| item | cost |
|---|---|
| Data | |
| Round-3 training data: writing (Luna $10.17, Gemini $55.30, Grok $49.06) + Astra batch judge (upper bound at list batch prices) | ≈ $280 |
LLM-evaluation data (llm_eval_data.md): 9.4K training questions after safety passes (writing $37.09 + Astra $37.95) + eval_llm test set (writing $15.55, judges $5.76, Jev $0.03) |
$96.38 |
llm_multilabel_v1 multilabel batch: writing $26.60 + Astra judge $42.36 |
$68.96 |
| Round-2 hard cases: writing $22.86 + blind Astra judge $36.24 + Jev second opinion $0.27 | $59.37 |
| eval2: writing, two judges, Jev second opinion | $32.72 |
mpos_distr_num_v1: 3,645 several-correct / distractor / number questions (Luna $1.87 + Astra $17.80 + Jev $0.14) |
$19.81 |
numdate_neg_v1: 549 double-negation / number / date questions (Luna $0.28 + Astra $2.78 + Jev $0.02) |
$3.08 |
Jev on every question of data/all.jsonl.gz (33,707 texts not cached) |
$1.95 |
Jev predictions for llm_multilabel_v1 |
$0.30 |
| Evaluation against Jev and GPT-6 Astra | |
| Jev + GPT-6 Astra on the 3,471 dev-benchmark questions (Jev $0.06, Astra $19.27) | $19.34 |
| Test-failure audit (Opus 5.5 relabel, 968 questions) | $4.92 |
| Jev on eval2 | $0.05 |
| AWS GPU (a few rows include the Jev calls of the same job) | |
| Jev soft targets on all data: RLCD (L40S, 7.2 h) + fine-tune (L40S, 6.0 h) | ≈ $29.70 |
| 10 learning-curve runs, 8 g5.xlarge | ≈ $23.00 |
selfjev-4b: retrain from scratch on everything with Jev targets, texts to 16K (L40S, 9.3 h) |
≈ $20.80 |
selfjev_4b_v2: the same recipe + batch mpos_distr_num_v1, worse, not kept (L40S, 9.6 h incl. 23 min idle) |
≈ $21.46 |
Engine check: selfjev-4b re-scored by TreeServer on the three test sets (A10G, 31 min) |
≈ $0.52 |
| End-to-end test of the product: deploy, SDK checks, two fine-tuning jobs (L40S g6e.xlarge, 26 min) | ≈ $0.81 |
selfjev_4b_repro: selfjev-4b's exact data rerun with the new code (L40S, 9.1 h) |
≈ $20.44 |
| Jev soft targets, C: RLCD with a confident-mistake cost from B (L40S, 7.1 h) | ≈ $15.90 |
qwen35_4b_combo: Qwen3.5 + option lists + round 3, full sequences (L40S, 7.0 h) |
≈ $15.80 |
| All-options, 27B teacher and distillation boxes | ≈ $12.00 |
tree_4b_combo: Qwen3 tree + option lists + round 3 (g5.xlarge, 11.7 h) |
≈ $11.70 |
qwen35_4b_tree: Qwen3.5 trained with the tree, the previous default (L40S, 4.8 h) |
≈ $10.80 |
tree_4b_instruct_r3: round-3 rerun (L40S) |
≈ $9.25 |
| Round-2 box: tree r2, stopped stock r2, tree r2b (g6e.4xlarge) | ≈ $6.70 |
| Round-3 training box (stopped) + scoring box | ≈ $6.25 |
tree_4b_combo_r2 and tree_4b_combo_ptr (2 g5.xlarge) |
≈ $6.22 |
| First RLCD test: RLCD and a fine-tune control on fresh questions (L40S, 2.8 h) | ≈ $6.20 |
| Stock 4B and 8B pipelines, one g6e.xlarge, 3.16 h | ≈ $5.89 |
| Combined levers: round-2b data + r64 / + MLP, 2 g5.xlarge | ≈ $5.26 |
qwen35_4b_r2x64: first Qwen3.5-4B run (g5.xlarge) |
≈ $4.75 |
| Instruct base + round-2 data + r64 | ≈ $2.45 |
| Qwen3.5 on vLLM: eval2 parity, latency sweep vs Jev, throughput (L40S, 50 min, + Jev $0.02) | ≈ $1.88 |
| Tree LoRA-capacity ablation, 2 g5.xlarge | ≈ $1.60 |
H100 latency sweep: tree_4b_combo on vLLM, bf16 and FP8 (p5.4xlarge spot, 28 min, + Jev $0.02) |
≈ $1.20 |
| eval2 scoring box | ≈ $1.12 |
| Latency sweep vs Jev (AWS $0.87 + Jev $0.04) | ≈ $0.91 |
Unit costs worth remembering:
| what | cost |
|---|---|
| g5.xlarge (A10G 24 GB) | $1.006/h: one 4B LoRA on 10K questions ≈ 40 min; the dev benchmark ≈ 4 min |
| g6e.2xlarge (L40S 48 GB) | $2.242/h: selfjev-4b (79.9K questions, texts to 16K) trained and scored in one 9.3 h box ≈ $21; qwen35_4b_tree (51.8K questions, texts to 8K) trained in 4.1 h; eval2 through vLLM in 2 min |
| g6e.4xlarge (L40S 48 GB), us-east-2 | $3.00424/h |
| g6.xlarge (L4 24 GB) | $0.805/h: the default machine of selfjev deploy aws; nothing timed on it yet |
| p5.4xlarge (H100 80 GB) | spot $2.544/h (obtained once, 2026-09-26, us-east-2b); on demand $6.88/h (never obtained) |
| GPT-6 Luna as a data writer | ≈ $0.001 per text ($1.62 for 1,676 in round 2, $1.87 for 1,671 in mpos_distr_num_v1) |
| Gemini 3.8 Flash as a data writer | ≈ $0.014 per text; through Google's Batch API ≈ $2.3 per 1,000 questions vs ≈ $7 through OpenRouter; Grok 4.7 ≈ $0.026 per text (mandatory hidden reasoning) |
| GPT-6 Astra as a blind judge, OpenAI Batch API | ≈ $3.40–4.60 per 1,000 questions |
| Jev | $0.042 per million input tokens |
Rules (AGENTS.md): paid resources need the user's explicit OK with a price, every time. AWS boxes
are tagged Project=personal-jev, get a shutdown -h cap with terminate-on-shutdown, and are terminated, with their
security group and key pair deleted, when done.