Skip to content
isotomaPublic
forked from firelex/jeff

About

Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

Jeff

New: v1.1 (29 September 2026). Jeff-Qwen3.5-0.8B and Jeff-Qwen3.5-2B now choose among up to 254 options (v1.0: 26), with better calibration. On our long-list test the 0.8B goes from 40% to 95%. The 2B's benchmark score dips from 83.1% to 82.0%. Details in the changelog; v1.0 stays available on Hugging Face as revision v1.0.

Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, fast decision models you slot into your code, with the same request format as Jev. You describe a situation and list the options in plain words; Jeff returns a calibrated probability for each option from a single forward pass. No generated text, no parsing: about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max (MLX).

Zero-shot means the options can be anything: support queues, user intents, moderation labels, voice commands, game moves. Your categories don't need to appear in the training data; you describe them, and Jeff picks.

What it is, and what it isn't. These are very small models. They make extremely fast, well-calibrated judgement calls between options, and they slot easily into your local code. On benchmarks they approach, and sometimes beat, Jev; but at this size their reasoning won't match Jev's, which runs on a much larger model. If zero-shot accuracy isn't good enough for your purposes, a short fine-tune on your own examples takes you much further: our voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU.

Built entirely on local hardware. Training on one RTX PRO 6000 workstation GPU (the 0.8B trains in about 2 hours, the 2B in about 3.5), all synthetic training data written by an open model (Qwen3.8-Flash-Next) on two DGX Sparks, testing on a MacBook. No cloud GPUs, and no closed-model output in the training data; a closed model was used only to spot-check the quality of a sample of the synthetic data.

Independent project. Jeff uses the same request format as Jev, but it is not affiliated with or endorsed by TypeSafe, the makers of Jev. Our training code starts from the open-source AutoJev recipe.

Models on Hugging Face: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B · chess fine-tune: Jeff-Qwen3.5-0.8B-Chess

Quick start

To serve a model you only need the serving install (--no-default-groups leaves out the training, data and evaluation packages; plain uv sync installs everything):

uv sync --no-default-groups                 # CPU
uv sync --no-default-groups --extra cuda    # NVIDIA GPU: adds the fast kernels (much slower without them)
uv sync --no-default-groups --extra mac     # Apple silicon: adds MLX
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b

# NVIDIA GPU or CPU (PyTorch)
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run --no-default-groups jeff-serve
# Apple silicon (MLX, much faster on a Mac; Qwen models only)
JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run --no-default-groups --extra mac jeff-serve
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "model": "jeff-latest",
  "state": "Refund request: the customer says the parcel arrived crushed and wants their money back.",
  "questions": {
    "route": {"type": "choice", "instructions": "Which team should handle this?",
              "criteria": {"1": "Refunds and payments", "2": "Damaged or lost parcels", "3": "Account and login problems"}},
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

Each answer has a probability per option, the chosen option and a confidence. Three question types: choice (pick one of up to 254 options with the v1.1 Qwen models, 26 with Jeff-Gemma4-E2B), noul (yes/no, returned as a probability) and score (a point on a scale you describe). Several independent questions in one request are answered together.

Benchmarks

4,599 questions from five public benchmarks, plus JevBench's public hard tier (105 items, scored separately):

Accuracy of Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B and Jeff-Gemma4-E2B against Jev's published figures, per benchmark

Benchmark Qwen3.5-0.8B untrained Jeff-Qwen3.5-0.8B Qwen3.5-2B untrained Jeff-Qwen3.5-2B Gemma 4 E2B untrained Jeff-Gemma4-E2B Jev (published) AutoJev-27B (published)
Overall (5 benchmarks) 45.3 79.1 46.5 82.0 62.5 81.6 83.0 84.9
BBH 39.5 64.9 46.0 68.7 51.3 66.4 94.3 82.8
Financial PhraseBank 36.0 95.7 53.4 94.7 86.0 96.1 77.0 84.2
JudgeBench 56.6 63.1 57.4 59.4 46.9 60.6 78.6 78.9
RAGTruth 49.1 85.6 35.9 87.7 63.8 87.4 77.3 88.9
WinoGrande 49.2 69.0 52.2 78.8 51.0 77.4 90.7 83.3
JevBench hard (separate) 36.2 46.7 45.7 57.1 41.0 48.6 73.3 70.3

Bold: the winner of Jeff against Jev in each row. Bold italic: AutoJev-27B where it is the best of all models in the row; it is shown for reference, since the head-to-head comparison is with Jev. The Qwen columns are v1.1; Jeff-Gemma4-E2B is v1.0. The published Jev and AutoJev figures were measured on a different sample of the same benchmarks. Jeff's overall score comes from classification and grounding, where it matches or beats the large models; on the reasoning-heavy benchmarks (BBH, JudgeBench, JevBench) it stays well below them, as you would expect at this size.

Games: a zero-shot test

Measured with v1.0 of the Jeff models.

To test zero-shot performance on tasks unlike anything in the benchmarks, we had Jeff play three games. Games aren't the ideal zero-shot test, since a game's state isn't typical unstructured data; but they are a common, and fun, way to test a System 1 model. Each turn, the code describes the situation and the legal moves in words, and the model picks one. The options state what each move leads to (Frogger: "you would be hit by a car and lose a life"; Doom: "the nearest monster is a little to your left"), but never which move is right. Each result is 20 episodes, seed 1234; ▶ opens a video of the run's first episode.

Jeff-Qwen3.5-0.8B playing, zero-shot (the bold row in the table below; click a clip for the full video):

Jeff-Qwen3.5-0.8B playing Doom
Doom
Jeff-Qwen3.5-0.8B playing Frogger
Frogger
Jeff-Qwen3.5-0.8B playing Pac-Man
Pac-Man
Model Doom, kills (monster's direction in words) Frogger, crossings (consequences) Pac-Man, pellets of 98 (consequences)
Random moves −0.05 0 11.2
Hand-coded rule bot 6.55 ▶ 10.25 ▶ 94.1 ▶
Qwen3.5-0.8B, untrained 5.0 ▶ 1.0 ▶ 25.8 ▶
Jeff-Qwen3.5-0.8B 6.55 ▶ 10.3 ▶ 57.0 ▶
Qwen3.5-2B, untrained 0.55 ▶ 0.05 ▶ 72.1 ▶
Jeff-Qwen3.5-2B −0.9 ▶ 6.0 ▶ 41.2 ▶
Gemma 4 E2B, untrained −0.55 ▶ 0 ▶ 3.2 ▶
Jeff-Gemma4-E2B 0.55 ▶ 0.15 ▶ 53.2 ▶
Jev (published, Doom) 6.55, told the aiming rule; −0.60 without it — —

Jeff-0.8B decides in 29–49 ms per move on an M4 Max; Jev's published Doom run took 212 ms per call over its API. The two times were not measured on the same hardware. To play them yourself:

uv sync --extra games
uv run python -m jeff.games --game doom --player jeff --criteria situation --url http://127.0.0.1:8765 --video --out runs/games/doom.json
uv run python -m jeff.games --game frogger --player jeff --criteria outcomes --url http://127.0.0.1:8765 --out runs/games/frogger.json
uv run python -m jeff.games --game pacman --player rule --out runs/games/pacman-rule.json

Fine-tuning example: chess

When zero-shot isn't enough, fine-tune. As a worked example we trained Jeff-Qwen3.5-0.8B on 600,000 Lichess positions, labelled by Stockfish, in about 3½ hours on one GPU. On 1,000 held-out chess puzzles:

Model Puzzles solved
Qwen3.5-0.8B, untrained 6.2%
Jeff-Qwen3.5-0.8B, zero-shot (no chess training) 15.5%
Jeff-Qwen3.5-0.8B-Chess 55.8%

It is not a strong player: about 1,000 Elo with no search, and it loses to Stockfish's weakest setting. The point is speed. Each move is one forward pass in tens of milliseconds, so one GPU keeps up with about 600 human blitz games at once.

Jeff-Qwen3.5-0.8B-Chess playing 100 blitz games at once; the featured game ends in checkmate

100 games at once in real time; the featured game is its one win of the 100, a nine-move checkmate. The full video, results and training details are on the model card. The code to reproduce it, step by step, is in examples/chess.

Speed and size

Median time per decision over the same 200 benchmark questions (about 200 input tokens each), one question at a time, from raw text to probabilities:

Model Parameters Weights (16-bit) NVIDIA RTX PRO 6000 Apple M4 Max (MLX) CPU (32 threads)
Jeff-Qwen3.5-0.8B 0.8B 1.7 GB 22 ms 28 ms 463 ms
Jeff-Qwen3.5-2B 2B 4.2 GB 24 ms 60 ms 708 ms
Jeff-Gemma4-E2B 2B effective (4.6B stored) 9.3 GB 29 ms — (MLX runs Qwen only) 1.0 s
AutoJev-27B 27B ~54 GB not published — —
Jev not disclosed API only 114–212 ms per call in published Doom runs, including the network

Using it well

  • Reason in code, decide with Jeff. It's a classifier, not a planner. State what each option leads to ("this move gets you hit by a car"); asked to forecast ("a car arrives in 2 turns"), it does no better than random.
  • Wording matters enormously. Describe options consistently: giving Frogger's goal option the same words as every other forward option took one episode from 15 crossings to 23.
  • Use short option keys and descriptive text: {"1": "Engagement letter"}, not long IDs, which cost time and add nothing.
  • Ask independent questions together in one request.
  • Fine-tune it if zero-shot isn't enough. A voice-navigation fine-tune on ~11k app-specific examples took about half an hour on one GPU and moved held-out accuracy from 31.7% to 95.8%, at about 40 ms per decision on an M4 Max: jeff-train --initial-checkpoint <jeff> --epochs 1 ....
  • Pick the size for the job. For fast option picking the 0.8B is the sweet spot: the 2B is more cautious and plays the games worse, despite scoring higher on the benchmarks.

Train your own

uv run jeff-mix ...          # build the training set (public data, synthetic data, leak filter)
scripts/train.sh RUN data/mix/public.jsonl data/mix 5e-6 40 Qwen/Qwen3.5-0.8B <revision> --epochs 1
uv run jeff-evaluate --data data/panel.jsonl --local --checkpoint checkpoints/RUN/final --output runs/eval/RUN.json

The full pipeline (synthetic data from a local teacher, leak filter, learning-rate sweeps, dashboard) is described in scripts/train_all.sh, and every training source with its licence in docs/data-sources.md. Training recipe: full-weight fine-tuning, one epoch, batches of 256, cross-entropy over the option letters, then one fitted temperature for calibration. From v1.1 we use the checkpoint at the end of the epoch; the benchmark panel is never used for selection or tuning. At least half of each training family follows the panel's layout conventions (formats only; no panel item is ever trained on).

Caveats

  • Option limits. The v1.1 Qwen models are trained on lists of 20 to 254 options and accept up to 254. Jeff-Gemma4-E2B is still v1.0: it picks reliably among at most 26, and the server refuses longer lists for it; shortlist first.
  • Small models don't reason. Expect fast, calibrated choices between the options you describe, not multi-step reasoning. At 0.8B–2B parameters this holds for every model, not just Jeff.
  • Jeff-2B is a weaker game player than Jeff-0.8B. The untrained 2B already appears more risk-averse than the untrained 0.8B, and our training seems to have made that worse. This needs more investigation.
  • Benchmark scores don't predict game play. The untrained Gemma 4 E2B beats the untrained Qwen models on the benchmarks yet plays the games worst: right most of the time, but not reliably. Training fixed its Pac-Man (3.2 → 53.2 pellets) but not its Doom or Frogger.
  • Prompts matter. Jev's own Doom prompt (a raw bearing number plus an aiming rule) does not work for any of our models; options that state consequences in words do.
  • English and text only.

Changelog

v1.1 (29 September 2026): Jeff-Qwen3.5-0.8B and Jeff-Qwen3.5-2B

  • Long lists. v1.0 never picked an option past the 26th, because no training question had more than 19 options (thanks to @puhuk for the report, #1). v1.1 adds 32,000 training questions with 20 to 254 options, built in code (src/jeff/longlists.py) and from the MASSIVE and CLINC150 training splits, and accepts up to 254 options. Long-list test (20 to 254 items): 0.8B 40.3% → 94.7%, 2B 95.2%.
  • Final checkpoint. v1.1 publishes the checkpoint at the end of the epoch. Picking the one with the lowest development loss, as v1.0 did, chose an early, less settled checkpoint in the new runs and cost 1–2 points.
  • Calibration error 0.049 → 0.021 (0.8B) and 0.028 → 0.026 (2B). JevBench hard tier: 47.6% → 46.7% (0.8B), 53.3% → 57.1% (2B).
  • Benchmarks: 0.8B unchanged at 79.1%; 2B 83.1% → 82.0%, mostly on JudgeBench (64.6% → 59.4%), a small benchmark where models of this size sit near chance. The 2B no longer edges past Jev's published 83.0%.
  • Not zero-shot any more: MASSIVE and CLINC150 (their training splits are now in the training data).
  • Serving-only install (uv sync --no-default-groups), thanks to @WavesMan (#2).
  • Jeff-Gemma4-E2B was not retrained and stays at v1.0. v1.0 of every model stays available on Hugging Face as revision v1.0.

v1.0 (28 September 2026): first release: Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B and Jeff-Gemma4-E2B.

History

Jeff began as a fork of AutoJev by Denis Yarats (MIT licence), an open recipe that fine-tunes Qwen3.8-27B to return Jev-style decisions. We kept its core design (one forward pass per decision, a trained answer readout, a fitted temperature for calibration) and built on it: small students (0.8B and 2B Qwen, Gemma 4 E2B), a local synthetic-data pipeline with a leak filter, prompt layouts for domain fine-tunes, MLX serving on Apple silicon, game tests and a training dashboard. The original copyright notice is kept in LICENSE.

Licence

Code: MIT (including AutoJev's). Model weights: Apache 2.0. Doom harness adapted from jev-plays-doom (MIT). Training data: see the dataset card; each source keeps its licence and is listed in docs/data-sources.md. We release the weights and code, not the training data; some sources are share-alike (CC BY-SA).

About

Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages