firelex/jeff: Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification · GitHub


Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, quick resolution fashions you fit into your
code, with the identical request format as Jev.
You describe a state of affairs and checklist the choices in plain phrases; Jeff returns a
calibrated likelihood for every choice from a single ahead go. No generated textual content, no parsing: about 22 ms per
resolution on an RTX PRO 6000 and 28 ms on an Apple M4 Max (MLX).

Zero-shot means the choices might be something: assist queues, person intents, moderation labels, voice instructions, sport
strikes. Your classes need not seem within the coaching information; you describe them, and Jeff picks.

What it’s, and what it is not. These are very small fashions. They make extraordinarily quick, well-calibrated judgement
calls between choices, they usually slot simply into your native code. On benchmarks they strategy, and generally beat, Jev;
however at this dimension their reasoning will not match Jev’s, which runs on a a lot bigger mannequin. If zero-shot accuracy is not
ok in your functions, a brief fine-tune by yourself examples takes you a lot additional: our
voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in underneath half an hour on one GPU.

Built solely on native {hardware}. Training on one RTX PRO 6000 workstation GPU (the 0.8B trains in about 2 hours,
the 2B in about 3.5), all artificial coaching information written by an open mannequin (Qwen3.8-Flash-Next) on two DGX Sparks,
testing on a MacE-book. No cloud GPUs, and no closed-model output within the coaching information; a closed mannequin was used solely to
spot-check the standard of a pattern of the artificial information.

Independent mission. Jeff makes use of the identical request format as Jev, however it isn’t affiliated with or endorsed by TypeSafe, the
makers of Jev. Our coaching code begins from the open-source AutoJev recipe.

Models on Hugging Face: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B

uv sync
uv run hf obtain mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b

# NVIDIA GPU or CPU (PyTorch)
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
# Apple silicon (MLX, a lot quicker on a Mac; Qwen fashions solely)
uv sync --extra mac
JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
curl -s localhost:8765/v1/systemone -H 'content-type: software/json' -d '{
  "mannequin": "jeff-latest",
  "state": "Refund request: the shopper says the parcel arrived crushed and desires their a refund.",
  "questions": {
    "route": {"sort": "alternative", "directions": "Which crew ought to deal with this?",
              "standards": {"1": "Refunds and funds", "2": "Damaged or misplaced parcels", "3": "Account and login issues"}},
    "indignant": {"sort": "noul", "directions": "Is the shopper indignant?"}
  }
}'

Each reply has a likelihood per choice, the chosen choice and a confidence. Three query sorts: alternative (choose one
of as much as 255 choices), noul (sure/no, returned as a likelihood) and rating (some extent on a scale you describe).
Several impartial questions in a single request are answered collectively.

4,599 questions from 5 public benchmarks, plus JevBench’s public arduous tier (105 objects, scored individually):

Benchmark Qwen3.5-0.8B untrained Jeff-Qwen3.5-0.8B Qwen3.5-2B untrained Jeff-Qwen3.5-2B Gemma 4 E2B untrained Jeff-Gemma4-E2B Jev (printed) AutoJev-27B (printed)
Overall (5 benchmarks) 45.3 79.1 46.5 83.1 62.5 81.6 83.0 84.9
BBH 39.5 64.0 46.0 68.0 51.3 66.4 94.3 82.8
Financial PhraseBank 36.0 96.4 53.4 96.3 86.0 96.1 77.0 84.2
JudgeBench 56.6 62.6 57.4 64.6 46.9 60.6 78.6 78.9
RAGTruth 49.1 86.1 35.9 88.9 63.8 87.4 77.3 88.9
WinoGrande 49.2 68.6 52.2 79.0 51.0 77.4 90.7 83.3
JevBench arduous (separate) 36.2 47.6 45.7 53.3 41.0 48.6 73.3 70.3

Bold: the winner of Jeff in opposition to Jev in every row. Bold italic: AutoJev-27B the place it’s the better of all fashions
within the row (on RAGTruth, tied with Jeff-Qwen3.5-2B); it’s proven for reference, because the head-to-head comparability is
with Jev. The printed Jev and AutoJev figures have been measured on a distinct pattern of the identical benchmarks. Jeff’s
general rating comes from classification and grounding, the place it matches or beats the big fashions; on the
reasoning-heavy benchmarks (BBH, JudgeBench, JevBench) it stays properly beneath them, as you’ll count on at this dimension.

To check zero-shot efficiency on duties in contrast to something within the benchmarks, we had Jeff play three video games. Games aren’t
the perfect zero-shot check, since a sport’s state is not typical unstructured information; however they’re a typical, and enjoyable, solution to
check a System 1 mannequin. Each flip, the code describes the state of affairs and the authorized strikes in phrases, and the mannequin picks
one. The choices state what every transfer results in (Frogger: “you’ll be hit by a automotive and lose a life”; Doom: “the
nearest monster is a little bit to your left”), however by no means which transfer is true. Each result’s 20 episodes, seed 1234; ▶
opens a video of the run’s first episode.

Jeff-Qwen3.5-0.8B enjoying, zero-shot (the daring row within the desk beneath; click on a clip for the complete video):


Model Doom, kills (monster’s course in phrases) Frogger, crossings (penalties) Pac-Man, pellets of 98 (penalties)
Random strikes −0.05 0 11.2
Hand-coded rule bot 6.55 ▶ 10.25 ▶ 94.1 ▶
Qwen3.5-0.8B, untrained 5.0 ▶ 1.0 ▶ 25.8 ▶
Jeff-Qwen3.5-0.8B 6.55 ▶ 10.3 ▶ 57.0 ▶
Qwen3.5-2B, untrained 0.55 ▶ 0.05 ▶ 72.1 ▶
Jeff-Qwen3.5-2B −0.9 ▶ 6.0 ▶ 41.2 ▶
Gemma 4 E2B, untrained −0.55 ▶ 0 ▶ 3.2 ▶
Jeff-Gemma4-E2B 0.55 ▶ 0.15 ▶ 53.2 ▶
Jev (printed, Doom) 6.55, informed the aiming rule; −0.60 with out it — —

Jeff-0.8B decides in 29–49 ms per transfer on an M4 Max; Jev’s printed Doom run took 212 ms per name over its API. The
two instances weren’t measured on the identical {hardware}. To play them your self:

uv sync --extra video games
uv run python -m jeff.video games --game doom --player jeff --criteria state of affairs --url http://127.0.0.1:8765 --video --out runs/video games/doom.json
uv run python -m jeff.video games --game frogger --player jeff --criteria outcomes --url http://127.0.0.1:8765 --out runs/video games/frogger.json
uv run python -m jeff.video games --game pacman --player rule --out runs/video games/pacman-rule.json

Median time per resolution over the identical 200 benchmark questions (about 200 enter tokens every), one query at a time,
from uncooked textual content to chances:

Model Parameters Weights (16-bit) NVIDIA RTX PRO 6000 Apple M4 Max (MLX) CPU (32 threads)
Jeff-Qwen3.5-0.8B 0.8B 1.7 GB 22 ms 28 ms 463 ms
Jeff-Qwen3.5-2B 2B 4.2 GB 24 ms 60 ms 708 ms
Jeff-Gemma4-E2B 2B efficient (4.6B saved) 9.3 GB 29 ms — (MLX runs Qwen solely) 1.0 s
AutoJev-27B 27B ~54 GB not printed — —
Jev not disclosed API solely 114–212 ms per name in printed Doom runs, together with the community

  • Reason in code, resolve with Jeff. It’s a classifier, not a planner. State what every choice results in (“this transfer
    will get you hit by a automotive”); requested to forecast (“a automotive arrives in 2 turns”), it does no higher than random.
  • Wording issues enormously. Describe choices constantly: giving Frogger’s aim choice the identical phrases as each
    different ahead choice took one episode from 15 crossings to 23.
  • Use brief choice keys and descriptive textual content: {"1": "Engagement letter"}, not lengthy IDs, which price time and add
    nothing.
  • Ask impartial questions collectively in a single request.
  • Fine-tune it if zero-shot is not sufficient. A voice-navigation fine-tune on ~11k app-specific examples took about half
    an hour on one GPU and moved held-out accuracy from 31.7% to 95.8%, at about 40 ms per resolution on an M4 Max:
    autojev-train --initial-checkpoint --epochs 1 ....
  • Pick the scale for the job. For quick choice selecting the 0.8B is the candy spot: the 2B is extra cautious and performs
    the video games worse, regardless of scoring larger on the benchmarks.
uv run autojev-mix ...          # construct the coaching set (public information, artificial information, leak filter)
scripts/practice.sh RUN information/combine/public.jsonl information/combine 5e-6 40 Qwen/Qwen3.5-0.8B <revision> --epochs 1
uv run autojev-evaluate --data information/panel.jsonl --local --checkpoint checkpoints/RUN/chosen --output runs/eval/RUN.json

The full pipeline (artificial information from a neighborhood trainer, leak filter, learning-rate sweeps, dashboard) is described in
scripts/train_all.sh, and each coaching supply with its licence in docs/data-sources.md. Training recipe: full-weight fine-tuning, one epoch, batches of 256, cross-entropy over the
choice letters, then one fitted temperature for calibration; checkpoints are chosen on a growth set, by no means on the
benchmark panel. At least half of every coaching household follows the panel’s structure conventions (codecs solely; no panel
merchandise is ever skilled on).

  • Small fashions do not cause. Expect quick, calibrated decisions between the choices you describe, not multi-step
    reasoning. At 0.8B–2B parameters this holds for each mannequin, not simply Jeff.
  • Jeff-2B is a weaker sport participant than Jeff-0.8B. The untrained 2B already seems extra risk-averse than the
    untrained 0.8B, and our coaching appears to have made that worse. This wants extra investigation.
  • Benchmark scores do not predict sport play. The untrained Gemma 4 E2B beats the untrained Qwen fashions on the
    benchmarks but performs the video games worst: proper more often than not, however not reliably. Training mounted its Pac-Man (3.2 → 53.2 pellets)
    however not its Doom or Frogger.
  • Prompts matter. Jev’s personal Doom immediate (a uncooked bearing quantity plus an aiming rule) doesn’t work for any of our
    fashions; choices that state penalties in phrases do.
  • English and textual content solely.

Jeff started as a fork of AutoJev by Denis Yarats (MIT licence), an open recipe
that fine-tunes Qwen3.8-27B to return Jev-style choices. We stored its core design (one ahead go per resolution, a
skilled reply readout, a fitted temperature for calibration) and constructed on it: small college students (0.8B and 2B Qwen, Gemma
4 E2B), a neighborhood synthetic-data pipeline with a leak filter, immediate layouts for area fine-tunes, MLX serving on Apple
silicon, sport assessments and a coaching dashboard. The authentic copyright discover is stored in LICENSE.

Code: MIT (together with AutoJev’s). Model weights: Apache 2.0. Doom harness tailored from
jev-plays-doom (MIT). Training information: see the dataset card; every
supply retains its licence and is listed in docs/data-sources.md. We launch the weights and code, not the coaching information; some sources are share-alike (CC BY-SA).



Source link