Trained at Home: The 0.8B Decision Models Are Ahead of Schedule
A 0.8B decision model trained at home answers in 30 ms. We checked the Hacker News wave against the project README and nine papers, from phi-1 to ThinkSLM.
Last weekend a project called Jeff hit the Hacker News front page: 518 points at the time of writing, posted September 28, 2026. The pitch sounds like a threshold being crossed. Decision models of 0.8B and 2B parameters, fine-tuned at home in about two hours, answering multiple-choice questions in roughly 30 ms, speaking the same request format as Jev, the hosted decision API from TypeSafe. No generated text, no parsing, one forward pass per decision.
Read the fine print and two things happen. First, the claims hold up: we checked every number against the project README and the benchmark table. Second, "trained at home" turns out to mean two DGX Sparks writing synthetic data and an RTX PRO 6000 doing the training, which one commenter summed up as "that's some local hardware." Still, the interesting part is not the hardware flex. It is that none of this should surprise anyone who read the last three years of the literature. The recipe was sitting in plain sight. Here is the verified rundown, and the research that predicted it.
What Jeff is, and is not
Every number below comes from the project README, accessed September 30, 2026. Jeff is a set of fine-tunes of Qwen3.5 (0.8B and 2B) and Gemma 4 E2B for zero-shot classification. You describe a situation and list the options in plain words; the model returns a calibrated probability for each option in a single forward pass. Latency on the 0.8B: 22 ms median on an RTX PRO 6000, 28 ms on an Apple M4 Max through MLX, 463 ms on a 32-thread CPU. The weights are 1.7 GB in 16-bit, Apache 2.0 licensed, with MIT training code forked from the open AutoJev recipe. The project states plainly that it is independent of TypeSafe.
On their five-benchmark panel, the 0.8B scores 79.1 overall against Jev's published 83.0, and the 2B reaches 83.1, though the README is careful to note the published figures were measured on a different sample of the same benchmarks. The shape of the wins and losses matters more than the averages. Where the task is pure classification, the small models win outright: 96.4 on Financial PhraseBank against Jev's 77.0. Where reasoning enters, the gap is brutal:
| Benchmark | Jeff 0.8B | Jev (published) |
|---|---|---|
| Overall (5 benchmarks) | 79.1 | 83.0 |
| Financial PhraseBank | 96.4 | 77.0 |
| RAGTruth | 86.1 | 77.3 |
| BBH | 64.0 | 94.3 |
| JevBench hard tier | 47.6 | 73.3 |
Their own caveats section is refreshingly blunt. At most 26 options per question, because the models never learned two-letter answer codes, so positions 27 and beyond are effectively never chosen. English text only. No multi-step reasoning. The README also demonstrates the ceiling honestly: a voice-navigation fine-tune took held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU, which is a great result that quietly concedes zero-shot was not enough for that task.
The research saw this coming
Synthetic data teaching small students is an established pattern. phi-1 (arXiv 2306.11644) showed a 1.3B model trained largely on filtered "textbook quality" web data plus a billion tokens of synthetic GPT-3.5 output reaching 50.6% on HumanEval after four days on eight A100s. Distilling Step-by-Step (arXiv 2305.02301) pushed the logic further: a 770M T5, trained on LLM rationales, outperformed the few-shot 540B PaLM using only 80% of the benchmark data. And MobileLLM (arXiv 2402.14905) argued that below a billion parameters, architecture outweighs data quantity; its 125M and 350M models came close to LLaMA-2 7B correctness on API calling. Jeff's recipe, an open teacher writing data and a small student learning a narrow skill, is that lineage applied to decisions instead of code or math.
Why one pass beats a paragraph
The calibration literature explains why a trained readout over fixed options is attractive. Calibrate Before Use (arXiv 2102.09690) showed GPT-3 few-shot accuracy swinging from near chance to near state of the art on prompt format and example order alone, with a simple bias correction recovering up to 30.0 points of absolute accuracy. Just Ask for Calibration (arXiv 2305.14975) found that RLHF-tuned models carry poorly calibrated conditional probabilities, and that explicit confidence statements cut expected calibration error by roughly half. A third study across 15 chat models (arXiv 2402.13213) concluded their max-softmax probabilities on multiple-choice questions are consistently miscalibrated, though they still tend to rank wrong answers lower than right ones.
Jeff's README rediscovers the first of these findings empirically: "wording matters enormously," including one rewording that moved a Frogger playthrough from 15 crossings to 23. The 2021 paper described the same fragility in GPT-3. Meanwhile, the generative alternative has its own documented failure mode. A CPU reliability benchmark for sub-2B tool calling (arXiv 2609.07370) found only 5 of 1,000 raw small-model generations parsed directly as JSON, with Qwen2.5-1.5B needing about 30.8 seconds mean latency per call on CPU. Generating text and then parsing it is the fragile, slow path. Scoring fixed option letters in one pass removes the parser entirely, and a fitted temperature on top addresses the calibration gap the other papers documented.
Where the ceiling is
ThinkSLM (arXiv 2502.11569) evaluated 72 small models across 17 reasoning benchmarks and found that training method and data quality, not scale alone, drive reasoning skill. But it also found larger models stayed steadier under adversarial perturbation, and Jeff's own table agrees: 64.0 on BBH against Jev's 94.3 is not a rounding error, it is a category line. On the agent side, PTC-Decoder (arXiv 2609.30836) showed small models almost entirely ignore plan-enforcement prompts; constrained decoding restored plan invocation (a statistically significant +1.21 mean score), while the authors keep final-answer accuracy listed as an open problem.
The practical limits stack up too. Twenty-six options maximum. English only. And a Hacker News commenter running it against their production workload reported 70% zero-shot accuracy where Jev scored 94%. Fine-tuning closes exactly this kind of gap, as Jeff's voice-navigation result shows, but needing a fine-tune changes the product category, as another commenter put it. A hosted general decision API and a fine-tuned domain model are different purchases with different maintenance contracts.
When the critics are right
Two pushes back from the same thread deserve their due. First, the classics refuse to die: one commenter benchmarked embeddings plus a logistic classifier against Jev and Laya on AG News, Emotion, MASSIVE Intent, and Banking77, and reports the old stack matching or beating both on plain classification. If your categories are static and you have labeled data, that is still the correct first experiment, and it is the same instinct behind our own work on compressed retrieval (4-bit recall on a CPU budget). Second, another commenter noted the benchmark panel leans reasoning-heavy (BBH, JudgeBench) precisely where 0.8B models die, and suggested classification-native sets like Banking77 as fairer ground for System-1 claims. The Jeff table supports both readings at once, which is the honest way to say it.
What this means for a local stack
We run the neighboring layers ourselves, on our own hardware, with published numbers. Our Jev-compatible sidecar on plain CPU serves typed decisions end to end in 281 to 671 ms per call on CPU, batching to 7.2 ms per question on a T4, with 16 of 16 contract checks passing. Our memory engine holds 98.2% recall_any@5 on our published benchmark, CPU-only and fully offline.
A trained-at-home decision model completes a picture we have been assembling all quarter. Recall runs in a CPU process of a few hundred megabytes. Decisions run in 1.7 GB of weights at 28 ms on Apple silicon. Retrieval runs in 4-bit vectors. Each layer has a credible offline implementation now, and the cloud becomes an optimization rather than a dependency. The layers stay complementary: memory (Uteke), decisions (Laya on CPU, or Jeff-class models when you can train), action (your code). Nothing in Jeff threatens the memory layer, and nothing in the memory layer pretends to decide.
The verdict: Jeff did not invent a new technique. It assembled a known recipe, synthetic data from an open teacher, a small student, a calibrated readout, at a scale where individuals can run the whole loop, and published its numbers together with its limits. That is how a category becomes commodity. The frontier stays with the hosted giants. The floor keeps rising underneath them.
References
- Jeff project README, github.com/firelex/jeff, accessed September 30, 2026
- Hacker News discussion, item 49883844, September 28, 2026, 518 points at time of writing
- Textbooks Are All You Need (phi-1), arXiv 2306.11644
- Distilling Step-by-Step, arXiv 2305.02301
- MobileLLM: Optimizing Sub-billion Parameter Language Models, arXiv 2402.14905
- Calibrate Before Use, arXiv 2102.09690
- Just Ask for Calibration, arXiv 2305.14975
- Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness, arXiv 2402.13213
- A CPU Reliability Benchmark for MCP-Style Tool Calling in Sub-2B Models, arXiv 2609.07370
- Towards Reasoning Ability of Small Language Models (ThinkSLM), arXiv 2502.11569
- PTC-Decoder: Towards Intelligent SLMs on Offline Edge Devices, arXiv 2609.30836
- Jev-compatible decision API on plain CPU, blog.codecora.dev, published September 2026
- Uteke vs ChromaDB vs Mem0 benchmark, blog.codecora.dev, published 2026