Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Source Entity
Hacker News

Jeff is a new series of small-scale, high-speed decision models derived from Qwen3.5 and Gemma 4. These models perform zero-shot classification with extreme latency efficiency, making them ideal for real-time automated decision-making tasks.
The Rise of Compact Decision Models
The introduction of 'Jeff,' a collection of 0.8B parameter decision models, marks a significant shift in how developers approach local machine learning. By fine-tuning established architectures like Qwen3.5 and Gemma 4, the creators have optimized these models for a singular purpose: zero-shot classification. Unlike massive language models designed for generative tasks, Jeff is engineered specifically to output calibrated probabilities, effectively turning a neural network into a high-speed, programmatic decision engine.
Unprecedented Latency and Local Performance
A primary breakthrough of the Jeff project is its operational speed. With performance metrics hitting approximately 22 ms on an RTX PRO 6000 and 28 ms on an Apple M4 Max via MLX, these models are designed for integration into latency-sensitive applications. By bypassing the need for text generation and subsequent parsing, the models deliver direct classification results, which is a critical improvement for systems requiring immediate feedback loops.
Zero-Shot Versatility
One of the most compelling aspects of the Jeff framework is its zero-shot capability. Because the models interpret options based on plain-language descriptions rather than being restricted to a fixed, pre-trained label set, they are highly flexible. This allows for dynamic implementation across diverse domains—from sorting support tickets and identifying user intents to moderating content or determining game moves—without the need for extensive retraining when new categories are introduced.
Technical Architecture and Integration
Jeff operates by mapping a described situation and a set of user-provided options to a probability distribution through a single forward pass. This architecture removes the common bottlenecks associated with LLM-based classification, such as token generation latency and the overhead of JSON parsing. By slotting directly into existing codebases, Jeff functions more like a traditional software utility than a complex AI service, lowering the barrier for developers to implement 'judgment' in their applications.
Broader Implications for AI Deployment
The development of 0.8B parameter models trained at home underscores a growing trend toward 'small AI.' As hardware constraints remain a hurdle for real-time deployment, specialized, small-scale models offer a viable pathway for running sophisticated logic on edge devices. This shift suggests a future where high-performance decision-making is no longer tethered to massive cloud infrastructure, but is instead embedded directly into the software stack.
Future Trends in Decision Logic
Looking ahead, the success of the Jeff models indicates that the industry may move toward hyper-specialized, small-parameter models for specific tasks. While generative models continue to dominate the headlines, the practical utility of these compact, fast, and well-calibrated decision models could prove more transformative for enterprise software, automation, and real-time user interface interactions.