Multimodal open d1 decision models for the edge
Source Entity
Hugging Face - Blog

The new d1 open decision models, built on Liquid Foundation Models, offer record-breaking performance for edge devices. By prioritizing single-pass decision-making over generative tokenization, these models deliver high efficiency on hardware like NVIDIA Jetson.
The Rise of Edge-Native Decision Models
The landscape of artificial intelligence is currently shifting from massive, cloud-dependent generative models toward specialized, efficient architectures designed for edge computing. The introduction of the d1 series—specifically the d1-3B model—marks a significant milestone in this transition. By achieving a score of 48.57 on the Decision Index 0.2.1, the d1-3B has demonstrated that smaller, optimized models can outperform significantly larger counterparts, including 4B and 9B parameter models, as well as the 35B-A3B Decider model. This benchmarks a new era where performance is measured by decision accuracy rather than mere parameter count.
Architectural Innovation: Beyond Generative AI
The core technological breakthrough behind the d1 series lies in its departure from traditional Large Language Model (LLM) architectures. While generative models are designed to predict and produce sequences of tokens, the d1 models are built on Liquid Foundation Models (LFMs) specifically engineered for decision-making. By answering in a single forward pass rather than generating output token-by-token, these models drastically reduce latency and computational overhead, making them ideal for real-time applications where every millisecond counts.
Multimodality at the Edge
Multimodality is no longer a luxury reserved for massive cloud servers; the d1 series brings this capability directly to the edge. The d1-3B model supports both text and image inputs, allowing for complex situational awareness in local environments. Complementing this is the d1-omni-600M, which offers even greater versatility by supporting text-to-image or text-to-audio modalities. This flexibility enables developers to deploy sophisticated AI systems that can interpret diverse sensory data without the need for constant internet connectivity.
Latency Benchmarks on NVIDIA Hardware
The real-world utility of the d1-3B is underscored by its impressive latency metrics on NVIDIA Jetson hardware. On the high-end NVIDIA Jetson AGX Thor, the model achieves a response time of just 16 ms. Even on more constrained hardware like the Jetson AGX Orin (26 ms) and the Jetson Orin Nano (50 ms), the model remains remarkably fast. These metrics prove that high-performance, intelligent decision-making is now viable on compact, power-efficient edge devices.
Broader Implications and Future Trends
This development signals a broader trend toward 'Edge Intelligence,' where the processing of critical decisions happens locally. By moving away from generative tokenization toward single-pass decision models, companies are reducing their reliance on expensive, latency-prone cloud infrastructure. As these models continue to evolve, we can expect to see them integrated into autonomous robotics, industrial IoT, and real-time monitoring systems, where the speed of a single forward pass provides a distinct competitive advantage over traditional generative approaches.