DFlash 2: Keep Drafting Parallel
Source Entity
Hacker News

Inco AI has announced DFlash 2, an advancement in inference technology designed to optimize token economics for autonomous AI agents. The tool aims to overcome current latency bottlenecks, building on the success and widespread industry adoption of the original DFlash architecture.
The Evolution of AI Inference: Introducing DFlash 2
In the rapidly maturing landscape of artificial intelligence, the transition from simple chatbot interactions to complex, autonomous agentic workflows has created a significant technical hurdle: inference latency. As agents are tasked with reading, planning, and executing tool calls over extended periods, the demand for computational resources has surged. Inco AI’s announcement of DFlash 2 addresses this specific bottleneck, aiming to align inference infrastructure with the unique token economics required for the next generation of agent-driven applications.
Scaling for the Agent Era
The fundamental challenge identified by Inco AI is that every token generated by an autonomous agent requires a complete forward pass over the underlying model. Unlike traditional chat interfaces, which are relatively lightweight, agentic workflows consume tokens at a rate that threatens to outpace current hardware capabilities. DFlash 2 is positioned as a foundational upgrade to the inference stack, designed to handle these intensive workloads without compromising on speed or efficiency.
Building on Proven Success
The predecessor to this technology, DFlash, released in January, has already achieved significant industry penetration. Its integration into major frameworks such as SGLang, vLLM, TensorRT-LLM, and llama.cpp demonstrates its versatility across diverse computing environments. By optimizing how tokens are processed, DFlash has become a critical component for infrastructure providers looking to maximize the utility of their hardware investments.
Benchmarking Performance Gains
The efficacy of the original DFlash architecture is well-documented through industry benchmarks. NVIDIA reported up to 15x throughput improvements on Blackwell GPUs, while Google observed a 3x increase in tokens per second on its TPU infrastructure. Furthermore, CoreWeave’s utilization of DFlash for its Kimi K2.7 Code endpoint—currently recognized as the fastest for that model on Artificial Analysis—serves as a testament to the real-world performance benefits of this technology.
Ecosystem Integration and Future Outlook
The widespread adoption of DFlash by major industry players like NVIDIA, Red Hat, and Modal highlights its status as an emerging standard in AI infrastructure. As the industry moves toward more complex agentic behaviors, the demand for efficient inference will only grow. DFlash 2 represents a strategic step forward, ensuring that the infrastructure layer can support the massive, sustained token consumption inherent in future AI agent deployments.
Conclusion
By focusing on the specific constraints of agentic inference, Inco AI is addressing one of the most critical limitations in current AI development. As DFlash 2 moves from a sneak peek to broader implementation, it is poised to further reduce the friction associated with long-running AI tasks, ultimately enabling more sophisticated and reliable autonomous agents across the technology sector.