DeepSeek V4 Flash 0731
Source Entity
Hacker News

DeepSeek has released the V4 Flash 0731 model, achieving significant performance metrics on the ARC-AGI benchmarks. The model offers tiered reasoning variants with highly competitive pricing structures for AI task execution.
Breakthrough Performance in ARC-AGI Benchmarking
DeepSeek’s release of the V4 Flash 0731 model on July 31, 2026, marks a significant milestone in the evolution of large language models focused on abstract reasoning. By achieving an 89.0% score on the ARC-AGI-1 Semi-Private benchmark, the model demonstrates a sophisticated capability to handle complex, unseen problem-solving tasks that move beyond simple pattern matching. This performance is particularly notable given the model’s intent to bridge the gap between high-level reasoning and cost-effective execution in automated environments.
The Economics of Reasoning
The introduction of a tiered pricing structure—specifically $0.02 per task for ARC-AGI-1 and $0.04 per task for ARC-AGI-2—underscores a strategic shift in the AI industry toward commoditized reasoning power. By providing these specific price points, DeepSeek is positioning itself as a leader in high-efficiency AI deployment. This economic model allows developers to scale complex reasoning tasks without the prohibitive costs traditionally associated with top-tier foundation models, potentially accelerating the adoption of AI in industrial and research settings.
Analyzing Reasoning Variants
The V4 Flash 0731 release includes three distinct reasoning variants—Max, High, and Low—which offer a spectrum of performance tailored to different computational requirements. The 'Max' variant, which captures the peak 89.0% score on ARC-AGI-1 and 61.4% on ARC-AGI-2, represents the current state-of-the-art for this specific architecture. Meanwhile, the 'High' and 'Low' variants provide lower benchmarks (87.0%/56.0% and 84.0%/46.0%, respectively), allowing users to optimize for latency or power consumption depending on the specific application needs.
Implications for General Intelligence
The ARC-AGI (Abstraction and Reasoning Corpus) benchmarks are widely considered the gold standard for evaluating progress toward Artificial General Intelligence (AGI). Unlike standard benchmarks that test knowledge retrieval, the ARC-AGI suite focuses on a model's ability to learn new skills from limited examples. DeepSeek’s ability to maintain high scores across these variants suggests that the V4 Flash architecture possesses a robust underlying engine capable of generalizing across diverse, non-repeating logic puzzles.
Future Trends and Industry Impact
Looking forward, the success of the V4 Flash 0731 series suggests that developers will increasingly prioritize 'reasoning density'—the amount of logic performed per dollar spent. As more models undergo rigorous verification on the ARC-AGI-2 leaderboard, we can expect a competitive race to push these percentages higher while simultaneously driving costs down. DeepSeek’s transparent reporting of pass/fail rates per reasoning level provides a roadmap for the future of model evaluation, where performance is not just a single number, but a customizable capability set.