GPT-6 Astra on robot arms
Source Entity
Hacker News

OpenAI's GPT-6 Astra was tested on robotic manipulation tasks, showing superior performance in block placement but struggling equally with complex puzzle insertions. These results highlight both the rapid advancement and the persistent limitations in AI-driven physical automation.
The Evolution of Embodied AI: Testing GPT-6 Astra
Recent benchmarking of OpenAI's GPT-6 Astra within the context of robotic manipulation marks a significant milestone in the evolution of embodied artificial intelligence. By utilizing YAM arms governed by the Inspect Robots agent policy, researchers have moved beyond theoretical language processing to evaluate how advanced models manage physical, real-world tasks. This shift is critical as the industry transitions from pure digital intelligence to systems capable of interacting with the physical environment.
Performance Metrics: Success and Speed
The comparative study between GPT-6 Astra and the Claude Fable 5 and 5.1 models reveals a clear trajectory in operational efficiency. In the relatively straightforward task of picking up a red block and placing it into a bowl, Astra achieved a success rate of 19 out of 20 trials. This stands in stark contrast to the Fable 5.1 model, which succeeded in only 8 trials, and the original Fable 5, which managed only a single success. Beyond success rates, Astra demonstrated superior temporal efficiency, completing the task in 2.5 minutes per trial compared to 6.8 minutes for Fable 5.1.
Economic Implications of AI Robotics
Beyond technical performance, the economic viability of these agents is a primary concern for industrial implementation. The study estimates the cost per run for Astra at $0.94, significantly lower than the $2.12 required for Fable 5.1. This reduction in operational expenditure, coupled with increased task speed, suggests that GPT-6 Astra is better positioned for scalable automation. If these cost-efficiency trends continue, the barrier to entry for integrating high-level AI into manufacturing and logistics hardware will drop precipitously.
The Persistence of Complex Manipulation Barriers
Despite the gains, the test results highlight a persistent plateau in AI capability regarding fine motor coordination. In the puzzle task—which requires picking up a round blue piece by a specific knob and inserting it into a matching groove—both GPT-6 Astra and Fable 5.1 struggled, with both models completing the insertion only 2 out of 20 times. The tendency for the models to stall at the final insertion step suggests that while vision-language models have improved at high-level planning, they still face significant challenges in tactile feedback and precision control.
Future Trajectories for Agent Policy
The failure of both models to overcome the final hurdle in the puzzle task indicates that current agent policies like 'Inspect Robots' may require deeper integration with low-level control systems. Future development will likely focus on closing the gap between 'intent'—the model understanding the goal—and 'execution'—the model managing the micro-adjustments needed for friction-heavy or precise mechanical tasks. As models like Astra evolve, the focus will shift from simple pick-and-place operations to more nuanced, high-dexterity manipulation.
Conclusion
The data provides a nuanced look at the current state of robotics. While GPT-6 Astra has set a new benchmark for speed and cost in basic manipulation, the shared inability to complete complex puzzles underscores that we are still in the early stages of true robotic autonomy. The next frontier in this field will be solving these remaining precision bottlenecks to move beyond basic tasks into more sophisticated, real-world applications.