Hand a frontier model tools instead of prompts and the results get strange fast. Fresh third-party runs put OpenAI’s GPT-6 Astra through drone control, shop keeping and desk robotics.
StationeryBench talks about spatial sense. Two models took turns steering the same dual-arm YAM robots through five chores, from uncapping a marker to passing a ruler between grips. Over 200 trials, Astra closed out 7 of 100 tasks on its own. Ai2’s MolmoAct2 closed out none. Median progress separated them at 46 against 12, on a scale that runs to 100.
Yoav Artzi, a Cornell researcher who also works at Google DeepMind, called that a step change in spatial reasoning, then cautioned that the model still trails human reliability.
The drone trial held five separate jobs. Astra beat the human-AI reference in every one of them, a first for a model. One job asked the system to write code letting a drone track a single chosen person. Reliability did not follow, with success rates described as uneven.
Andon Labs ran the commercial side. Its Vending-Bench hands a model a simulated machine to stock, price and run. Astra banked $15,515 on average, which the lab describes as close to three times what Anthropic’s Claude Fable 5.1 collected. The OpenAI system also refused illegal price-fixing offers that Fable took.