
GPT-6 Astra completed 7 of 100 tasks in the StationeryBench benchmark with two-handed robots. Competing model MolmoAct2 did not complete any tasks, according to The Decoder.
The publication conveys the researcher's assessment, who described the result as a “leap” in spatial reasoning. However, the available data are presented only as a brief description of the source, with no details about the methodology and test conditions.
The practical significance of the result is currently limited to one early comparison. To assess the robustness of the advantage, information on other task sets, repeatability of the experiment, and model behavior in real robotic scenarios are needed.
editorial commentary
Why it matters
Possible consequence is increased interest in checking the spatial abilities of models for robot control. The next observable signal will be full methodological data, repeated tests, and results on other tasks. Substantial uncertainty is due to the available description being limited to metadata from a single source.