
GPT-6 Astra received conflicting scores in testing. Epoch AI ranked the model ahead with a score of 169 points, whereas Artificial Analysis rated it no better than its predecessor and below Claude Fable 5.1.
The most attention focused on ARC-AGI-3: according to the provided description, Astra for the first time operated more efficiently than the average human. ARC Prize head François Chollet does not consider this proof of AGI, but stated that progress is occurring roughly twice as fast as he expected and has revised his forecast regarding AGI timelines.
Interpreting the result requires accounting for the limited available data: the package contains only The Decoder's synopsis, not full testing methodologies or primary statements from participants. Therefore, directly comparing figures from different evaluations is not yet possible.
editorial commentary
Why it matters
The likely consequence is increased focus on model efficiency in reasoning tasks, not just their final scores. The next observable signals will be full methodologies and independent verifications of the ARC-AGI-3 results. Substantial uncertainty remains due to the absence of primary data and diverging assessments from different analysts.