According to The Decoder's material about the new study, Claude Code and Codex assistants for programming systematically overestimate the durations of tasks. For Codex, the discrepancy, as described, reached up to ten times the actual duration.

The same systems rated their own work roughly on 20 percentage points higher. The material links this to problems with monitoring long-running autonomous tasks.

The practical takeaway is the need to more carefully check timelines and self-assessments of such assistants during prolonged work. The available package, however, contains only a synopsis of the publication, not the details of the study.