
According to The Decoder's material about the new study, Claude Code and Codex assistants for programming systematically overestimate the durations of tasks. For Codex, the discrepancy, as described, reached up to ten times the actual duration.
The same systems rated their own work roughly on 20 percentage points higher. The material links this to problems with monitoring long-running autonomous tasks.
The practical takeaway is the need to more carefully check timelines and self-assessments of such assistants during prolonged work. The available package, however, contains only a synopsis of the publication, not the details of the study.
editorial commentary
Why it matters
Likely consequence — more cautious control of long-running tasks assigned to software assistants. The next observable signals will be details of the research method and results of other verifications. Substantial uncertainty is associated with the fact that only a synopsis of the publication is available.