
Former pre-training researcher at OpenAI and Anthropic, Jacob Coxon, left the companies and accused them of consciously exposing humanity to the risk of extinction. The Decoder reports this.
Coxon's colleague at Anthropic, Evan Hubinger, estimates the probability that an unaligned superintelligent AI will destroy humanity within the next decade to be more than 10%.
The significance of this story lies not in a confirmed forecast, but in the publicly attributed assessment of extreme risk by a researcher and criticism of two prominent companies. Available confirmation is limited to a brief description by The Decoder, rather than the full text of the primary source.
editorial commentary
Why it matters
The likely consequence is an intensification of debates regarding AI developer safety and responsibility, but this is not a confirmed forecast. The next observable signals will be responses from OpenAI and Anthropic, as well as the publication of methodologies or primary materials underlying the statements. Significant uncertainty remains due to the absence of the full text and independent confirmation in the source package.