Researchers developed a four-phase methodology to test whether language models use signals of their own confidence when choosing between answering and abstaining. In the first phase, model confidence was measured without the option to refuse an answer.

The study was published in Nature Machine Intelligence under the title "Causal evidence that language models use confidence to drive behaviour." In the work's description, metacognition is considered as an evaluation of the quality of one's own cognitive actions, linked to adaptive behavior across different species.

The available source contains only a publisher's summary and does not report the results of the remaining phases, the names of the models studied, or quantitative metrics. Therefore, while the setup of the causal test is confirmed, its outcome is not.