Perplexity Research published a post-training study of the system inside Perplexity Computer. Real user sessions, including unsuccessful ones, were used for training. The method combines fine-tuning on selected successful examples with hint-guided distillation.

In a live A/B test between two trained versions, the share of tool invocation failures decreased from 2,24% to 1,77%. The source reports this as a result of comparing two checkpoints.

Practical meaning of the approach is to use real errors as material for subsequent improvement. The published data package is based on metadata and a summary of MarkTechPost, not the full text of the primary study, so details of the experiment and the robustness of the result remain unclear.