
Perplexity Research published a post-training study of the system inside Perplexity Computer. Real user sessions, including unsuccessful ones, were used for training. The method combines fine-tuning on selected successful examples with hint-guided distillation.
In a live A/B test between two trained versions, the share of tool invocation failures decreased from 2,24% to 1,77%. The source reports this as a result of comparing two checkpoints.
Practical meaning of the approach is to use real errors as material for subsequent improvement. The published data package is based on metadata and a summary of MarkTechPost, not the full text of the primary study, so details of the experiment and the robustness of the result remain unclear.
editorial commentary
Why it matters
Likely consequence — further development of training software systems on real errors and subsequent verification of such methods in deployment. The next observable signal will be the publication of a full study with sample size, test methodology, and results on new scenarios. A substantial uncertainty is that only metadata from one source's summary is currently available.