Story thread · 2 reports / 2 sources

Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation

marktechpost.com · 1h

Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% […] The post Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation appeared first on MarkTechPost .

First report: Perplexity.AI reduces tool-call failures by 21% with new model training approach — cryptobriefing.com, 2d

The conversation · 0

Sign in to join the conversation.

No comments yet — start the thread.