Story thread · 2 reports / 2 sources
Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation
marktechpost.com · 1h
Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% […] The post Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation appeared first on MarkTechPost .
First report: Perplexity.AI reduces tool-call failures by 21% with new model training approach — cryptobriefing.com, 2d
The conversation · 0
Sign in to join the conversation.
No comments yet — start the thread.