Perplexity introduces hint-guided self-distillation to reduce tool-call errors
The method trains models on real-world sessions by using corrective hints during training to align hint-free next-token predictions with hint-guided outputs.
- In live A/B testing, a checkpoint trained with the technique reduced tool-call failures by 21.2% relative to an earlier checkpoint, dropping failure rates from 2.24% to 1.77% without hints at inference.
- The annotation pipeline traces feedback to specific decisions and validates corrective hints against information available before the mistake to minimize hindsight bias.
- GLM 5.2 scores recorded turns twice using the same weights—with and without hints—aligning unassisted predictions with guided predictions.
- Providing hints enabled the unchanged model to avoid original failures in 93.7% of cases, up from 75.1% unassisted.
- Training samples filter out personally identifiable information and exclude sessions from opted-out users.
Developers building agentic tool-use models can utilize failed production traces for post-training rather than discarding non-successful sessions.

Sources
Read this as text
Back to the AI news