Perplexity releases pplx-embed-v2-context-9b-preview contextual embedding model
The model trains contextual chunk retrieval by distilling relevance scores from a query-aware context compression model instead of single gold chunks.
- The preview model is available on Hugging Face.
- Uses 1 KB per vector with 1,024 dimensions in int8 format, compared to 8 KB for voyage-context-4.
- Achieves the highest average nDCG@10 on ConTEB among tested models and leads turbopuffer's context-bench in answer and evidence retrieval.
Engineers building RAG pipelines get higher chunk retrieval accuracy with lower storage overhead by capturing full-document context.

Sources
Read this as text
Back to the AI news