# Perplexity releases pplx-embed-v2-context-9b-preview contextual embedding model

The model trains contextual chunk retrieval by distilling relevance scores from a query-aware context compression model instead of single gold chunks.

- The preview model is available on Hugging Face.
- Uses 1 KB per vector with 1,024 dimensions in int8 format, compared to 8 KB for voyage-context-4.
- Achieves the highest average nDCG@10 on ConTEB among tested models and leads turbopuffer's context-bench in answer and evidence retrieval.

## Why it matters

Engineers building RAG pipelines get higher chunk retrieval accuracy with lower storage overhead by capturing full-document context.

## Sources

- [Perplexity: Perplexity releases pplx-embed-v2-context-9b-preview contextual embedding model](https://x.com/perplexity_ai/status/2105373989262827915)

---

Summarized by dstilled on 2026-09-30. https://dstilled.ai/story/2e5c9e32-cd85-40a3-928f-a3257bd91605
