# Perplexity releases Q2D-Web retrieval benchmark for agentic RAG

The dataset evaluates embedding models on web search using 190M documents and 69,721 agent-reformulated queries across ten languages.

- Draws on 23,000 PII-free production searches collected over nine months across fields including law, health, finance, and programming.
- Features three relevance sets: agent citations, production web rankings, and a combined set expanded with LLM judgments averaging 99.6 positive matches per query.
- RRF subsampling uses 31.7% of the corpus to preserve model rankings while reducing pplx-embed-v1-4b evaluation from 4,608 to roughly 1,500 H200 GPU-hours.
- Across 13 baseline models evaluated on Recall@1000, pplx-embed-v1-4b leads the Combined set at 69.11, while Nemotron-3-Embed-8B leads Citations at 61.68.
- Perplexity opened an evaluation request form for researchers to submit public Hugging Face retrieval models to the leaderboard.

## Why it matters

Engineers and researchers training search models get a large-scale, production-grounded benchmark specifically tailored to multi-query agentic retrieval.

## Sources

- [Perplexity: Perplexity releases Q2D-Web retrieval benchmark for agentic RAG](https://x.com/perplexity_ai/status/2097782467210166601)

---

Summarized by dstilled on 2026-09-09. https://dstilled.ai/story/88484e66-68cd-4117-a5ef-6b07c71f56ce
