Perplexity releases Q2D-Web retrieval benchmark for agentic RAG
The dataset evaluates embedding models on web search using 190M documents and 69,721 agent-reformulated queries across ten languages.
- Draws on 23,000 PII-free production searches collected over nine months across fields including law, health, finance, and programming.
- Features three relevance sets: agent citations, production web rankings, and a combined set expanded with LLM judgments averaging 99.6 positive matches per query.
- RRF subsampling uses 31.7% of the corpus to preserve model rankings while reducing pplx-embed-v1-4b evaluation from 4,608 to roughly 1,500 H200 GPU-hours.
- Across 13 baseline models evaluated on Recall@1000, pplx-embed-v1-4b leads the Combined set at 69.11, while Nemotron-3-Embed-8B leads Citations at 61.68.
- Perplexity opened an evaluation request form for researchers to submit public Hugging Face retrieval models to the leaderboard.
Engineers and researchers training search models get a large-scale, production-grounded benchmark specifically tailored to multi-query agentic retrieval.

Sources
Read this as text
Back to the AI news