Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep
Perplexity AI has released WANDR, an open-source evaluation benchmark designed to test the ability of research agents to gather evidence-based information across complex topics. The tool consists of 500 tasks that require agents to identify multiple entities and provide verifiable citations for each. By creating a standardized way to measure retrieval accuracy and depth, this benchmark aims to improve how developers evaluate the reliability of automated research systems.
Covered by 1 source
- MMarkTechPost↗Asif Razzaq2d ago