DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
Researchers have introduced DeepAmbigQA, a new benchmarking dataset designed to evaluate how effectively large language models handle complex, multi-hop questions with ambiguous components. By requiring systems to identify multiple distinct answers for nuanced queries, this framework addresses a common limitation where models fail to provide exhaustive responses during information retrieval. This development aims to improve the reliability of retrieval-augmented generation systems when tasked with parsing intricate, open-domain requests that demand both synthesis of external data and logical reasoning.
Covered by 2 sources · 5 articles
- AApple Machine Learning Blog↗Aug 6
- AarXiv CS.AI↗Maodong Li, Xinyue Kang, Yuanchen Shi, Fang KongAug 4
- AarXiv CS.AI↗Jiaoyang Li, Junhao Ruan, Shengwei Tang, Kaiyan Chang, Zhengtao Yu, Tong Xiao, Jingbo ZhuAug 6
- AarXiv CS.AI↗Can Wang, Haoran Chen, Haowen Gao, Hao Ding, Zhaoyang Liu, Zhiying TuAug 4
- AarXiv CS.AI↗Weidong Bao, Yingying Sun, Jun Yang, Yilin Wang, Zili Wei, Yubin Bao, Fangling Leng, Minghe Yu, Tiancheng Zhang, Ge YuAug 4