LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis
Researchers have introduced LongDS-Bench to evaluate how AI agents perform during complex, multi-step data analysis tasks that require tracking evolving context over time. By shifting focus from isolated interactions to long-horizon workflows, this benchmark aims to address the limitations of existing testing methods that currently fail to measure how well agents manage iterative analytical processes.
Covered by 1 source
- AarXiv CS.AI↗Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding, Haoming Xu, Lei Liang, Ningyu ZhangAug 31