ORCA-bench: How Ready Are Language Model Agents for Oncall?
Researchers introduced ORCA-bench, a new evaluation framework designed to measure how effectively language model agents perform root cause analysis for software infrastructure incidents. Unlike standard coding tasks, this benchmark tests a model's ability to interpret noisy system logs, metrics, and traces to diagnose service failures from ambiguous reports. This effort aims to determine whether automated systems are reliable enough to assist engineers in real-world, high-pressure troubleshooting scenarios.
Covered by 1 source
- AarXiv CS.AI↗Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi1d ago