Submitted by Gabriel Tomitsuka 33 Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows TextQL 8 2