SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? Paper • 2609.09113 • Published 11 days ago • 19
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published Aug 3 • 25