FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 19 days ago • 148
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 20 days ago • 207
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published Aug 11 • 69
SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs Paper • 2509.20758 • Published Sep 25, 2025 • 2
Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning Paper • 2602.01058 • Published Feb 1 • 45