R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets Paper • 2608.16033 • Published about 1 month ago • 17
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 23 days ago • 207
PatchWorld: Gradient-Free Optimization of Executable World Models Paper • 2605.30880 • Published May 29 • 12
RubricBench: Aligning Model-Generated Rubrics with Human Standards Paper • 2603.01562 • Published Mar 2 • 64