Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? Paper • 2608.04828 • Published 24 days ago • 1
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration Paper • 2605.03042 • Published May 4 • 148
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 4 days ago • 139
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Paper • 2406.01574 • Published Jun 3, 2024 • 57
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments Paper • 2509.14233 • Published Sep 17, 2025 • 24
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 32
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published 26 days ago • 182