Don't Train the Model, Evolve the Harness
🌿
66
Evolving an agent's harness, not its model, on Harvey's LAB
DABstep Reasoning Benchmark Leaderboard
Explore and compare model scores on RewardBench benchmarks
Track, rank and evaluate open LLMs and chatbots
Find upcoming AI conference deadlines instantly
The ultimate guide to training LLM on large GPU Clusters