Running Agents 11 Official Benchmarks Leaderboard 2026 🏆 11 Explore and compare AI model scores across official benchmarks
SWE-rebench-V2 Collection SWE-rebench-V2 is a curated dataset of software-engineering tasks derived from real GitHub issues and pull requests. • 3 items • Updated Mar 3 • 12
mistralai/Voxtral-Mini-4B-Realtime-2602 Automatic Speech Recognition • 4B • Updated Mar 11 • 1.33M • 859
Ministral 3 Collection A collection of edge models, with Base, Instruct and Reasoning variants, in 3 different sizes: 3B, 8B and 14B. All with vision capabilities. • 9 items • Updated Dec 2, 2025 • 167
Running Featured 598 Image Arena Leaderboard 📊 598 Image Generation and Image Editing Arena & Leaderboard
Running Featured 459 LLM Performance Leaderboard 🐨 459 View the latest LLM performance leaderboard online
Running Agents Featured 135 Open VLM Video Leaderboard 🌎 135 VLMEvalKit Eval Results in video understanding benchmark