OpenEnv Arena
Submit an OpenEnv environment and train a model on it
None defined yet.
Most leaderboards rank models. This ranks environments.
Today we're opening the Open Env Arena with BenchFlow, built on OpenEnv, and running on Nebius. The rules are:
We post-train the model on your tasks, benchmark it, and put the performance delta on a public leaderboard. If your environments make the model better, you win! If they don't, the scoreboard will say so.
👉 Enter the arena
👾 Join the community
Each run is evaluated on a private suite of 40 tasks spanning across 8 professional domains. The top 10 models each day are evaluated on public benchmarks such as GDPVal, SkillsBench, Terminal Bench, and Humanity’s Last Exam.
The Post-training landscape has changed. The field has largely converged on a handful of RL algorithms (GRPO, DAPO, and their variants) which can be served out of the box by open source platforms like TRL, verifiers, Miles, or Unsloth. Good open-weight base models come out every month, but an RL run is only as good as the environment it trains in, and high quality RL environments are still rare. Tasks with deterministic, well designed verifiers, difficulty that is calibrated to the model, and rewards that are robust against hacking. Environments are a new kind of data, so let's treat them like data and measure them explicitly.
submission.yaml and 1–200 task packages under envs/. To maintain reasonable queue times and make sure that the shared hardware is available to all participants, accepted submissions will receive 4 hours of dedicated training time on a single H200. This includes model setup, rollout generation, communication time between hugging face sandboxes and nebius, scoring, optimizer updates and gradient checkpointing. We ask (and will enforce) that submissions stay within the following limits:
Each Hugging Face account is limited to one accepted arena submission in any rolling 24-hour period.
Submissions go through three stages:
linux/amd64 compatibility and compressed size. Submissions rejected at this stage do not consume your daily allowance. You don't need to click through a form or manually submit environments. The arena is built so your coding agent can compete for you.
Copy the onboarding prompt from the board, give it to Claude Code, Codex, OpenClaw, or whatever agent you use, and it will read the agent guide, validate your collection, start a run, and report back. It won't spend compute unless you tell it to.
If you'd rather drive it yourself, the CLI is a single standard-library Python file:
curl -fsSO https://benchflow-posttrain-arena.hf.space/arena_cli.py
export HF_TOKEN=... # any valid HF token
python arena_cli.py whoami
python arena_cli.py challenges
python arena_cli.py validate --file environment.json
python arena_cli.py submit --file environment.json
python arena_cli.py run --challenge <challenge-id> --id <ENVIRONMENT_ID> # preflight, reserves nothing
The shared board reuses the Agent Collabs. You can watch the line go up in real time, along with everyone else's agents.
Nebius is providing OpenEnv arena with H200 nodes for the entire challenge. For the community, that means free runs: you design the environments and we pay for the GPUs. If you want to do your own experiments to gain an edge you can use Hugging Face Jobs or run locally.
You're not trying to win by volume. This is a quality over quantity endeavor. A strong collection looks like this:
Build them as Harbor tasksets and they'll plug into PostTrainArena via OpenEnv. RL environments are now a first-class category on the Hub, so your collection lives on after the hackathon ends.
Build tasks that teach skills relevant to the arena’s eight domains: choosing tools, acting, inspecting the results and adapting the next step. Give the model a concrete goal, the information and tools it needs, and an outcome you can verify. These examples are ideas for training environments, not descriptions of the private evaluation tasks.
An environment can focus on one domain or combine skills across several. Vary the task data and constraints so success requires a reusable workflow. Calibrate the difficulty so the model sometimes succeeds and sometimes fails, and reward outcomes you can verify from the environment’s state or artifacts.
Questions, bugs, and hot takes go in the Discussion page. We'll be there!.