AI & ML interests

Open RL Environments at Scale

Recent Activity

AdithyaSKย  updated a Space about 5 hours ago
FineEnvs/geoguesser-article
AdithyaSKย  updated a Space about 6 hours ago
FineEnvs/README
AdithyaSKย  updated a collection about 6 hours ago
Multilingual Multimodal Envs
View all activity

Organization Card

FineEnv_banner

๐Ÿ’ป Code ๐Ÿ“– Guide ๐ŸŽฅ Slides

๐Ÿค— FineEnvs: Open RL Environments

FineEnvs is a home for end-to-end RL environment recipes, built to make it easier to explore, reproduce, train, and evaluate agent systems.

Explore complete and reproducible environment projects from us and the community, including:

  • ๐ŸŒ Open RL environments
  • ๐Ÿงฉ End-to-end environment recipes
  • ๐Ÿ’ป Complete implementations
  • ๐Ÿ“ฆ Models, datasets, and artifacts
  • ๐Ÿงช Training and evaluation setups
  • ๐Ÿš€ Demos and Spaces
  • ๐Ÿ“š Tutorials and guides

All the reproducible code โ€” environments, rollouts, training configs, notebooks, article and slide sources โ€” lives in one repo: github.com/adithya-s-k/FineEnvs. The artifacts those produce live here on the Hub.

FineEnvs Projects

A growing collection of open projects, environments, resources, and artifacts.

Project What it is Explore
FineEnvs Academy Articles, guides, tutorials, slides, and hands-on resources for learning how to build RL environments and agent systems. Explore โ†’
Data Agent Training SLMs for data science with multi-harness RL environments. Explore โ†’
MiMo-V2.6-RL in Harbor All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, plus an explorer to browse them and run graded rollouts. Explore โ†’ ยท Explorer โ†’
Repo2RLEnv Verifiable coding and terminal RL environments in Harbor format, with per-task quality labels and provenance. Explore โ†’

Each numbered project below is a self-contained recipe: an environment, a training run, and every artifact it produced.

# Project What it is Explore
00 RL Environments 101 Three environments implemented six times over, one per framework. Same logic, six dialects. Source โ†’ ยท Collection โ†’
01 LaTeX OCR Qwen3-VL-2B trained to read rendered math into LaTeX, scored by a reward served from a live Space. Collection โ†’
02 Watercolour Qwen3.5-35B-A3B trained to paint watercolours by writing p5.brush sketches, rewarded by taste rather than correctness. Collection โ†’
03 GeoGuesser A multi-turn visual geolocation environment, and the 4B trained on it until it outscored gpt-5.4-mini and claude-haiku-4.5. Collection โ†’
04 SmolDataEnvs 5.5K+ data-analysis tasks for hill-climbing small models, graded deterministically with no LLM judge: plain prompts, verified SFT traces, and Harbor task suites. Collection โ†’
05 SmolDataEnvs: Multi-harness RL One small model trained with GRPO inside four unmodified coding agents (OpenCode, Claude Code, Codex, Mini-SWE-Agent), with LFM2.5-2.6B and Qwen3.5-2B checkpoints. Article โ†’ ยท Collection โ†’
06 Multilingual OCR A million document pages in 22 languages behind one OpenEnv server, plus Sarvam Indic OCR Bench, and Gemma 4 trained to read Kannada. Collection โ†’
07 Multilingual ASR All 102 FLEURS languages behind one OpenEnv server, and Gemma 4 trained to hear Kannada: held-out character error โˆ’45%. Collection โ†’

Articles & Talks

What it covers Read / Watch
๐Ÿ“– The Ultimate Guide to RL Environments Building and scaling RL environments in the LLM era โ€” how frameworks are built, how rewards are wired, how they scale to thousands of concurrent sessions. Read โ†’
๐ŸŽž๏ธ RL Environments 101 From "what is an env?" to training your own: RL fundamentals โ†’ environment anatomy โ†’ OpenEnv โ†’ training with TRL. Watch โ†’
๐Ÿ“ˆ Scaling RL for LLMs RL environments and RL training โ€” what an environment is, how reward hacking happens, how to train against your own. AMD AI Dev Day. Watch โ†’
๐Ÿ”€ Multi-Harness Training OpenEnv ร— Harbor โ€” why an environment's failure model decides whether it can be trained against. Watch โ†’
๐Ÿงญ The Ultimate Guide to Multi-Harness RL Training small models on SmolDataEnvs with the same tasks and reward but a different tool loop each time (TRL, native OpenCode, Harbor), and what changes. Read โ†’
๐ŸŽค Training a Coding Agent Through a Harness You Did Not Write Multi-harness RL talk by Sergio Paniego Blanco: one model, four unmodified coding agents, GRPO. Watch โ†’
๐ŸŒ How to turn a game into an RL environment The technical intuition, end to end: curating the data, designing the environment, shipping it with OpenEnv, and training a 4B against it with TRL. Read โ†’

Environments

Three reference environments, each implemented across six frameworks โ€” openenv, ors, nemo_gym, verifiers, skyrl_gym, gem. Same logic, six dialects. Source โ†’ ยท RL Envs 101 collection โ†’

Environment Tools OpenEnv ORS NeMo Gym
Jupyter agent โ€” real code execution in an E2B sandbox 4 Space Space Space
Wordle โ€” multi-turn, pure Python, no backend 1 Space Space Space
Desktop โ€” computer-use, vision-driven Linux desktop 19 Space Space โ€”

Build your own

Five agent skills turn a plain-English description into a runnable RL environment across four frameworks โ€” works with Claude Code, Cursor, Codex, OpenCode, Gemini CLI and others.

npx skills add adithya-s-k/FineEnvs

We're looking for new end-to-end recipes โ€” a task, an environment, a training run, and honest results. Contributing guide โ†’

Citation

@misc{fineenvs,
  author = {Kolavi, Adithya S},
  title  = {FineEnvs: Open Source RL Environments for LLM Agents},
  year   = {2026},
  url    = {https://github.com/adithya-s-k/FineEnvs}
}