AI & ML interests
Open RL Environments at Scale
Recent Activity
๐ค FineEnvs: Open RL Environments
FineEnvs is a home for end-to-end RL environment recipes, built to make it easier to explore, reproduce, train, and evaluate agent systems.
Explore complete and reproducible environment projects from us and the community, including:
- ๐ Open RL environments
- ๐งฉ End-to-end environment recipes
- ๐ป Complete implementations
- ๐ฆ Models, datasets, and artifacts
- ๐งช Training and evaluation setups
- ๐ Demos and Spaces
- ๐ Tutorials and guides
All the reproducible code โ environments, rollouts, training configs, notebooks, article and slide sources โ lives in one repo: github.com/adithya-s-k/FineEnvs. The artifacts those produce live here on the Hub.
FineEnvs Projects
A growing collection of open projects, environments, resources, and artifacts.
| Project | What it is | Explore |
|---|---|---|
| FineEnvs Academy | Articles, guides, tutorials, slides, and hands-on resources for learning how to build RL environments and agent systems. | Explore โ |
| Data Agent | Training SLMs for data science with multi-harness RL environments. | Explore โ |
| MiMo-V2.6-RL in Harbor | All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, plus an explorer to browse them and run graded rollouts. | Explore โ ยท Explorer โ |
| Repo2RLEnv | Verifiable coding and terminal RL environments in Harbor format, with per-task quality labels and provenance. | Explore โ |
Each numbered project below is a self-contained recipe: an environment, a training run, and every artifact it produced.
| # | Project | What it is | Explore |
|---|---|---|---|
| 00 | RL Environments 101 | Three environments implemented six times over, one per framework. Same logic, six dialects. | Source โ ยท Collection โ |
| 01 | LaTeX OCR | Qwen3-VL-2B trained to read rendered math into LaTeX, scored by a reward served from a live Space. | Collection โ |
| 02 | Watercolour | Qwen3.5-35B-A3B trained to paint watercolours by writing p5.brush sketches, rewarded by taste rather than correctness. | Collection โ |
| 03 | GeoGuesser | A multi-turn visual geolocation environment, and the 4B trained on it until it outscored gpt-5.4-mini and claude-haiku-4.5. |
Collection โ |
| 04 | SmolDataEnvs | 5.5K+ data-analysis tasks for hill-climbing small models, graded deterministically with no LLM judge: plain prompts, verified SFT traces, and Harbor task suites. | Collection โ |
| 05 | SmolDataEnvs: Multi-harness RL | One small model trained with GRPO inside four unmodified coding agents (OpenCode, Claude Code, Codex, Mini-SWE-Agent), with LFM2.5-2.6B and Qwen3.5-2B checkpoints. | Article โ ยท Collection โ |
| 06 | Multilingual OCR | A million document pages in 22 languages behind one OpenEnv server, plus Sarvam Indic OCR Bench, and Gemma 4 trained to read Kannada. | Collection โ |
| 07 | Multilingual ASR | All 102 FLEURS languages behind one OpenEnv server, and Gemma 4 trained to hear Kannada: held-out character error โ45%. | Collection โ |
Articles & Talks
| What it covers | Read / Watch | |
|---|---|---|
| ๐ The Ultimate Guide to RL Environments | Building and scaling RL environments in the LLM era โ how frameworks are built, how rewards are wired, how they scale to thousands of concurrent sessions. | Read โ |
| ๐๏ธ RL Environments 101 | From "what is an env?" to training your own: RL fundamentals โ environment anatomy โ OpenEnv โ training with TRL. | Watch โ |
| ๐ Scaling RL for LLMs | RL environments and RL training โ what an environment is, how reward hacking happens, how to train against your own. AMD AI Dev Day. | Watch โ |
| ๐ Multi-Harness Training | OpenEnv ร Harbor โ why an environment's failure model decides whether it can be trained against. | Watch โ |
| ๐งญ The Ultimate Guide to Multi-Harness RL | Training small models on SmolDataEnvs with the same tasks and reward but a different tool loop each time (TRL, native OpenCode, Harbor), and what changes. | Read โ |
| ๐ค Training a Coding Agent Through a Harness You Did Not Write | Multi-harness RL talk by Sergio Paniego Blanco: one model, four unmodified coding agents, GRPO. | Watch โ |
| ๐ How to turn a game into an RL environment | The technical intuition, end to end: curating the data, designing the environment, shipping it with OpenEnv, and training a 4B against it with TRL. | Read โ |
Environments
Three reference environments, each implemented across six frameworks โ openenv, ors, nemo_gym, verifiers, skyrl_gym, gem. Same logic, six dialects. Source โ ยท RL Envs 101 collection โ
| Environment | Tools | OpenEnv | ORS | NeMo Gym |
|---|---|---|---|---|
| Jupyter agent โ real code execution in an E2B sandbox | 4 | Space | Space | Space |
| Wordle โ multi-turn, pure Python, no backend | 1 | Space | Space | Space |
| Desktop โ computer-use, vision-driven Linux desktop | 19 | Space | Space | โ |
Build your own
Five agent skills turn a plain-English description into a runnable RL environment across four frameworks โ works with Claude Code, Cursor, Codex, OpenCode, Gemini CLI and others.
npx skills add adithya-s-k/FineEnvs
We're looking for new end-to-end recipes โ a task, an environment, a training run, and honest results. Contributing guide โ
Citation
@misc{fineenvs,
author = {Kolavi, Adithya S},
title = {FineEnvs: Open Source RL Environments for LLM Agents},
year = {2026},
url = {https://github.com/adithya-s-k/FineEnvs}
}
-
Nayana Multilingual OCR
๐OpenEnv document OCR, layout and VQA over 22 languages
-
FLEURS Multilingual ASR
๐OpenEnv speech recognition over 102 FLEURS languages
-
FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo
Image-Text-to-Text โข Updated -
FineEnvs/gemma-4-E4B-it-kannada-asr-grpo
Automatic Speech Recognition โข Updated
-
The ultimate guide to multi-harness RL
๐78Train open models with RL inside real agent harnesses
-
SmolDataEnvs Multi-harness | SETA Whitebox
๐งชAnswer dataset questions and view scoring results
-
SmolDataEnvs Multi-harness | Native OpenCode
๐งชRun OpenCode tasks and view grading results
-
SmolDataEnv RL
๐Visualize RL agent comparison metrics in a web dashboard
-
Nayana Multilingual OCR
๐OpenEnv document OCR, layout and VQA over 22 languages
-
FLEURS Multilingual ASR
๐OpenEnv speech recognition over 102 FLEURS languages
-
FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo
Image-Text-to-Text โข Updated -
FineEnvs/gemma-4-E4B-it-kannada-asr-grpo
Automatic Speech Recognition โข Updated
-
The ultimate guide to multi-harness RL
๐78Train open models with RL inside real agent harnesses
-
SmolDataEnvs Multi-harness | SETA Whitebox
๐งชAnswer dataset questions and view scoring results
-
SmolDataEnvs Multi-harness | Native OpenCode
๐งชRun OpenCode tasks and view grading results
-
SmolDataEnv RL
๐Visualize RL agent comparison metrics in a web dashboard
spaces 29
How to turn a game into an RL environment
From an idea to a trained 4B, with the dead ends left in
Nayana Multilingual OCR
OpenEnv document OCR, layout and VQA over 22 languages
FLEURS Multilingual ASR
OpenEnv speech recognition over 102 FLEURS languages
Multilingual Multimodal Trackio
Kannada ASR and OCR GRPO runs, with live held-out evals
Multi-harness RL slides
Explore a tutorial on training coding agents with RL
