5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
AI & ML interests
RL Environments at Scale
Recent Activity
View all activity
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
FineEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 42 โข 1 -
FineEnvs/watercolour-reference-pool
RL Environment โข Updated โข 178 โข 1.77k โข 2 -
FineEnvs/watercolour-rollouts-hps-only
RL Environment โข Updated โข 470 โข 980 -
FineEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated โข 35 โข 1
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
-
The ultimate guide to RL environments: building and scaling them in the LLM era
๐243Building and scaling RL environments for LLM training
-
RL Environments 101 โ Slides
๐40Explore RL Environments 101 slide presentation
-
How to turn a game into an RL environment
๐68From an idea to a trained 4B, with the dead ends left in
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
-
Geoguesser Environment
๐1Interact with a GeoGuessrโstyle environment through custom actions
-
FineEnvs/geoguesser-tasks
RL Environment โข Updated โข 3.65k โข 185 -
FineEnvs/geoguesser-qwen3.5-4b-grpo
Image-Text-to-Text โข Updated โข 48 -
FineEnvs/geoguesser-qwen3.5-4b-grpo-v3
Image-Text-to-Text โข Updated โข 79
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
-
Geoguesser Environment
๐1Interact with a GeoGuessrโstyle environment through custom actions
-
FineEnvs/geoguesser-tasks
RL Environment โข Updated โข 3.65k โข 185 -
FineEnvs/geoguesser-qwen3.5-4b-grpo
Image-Text-to-Text โข Updated โข 48 -
FineEnvs/geoguesser-qwen3.5-4b-grpo-v3
Image-Text-to-Text โข Updated โข 79
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
FineEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 42 โข 1 -
FineEnvs/watercolour-reference-pool
RL Environment โข Updated โข 178 โข 1.77k โข 2 -
FineEnvs/watercolour-rollouts-hps-only
RL Environment โข Updated โข 470 โข 980 -
FineEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated โข 35 โข 1
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
-
The ultimate guide to RL environments: building and scaling them in the LLM era
๐243Building and scaling RL environments for LLM training
-
RL Environments 101 โ Slides
๐40Explore RL Environments 101 slide presentation
-
How to turn a game into an RL environment
๐68From an idea to a trained 4B, with the dead ends left in