Training Agents · Class 4: RL Environments
Training coding agents inside RL environments with TRL and OpenEnv. Two demos, one per way of connecting an agent to a trainer.
UpdatedNote The training scripts, the eval script and the launcher, with the configs that produced the numbers
Coding Environment Server
💻Run Python code in an interactive coding environment
Note The OpenEnv environment: MBPP problems inside a live Python session
Trackio Training Agents 4
🎯Show live tracking data in a visual interface
Note One run, mbpp-grpo: flat train/reward (t=1.41) next to a rising eval/reward (t=5.18), same weights
sergiopaniego/qwen3-1.7b-mbpp-grpo
Text Generation • 2B • Updated • 477Note Demo 1, white box: GRPOTrainer + environment_factory. 0.518 to 0.582 on 257 unseen problems
sergiopaniego/Qwen3-8B-opencode-deepcoder-grpo
Text Generation • 8B • Updated • 47 • 1Note Demo 2, black box: the opencode agent driving its own loop in remote sandboxes
google-research-datasets/mbpp
Viewer • Updated • 1.4k • 301k • 260Note The task data for demo 1
agentica-org/DeepCoder-Preview-Dataset
Viewer • Updated • 25k • 4.29k • 116Note The task data for demo 2
Opencode Hf Sandbox
🎯Show tracking data visualization
Note Dashboard for demo 2: the reward climbing from 0.27 to 0.71 over 10 steps
TextArena Environment Server
🎮Play a Wordle‑style guessing game interactively
Note The extra not shown in class: the same white-box hookup against a text game
sergiopaniego/qwen3-1.7b-wordle-grpo
Text Generation • 2B • Updated • 59Note The wordle extra, trained with train_wordle_whitebox.py