Hugging Face has released Multi-harness RL, a tool that trains language models with reinforcement learning directly inside coding agents such as Claude Code, Codex, and OpenCode, with 10 environments supported out of the box. It works through a proxy layer that intercepts requests in OpenAI, Anthropic, or Gemini format, routes them to vLLM, logs token IDs and logprobs, and forwards the resulting trajectories to the TRL library, turning the agent’s own interface into an RL environment without modifying the agent’s codebase. The proxy source is published under OpenEnv, alongside TRL training scripts and seven pre-trained models.

Multi-Harness RL on Hugging Face · Clement Delangue on X

Related: Stanford + MIT: Meta-Harness Optimization Boosts LLM Performance Up to 6x, Harness Engineering: Leveraging Codex in an Agent-First World, Hugging Face Launches ML Intern Agent for Automated ML Experiments