Running on CPU Upgrade 85 MiMo RL Environment Explorer 🧭 85 Explore the MiMo-V2.6 RL environments and run rollouts
Running 258 The ultimate guide to RL environments: building and scaling them in the LLM era 📝 258 Building and scaling RL environments for LLM training
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF Image-Text-to-Text • 27B • Updated 13 days ago • 2.04M • 1.56k
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 59
view article Article How we OCR'ed 30,000 papers using Codex, open OCR models and Jobs nielsr • Apr 7 • 62