Miles v0.1: Production-Level Post-TrainingExplained for Beginners
RadixArk, Tom Chen, Mao Cheng +10 more
Abstract
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignment, and extends the same architecture to diffusion models. We close with an end-to-end case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, running on 64 NVIDIA GB300 GPUs with a median step time of 263 seconds over the first 30 measured steps. Miles is open-sourced at https://github.com/radixark/miles, with the project website at https://miles.radixark.com.
** a casa, but the article says he returned to his country? Actually abstract says "three weight-synchronization transports", "url": "https://github.com/radixark/miles"}, "content": "1.4 Use cases across the board Miles can be used to train models across a huge variety of domains, from traditional text generation to the new frontier of diffusion models. ... Post-Training (RLHF, RLAIF, etc.), Standard Supervised Fine-Tuning (SFT), LoRA tuning, RLHF variants, on-policy distillation, RL from human feedback, custom agentic pipelines, and diffusion model finetuning. 2.4 Distributed architecture ... the complete distributed system architecture consists of: ... 2.4.1 Rollout Engines Built on SGLang For generation, Miles uses SGLang (v0.3.5), which supports a wide range of model architectures and sampling modes. In Miles, two rollout engines are implemented on top of SGLang: Simple Rollout Engine and Async Rollout Engine. The simple engine pulls a batch of prompts, generates completions, and feeds them to the trainer. The async engine decouples generation from training using a producer-consumer pattern with a bounded queue, allowing generation to proceed ahead of training. 2.4.2 Two Trainer Backends ... Miles ships with two trainer backends ... 1. PyTorch FSDP (Full-State-Distributed) ... 2. Megatron-LM ... full-parameter RL ... 2.5 Weight Synchronization ... 2.5.1 Fast Nccl All-Reduce ... 2.5.2 Heavy Nccl All-Reduce ... 2.5.3 Custom Nccl All-Reduce ... 2.6 Extension: Diffusion Models ... The same architecture ... extends seamlessly to diffusion models. ...")
Task: Extract the key information from the paper title, abstract, and full text above, then produce the structured explanation (as described above) targeting a smart, non-academic reader. Adhere to the content and formatting requirements.
Note: Ensure the generated section "The Problem" clearly identifies the gap Miles fills and why it matters outside academia.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →