Papers
arxiv:2609.34645

Nereus: Adaptive Parallelism for LLM Post-Training

Published on Sep 28
ยท Submitted by
Songlin Jiang
on Sep 29
Authors:
,
,
,
,
,

Abstract

Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27times over OpenRLHF and by 1.10--1.47times over Verl across diverse clusters.

Community

Paper submitter

Nereus adapts the parallel execution plan of an RL post-training job while the job runs.
Resource availability, sequence length, memory pressure, and stage bottlenecks change during a run, so a plan that was good at the start can become slow or even infeasible. Nereus's controller selects a memory-feasible global plan and admits a transition only when the current plan is infeasible or the savings repay the transition cost. It represents each model-stage replica as an Elastic Model Unit and uses a global transition graph to order the transformations and GPU transfers across all models and stages. It achieves 27.7% lower average step latency than a fixed TP/PP layout with DP scaling, on a trace built from real data

Deploying this means the scheduler has to decide whether the remaining run is long enough to pay for a reconfiguration, and that's a bet on the horizon, not on throughput. Most auto-parallelism numbers get reported after the new plan is warm, which quietly prices the transition at zero. What I'd want on the plot is time-to-target including resharding, against how far into the run the re-plan fires โ€” a switch at 80% through a job is almost always a loss, and the scheduler needs to know that before it says yes. The other thing I'd poke at is optimizer state: changing the TP/PP split means resharding Adam's two moments plus the master weights, so if "reuse the job's distributed state" means re-sharding those tensors over the new mesh, the transition cost is a full extra checkpoint round-trip and the decision threshold moves a lot. If it's something cheaper than that, that's the interesting part and I'd want it spelled out. A run where sequence length grows monotonically would be the clean test โ€” static plans are provably wrong there, so adaptation should show up clearly or not at all.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.34645
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.34645 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.34645 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.34645 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.