Bump verl from 0.2.0.dev0 to 0.4.0 #39

dependabot · 2025-06-09T20:16:32Z

Bumps verl from 0.2.0.dev0 to 0.4.0.

Release notes

v0.4.0 release: large MoEs, tool calling, and low resource friendly

Highlights

Large MoE models support: DeepSeek 671b & Qwen3 235b

Preview features are provided to enable large MoE RL training with Megatron backend, such as DeepSeek 671b documentation. The Megatron backend now supports:

expert parallelism, context parallelism, gradient checkpointing

DeepSeek-V3, Qwen3-235b, Mixtral, Moonlight

dist-ckpt support

Tool-calling, multi-turn RL, SGLang rollout

Sample-level rollout with tool calling and multi-turn RL is supported via SGLang. We provide the Search-R1 recipe built on top of that. A prototype for sample-level async tool calling is also available with vllm AsyncLLM server. Multiple enhancements and improvements are made to SGLang rollout, supporting multi-node and multimodal. Sandbox fusion is integrated.

Low resource friendly

LoRA support is available, enabling 70B+ models on a single node with A100x8 GPUs. Fused cross entropy kernel to drastically reduce peak memory: actor_rollout_ref.model.use_fused_kernels=True

New models, algorithms and recipes

Documentation for PPO and GRPO

Recipe: DAPO

Recipe: Self-Play Fine-Tuning (SPIN)

Recipe: Self-Play Preference Optimization (SPPO)

OPO: On-Policy RL with Optimal Reward Baseline, DrGRPO, REINFORCE++, Dual-Clip PPO

New models and training utils include:

kimi_vl example

qwen3 example

video inputs support

Warmup-Stable-Decay scheduler

rope scaling

evals for GPQA, livecodebench

logging to ClearML

FSDP2 and training optimizations

FSDP2 is recommended to replace FSDP1, providing better throughput and memory usage, and is composable with other features (e.g. torch.compile):
actor_rollout_ref.ref.strategy=fsdp2
actor_rollout_ref.actor.strategy=fsdp2
critic.strategy=fsdp2 
reward_model.strategy=fsdp2 
Furthermore, FSDP2 cpu offloading is compatible with gradient accumulation. You can turn it on to save memory with actor_rollout_ref.actor.offload_policy=True.

... (truncated)

Commits

See full diff in compare view

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.

Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

@dependabot rebase will rebase this PR
@dependabot recreate will recreate this PR, overwriting any edits that have been made to it
@dependabot merge will merge this PR after your CI passes on it
@dependabot squash and merge will squash and merge this PR after your CI passes on it
@dependabot cancel merge will cancel a previously requested merge and block automerging
@dependabot reopen will reopen this PR if it is closed
@dependabot close will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually
@dependabot show <dependency name> ignore conditions will show all of the ignore conditions of the specified dependency
@dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
@dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
@dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

Bumps [verl](https://github.com/volcengine/verl) from 0.2.0.dev0 to 0.4.0. - [Release notes](https://github.com/volcengine/verl/releases) - [Commits](https://github.com/volcengine/verl/commits/v0.4.0) --- updated-dependencies: - dependency-name: verl dependency-version: 0.4.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com>

dependabot · 2025-06-30T23:37:32Z

Superseded by #74.

dependabot bot added dependencies Pull requests that update a dependency file python Pull requests that update python code labels Jun 9, 2025

dependabot bot closed this Jun 30, 2025

dependabot bot deleted the dependabot/pip/verl-0.4.0 branch June 30, 2025 23:37

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Bump verl from 0.2.0.dev0 to 0.4.0 #39

Bump verl from 0.2.0.dev0 to 0.4.0 #39

Uh oh!

dependabot bot commented on behalf of github Jun 9, 2025

Uh oh!

dependabot bot commented on behalf of github Jun 30, 2025

Uh oh!

Uh oh!

Bump verl from 0.2.0.dev0 to 0.4.0 #39

Bump verl from 0.2.0.dev0 to 0.4.0 #39

Uh oh!

Conversation

dependabot bot commented on behalf of github Jun 9, 2025

v0.4.0 release: large MoEs, tool calling, and low resource friendly

Highlights

Large MoE models support: DeepSeek 671b & Qwen3 235b

Tool-calling, multi-turn RL, SGLang rollout

Low resource friendly

New models, algorithms and recipes

FSDP2 and training optimizations

Uh oh!

dependabot bot commented on behalf of github Jun 30, 2025

Uh oh!

Uh oh!