Rollout-as-a-service for multi-turn agent RL¶
Scope: moving the entire agentic rollout lifecycle (sandbox creation, the multi-turn agent loop, tool execution, and reward scoring) out of the RL trainer process and behind an HTTP service, so the trainer submits a task instance and receives a scored token-level trajectory. Covers the service boundary, the three-stage worker pipeline, the rootless HPC sandbox layer that makes it run on Slurm without Docker, the backend router, and what the anchor paper did and did not measure. The GPU-side pool ratio is rollout fleet sizing, the CPU fleet arithmetic is the agentic rollout sandbox fleet, the staleness gate between the two halves is the RL orchestrator control loop, and the training algorithm is agentic and tool-use RL. This page is the transport and lifecycle boundary those pages assume, not a second copy of them.
Evidence status, verified 2026-08-30. Design and result claims come from arXiv 2603.18815v1 (Zhang, Liu, Zhang, Han, Hu, Jin, Zhang, Diao, Lu, Xu, Yu, Kautz, Dong), submitted 19 March 2026, CC BY 4.0, fetched as HTML and PDF and read in full. The paper states no author affiliations; the only attribution is on page 1, where "* Core contribution." and "© 2026 NVIDIA. All rights reserved." appear as separate lines. Code claims were read from the repository itself at the paper-era commit
0a3edf40and at HEAD6a1ead6b, not from the paper's description of it. No benchmark on this page was reproduced here, no cluster was run, and no accuracy number was measured. What was computed here: the NumPy block re-derives the routing policy the shipped source implements and prices it against the one the paper's own listing describes; its assertions pass. Figure 5 publishes no numeric table, so the scaling figures below were extracted from the published PNG by pixel, and the extraction method is stated where they appear.
What it is¶
ProRL Agent is an infrastructure paper with a single architectural claim: the agentic rollout should be a service, not a library inside the trainer. The abstract states the principle directly, that "existing infrastructures often couple rollout orchestration with the training loop, making systems hard to migrate and maintain", and that under "the rollout-as-a-service philosophy" the system "serves the full agentic rollout lifecycle through an API service".
Three components implement it.1
- A pluggable task abstraction. Every task domain subclasses one interface,
AgentHandler, with three lifecycle methods:initprovisions the sandbox and configures the toolset,rundrives the multi-turn agent loop and collects the trajectory,evalscores the output and returns a scalar reward. Each handler also exposesinit_exception,run_exception,eval_exception, and afinal_resultserializer, so a rollout that dies partway still emits a well-formed response instead of stalling a worker. - A rootless container runtime.
SingularityRuntimelaunches each sandbox as an unprivileged child process in its own session, with no daemon. Images are.siffiles built from Jinja2 templates with three cache modes (Scratch always rebuilds, Versioned reuses when base image and framework version are unchanged, Lock reuses when the dependency lockfile is identical). Each container gets a unique loopback address in the127.x.x.xrange from a thread-safe allocator, which removes port-collision handling for many concurrent sandboxes on one node. - An HTTP server. Three independent worker pools drain three FIFO queues (INIT, RUN, EVAL), so the three phases overlap across the job population instead of serializing inside each job. A management API exposes job submission, per-job cancellation, LLM backend registration, and server lifecycle. Trajectories cross the boundary as token IDs, not text.
The trainer, in the paper's design, is "any training framework (e.g., veRL, NeMo RL)" that "interacts with the server solely via HTTP".
The name collides with an earlier, unrelated result. ProRL (arXiv 2505.24864) is a training-recipe paper about prolonged RL, cited on RLVR. ProRL Agent is this infrastructure paper. They share an author in Mingjie Liu and the recipe is reused for the math and STEM runs here, but the contributions are different in kind.
Why use it¶
- The two workloads want different machines. The paper's framing is that rollout is "I/O-intensive, involving sandbox creation, long-lived tool sessions, and asynchronous coordination across hundreds of concurrent instances", while training is "GPU-intensive". Once rollout is a service, the CPU fleet and the GPU fleet scale on separate axes, which is the same argument the sandbox fleet page reaches from measured trajectories.
- Trainer portability becomes real. With rollout in-process, "migrating to a different training backend often requires re-implementing the entire agent execution pipeline". Behind an HTTP boundary, adding a task is a handler plugin and changing trainers touches no rollout code.
- It runs where Docker does not. This is the paper's strongest and least-contested contribution. HPC clusters "typically forbid Docker daemons for security reasons"; Singularity
.sifimages plus--fakerootplus per-container loopback addressing give rootless, Slurm-native sandboxes. Nothing else in the compared set does this. - Token-in/token-out removes a silent correctness bug. Trajectories carry
input_ids,output_ids, andlogprobsend to end; prior assistant turns keep their original token IDs and only new observations are tokenized. This is the same defect token-in, token-out covers: text round-tripping re-tokenizes, the resulting sequence can differ from the one generated, and the update silently goes off-policy. - One measurement is clean and portable. Replacing a tmux-mediated shell with a direct
ptyprocesspseudo-terminal cut mean shell action time from 0.78 s to 0.42 s, a 46% reduction, and lifted rollout throughput from 0.29 to 0.37 instances/s.2 That change is worth making in any agent harness, whether or not this system is adopted.
When to use it (and when not)¶
Use it when:
- Rollouts run on a shared Slurm or HPC cluster where a Docker daemon and root are unavailable. This is the case the design is actually built for.
- More than one trainer is in play, or the trainer is expected to change, and the rollout stack should outlive that choice.
- Task domains are heterogeneous (software engineering, math, STEM, computer use) and each needs its own environment, toolset, and reward.
Do not use it when:
- A throughput win over an incumbent framework is the reason for switching. The paper contains no head-to-head system benchmark against SkyRL-Agent, VeRL-Tool, rLLM, GEM, or Agent Lightning. Table 1 is a three-column qualitative matrix on which this system wins every column, Table 3 is a self-ablation, and Figure 5 is self-scaling. The claim that decoupling is faster is argued, not measured against anything.
- Capacity planning assumes linear rollout scaling. See the scaling numbers below: 8x the nodes buys roughly 4.3x to 4.9x the throughput.
- The intent is to run the paper's code from the repository's default branch. It is not there any more; see maintenance below.
- The cluster already runs Docker comfortably and the rollout stack is not changing. The service boundary is real engineering, and its benefit is portability, not speed.
Architecture¶
The service boundary is the whole design. The trainer holds no container, no tool session, and no grader; it holds an HTTP client.
flowchart LR
T["RL trainer<br/>(veRL / NeMo RL)"] -->|"submit instance + sampling params"| S["ProRL Agent server"]
S -->|"token IDs, logprobs, reward"| T
subgraph S2["Server: three pools, three queues"]
Q1["INIT queue"] --> W1["init workers<br/>(I/O bound)"]
W1 --> Q2["RUN queue"]
Q2 --> W2["run workers<br/>(LLM bound)"]
W2 --> Q3["EVAL queue"]
Q3 --> W3["eval workers<br/>(ms to minutes)"]
end
S --- S2
W1 --> C["SingularityRuntime sandbox<br/>rootless .sif, 127.x.x.x, --fakeroot"]
W2 --> C
W2 --> B["LLM backend heap"]
B --> V["vLLM / SGLang replicas"]
W3 --> R["reward scoring"]
Three details carry the design. The stages are separately sizable because their resource profiles differ: container startup waits on disk and network, the agent loop waits on GPU, and evaluation "ranges from a few milliseconds for direct scoring to several minutes for full test-suite execution". A single worker walking all three phases idles through whichever is slow. Timeouts are phase-aware: a PausableTimer accumulates only during active stages, "while excluding time spent waiting in inter-stage queues", so queue backpressure cannot masquerade as a rollout timeout. The container is freed before evaluation starts, which matters because eval is the longest-tailed stage and holding a sandbox through it wastes the fleet.
The stage decomposition has prior art the pipeline subsection does not acknowledge. SkyRL-v0, which this paper cites elsewhere, already published "Three-stage Producer-Consumer Pipelining", naming the same three stages (runtime initialization, trajectory generation, reward calculation) with the same stated motivation of reducing GPU idle time.3 A producer-consumer split is a generic pattern and nothing here establishes copying; what it does establish is that the decomposition predates the paper and is presented without that context. SkyRL-v0 also already describes "a scalable Remote Sandbox Server that separates environment execution from training", yet Table 1 marks SkyRL-Agent with a cross on "Training-Rollout Decoupled?". The paper's narrow reading is defensible, because rollout control does remain in SkyRL's trainer and the appendix argues that case per framework, but a binary cross in the headline table overstates the gap.
How to use it¶
The management API is seven endpoints. This is the paper's Listing, cross-checked against scripts/start_server.py at the paper-era commit 0a3edf40, where every route below is registered.5
POST /add_llm_server {"address": "http://host:port/v1"} register a backend
POST /clear_llm_server flush all backends
POST /process {"instance": {...}, "sampling_params": {...}} submit a rollout
POST /cancel {"job_id": "..."} abort a running job
POST /start | POST /stop server lifecycle
GET /status queue depths
The checkpoint-swap protocol is the part worth internalizing. When the policy advances, the old inference weights are invalid, but the server does not restart: the trainer calls POST /clear_llm_server, then re-registers the reloaded endpoints, and "all subsequent rollouts automatically use the updated model, with no interruption to jobs already in the pipeline". That is a deliberate staleness decision, not a race. Jobs already mid-rollout keep finishing against the policy version they started on, which is exactly the off-policy age that the orchestrator control loop bounds with an admission gate. This service does not bound it; the trainer must.
POST /process is a blocking request. The paper's own pseudocode ends the pipeline with job.done.set() # unblock the waiting HTTP handler, so the connection is held for the life of the rollout, which for a trajectory spanning "dozens of turns" is minutes to hours. Size proxy, keepalive, and connection limits accordingly; the paper does not discuss them.
How to develop with it¶
The routing rule is described three different ways, and while they are not pairwise contradictory they cannot all hold at once. It is worth settling, because it decides whether the backend pool is load-aware or not.
- The architecture listing comments the structure
# min-heap keyed by in-flight count, which implies a decrement when a rollout finishes. - The prose says "The counter is incremented once per task (rather than per call), ensuring that all subsequent calls within the same task are consistently routed to the same backend to maximize prefix cache reuse."
- The very next sentence, defining the same symbol, says "where $w_s$ counts the total number of inference calls assigned to server $s$ since it was registered."
The first two are compatible with each other, since a key that counts active tasks would be incremented once per task and decremented on completion. Neither is compatible with the third, which is cumulative and never falls. The source settles which one shipped. At 0a3edf40, openhands/nvidia/async_server.py holds the entire mutation of the key in create_llm_config:
address = self.weighted_addresses[0][1]
self.weighted_addresses[0][0] += 1 # type: ignore
heapq.heapreplace(self.weighted_addresses, self.weighted_addresses[0])
git grep weighted_addresses over that commit shows the counter initialized to zero, incremented here, and never decremented anywhere. create_llm_config is called once per job, at job creation, and the result is stored on job_details.llm_config, so the "once per task" reading is the shipped one and the "inference calls" sentence is wrong about its own code. The project's own unit test agrees on the character of the result: test_weighted_addresses_load_balancing asserts only that two addresses are both used across four configs, and its comment reads "ensure addresses are rotated".
A min-heap on a never-decremented per-task counter is round-robin over a fixed backend set: every backend is held to an equal cumulative count, with ties broken deterministically by heap order. It is not round-robin across registration changes, since a newly added backend starts at zero and takes consecutive jobs until its count catches up, and clear_llm_server resets every counter. The paper half-concedes the steady-state behaviour ("achieving a round-robin-like balance"), but not the consequence: the router cannot see that a backend is still busy, because its key has already moved on. The block below is a toy discrete-event model of that one decision. It counts a rollout as one unit of backend load for its whole life and models nothing else: no token volume, no per-turn call pattern, no tool-wait time, no queueing, no cache effects, no backend throughput. What it can therefore show is concurrency imbalance, not GPU utilization or latency. On that narrow question the ambiguity is free when rollout durations are uniform and costly when they are not, and the paper's own motivation notes that agentic rollouts "may incur highly variable latency".
# prorl_backend_routing.py - the paper specifies the LLM-backend min-heap key two
# incompatible ways: the architecture listing comments it "min-heap keyed by
# in-flight count", while the prose defines w_s as "the total number of inference
# calls assigned to server s since it was registered" and concedes the result is
# "round-robin-like". The paper-era source settles it: the counter is incremented
# once per job and never decremented, so over a fixed backend set the shipped
# policy is round-robin. This
# block measures what that costs against the least-in-flight policy the listing
# describes. Discrete-event model in numpy; it tests the two specifications
# against each other and measures nothing about the shipped server.
import numpy as np
def route(n_backends, durations, key):
"""One rollout arrives per tick, assigned to the min-key backend (ties to the
lowest index, as a heap pop would). key='cumulative' never decrements the
counter, matching the shipped code; key='inflight' decrements on completion,
matching the listing's comment. Returns (assignment, ticks x backends
concurrency trace)."""
w = np.zeros(n_backends, dtype=np.int64)
inflight = np.zeros(n_backends, dtype=np.int64)
free_at = [[] for _ in range(n_backends)]
assign = np.empty(len(durations), dtype=np.int64)
trace = np.empty((len(durations), n_backends), dtype=np.int64)
for t, d in enumerate(durations):
for b in range(n_backends):
keep = [c for c in free_at[b] if c > t]
retired = len(free_at[b]) - len(keep)
free_at[b] = keep
inflight[b] -= retired
if key == "inflight":
w[b] -= retired
s = int(np.argmin(w))
w[s] += 1
inflight[s] += 1
free_at[s].append(t + d)
assign[t] = s
trace[t] = inflight
return assign, trace
def imbalance(trace):
"""Mean over time of (busiest backend - idlest backend), and the mean spatial
coefficient of variation of concurrency across backends."""
gap = float((trace.max(axis=1) - trace.min(axis=1)).mean())
cv = float((trace.std(axis=1) / np.maximum(trace.mean(axis=1), 1e-9)).mean())
return gap, cv
N, M = 8, 4000
rng = np.random.default_rng(0)
uniform = np.full(M, 20)
heavy = np.clip(rng.lognormal(2.5, 1.4, M), 1, None).astype(int) # median 11, max 1164
aligned = np.where(np.arange(M) % N == 0, 400, 10) # skew on the RR period
# 1) Over a fixed backend set the never-decremented counter is exactly round-robin,
# for every duration profile. Duration never enters the decision at all.
for durations in (uniform, heavy, aligned):
a, _ = route(N, durations, "cumulative")
assert np.array_equal(a, np.arange(M) % N), "monotone counter must be exact round-robin"
# 2) Uniform durations: both readings hold concurrency within one rollout, so the
# ambiguity in the paper is invisible in this regime and costs nothing.
gap_uc, _ = imbalance(route(N, uniform, "cumulative")[1])
gap_ui, _ = imbalance(route(N, uniform, "inflight")[1])
assert gap_uc <= 1.0 and gap_ui <= 1.0
# 3) Heavy-tailed durations, which is the regime the paper's own motivation
# describes: the two readings separate by several-fold.
gap_hc, cv_hc = imbalance(route(N, heavy, "cumulative")[1])
gap_hi, cv_hi = imbalance(route(N, heavy, "inflight")[1])
assert gap_hi < gap_hc / 3 and cv_hi < cv_hc / 3
# 4) Adversarial: long rollouts landing on the round-robin period pin every one of
# them onto backend 0. Least-in-flight routes around it; round-robin cannot see
# it, because its counter has already moved on.
_, tr_ac = route(N, aligned, "cumulative")
_, tr_ai = route(N, aligned, "inflight")
gap_ac, _ = imbalance(tr_ac)
gap_ai, _ = imbalance(tr_ai)
assert tr_ac[:, 0].max() >= 20 * tr_ac[:, 1:].max(), "round-robin must pile onto backend 0"
assert gap_ac > 30 * gap_ai
assert tr_ai.max() <= tr_ac.max() / 5
# 5) Edge: one backend leaves both readings identical and the metric well-defined.
a1c, tr1 = route(1, heavy, "cumulative")
a1i, _ = route(1, heavy, "inflight")
assert np.array_equal(a1c, a1i) and int(a1c.max()) == 0
assert imbalance(tr1) == (0.0, 0.0)
print(f"1) monotone counter == round-robin (rollout t -> backend t % {N}), all three profiles: True")
print(f"2) uniform durations mean busiest-idlest gap round-robin={gap_uc:.3f} least-in-flight={gap_ui:.3f}")
print(f"3) heavy-tailed mean busiest-idlest gap round-robin={gap_hc:.3f} least-in-flight={gap_hi:.3f}")
print(f" heavy-tailed mean concurrency CV round-robin={cv_hc:.4f} least-in-flight={cv_hi:.4f}")
print(f"4) period-aligned skew peak concurrency round-robin={int(tr_ac.max())} least-in-flight={int(tr_ai.max())}")
print(f" round-robin's peak is backend 0 at {int(tr_ac[:, 0].max())}; every other backend peaks at {int(tr_ac[:, 1:].max())}")
print(f" period-aligned skew mean gap round-robin={gap_ac:.3f} least-in-flight={gap_ai:.3f}")
print(f"5) single backend: both readings identical, imbalance exactly zero: True")
Executed output (.venv/bin/python, Python 3.12.3, numpy 2.5.1, pasted verbatim):
1) monotone counter == round-robin (rollout t -> backend t % 8), all three profiles: True
2) uniform durations mean busiest-idlest gap round-robin=1.000 least-in-flight=1.000
3) heavy-tailed mean busiest-idlest gap round-robin=4.865 least-in-flight=1.093
heavy-tailed mean concurrency CV round-robin=0.4266 least-in-flight=0.1208
4) period-aligned skew peak concurrency round-robin=50 least-in-flight=8
round-robin's peak is backend 0 at 50; every other backend peaks at 2
period-aligned skew mean gap round-robin=46.552 least-in-flight=0.988
5) single backend: both readings identical, imbalance exactly zero: True
The practical readings: under uniform rollout durations the two specifications are indistinguishable, which is why the ambiguity survived review. Under a lognormal duration profile with median 11 and maximum 1164, round-robin's mean busiest-to-idlest concurrency gap is 4.9 rollouts against 1.1, a 4.4x worse imbalance. Under skew aligned to the rotation period, every long rollout lands on the same backend and peak concurrency there reaches 50 while every other backend peaks at 2. Prefix-cache affinity, which is the stated reason for pinning a whole task to one backend, is a real and correct goal; the defect is that nothing bounds the resulting imbalance. If the backend pool is uneven in speed or the task mix is long-tailed, add a load term rather than trusting the counter.
This also complicates the paper's own ablation. The "without Load Balancing" arm is described as "a simple baseline assignment strategy that distributes an equal number of instances to each LLM server", which is round-robin, and the method under test is a counter the paper itself calls "round-robin-like". Two near-identical static policies should not move GPU utilization from 42% to 78%. Either the effect belongs to the separate node-locality phase described in section 3.4, which Table 3 does not name, or the ablated baseline differs from its description.
How to maintain it¶
- Pin the commit, and know which system the pin points at. The repository was rewritten after publication.
75de6df3("Merging Polar to Main", 2026-05-14) is a single commit removing about 212,000 lines and adding about 20,700; the paper'sopenhands/nvidia/andtrainer_integration/verl/trees are gone. HEAD (6a1ead6b, 2026-08-13) implements a successor project, Polar, whose own report is arXiv 2605.24220.AgentHandler,SingularityRuntime,POST /process,POST /add_llm_server, and the backend heap are absent at HEAD; the readable implementation of this paper is0a3edf40(2026-04-02) or earlier. - Expect the interfaces to have moved, not merely been renamed. At HEAD the submission path is
POST /rollout/task/submitreturning a task id plusGET /rollout/task/{id}, so the blocking call is gone; cancellation isDELETE /sessions/{session_id}; the handler contract isBaseHarnesswithsetup/run_steps/postrun_steps/postprocess, which emits shell commands rather than owning the lifecycle; the stage machine is four stages (INIT, READY, RUNNING, POSTRUN) with evaluation folded into POSTRUN; and there is no backend heap at all, because each gateway node binds one upstream from static config. Token-in/token-out is the one named component that survived intact. - Treat the default branch carefully. Real work is on
stable;origin/mainis a stub. The project's ancestry is an OpenHands fork, and OpenHands is now one of eleven harness presets rather than the substrate. - Re-verify the rootless story on your own cluster.
--fakerootdoes not remove the privilege dependency, it relocates it: Apptainer requires either a setuid installation or unprivileged user namespaces plus admin-provisioned/etc/subuidand/etc/subgidmappings. On a hardened site that is an administrator ticket, not a user action. The paper presents the flag as though it settles the question. - Pin inference-engine versions hard if using the current code. Token metadata requires patched servers; the repository's patch script hard-pins SGLang
0.5.13and fails on mismatch, while vLLM is left unpinned.
How to run it in production¶
- Size INIT, RUN, and EVAL pools independently, and measure before choosing. The design's whole point is that the three have different profiles, so more init workers absorb slow I/O startup and more eval workers absorb long test suites. The paper does not publish the pool sizes used for any experiment. For reference, the paper-era defaults were 6 init and 5 run with eval mirroring run, while the current code defaults to 4 init, 2 run, and 4 postrun; neither is a recommendation for your workload.
- Adopt phase-aware timeouts even if nothing else here is adopted. Excluding inter-stage queue wait from the timeout budget costs nothing and stops backpressure from being misread as rollout failure.
- Pair rollout servers with node-local inference. The trainer-side scheduler assigns LLM servers preferentially to servers on the same physical node by IP match, then round-robins the remainder. That locality phase is cheap and is plausibly where the ablation's utilization gain actually comes from.
- Plan capacity at roughly 55% to 61% parallel efficiency, not linear. Section 4.4.1 claims throughput "increases nearly linearly with the number of nodes" with "minimal scaling overhead". The paper publishes no table for Figure 5, so the figure PNG was calibrated against its own gridlines and the three series extracted by colour: 4B goes 0.386 to 1.667 instances/s from 1 to 8 nodes (4.32x, 54%), 8B goes 0.350 to 1.717 (4.91x, 61%), and 14B goes 0.260 to 1.263 (4.85x, 61%). The plotted 8B value also falls from 0.653 at 2 nodes to 0.621 at 4 nodes, a 4.8% drop inside a figure whose caption asserts a "near-linear increase"; because the paper reports neither a repeated-run count nor error bars, that is a decline in the plotted value rather than a measured regression, but the endpoint efficiencies do not depend on it.6 The paper never states GPUs per node, so the x-axis cannot be converted to GPUs.
- Bound off-policy age in the trainer.
clear_llm_serverswaps weights without draining in-flight rollouts by design. Nothing in the service caps how stale a returned trajectory is. - Do not treat the accuracy table as evidence about the transport. Table 2 measures model quality, which is a property of the recipe (DAPO, data, thinking mode), not of the rollout boundary.
Failure modes¶
- Backend pile-up under long-tailed rollouts. The router is round-robin in effect, so a backend holding several long trajectories keeps receiving new ones. Symptom: uneven queue depth and GPU utilization across otherwise identical inference replicas. Add a load term or cap per-backend concurrency.
- Held connections exhaust the proxy, not the server.
POST /processblocks for the whole rollout. At hundreds of concurrent instances this is hundreds of multi-minute connections through whatever sits in front of the service. - Rollout service death loses every in-flight trajectory. Making rollout a single service centralizes the failure the paper criticizes Agent Lightning for. The paper offers no fault-tolerance story for the server itself.
--fakerootunavailable on the cluster. Image builds and in-container package installs fail at submission time, not at design time. Verify subuid/subgid mappings during bring-up.- Silent off-policy drift after a checkpoint swap. Jobs in flight during
clear_llm_servercomplete against the old weights and are returned as if current. - Citing the repository HEAD as the paper's implementation. It is a different system with a different API. Pin
0a3edf40. - Reading Table 3 as three independent wins. Only Efficient Bash has an isolated, quantified mechanism (action time 0.78 s to 0.42 s). The IPython in-process kernel and the Unix-domain-socket transport, two of the three claimed tool-backend optimizations, are never measured anywhere in the paper.
Two further cautions about the evidence itself, because the numbers are easy to over-read. The SWE-Bench comparison is not like-for-like: SkyRL-v0's own writeup records "SkyRL-Agent-8B-v0 from Qwen3-8B (no thinking): 3.6% → 9.4%" and "SkyRL-Agent-14B-v0 from Qwen3-14B (thinking-enabled): 18.0% → 21.6%", while this paper enables thinking mode for its 8B and 14B runs.4 So the headline "nearly a 2x improvement" at 8B compares a thinking-mode run against a no-thinking one. Measured instead as the final-to-base score multiple each system achieved over its own starting checkpoint, SkyRL's 8B is 2.61x (3.6 to 9.4) against this paper's 1.88x (9.6 to 18.0), which is the comparison the table's layout makes hard to see. The 14B head-to-head margin is 23.6 against 21.6, which on the 500-instance SWE-Bench Verified set is ten instances, reported without seeds or error bars, and this paper's reproduced Qwen3-14B base (15.4) sits 2.6 points below the base SkyRL reports for the same model in the same mode (18.0). Separately, the ablation names Qwen3-14B-Instruct-2507; that repository does not resolve on the Hugging Face API (HTTP 401, the response for an absent or private repo) while Qwen3-4B-Instruct-2507, Qwen3-8B, and Qwen3-14B all return 200, so the ablation's checkpoint has no public counterpart. The paper has no limitations section.
References¶
- Zhang, Liu, Zhang, Han, Hu, Jin, Zhang, Diao, Lu, Xu, Yu, Kautz, Dong, ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents (arXiv 2603.18815, submitted 19 March 2026, CC BY 4.0): https://arxiv.org/abs/2603.18815
- ProRL Agent Server repository (Apache-2.0; paper-era code at commit
0a3edf40, HEAD is the successor project Polar): https://github.com/NVIDIA-NeMo/ProRL-Agent-Server - Xu et al., Polar: Agentic RL on Any Harness at Scale (arXiv 2605.24220, 22 May 2026; the report the repository's current HEAD implements): https://arxiv.org/abs/2605.24220
- NVIDIA NeMo Gym (the library the paper states ProRL Agent is integrated into): https://github.com/NVIDIA-NeMo/Gym
- NVIDIA NeMo RL (one of the two trainers the paper names): https://github.com/NVIDIA-NeMo/RL
- Cao et al., SkyRL-Agent (arXiv 2511.16108); SkyRL-v0 writeup, the source of the baseline numbers Table 2 cites without a URL: https://arxiv.org/abs/2511.16108 and https://novasky-ai.notion.site/skyrl-v0
- Yu et al., DAPO: An Open-Source LLM Reinforcement Learning System at Scale (the paper's default RL algorithm): https://arxiv.org/abs/2503.14476
- Jiang et al., VeRL-Tool (arXiv 2509.01055), Luo et al., Agent Lightning (arXiv 2508.03680), Liu et al., GEM (arXiv 2510.01051): the frameworks compared in Table 1: https://arxiv.org/abs/2509.01055 https://arxiv.org/abs/2508.03680 https://arxiv.org/abs/2510.01051
- vLLM blog, Agent Lightning (the re-tokenization drift the paper cites for token-in/token-out): https://blog.vllm.ai/2025/10/22/agent-lightning.html
- Apptainer fakeroot documentation (what
--fakerootactually requires from the cluster): https://apptainer.org/docs/user/main/fakeroot.html
Related: Sizing the agentic rollout sandbox fleet, The RL orchestrator control loop, Rollout fleet sizing, Agentic and tool-use RL, Async and disaggregated RL systems, Token-in, token-out, Agent Lightning, NeMo-RL, verl, RL libraries overview, Agent sandboxing and isolation, RL environment sandbox escape
-
Figure 2 caption, verbatim: "(1) Sandbox Environment: each rollout is executed inside a SingularityRuntime container and orchestrated via AgentHandler, which exposes three lifecycle methods including init(), run(), and eval() ... (2) ProRL Agent Server: an HTTP service that manages rollouts through a three-stage asynchronous pipeline (INIT → RUN → EVAL) with independent worker pools, and maintains a min-heap LLM backend pool supporting dynamic registration and checkpoint swapping. (3) RL Trainer: any training framework (e.g., veRL, NeMo RL) interacts with the server solely via HTTP". Sections 3.2.2 and 3.2.3 supply the
.sifbuild cache modes, the127.x.x.xallocator,--fakeroot,--network none, and the three tool-backend optimizations (ptyprocess PTY, in-process IPython kernel, Unix domain sockets). ↩ -
Table 3, all four rows, on Qwen3-14B-Instruct-2507 with 8 H100 GPUs under DAPO. Full system: action time 0.42 s, GPU util 78%, throughput 0.37 instance/s. Without Load Balancing: 0.42 s, 42%, 0.25. Without Efficient Bash: 0.78 s, 68%, 0.29. Without Stale Job Cleanup: 0.42 s, 65%, 0.30. Section 4.1 gives the shared setup: DAPO, batch size 32, mini-batch size 8, 8 rollouts per instance, KL coefficient 1e-4, learning rate 1e-6, and "All RL training is performed on 32 NVIDIA H100 GPUs" (the ablation separately uses 8). ↩
-
SkyRL-v0, cited in the paper as Cao et al. 2025a with no URL in its bibliography: "Three-stage Producer-Consumer Pipelining – further, to reduce GPU idle time, we can decouple and overlap the stages of (i) runtime initialization (e.g., building images, launching and connecting containers), (ii) trajectory generation, and (iii) reward calculation (e.g., sending patches, running tests), to maximize parallelism and system throughput, as shown in the bottom part of Figure 4", and "we implemented a scalable Remote Sandbox Server that separates environment execution from training. This disaggregated design enables independent scaling of environment workers to maintain a high rate of environment interaction, ensuring high GPU utilization in the rollout stage". Retrieved 2026-08-30 from the page's Notion content API, since the published page is a JavaScript shell. ↩
-
Same source, two list items quoted with their leading marker glyphs cut: "... SkyRL-Agent-8B-v0 from Qwen3-8B (no thinking): 3.6% → 9.4%" and "... SkyRL-Agent-14B-v0 from Qwen3-14B (thinking-enabled): 18.0% → 21.6%". Table 2 of the paper reports only the 9.4 and 21.6 endpoints, in a column headed "Reported", alongside its own reproduced bases of 9.6 (Qwen3-8B) and 15.4 (Qwen3-14B) and finals of 18.0 and 23.6. ↩
-
Section 3.3.2 Listing. Confirmed against
0a3edf40:scripts/start_server.py, whereFastAPI(title='OpenHands Async Server API')registers/start,/stop,/status,/cancel,/add_llm_server,/clear_llm_server, and/process. The blocking behaviour is from the section 3.3 Listing, whose final line isjob.done.set() # unblock the waiting HTTP handler. ↩ -
Figure 5 publishes no numeric table, so values were recovered from the published PNG (
https://arxiv.org/html/2603.18815v1/figures/fig4_throughput.png, 1622x1202). Calibration used the figure's own gridlines: horizontal rules at pixel rows 1008 and 209 correspond to 0.2 and 1.6 instances/s, and vertical rules at columns 382 and 1535 to 1 and 8 nodes, both linear. Series were separated by their exact plot colours, RGB (31, 59, 115) for 4B, (217, 95, 2) for 8B, and (27, 158, 119) for 14B, and each marker's value taken as the mean row of matching pixels within nine columns of the tick. One extra step is needed and is easy to miss: the legend sits over the 1-node and 2-node columns and its swatches are drawn in the same three colours, so for tick columns left of x=560 only rows below y=420 are counted. Without that exclusion the 1-node values are contaminated by legend pixels. The figure publishes no error bars and the paper states no repeated-run count, so what declines between 2 and 4 nodes is the plotted value for 8B, not a quantity with a stated uncertainty; the decline is visible by eye. ↩