突破舒适区:面向RLVR的高效策略引导探索框架NudgeRL

HuggingFace Daily Papers(社区热门论文)·2026-05-15 08:00·128天前
AI 导读

强化学习与可验证奖励范式面临探索效率瓶颈。为此,研究团队提出NudgeRL框架,其核心是“策略助推”技术,通过为每次策略采样注入轻量级策略级上下文,引导模型产生多样化推理轨迹,无需依赖昂贵的外部监督。该框架进一步提出一个统一目标,将奖励分解为上下文间与上下文内组件,并通过蒸馏目标将有效行为迁移回基础策略。在五个高难度数学基准测试中,NudgeRL的表现优于标准GRPO方法,其效果相当于后者使用高达8倍采样预算的结果,且平均表现超过了依赖特权信息的Oracle引导基线,证明了结构化探索的高效性与可扩展性。

HuggingFace Daily Papers(社区热门论文)
精选
71AI 编辑部评分,满分 100

突破舒适区:面向RLVR的高效策略引导探索框架NudgeRL

2026-05-15 08:00· 128天前
AI 导读

强化学习与可验证奖励范式面临探索效率瓶颈。为此,研究团队提出NudgeRL框架,其核心是“策略助推”技术,通过为每次策略采样注入轻量级策略级上下文,引导模型产生多样化推理轨迹,无需依赖昂贵的外部监督。该框架进一步提出一个统一目标,将奖励分解为上下文间与上下文内组件,并通过蒸馏目标将有效行为迁移回基础策略。在五个高难度数学基准测试中,NudgeRL的表现优于标准GRPO方法,其效果相当于后者使用高达8倍采样预算的结果,且平均表现超过了依赖特权信息的Oracle引导基线,证明了结构化探索的高效性与可扩展性。

推荐理由

NudgeRL 首次把结构化探索引入 RLVR,比 GRPO 节省 8 倍 rollout 预算,数学推理效果还更好。做 LLM 推理优化的团队,这篇值得复现。

Abstract

Reinforcement learning with verifiable rewards (RLVR) has emerged as a scalable paradigm for improving the reasoning capabilities of large language models. However, its effectiveness is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. While increasing the number of rollouts alleviates this issue, such brute-force scaling is computationally expensive, and existing approaches that modify the optimization objective provide limited control over what is explored. In this work, we propose NudgeRL, a framework for structured and diversity-driven exploration in RLVR. Our approach introduces Strategy Nudging, which conditions each rollout on lightweight, strategy-level contexts to induce diverse reasoning trajectories without relying on expensive oracle supervision. To effectively learn from such structured exploration, we further propose a unified objective, which decomposes the reward signal into inter- and intra-context components and incorporates a distillation objective to transfer discovered behaviors back to the base policy. Empirically, NudgeRL outperforms standard GRPO with up to 8 larger rollout budgets, while outperforming oracle-guided RL baseline on average across five challenging math benchmarks. These results demonstrate that structured, context-driven exploration can serve as an efficient and scalable alternative to both brute-force rollout scaling and feasibility-oriented methods based on privileged information. Our code is available at https://github.com/tally0818/NudgeRL.

1 Introduction

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for improving the reasoning capabilities of large language models (LLMs) [20, 7]. By leveraging verifiable rewards, methods such as Group-Relative Policy Optimization (GRPO) [18] enable scalable post-training without requiring dense supervision. This paradigm has been successfully applied across a wide range of domains.

Despite its success, RLVR remains fundamentally limited by its ability to explore the space of reasoning trajectories. A natural approach is to scale the number of sampled rollouts, which increases the probability of discovering rare trajectories [5]. However, such brute-force scaling quickly becomes computationally prohibitive, motivating alternative approaches that improve exploration efficiency.

Refer to caption
Figure 1: Concept: Improving exploration diversity through Strategy Nudging. (a) Naive sampling methods (e.g., GRPO) often collapse to a dominant reasoning mode, limiting the exploration of the reasoning space. (b) NudgeRL introduces Strategy Nudging, which appends lightweight strategy to the input, forcing the model to traverse diverse reasoning modes. (c) As a result, Strategy Nudging significantly increases the number of distinct reasoning approaches discovered compared to the baseline, effectively mitigating the exploration bottleneck. Additional details are in Appendix˜B

Recent work has sought to address this limitation by modifying the optimization objective, for example through entropy regularization or decoupled clipping [26, 24]. While these methods encourage broader exploration at the distribution level, they provide limited control over what is explored, and often fail to ensure coverage of semantically meaningful reasoning strategies. Another line of work leverages privileged information, such as oracle solutions or intermediate reasoning steps, to improve the feasibility of discovering correct trajectories [27, 16, 8, 19]. Although effective, these approaches are primarily feasibility-oriented and rely on strong supervision signals that are expensive to obtain and difficult to scale. Moreover, by guiding the policy toward a narrow set of predefined successful trajectories, they may limit exploration diversity and hinder the discovery of alternative reasoning strategies [25, 23].

In this work, we address the exploration bottleneck by explicitly structuring the reasoning space in a scalable manner. We propose NudgeRL, a framework that introduces Strategy Nudging during the exploration phase. Instead of relying on expensive oracle data, Strategy Nudging appends lightweight, heuristic text prompts (e.g., specific strategies for math problems or reasoning keywords) to the input. This deliberately forces the model to traverse distinct, diverse reasoning modes that it might otherwise ignore under purely naive sampling.

However, learning from such context-conditioned exploration introduces new challenges. Since rollouts are generated under different context-conditioned prompts, the samples are naturally partitioned into multiple distinct groups, where reward variation reflects both the intrinsic trajectory quality and context-specific biases, making standard group-wise advantage estimation unreliable. Furthermore, context forcing creates a mismatch between how trajectories are sampled and how the policy is finally used at inference time. Without intervention, improvements discovered under context-forced exploration may not transfer directly to the base policy. To address these challenges, we further introduce (i) an Inter-Intra group advantage to enable meaningful credit assignment across context-induced groups, and (ii) a distillation-augmented objective that explicitly transfers effective behaviors discovered during context-forced exploration back to the base policy.

Our approach enables structured and diversity-driven exploration while remaining fully compatible with standard RLVR pipelines. Empirically, NudgeRL achieves performance surpassing GRPO even when GRPO is given an larger rollout budget, while outperforming oracle-guided baselines. This demonstrates that scalable, diversity-oriented exploration can serve as an effective alternative to both brute-force rollout scaling and feasibility-driven privileged information.

2 Preliminaries

2.1 Group-Relative Policy Optimization (GRPO)

We consider an empirical distribution of prompts . For each prompt , a policy generates a group of rollouts , where each rollout is sampled as . Each rollout is evaluated by a verifiable reward function .

Unlike standard PPO [17], which typically estimates advantages using a learned value function, GRPO [18] derives advantages from group-wise rewards. For rollouts sampled from the same prompt , let denote the reward of rollout . The group-wise advantage is then defined as:

(1)

where and are the reward mean and standard deviation within the group, and is used for numerical stability. This yields a relative advantage estimate without training a value function.

The policy is then optimized with a PPO-style clipped objective:

(2)

Thus, GRPO retains PPO’s clipped objective while using group-relative advantages.

2.2 Motivation: From Exploration to Performance Gain

To understand why exploration is a fundamental bottleneck in RLVR, we look beyond trajectory-level rewards and examine how the probability mass of generated tokens shifts during training. Hu et al. [5] characterizes the expected one-step performance improvement () in RLVR as:

(3)

where and denote the total probability mass of correct and incorrect tokens, is the learning rate, and is the number of rollouts. and are the second moments of sampled correct and incorrect tokens, while and are those of unsampled correct and incorrect tokens. represents the net reward contribution from sampled tokens.

Since , the first two terms in Eq.˜3 are non-negative and drive learning forward. The third term, however, acts as a potential penalty. Because incorrect tokens typically dominate the probability mass (), a large , meaning the model has significant probability mass on correct trajectories that it simply fails to explore, creates a dominant negative force that hinders performance gain. Therefore, the core bottleneck of RLVR lies in the unexplored correct regions.

Limitations of rollout scaling.

To mitigate this penalty, a naive solution is to increase the rollout size . Hu et al. [5] shows that for a collection of tokens with probabilities , the expected unsampled second moment after draws is:

(4)

which decreases monotonically with . However, tokens with small decay slowly, so fully covering long-tail correct trajectories requires prohibitively large rollout budgets.

This highlights the limitation of blindly scaling to reduce the unexplored correct mass (). Long-tail correct trajectories remain unlikely to be sampled even under large , suggesting the need for a structured exploration mechanism that can efficiently expose such latent trajectories.

3 NudgeRL

We introduce NudgeRL, a framework for structured exploration and learning in RLVR. NudgeRL consists of three components: (i) Strategy Nudging, which conditions rollout generation on strategy-level contexts to induce diverse reasoning trajectories; and (ii) Inter-intra Group Advantage, a credit assignment method that enables controlled exploration and exploitation of strategies; and (iii) Distillation augmented RL objective to learn from context-conditioned rollouts and distill effective strategies into the policy under the original prompt for inference without external context.

3.1 Strategy Nudging: Structured Exploration via Strategy-Level Contexts

Given that prior work [5] alleviates the exploration bottleneck by reducing unsampled probability mass through larger rollout budgets, a natural question arises: how many rollouts are required to reliably discover a rare trajectory? To quantify this discovery cost, consider a rare trajectory with . The expected number of rollouts required to observe at least once is:

(5)

This implies that for low-probability trajectories, the required rollout budget grows prohibitively large. In practice, naive rollout scaling repeatedly samples from high-probability modes of the current policy, leading to diminishing returns in covering rare trajectories.

This motivates conditioning generation on a context that can shift the sampling distribution toward otherwise rare trajectories. If such a context increases the probability of a trajectory , i.e., , then its expected number of rollouts becomes:

(6)

Thus, contexts need not provide a solution; they can serve as lightweight controls that alter the sampling distribution and reduce the cost of discovering rare trajectories.

Strategy Nudging.

Even though context conditioning can improve exploration efficiency in principle, simply placing multiple contexts in a single prompt leaves the choice of strategy to the policy, which may ignore some contexts and repeatedly follow dominant reasoning patterns. To enforce coverage over contexts, we instead assign a single sampled context to each rollout before generation.

Let denote a pool of Strategy-level contexts for the original prompt . For each rollout index , we begin with sampling . To avoid relying exclusively on the context pool and to retain compatibility with the original prompt, we further apply context dropout. Specifically, we sample a mask and define the context as:

(7)

We then construct the final prompt , and generate . By varying across rollout indices, Strategy Nudging induces diversity at the input-conditioning level, rather than relying solely on sampling from a single prompt. Details on generating are in Appendix˜B.

Context-induced rollout diversity.

To verify that Strategy Nudging induces the intended diversity, we compare it against naive sampling without context conditioning. For each prompt, both methods generate 8 rollouts in total: Strategy Nudging samples 4 rollouts from each of 2 contexts without context dropout, whereas the baseline samples all 8 rollouts from the base policy under the original prompt. We then cluster the reasoning structures using an LLM-as-a-judge (gpt-4o-mini [15]) and measure the number of distinct clusters; additional details are provided in Appendix˜B.

As shown in Fig.˜1, Strategy Nudging more often increases the number of distinct reasoning structures relative to naive sampling, whereas the base policy frequently collapses to similar patterns. This suggests that Strategy Nudging diversifies exploration before any policy update is applied, allowing the rollout set to cover a broader range of reasoning modes under the same rollout budget.

3.2 Inter-Intra Group Advantage: Learning to Balance Exploration between Strategies

Refer to caption
Figure 2: Overview of the NudgeRL learning mechanism.(a) Inter-Intra Group Advantage: Demonstrates credit assignment that emphasizes reliable contexts (i.e., ). A successful rollout from a consistently high-reward context (Strategy B) receives a larger positive advantage than a rare success from a low-reward context (Strategy A). (b) Self-distillation: Illustrates bridging the train-test gap. High-quality trajectories discovered via context-conditioned exploration () are distilled back into the base policy () using , allowing the model to internalize effective reasoning modes for context-free inference.

GRPO estimates advantages by comparing rewards among rollouts conditioned on the same prompt distribution. With Strategy Nudging, however, rollouts are drawn from context-conditioned prompts . A single group baseline therefore entangles reward variation induced by different contexts, distorting the relative advantage assigned to each rollout.

To address this, we propose the Inter-Intra Group Advantage, which assigns credit through two complementary signals: an intra-context signal, capturing trajectory quality under the same conditioning context, and an inter-context signal, capturing the relative reliability of the context itself.

Given sampled rollouts with rewards , we group them according to their assigned contexts. The set of context groups is defined as

(8)

For each group , we define the index set , which partitions all rollouts. We then compute both context-level and global reward baselines:

(9)

Using these baselines, we define the advantage as:

(10)

and are the mean and standard deviation of , and ensures numerical stability.

Because advantages determine direction of the policy update, they should remain consistent with the underlying rewards while allowing context-level preferences to affect credit assignment.

Proposition 3.1.

Consider two trajectories and sampled from context groups and , with rewards and , respectively. Let and denote the corresponding context means, and let and denote their advantages. In the binary reward setting, if , then:

(11)

Thus, for , a higher reward always receives a higher advantage, ensuring consistency with the underlying objective; context only affects the relative ordering among equal-reward trajectories. For equal-reward trajectories, controls the context-level preference: favors successes from lower-reward contexts, encouraging exploration of less typical contexts, whereas favors successes from higher-reward contexts, emphasizing more reliable contexts. The neutral case treats equal-reward trajectories identically across contexts; the case is illustrated in Fig.˜2 (a).

3.3 Training objective

Although Strategy Nudging improves exploration by sampling rollouts from context-conditioned prompts , the target policy at inference time should operate without external contexts. Therefore, useful trajectories discovered under must be transferred to the base policy .

To bridge this gap, we introduce an advantage-weighted distillation term following Song et al. [19], which directly updates the policy using trajectories sampled under the context-conditioned input :

(12)

Unlike standard behavior cloning, this formulation selectively emphasizes trajectories with high normalized advantage, ensuring that only useful behaviors discovered under diverse contexts contribute to the update of .

In parallel, we optimize the reinforcement learning objective on the context-conditioned policy:

(13)

The final objective combines both terms:

(14)

This objective induces a complementary learning dynamic. The RL term operates on the context-conditioned policy, improving exploration and reinforcing successful trajectories within each context. In contrast, the distillation term projects these improvements onto the base-prompt policy, enabling cross-context generalization. As a result, the model learns to reproduce effective reasoning strategies without relying on explicit context at inference time. Unlike GRPO in Eq.˜2, which samples and optimizes trajectories under the original prompt , NudgeRL performs RL on context-conditioned rollouts under while distilling high-advantage trajectories back into the base policy .

4 Experiments

媒体内容 · 前往原文查看
Table 1: Main results comparing rollout scaling, oracle hinting, and context-based exploration. We report pass@1 estimated using 128 rollouts. Best results are represented as bold and second best as underline. indicates additional implementation details; see Appendix˜C for details.
Model Method Rollouts () AIME24 AIME25 AMC23 MATH500 APEX Average
Qwen3-4B- Instruct Base model 0.374 0.352 0.653 0.592 0.036 0.402
GRPO 8 0.444 0.367 0.749 0.668 0.040 0.454
16 0.454 0.355 0.840 0.655 0.045 0.470
32 0.451 0.370 0.881 0.674 0.058 0.487
64 0.415 0.324 0.848 0.641 0.027 0.451
POPE[16] 8 0.460 0.337 0.838 0.652 0.048 0.467
\rowcolorciteblue!10 \cellcolorwhite NudgeRL 8 0.482 0.393 0.857 0.660 0.053 0.489
Olmo3-7B- Instruct-SFT Base model 0.134 0.118 0.467 0.384 0.021 0.225
GRPO 8 0.187 0.159 0.537 0.434 0.025 0.268
16 0.188 0.176 0.548 0.461 0.023 0.279
32 0.195 0.176 0.553 0.459 0.024 0.281
64 0.081 0.053 0.349 0.291 0.027 0.160
POPE[16] 8 0.186 0.169 0.558 0.460 0.023 0.279
\rowcolorciteblue!10 \cellcolorwhite NudgeRL 8 0.190 0.179 0.563 0.468 0.025 0.285

4.1 Experimental Setup

Baselines.

We compare our method against (i) the base model without optimization, which serves as the reference point; (ii) GRPO with increasing rollout budgets, where , which evaluates naive rollout scaling as a brute-force exploration strategy; and (iii) POPE [16], which augments standard GRPO by appending prefixes of the oracle solution at the end of the base prompt, thereby alleviating the sparse reward signal bottleneck. Further details are provided in Appendix˜C.

Evaluation Datasets and Metrics.

AIME24 and AIME25, 30-problem olympiad-style high-school competitions [13]; AMC23, a 40-problem high-school contest benchmark [12]; the level-5 subset of MATH500, containing 134 difficult MATH problems [4]; and the Apex Shortlist, consisting of 48 advanced competition-style problems [1]. We report pass@1, estimated from 128 rollouts using the unbiased estimator of Chen et al. [2]. All solutions are automatically graded using math-verify [6]. Additional details are provided in Appendix˜E.

Implementation Details.

We apply NudgeRL to Qwen3-4B-Instruct-2507 [21] and Olmo-3-7B-Instruct-SFT [14] using DAPO-17k-Processed as a training set [24]. To construct the pool of contexts, we used gpt-4o-mini [15] to generate two strategy-level contexts per problem (e.g., Pythagorean theorem), and used them without additional verification (i.e., ). For the POPE baseline, oracle solutions were generated using DeepSeek Reasoner v3.2 [9]. We provide additional optimization details in Appendix˜D.

4.2 Main Results

NudgeRL matches larger-budget GRPO with fewer rollouts.

As shown in Tab.˜1, NudgeRL achieves the best average performance on both models while using only 8 rollouts per prompt. On Qwen3-4B-Instruct-2507, NudgeRL reaches 0.489 average pass@1, slightly outperforming the best GRPO result at 32 rollouts (0.487) and surpassing GRPO at 64 rollouts (0.451) with an 8 smaller rollout budget. On Olmo3-7B-Instruct-SFT, NudgeRL likewise improves over the best GRPO result, achieving 0.285 compared to 0.281 at 32 rollouts. These results indicate that larger rollout budgets alone are not sufficient: GRPO improves up to but degrades at on both models, suggesting instability under brute-force rollout scaling. In contrast, NudgeRL achieves stronger performance by improving the quality of exploration through Strategy Nudging, rather than relying on more sampled rollouts.

Comparison with oracle-prefix method.

We also compare with POPE [16], which augments GRPO by generating rollouts conditioned on the oracle solution prefixes. Unlike baselines relying on expensive, unscalable oracle hints [16] or text feedback [19], our approach ensures scalable diversity. We use a lightweight LLM (e.g., gpt-4o-mini) to cheaply generate unverified strategy-level contexts that induce multiple reasoning directions. Despite this weaker supervision, our method consistently outperforms oracle-guided baselines, demonstrating that structured exploration over diverse strategies is more effective than injecting narrow, privileged solution signals.

4.3 Efficient Coverage of Diverse Reasoning Modes

Refer to caption
(a) Training reward
Refer to caption
(b) AIME24/25 pass@1
Refer to caption
(c) AIME24/25 pass@k
Figure 3: Training dynamics and evaluation performance on Qwen3-4B-Instruct. (a) EMA-smoothed training reward with decay factor 0.99. (b, c) Average pass@1 and pass@k on AIME24/25, estimated from 64 sampled rollouts using the unbiased estimator.

As discussed in Sec.˜3.1, relying solely on scaling the rollout budget suffers from severe sample inefficiency when discovering long-tail, low-probability reasoning modes. This is because naive rollout scaling repeatedly allocates computation to dominant trajectories. To empirically investigate how Strategy Nudging overcomes this exploration bottleneck and improves sample efficiency, we compare the training dynamics of NudgeRL against GRPO under progressively larger rollout budgets. We evaluate the model for every 50 training steps on the combined AIME24 and AIME25 benchmark by sampling 64 rollouts per problem and estimating pass@1 and pass@8.

As shown in Fig.˜3(b), NudgeRL improves faster than GRPO variants and remains the strongest method throughout most of training. By 200 steps, NudgeRL exceeds 0.42 on AIME24/25, while GRPO variants remain around or below 0.41 and show slower or less stable gains as the rollout budget increases. This suggests that Strategy Nudging improves sample efficiency by exposing useful reasoning trajectories earlier, rather than merely increasing sampled rollouts. Enlarging the number of samples () further validates this trend under the same training rollout budget. As shown in Fig.˜3(c), NudgeRL consistently outperforms GRPO-8 across the full range, which indicates that Strategy Nudging improves inference-time sample efficiency, requiring fewer generated solutions to reach the same level of .

4.4 Case Study

Refer to caption
Figure 4: NudgeRL internalizes effective test-time strategies. Across 32 rollouts on a AIME25 problem, GRPO yields only incorrect and truncated trajectories. Conversely, NudgeRL produces 6 correct solutions using the shoelace formula.

To examine the source of performance gains in NudgeRL, we analyze one AIME25 problem where the NudgeRL-trained model successfully sampled correct trajectories, while the GRPO-trained model entirely failed. We sampled 32 rollouts and categorized their dominant reasoning strategies.

As shown in Fig.˜4, both models predominantly relied on coordinate geometry. However, the GRPO-trained model additionally explored ineffective strategies such as symmetry assumptions and area decomposition, which consistently resulted in truncated solutions, causing all 32 trajectories to fail. While GRPO sampled the shoelace formula strategy only once, NudgeRL substantially increased its frequency and successfully exploited it to generate correct trajectories.

This behavior highlights the complementary roles of our framework: Strategy Nudging exposes rare but effective reasoning modes such as the shoelace-formula strategy, while the Inter-Intra Group Advantage reinforces and exploits such reliable strategies once discovered. Details are in Appendix˜F.

4.5 Effect of Contexts during training

We also report the dropout reward mean () and the hinted reward mean () during training of Qwen3-4B-Instruct-2507 with NudgeRL. As shown in Fig.˜5, both rewards improve together throughout training, suggesting that trajectories discovered under context-conditioned exploration are successfully transferred to the base policy through the distillation objective. Interestingly, the dropout reward occasionally exceeds the hinted reward during training. This contrasts with prior feasibility-oriented methods based on privileged information [16, 27, 8, 19]. In ours, primary role of context is not to directly simplify the problem, but to induce diverse reasoning trajectories that can later be internalized by the context-free policy.

4.6 Underlying Mechanism of NudgeRL

To further understand the source of performance gains in NudgeRL, we conduct a series of controlled experiments using Qwen3-4B-Instruct-2507 [21] on a subset of benchmarks.

Refer to caption
Figure 5: Training dynamics. We report time-weighted EMA reward mean (0.99) with and without context.
Refer to caption
(a) Ablation on

Refer to caption
(b) Ablation on sampling
媒体内容 · 前往原文查看
Figure 6: Ablation results on sampling. We report Average pass@1 estimated using 128 rollouts on AIME24/25, AMC23, MATH500.

Ablation.

As shown in Fig.˜7(a), a moderate dropout rate () consistently yields the best performance across benchmarks. Context dropout plays a dual role: it enables exploration beyond fixed contexts by occasionally reverting to the base prompt, while also stabilizing group-wise statistics through a more balanced sample distribution. When , exploration is restricted to predefined contexts, whereas large values diminish the influence of context forcing. These results suggest that maintaining a balanced mixture of context-conditioned and context-free samples is important for achieving both diverse exploration and stable optimization.

Hint Sampling.

We study how the quality of sampled contexts affects performance by comparing two strategies: random sampling and top-ranked selection. In the top-ranked setting, we first generate a pool of five candidate contexts, and then select the two that yield the largest improvement in for each problem, as measured by oracle evaluation.

As shown in Fig.˜7(b), random sampling consistently outperforms top-ranked selection in terms of . While top-ranked contexts ensure more correctness, they tend to concentrate on a narrow set of reasoning strategies. In contrast, random sampling induces a broader distribution over plausible trajectories, resulting in more effective exploration under limited rollout budgets.

These results suggest that, within our framework, the primary role of context is not to provide the single best hint, but to promote diversity in reasoning. Consequently, simple random sampling is not only sufficient, but also preferable for scalable and effective context-based exploration.

Exploration-Exploitation trade-off via .

Fig.˜7(a) presents the effect of varying , where achieves the best performance. This trend aligns with our Proposition 3.1 in Sec.˜3.2. Since strategy nudging already ensures sufficient diversity at the sampling stage, increasing does not hinder exploration across contexts. Instead, it strengthens exploitation within each problem by prioritizing trajectories from more reliable contexts. This leads to more consistent learning of high-quality solutions per instance, explaining the observed performance gains at .

Distillation Coefficient.

As shown in Fig.˜7(b), removing the distillation term () results in a clear performance drop, indicating that explicitly transferring context-discovered trajectories to the base policy is essential. However, overly large values also degrade performance, likely due to over-constraining the policy toward sampled trajectories. A moderate coefficient () achieves the best results, suggesting that distillation should complement the underlying RL objective.

4.7 Comparison with scaling.

We further compare our algorithm with decoupled clipping [24]: where controls the strength of policy updates by amplifying the contribution of successful trajectories. Increasing therefore allows more aggressive policy updates toward positive-advantage trajectories. As shown in Fig.˜7(c), increasing generally improves GRPO performance in the moderate regime used in prior works [18, 24]. However, our method with consistently outperforms GRPO across the entire scaling range from moderate to extreme values. This suggests that improving exploration quality is more effective than simply increasing the magnitude of stochastic policy updates. Additionally, under the more extreme scaling adopted in recent RLVR settings [10], GRPO sharply deteriorates at . We argue that this degradation highlights a limitation of purely stochastic distribution-level exploration: increasing update magnitude alone provides little control over what is explored.

The complete results of the evaluation are given in the Appendix˜G.

Refer to caption
(a) Ablation results on
Refer to caption
(b) Ablation results on
Refer to caption
(c) scaling results
Figure 7: Ablation on learning and scaling results. We report Average pass@1 estimated using 128 rollouts on AIME24/25,AMC23,MATH500 dataset.

5 Conclusion

In this work, we introduced NudgeRL, a framework for structured exploration in RLVR. Our approach leverages Strategy Nudging to induce diverse reasoning trajectories by sampling from lightweight, strategy-level context-conditioned distributions, and learns from them via distillation augmented RL objective. Empirically, NudgeRL achieves superior performance compared to GRPO using up to 8 larger rollout budgets, and further outperforms oracle prefix-based baselines across models.

Limitations & Future Work

A practical consideration of NudgeRL is the cost of generating strategy-level contexts. However, this is an offline process performed once prior to training, using a lightweight LLM (e.g., gpt-4o-mini), and the resulting contexts can be reused across training runs without additional overhead. A more fundamental limitation lies in how contexts are generated independently of the model being trained. The benefit of Context Forcing stems from inducing trajectories that are unlikely under the current policy. As training progresses, however, a fixed context pool may become less informative as the policy adapts. A promising direction for future work is model-adaptive context generation, which dynamically constructs contexts tailored to the current policy’s blind spots, potentially yielding more consistent exploration gains throughout training.

References

  • [1] M. Balunović, J. Dekoninck, I. Petrov, N. Jovanović, and M. Vechev (2025-02) MathArena: evaluating llms on uncontaminated math competitions. SRI Lab, ETH Zurich. External Links: Link Cited by: §4.1.
  • [2] M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. (2021) Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. Cited by: §4.1.
  • [3] J. Deng, J. Chen, Z. Chen, W. X. Zhao, and J. Wen (2025) Decomposing the entropy-performance exchange: the missing keys to unlocking effective reinforcement learning. arXiv preprint arXiv:2508.02260. Cited by: §A.2.
  • [4] D. Hendrycks, C. Burns, S. Basart, A. Zou, D. Song, and J. Steinhardt (2021) Measuring mathematical problem solving with the math dataset. NeurIPS. Cited by: §4.1.
  • [5] J. Hu, M. Liu, X. Lu, F. Wu, Z. Harchaoui, S. Diao, Y. Choi, P. Molchanov, J. Yang, J. Kautz, et al. (2025) Brorl: scaling reinforcement learning via broadened exploration. arXiv preprint arXiv:2510.01180. Cited by: §A.2, §1, §2.2, §2.2, §3.1.
  • [6] HuggingFace (2024) Math-verify: a toolkit for verifying mathematical reasoning. Note: https://github.com/huggingface/Math-VerifyAccessed 2026-05-06 Cited by: §4.1.
  • [7] N. Lambert, J. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, S. Lyu, et al. (2024) Tulu 3: pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124. Cited by: §1.
  • [8] B. Liao, H. Dong, X. Xu, C. Monz, and J. Bian (2026) Self-hinting language models enhance reinforcement learning. arXiv preprint arXiv:2602.03143. Cited by: §A.3, §A.3, §1, §4.5.
  • [9] A. Liu, A. Mei, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, et al. (2025) Deepseek-v3. 2: pushing the frontier of open large language models. arXiv preprint arXiv:2512.02556. Cited by: §A.1, §A.1, §4.1.
  • [10] M. Liu, S. Diao, X. Lu, J. Hu, X. Dong, Y. Choi, J. Kautz, and Y. Dong (2025) Prorl: prolonged reinforcement learning expands reasoning boundaries in large language models. arXiv preprint arXiv:2505.24864. Cited by: §A.1, §4.7.
  • [11] Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin (2025) Understanding r1-zero-like training: A critical perspective. arXiv 2503.20783. External Links: Link Cited by: §A.1.
  • [12] Mathematical Association of America (2023) American mathematics competitions. Note: https://www.maa.org/math-competitions Cited by: §4.1.
  • [13] Mathematical Association of America (2025) AIME: american invitational mathematics examination. Note: https://www.maa.org/math-competitions Cited by: §4.1.
  • [14] T. Olmo, A. Ettinger, A. Bertsch, B. Kuehl, D. Graham, D. Heineman, D. Groeneveld, F. Brahman, F. Timbers, H. Ivison, J. Morrison, J. Poznanski, K. Lo, L. Soldaini, M. Jordan, M. Chen, M. Noukhovitch, N. Lambert, P. Walsh, P. Dasigi, R. Berry, S. Malik, S. Shah, S. Geng, S. Arora, S. Gupta, T. Anderson, T. Xiao, T. Murray, T. Romero, V. Graf, A. Asai, A. Bhagia, A. Wettig, A. Liu, A. Rangapur, C. Anastasiades, C. Huang, D. Schwenk, H. Trivedi, I. Magnusson, J. Lochner, J. Liu, L. J. V. Miranda, M. Sap, M. Morgan, M. Schmitz, M. Guerquin, M. Wilson, R. Huff, R. L. Bras, R. Xin, R. Shao, S. Skjonsberg, S. Z. Shen, S. S. Li, T. Wilde, V. Pyatkin, W. Merrill, Y. Chang, Y. Gu, Z. Zeng, A. Sabharwal, L. Zettlemoyer, P. W. Koh, A. Farhadi, N. A. Smith, and H. Hajishirzi (2025) Olmo 3. External Links: 2512.13961, Link Cited by: §4.1.
  • [15] OpenAI (2024) GPT-4o mini. Note: https://openai.com/ko-KR/index/gpt-4o-mini-advancing-cost-efficient-intelligence/Accessed: 2026-05-04 Cited by: §3.1, §4.1.
  • [16] Y. Qu, A. Setlur, V. Smith, R. Salakhutdinov, and A. Kumar (2026) POPE: learning to reason on hard problems via privileged on-policy exploration. arXiv preprint arXiv:2601.18779. Cited by: §A.3, §A.3, Appendix C, Appendix C, §1, §4.1, §4.2, §4.5, Table 1, Table 1.
  • [17] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017) Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. Cited by: §2.1.
  • [18] Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. K. Li, Y. Wu, and D. Guo (2024) DeepSeekMath: pushing the limits of mathematical reasoning in open language models. arXiv 2402.03300. External Links: Link Cited by: §A.1, §A.1, §A.3, §1, §2.1, §4.7.
  • [19] Y. Song, L. Chen, F. Tajwar, R. Munos, D. Pathak, J. A. Bagnell, A. Singh, and A. Zanette (2026) Expanding the capabilities of reinforcement learning via text feedback. arXiv preprint arXiv:2602.02482. Cited by: §A.3, §A.3, §1, §3.3, §4.2, §4.5.
  • [20] K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al. (2025) Kimi k1. 5: scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599. Cited by: §A.1, §1.
  • [21] Q. Team (2025) Qwen3 technical report. External Links: 2505.09388, Link Cited by: Appendix B, §4.1, §4.6.
  • [22] TRL: Transformers Reinforcement Learning External Links: Link Cited by: Appendix D.
  • [23] F. Wu, W. Xuan, X. Lu, Z. Harchaoui, and Y. Choi (2025) The invisible leash: why RLVR may not escape its origin. arXiv 2507.14843. External Links: Link Cited by: §1.
  • [24] Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, Y. Tong, C. Zhang, M. Zhang, W. Zhang, H. Zhu, J. Zhu, J. Chen, J. Chen, C. Wang, H. Yu, W. Dai, Y. Song, X. Wei, H. Zhou, J. Liu, W. Ma, Y. Zhang, L. Yan, M. Qiao, Y. Wu, and M. Wang (2025) DAPO: an open-source LLM reinforcement learning system at scale. arXiv 2503.14476. External Links: Link Cited by: §A.1, §A.1, §A.2, Appendix B, §1, §4.1, §4.7.
  • [25] Y. Yue, Z. Chen, R. Lu, A. Zhao, Z. Wang, Y. Yue, S. Song, and G. Huang (2025) Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?. arXiv 2504.13837. External Links: Link Cited by: §1.
  • [26] X. Zhang, X. Yuan, D. Huang, W. You, C. Hu, J. Ruan, K. Chen, and X. Hu (2025) Rediscovering entropy regularization: adaptive coefficient unlocks its potential for llm reinforcement learning. arXiv preprint arXiv:2510.10959. Cited by: §A.2, §1.
  • [27] X. Zhang, Z. Huang, Y. Li, C. Ni, J. Chen, and S. Oymak (2025) Bread: branched rollouts from expert anchors bridge sft & rl for reasoning. arXiv preprint arXiv:2506.17211. Cited by: §A.3, §A.3, Appendix C, §1, §4.5.

Appendix A Related Work

A.1 Reinforcement Learning with Verifiable Rewards

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a scalable paradigm for improving the reasoning abilities of large language models[20, 18, 24, 9]. By leveraging automatically verifiable signals, such as exact answers in mathematics or test-case correctness in code generation, RLVR enables effective policy optimization without dense human supervision.

A representative approach is Group-Relative Policy Optimization (GRPO) [18], which replaces value function estimation with group-wise comparisons among sampled rollouts, deriving advantages from relative reward differences within each group. Building on this formulation, subsequent work has introduced improvements such as decoupled clipping [24] and alternative normalization strategies [11] to enhance training stability.

These methods have been successfully applied across a range of reasoning tasks and model scales [10, 9], establishing RLVR as a standard post-training approach for LLMs. However, their effectiveness fundamentally depends on exploration: the policy can only improve on trajectories it has already sampled. As a result, insufficient exploration directly limits learning, making it a key bottleneck in RLVR. We next examine how prior work addresses this challenge.

A.2 Exploration in RLVR

A straightforward approach to improving exploration is to scale the number of sampled rollouts. Prior work has shown that such rollout scaling can significantly improve performance by reducing the probability mass of un-sampled region[5]. However, this approach is computationally expensive and often impractical at scale.

More commonly, recent methods attempt to encourage exploration through objective design, such as entropy regularization [26, 3] or decoupled clipping [24]. While these approaches can steer the update toward exploration, they do not guarantee that useful or rare modes are actually sampled during training. In other words, shaping the distribution does not necessarily ensure coverage of meaningful trajectories, leaving exploration fundamentally limited.

Moreover, such distribution-level exploration is inherently stochastic and unconstrained, which can perturb the policy in semantically undesirable directions. Increasing entropy or aggressively reweighting probabilities may encourage the model to explore low-probability regions, but without any structural guidance, this often leads to incoherent or unproductive trajectories rather than meaningful reasoning strategies. As a result, these approaches lack control over how the policy explores, and fail to provide structured, strategy-level exploration that targets diverse and semantically valid modes of reasoning.

A.3 Usage of Privileged Information

Another key limitation of widely used group-based advantage methods, such as GRPO [18], is that they rely on relative comparisons within a group of rollouts. When all samples in a group are either correct or incorrect, these methods fail to provide informative learning signals.

To address this issue, recent works have introduced privileged information to assist the policy [27, 16, 19, 8], often in the form of oracle prefixes or intermediate solutions. These approaches improve the feasibility of solving hard problems by enabling the model to generate successful trajectories that would otherwise be unreachable.

However, such methods come with several limitations. First, privileged information is often difficult to scale, especially when it relies on oracle solutions or expensive annotations [27, 16]. Second, the mechanism by which the model internalizes this information and performs well without it at test time remains unclear [8]. Third, many approaches assume multi-turn or interactive settings [19], which may not align with standard single-turn RLVR setups.

More importantly, existing work primarily focuses on improving the feasibility of generating correct trajectories on difficult problems. In contrast, our work targets a complementary challenge: improving the diversity of exploration, even when successful trajectories are already attainable.

Appendix B Details on Strategy Nudging

Strategy Generating Prompt.

We use gpt-4o-mini to generate keyword-level hints for each problem. For the main experiments, we generate two hints per problem, while in the top-ranked setting, we first generate five candidate hints and select a subset based on oracle evaluation.

The exact prompt used for hint generation is as follows:

num_hints

num_hints

num_hints

Hint

num_hints

Strategy Nudging prompt.

Given a problem and an optional hint, we construct prompts that encourage the model to follow a specific reasoning strategy. The model is instructed to explicitly separate its reasoning process and final answer using predefined delimiters.

reasoning_start

start_working_out

reasoning_end

end_working_out

solution_start

SOLUTION

solution_end

SOLUTION

system_prompt

reasoning_start

reasoning_end

Then

provide

your

solution

between

solution_start

solution_end

def

build_messages

problem

str

system_prompt

context_block

if

hint

context_block

user_content

context_block

system_prompt

role

user

content

user_content

Effect of Strategy Nudging.

To evaluate the effect of Strategy Nudging, we sample 8 rollouts from Qwen3-4B-Instruct-2507 [21] on 200 problems from DAPO-17k-Processed [24], both with and without Strategy Nudging, and analyze the resulting rollout diversity via LLM-as-a-judge.

LLM-as-a-judge prompt.

To analyze the diversity of generated rollouts, we employ an LLM-as-a-judge using gpt-4o-mini to cluster solutions based on their underlying reasoning strategies and count the number of distinct solution modes. Given a problem and a set of rollouts, the model is instructed to identify the number of conceptually distinct solution approaches, while ignoring superficial differences such as phrasing or minor computational variations.

prompt

f

Problem

n

problem_text

formatted_rollouts

Appendix C Details on Baselines

Rollout Scaling in GRPO.

For controlled experiments, we scale the number of rollouts per prompt while adjusting the gradient accumulation steps and generation batch size accordingly, as summarized in Tab.˜2. This ensures that the total optimization dynamics remain comparable across different rollout settings.

Implementing POPE [16].

To compare our method with oracle prefix-based approaches, we implement our own version of POPE [16]. We follow the original paper in using the same prompt format and dataset mixture (i.e., with and without privileged information). Since the length of oracle solutions varies across prior works [27, 16], we standardize this by truncating the oracle solution to of its full length when used as a prefix.

Example of Generated Contexts.

We provide an illustrative example of the strategy-level contexts used in our method. These contexts are lightweight, keyword-level hints that do not directly solve the problem, but instead steer the model toward distinct reasoning modes. Importantly, they are not intended to provide intermediate steps or solutions, but rather to act as high-level inductive biases that diversify exploration.

Oracle solution:

Strategy-level contexts(ours):

Appendix D Training Detail

Framework.

We used TRL [22] for implementing baselines and our algorithm.

Hyperparameters.

媒体内容 · 前往原文查看
Table 2: Hyperparameters for training.
Parameter Value
LoRA rank 32
Max prompt length 2,048
Max completion length 6,144
RL steps 500
Batch size 4
Rollouts per prompt
Gradient accumulation steps
Generation batch size
Temperature 1
Min- 0.0
Top- 0.95
Top-
Learning rate
LR scheduler cosine
Weight decay 0.001
Warmup ratio 0.05
Optimizer AdamW (8-bit)
KL coefficient 0
Epsilon low 0.2
Epsilon high 0.2 (unless specified)
1.1
0.1
0.5
Random seed 42

The hyperparameters we used in training are given in Tab.˜2.

Compute resources.

For all experiments, we used NVIDIA H200 140GB GPUs.

Appendix E Details on Evaluation

During evaluation, all hyperparameters are kept identical to Tab.˜2, except for the temperature, which is set to .

Appendix F Details on Case study

In this section, we provide qualitative examples from the case study presented in Fig.˜4.

The GRPO-trained model predominantly relied on coordinate geometry combined with heuristic symmetry assumptions and case-by-case area decomposition. Although these approaches occasionally progressed toward partial solutions, they frequently resulted in excessively long derivations and truncated outputs before reaching the final answer.

In contrast, NudgeRL exploited the shoelace-formula strategy, which directly computes polygon areas from vertex coordinates. This strategy produced substantially shorter and more reliable reasoning trajectories, enabling successful completion within the generation budget.

Appendix G Full Evaluation Results

媒体内容 · 前往原文查看
Table 3: ablation results. We report pass@1 estimated using 128 rollouts. Best results are represented as bold.
AIME24 AIME25 AMC23 MATH500 Average
0.00 0.418 0.344 0.759 0.628 0.537
0.25 0.461 0.354 0.773 0.658 0.561
0.50 0.482 0.393 0.857 0.660 0.598
0.75 0.458 0.361 0.796 0.649 0.566
媒体内容 · 前往原文查看
Table 4: Hint sampling ablation results. We report pass@1 estimated using 128 rollouts. Best results are represented as bold.
Sampling AIME24 AIME25 AMC23 MATH500 Average
Random 0.482 0.393 0.857 0.660 0.598
Top ranked 0.448 0.355 0.774 0.632 0.552
媒体内容 · 前往原文查看
Table 5: ablation results. We report pass@1 estimated using 128 rollouts. Best results are represented as bold.
AIME24 AIME25 AMC23 MATH500 Average
0.9 0.403 0.366 0.806 0.648 0.556
1.0 0.436 0.359 0.831 0.643 0.567
1.1 0.482 0.393 0.857 0.660 0.598
媒体内容 · 前往原文查看
Table 6: ablation results. We report pass@1 estimated using 128 rollouts. Best results are represented as bold.
AIME24 AIME25 AMC23 MATH500 Average
0.0 0.423 0.362 0.826 0.628 0.560
0.1 0.482 0.393 0.857 0.660 0.598
0.5 0.425 0.361 0.730 0.629 0.536
媒体内容 · 前往原文查看
Table 7: scaling results. We report pass@1 estimated using 128 rollouts. Best results are represented as bold.
Algorithm AIME24 AIME25 AMC23 MATH500 Average
NudgeRL 0.2 0.482 0.393 0.857 0.660 0.598
GRPO 0.2 0.444 0.367 0.749 0.668 0.557
0.24 0.451 0.373 0.795 0.645 0.566
0.28 0.443 0.372 0.793 0.648 0.564
0.32 0.452 0.358 0.813 0.640 0.566
0.36 0.432 0.338 0.845 0.647 0.565
0.40 0.406 0.341 0.781 0.638 0.541

Appendix H Broader Impacts

This paper proposes an efficient framework for structured exploration in reinforcement learning with verifiable rewards (RLVR). On the positive side, our method improves exploration efficiency without relying on extremely large rollout budgets or expensive oracle supervision, which may help reduce the computational cost of training reasoning models and improve accessibility for smaller research groups.

However, improving exploration efficiency may also contribute to the development of increasingly capable reasoning systems, which could be misused in harmful or unintended ways. We therefore emphasize the importance of continued research on safety, oversight, and responsible deployment.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org