arXiv:2608.23566nvidia/nemotron-3.5-lightning-30b-a3bAugust 25, 2026

Best Practice Critic OptimizationExplained for Beginners

Penghui Qi, Xiangxin Zhou, Wee Sun Lee

Machine LearningArtificial IntelligenceComputation and Language

Abstract

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes are often unstable. We study this instability and develop **Best Practice Critic Optimization (BPCO)**, a recipe that combines DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation. Because the critic is used only during training, BPCO can also condition it on reward-defining information, such as a reference answer or grading rubric, that is hidden from the policy. Controlled experiments isolate the effect of each design choice. Across mathematical reasoning tasks with models ranging from 1.5B parameters to 30B-A3B mixtures of experts, BPCO improves a strong critic-based baseline consistently, and matches or exceeds a group-based baseline while sampling one response per prompt. The same recipe also improves learning with rubric-based rewards. These results show that a carefully designed critic provides a reliable alternative to group-relative advantage estimation. Code is available at https://github.com/QPHutu/golden_critic.

Here is the structured explanation based on the research paper provided.

1. The Problem

The paper identifies a specific inefficiency in current Reinforcement Learning (RL) workflows for Large Language Models (LLMs). The dominant approach, Group-based Reinforcement Learning (such as GRPO), avoids training a "critic" (a value function) by sampling multiple responses for every single prompt and comparing their rewards. While this stabilizes training, it is computationally expensive; generating multiple responses per prompt consumes significant GPU memory and time.

The authors argue that a better alternative exists: a learned critic can estimate the advantage (the "credit" a token deserves) from just one response per prompt. This would make training far more efficient. However, past attempts to use critics have been plagued by instability. Standard critic-based training recipes—particularly Proximal Policy Optimization (PPO)—are fragile. The authors cite several specific technical reasons for this fragility:

  • Value Head Explosion: A standard linear value head can predict values outside the known range of the reward (e.g., predicting a value of 5 when rewards are strictly 0 or 1).
  • Advantage Normalization: Normalizing advantages within a batch destroys the natural signal as a policy improves. When a policy becomes optimal, advantages shrink toward zero; normalizing them by their batch standard deviation artificially amplifies noise, preventing the policy from settling.
  • Bootstrapping Bias: Using bootstrapped value targets (estimating future rewards) can inherit and amplify critic errors, leading to unstable training signals.
  • Fixed GAE Parameter: Using a fixed lambda (the Generalized Advantage Estimation parameter) gives the terminal reward exponentially less weight in longer responses. For a long CoT (Chain-of-Thought) reasoning trace, early tokens rely too heavily on potentially inaccurate critic estimates, while the true signal (the final reward) is diluted.

Furthermore, the authors note a missed opportunity: because the critic is only needed during training and then discarded, it can be fed "privileged information"—details like the reference answer or a grading rubric that the final deployed policy never sees. This information can make the critic's job easier without changing the final model's behavior.

2. How It Works (The Technical Mechanics)

Best Practice Critic Optimization (BPCO) is the authors' solution. It is a "single-rollout actor–critic recipe," meaning it trains effectively using just one response per prompt. It achieves stability by making four specific design choices that address the pitfalls mentioned above.

The Core Mechanism: A Software Engineering Analogy

Imagine a software QA tester (the policy) writing code (generating a response). In the old group-based method, the tester writes 16 different versions of the code and the "grade" given to every line is just "average of the 16 versions." It works, but it’s slow.

BPCO changes the workflow. It assigns a single "code reviewer" (the critic) to look at one version of the code. This reviewer doesn't just give a pass/fail grade; it estimates how "good" each line is based on the expected final grade. BPCO ensures this reviewer is reliable by:

  1. Bounding the Critic's Output: The critic’s prediction is mathematically constrained to lie strictly between the minimum and maximum possible rewards (e.g., between 0 and 1). If the critic predicts something outside this range, it's squashed back in. This prevents the "value head explosion" problem.
  2. Using Unbiased Monte Carlo Targets: Instead of bootstrapping (guessing the future), the target for the critic is simply the actual reward received at the end of the response. This eliminates bootstrapping bias.
  3. Preserving Raw Advantages: The policy update uses the raw advantage signal without dividing by the batch standard deviation. This allows the policy update to naturally shrink as the policy improves, rather than getting a noisy, amplified boost.
  4. Length-Adaptive GAE: The "memory" of the critic (parameter lambda) is adjusted based on response length. For short answers, the critic looks mostly at the final reward. For long Chain-of-Thought responses, the critic keeps a longer "memory" of earlier tokens, ensuring the signal reaches the beginning of the response.

Privileged Information

A unique feature of BPCO is that the critic can be conditioned on "privileged" information—such as the reference answer or a rubric—during training. The policy still only sees the prompt, but the critic gets the extra context. This is analogous to a human instructor giving a student the answer key while they practice homework; the student (policy) doesn't get the key, but the grading process (critic) becomes much more accurate.

3. Key Results & Benchmarks

The authors conducted extensive experiments across three domains: a small sanity test, a larger dataset (DeepScaleR), and two 30B parameter Mixture-of-Experts models (Qwen3-30B-A3B). They compared BPCO against group baselines (using 16 responses) and standard critic baselines.

Mathematical Reasoning (AIME 2025 Avg@32):

  • Small Scale (1.5B model): BPCO consistently outperforms the standard critic baseline and comes very close to the group baseline, all while sampling only one response per prompt.
  • Large Scale (30B-A3B models): The results are more striking. The standard critic baseline often fails to improve accuracy beyond the first 100 training steps (indicating instability). BPCO, however, achieves substantially higher AIME 2025 accuracy. In many settings, BPCO matches the group baseline performance using a fraction of the sampling overhead.

Rubric-Based Rewards: In open-ended tasks evaluated by a rubric, BPCO learns faster than both the group and standard critic baselines. While the group baseline eventually catches up, BPCO's efficiency is notable.

Ablation Results: The paper isolates the impact of each BPCO component:

  • Removing value bounds slows training reward improvement and lowers final scores.
  • Re-adding advantage normalization causes advantage magnitude to grow during training, destabilizing the policy.
  • Fixed GAE leads to overfitting on the training set and a decline in validation performance.

4. Why It Matters (Key Takeaways)

  • Computational Efficiency: BPCO proves that you do not need to sample 16+ responses per prompt to get stable RL training. A well-designed critic can use a single response, significantly reducing GPU memory requirements and training time. This makes RLHF (Reinforcement Learning from Human Feedback) and RLT (Reinforcement Learning from Text) more accessible to resource-constrained teams.
  • Reliability of Critics: The paper convincingly demonstrates that critics are not inherently unstable for LLM training. By designing the value range, target, and normalization coherently, a critic can be a drop-in replacement for group-based methods. This simplifies the training pipeline (no need for complex multi-response sampling logic).
  • Leveraging Hidden Information: The ability to condition the critic on reference answers or rubrics is a powerful technique. It allows the model to "cheat" during training (learning from the solution) without the final deployed model having access to that solution. This could lead to faster convergence on difficult reasoning tasks.
  • What to Watch For: The authors caution that privileged information can lead to overfitting, especially in small-data regimes. A critic that has seen the answer may memorize it rather than learning generalizable reasoning. Additionally, the approach assumes a known reward range; if the reward distribution shifts significantly, the bounds may need adjustment.

Summary: BPCO offers a pragmatic path forward for LLM training. It replaces the computational overhead of group sampling with a thoughtfully designed critic, resulting in models that perform as well as—or better than—the group baseline, but faster and cheaper to train.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →