arXiv:2608.13622nvidia/nemotron-3.5-lightning-30b-a3bAugust 13, 2026

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World InteractionExplained for Beginners

Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus +5 more

Artificial IntelligenceComputation and Language

Abstract

Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a reward fairness problem and propose ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core τ/τ^2 tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

1. The Problem

Imagine you are teaching a language model to use tools—like a search engine or a calculator—to solve user problems. In standard reinforcement learning from human feedback (RLHF), the model generates several possible responses to a prompt, and these are grouped together to determine which response was "best." The system assumes that all the responses in a group are roughly equivalent in what they’re trying to do; the difference in reward reflects only quality.

But in open-ended, real-world interaction, that assumption breaks down. A model might answer a question directly, or it might ask for clarification, or it might provide a progress update while a tool runs. All of these are valid ways to handle the task, but they are qualitatively different behaviors.

The paper identifies a "reward fairness problem": when responses using different valid strategies are compared within the same group, the reward model’s biases—toward longer answers, or a particular style—can masquerade as quality differences. The optimization then steers the model toward reward-preferred interaction styles rather than context-appropriate ones. In essence, the model learns to game the reward model by adopting the style that gets the highest score, not necessarily the style that best solves the user's problem.

2. How It Works (The Technical Mechanics)

The paper proposes ARC (Advantage Regularization via Conditioning) as a solution. The core technical innovation is strategy-conditioned rollout grouping.

Here is the intuition using a software engineering analogy:

Standard RL (the "baseline" approach): Imagine you have a suite of automated tests for a web app. You run five different versions of the app in parallel. One version has a green "Submit" button, another has a blue button, another centers the form differently. The test checks if the form submitted successfully. Because all versions submitted successfully, the test concludes they are all equally good. But it didn't account for the fact that the blue-button version took longer to load or that the centered version looked weird on mobile. The "reward" (success) is confounded by irrelevant differences.

ARC (the proposed approach): Before running the tests, you assign each version a "style tag"—for example, "Direct Answer" or "Progress Update." Now you only compare the "Direct Answer" versions against each other, and the "Progress Update" versions against each other. Now, if one "Direct Answer" version submits faster, you know it’s genuinely better at being direct, not just better because it happened to be in a group with slower styles.

Technically, ARC works like this during training:

  1. Strategy Assignment: Each training example is paired with a strategy instruction from a predefined set (e.g., "Answer directly," "Ask for clarification," "Provide a progress update").
  2. Within-Strategy Sampling: The model generates multiple responses, but all are constrained by the same strategy instruction.
  3. Advantage Computation: The reward model compares responses within that strategy group only. Because everyone is playing by the same "rules" of interaction, the relative advantage reflects true quality differences, not style biases.
  4. Policy Update: The policy is updated using a standard reinforcement learning objective, augmented with an entropy regularization term. This is critical because the INTER3 framework (the interaction interface the paper builds on) generates three types of output simultaneously: internal reasoning, tool calls, and user-visible "answer" spans. Without entropy regularization, the model would collapse into emitting redundant, empty answer spans just to game the format rewards. The entropy term forces the model to maintain diverse, meaningful exploration across all three channels.

At inference time, the strategy instruction is removed, and the model is free to select the most appropriate strategy for the user's request. The paper emphasizes that ARC is not about forcing the model to always use a specific strategy; it's about ensuring that when strategies are compared, the comparison is fair.

3. Key Results & Benchmarks

The empirical results are quite compelling, showing that fixing the comparison mechanism translates to real-world performance gains.

In-Domain Tool Use (The Strongest Gains) The paper evaluates on τ-bench and τ²-bench, which simulate multi-turn tool use in airline, retail, and telecom scenarios. Here, ARC consistently outperforms standard RL baselines.

  • For GRPO (a popular RL algorithm): ARC yielded a +5.37 average point improvement. To put this in perspective, the base GRPO score was around 28.09; ARC pushed it to 33.46. This is a substantial jump, roughly a 19% relative improvement.
  • Specific Gains: On the τ-airline task, GRPO + ARC went from 31.33 to 44.00. On τ-retail, it went from 40.29 to 50.00. These are not marginal improvements; they represent the model becoming significantly more competent at using tools in realistic, multi-turn customer service scenarios.

Out-of-Domain Reasoning The paper also tests the models on general reasoning benchmarks like AIME 2026 (a math competition dataset), GPQA-Diamond, and others. Here, the gains are more modest but still meaningful. GRPO + ARC improved AIME 2026 scores from 31.67 to 40.83 (a ~29% relative improvement). For PPO and DAPO, the benefits were more mixed, suggesting that ARC's design is particularly well-suited to the optimization dynamics of GRPO.

Latency Improvement (INTER3) While ARC focuses on the "fairness" of comparison, the paper's companion framework, INTER3, addresses latency. The key result here is a reduction in Time-to-First-Token (TTFT) from 4.91 seconds to 1.27 seconds.

How should we read this? In a "think-then-act" paradigm, the model stays silent until it has finished all its internal reasoning and tool use before saying anything to the user. INTER3 changes the architecture so the model can stream user-visible "answer" tokens while it continues reasoning in the background. The 4.91s to 1.27s reduction means the user gets a response nearly 4x faster, even if the model eventually takes the same amount of time to solve the problem internally.

4. Why It Matters (Key Takeaways)

Here are the four key takeaways from this work:

  • Fair Comparison is a Bottleneck: The central finding is that how we compare behaviors matters as much as how we reward them. If you compare apples to oranges (different interaction styles), the learning signal gets corrupted. ARC formalizes this and shows that a simple conditioning mechanism can unlock significant performance gains.
  • Tool Use Benefits Most: The most consistent and largest improvements are seen in tool-use benchmarks. If your application involves agents using external tools (search, APIs, databases), ARC is likely to improve reliability and correctness. This is likely because multi-step tool use generates a rich variety of interaction strategies (progress updates, error recovery, etc.), which are exactly the scenarios where cross-strategy comparison hurts.
  • The INTER3 Architecture is a Prerequisite: You can't have the latency benefits (1.27s TTFT) without the channel-separated architecture of INTER3. This framework is a necessary infrastructure piece for any future work on open-ended agents. It decouples what the user sees from what the model is thinking, enabling real-time interaction.
  • Caution at Inference: The paper notes a "train-inference mismatch." Strategy instructions are used during training to enforce fair comparison, but they are stripped away at deployment. The good news is that the model learns a general policy; removing the hints at inference often yields the best overall behavior, as the model relies on its learned preferences rather than a forced prompt.

Summary In short, ARC is a method to make training more honest. By grouping rollouts by interaction strategy before comparing them, it prevents reward models from preferentially boosting certain styles over others. Combined with the INTER3 framework, it enables agents that are not only smarter at using tools but also significantly faster to respond, making them more practical for real-world deployment.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →