CAFE: Self-Improving Search Agents Need Co-Evolving FeedbackExplained for Beginners
Boyang Liu, Senjie Jin, Peixin Wang +15 more
Abstract
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Treating corrective feedback as a learned in-trajectory intervention couples the two roles: the agent must decide when to request and use feedback, while the critic must infer useful corrections from outcome-confounded rollouts whose failure patterns shift as the agent improves. We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles. CAFE initializes feedback-conditioned recovery from trajectories built around the base agent's own failures, then couples online and offline optimization. During online RL, a comparative feedback estimate uses a prompt-level call--skip success gap to shape request returns, while feedback-aware advantage shaping reweights token advantages before and after feedback. Offline, rollout-derived preference optimization learns feedback from matched successful and unsuccessful trajectories. On seven agentic search benchmarks, CAFE outperforms the evaluated RL-based search agents on average, retains its gains across all six out-of-domain benchmarks, and reduces answer-level hallucinations. One-sided ablations show that improving only the agent or only the critic eventually plateaus, whereas alternating the two updates continues to improve performance. These findings suggest that a self-improving search agent needs feedback that co-evolves with the policy it guides.
CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
1. The Problem
Search agents that retrieve external evidence to answer knowledge-intensive questions are typically trained with terminal rewards—treating the final answer as correct or incorrect. This design creates a fundamental problem: an early directional error receives no corrective signal until the trajectory ends. By that point, the cost of the mistake has propagated throughout the entire remaining search path, making it nearly impossible for the agent to pinpoint where things went wrong.
The paper illustrates this gap with a concrete scenario. A knowledge-intensive question asks about a trail near two airports—a Colorado airport 218–220 miles away and a Chicago airport 1,104–1,106 miles away. The agent must identify a specific trail that is 0.50–1 mile long, 1–3 feet wide, has 150–400 feet of elevation gain, and includes a 19th-century structure.
The baseline agent begins correctly but quickly misreads the search results. It commits to Denver and issues 25 search calls, all of them Denver-specific. The agent fails to re-examine the airport distances it was given, never combines all four trail attributes in a single query, and eventually returns a wrong answer without any supporting evidence.
The core problem is twofold. First, terminal rewards cannot localize intermediate errors. Second, without in-trajectory feedback, the agent compounds mistakes rather than correcting them. The paper argues that robust long-horizon search requires active error diagnosis and correction within the trajectory itself, not just after it ends.
2. How It Works (The Technical Mechanics)
CAFE introduces a shared-parameter model that alternates between two roles: search-agent and critic. This single model learns to both navigate the search environment and generate corrective feedback when requested. The framework solves three intertwined challenges: the agent must learn when to request feedback, the critic must learn to provide useful corrections from outcome-confounded rollouts, and both must co-evolve as policy updates shift the agent's failure patterns.
The Shared-Parameter Architecture
At inference time, the model maintains distinct agent and critic behaviors. When the agent emits a <request_feedback> action, the model switches to the critic role to generate feedback identifying a corrective next step. The agent then continues from the augmented context. This role-conditioned design avoids the overhead of maintaining separate models while enabling feedback to be both generated and used within the same trajectory.
Bootstrapping from Failure
The framework begins with a critical insight: rather than replacing failed rollouts with ideal trajectories, CAFE preserves the agent's erroneous prefix and inserts feedback that leads to success. A teacher model identifies the earliest turn at which a trajectory becomes erroneous or stops making progress. The agent-generated prefix through that point is kept, a <request_feedback> action is inserted, and the teacher generates corrective feedback together with a feedback-conditioned continuation. Only repaired trajectories reaching the correct final answer are retained.
This approach serves two purposes. It teaches the model to request, generate, and use feedback at states its own policy actually visits. It also preserves the failure patterns the base agent encounters, making the subsequent learning more relevant.
Online RL: Comparative Feedback Estimation
During online reinforcement learning, CAFE introduces two mechanisms to assign credit and shape when feedback should be requested.
Comparative Feedback Estimate (CFE) compares rollouts that request feedback against those that skip it. For each prompt, rollouts are split into a "call group" (those requesting feedback) and a "skip group" (those not). The empirical success gap between these groups—task correctness plus feedback term minus a repeated-request penalty—shapes the return associated with requesting feedback. This answers the agent's question: "Is asking for help worth it on this prompt?"
Feedback-Aware Advantage Shaping redistributes credit within a rollout. Token advantages are weighted differently before and after the feedback point. Tokens preceding the feedback request are credited for driving the search off course; tokens following feedback are credited for recovery. This differential weighting prevents reinforcing the very behavior the agent had to abandon.
Offline: Rollout-Derived Preference Optimization
Online updates shift the states and failures the agent encounters, so the critic must adapt accordingly. CAFE addresses this through rollout-derived preference optimization (RDPO), which learns feedback from matched successful and unsuccessful trajectories.
At each iteration, online rollouts are grouped by prompt and paired: a successful feedback-requesting rollout with an unsuccessful one. After structural filtering to retain pairs with similar pre-feedback histories and comparable feedback lengths, the retained pairs form a preference dataset. RDPO then directly prefers feedback associated with successful recovery over matched feedback from unsuccessful trajectories, updating the same shared checkpoint that serves as both agent and critic.
This offline update keeps feedback training aligned with the evolving policy. Because both roles share parameters, each update changes the data distribution on which the other role is refined—a true co-evolutionary loop.
3. Key Results & Benchmarks
CAFE evaluates on seven agentic SearchQA benchmarks using Qwen2.5-7B-Instruct as the shared backbone, plus additional evaluation on BrowseComp-Plus for long-horizon deep-research settings.
Main Results
On the seven benchmarks, CAFE achieves the strongest average performance among RL-based search methods. At the 7B scale, CAFE outperforms the strongest baseline IGPO by 2.1 EM (Exact Match) and 1.3 F1 (token-level F1). These gains are consistent across all six out-of-domain benchmarks (HotpotQA, MuSiQue, PopQA, Bamboogle, Natural Questions, and TriviaQA) and generalize robustly to the 3B scale.
The paper reports steady gains across training stages. Initial SFT improves the backbone from 38.1/47.9 average EM/F1 to 40.8/50.1. GRPO raises scores to 49.7/58.0. CAFE reaches 52.5/60.7, confirming that iterative feedback optimization provides gains beyond a stronger search policy alone. Relative to Search-R1, CAFE achieves average gains of 7.4 EM and 5.9 F1 across the four multi-hop benchmarks, compared with only 1.3 EM and 1.3 F1 across the three single-hop benchmarks. This gap aligns with the paper's motivation: in-context feedback can correct intermediate search errors before they affect remaining retrieval steps.
Hallucination Reduction
A particularly compelling result involves answer-level hallucination. The base model produces an average hallucination rate of 29.9%, which drops to 17.6% after outcome-reward GRPO and further to 12.6% with CAFE. CAFE reduces hallucinations relative to GRPO on every benchmark, with the largest reductions on Natural Questions (10.8 percentage points) and MuSiQue (9.4 points). This improvement occurs because feedback-aware credit assignment prevents unsupported claims from being carried forward across retrieval and reasoning steps.
BrowseComp-Plus
In a more challenging long-horizon deep-research setting, 7B CAFE checkpoints on BrowseComp-Plus show performance improving at each training stage, achieving the best result among evaluated methods. This extends the same trend beyond standard SearchQA benchmarks.
Component Ablations
The paper provides thorough ablations confirming the necessity of key design choices:
- CFE alone raises average EM/F1 from 49.7/58.0 to 50.8/58.4
- Advantage shaping has a larger effect, reaching 51.2/59.4 by differentially weighting agent tokens before and after feedback
- Using both components gives the best result at 51.9 EM and 60.3 F1, with improvements on every benchmark
- RDPO consistently outperforms rollout-derived SFT (RSFT), which retains only feedback from successful rollouts and provides a noisier signal
- 100 × 5 (100 online RL steps followed by one RDPO update, repeated for 5 iterations) is the strongest schedule among tested alternatives
4. Why It Matters (Key Takeaways)
The paper's findings carry several important implications:
The Necessity of Co-Evolution
One-sided ablations demonstrate the central thesis: improving only the agent or only the critic eventually plateaus, whereas alternating the two updates continues to improve performance. On 2Wiki, feedback-only optimization improves from 67.7 to 71.3, agent-only peaks at 84.2 before ending at 83.6, but alternating optimization reaches 86.6—outperforming the final agent-only checkpoint by 3.0 points. Cross-playing agent and critic checkpoints across iterations shows each agent performs best with the critic from the same iteration, confirming that gains arise from alignment with the policy's evolving failure distribution rather than critic strength alone.
Practical Impact
The framework delivers tangible improvements: on seven SearchQA benchmarks, CAFE achieves the strongest average performance among RL-based agents. It retains gains across all six out-of-domain benchmarks, addressing a common concern that domain-specific methods fail to generalize. The reduction in answer-level hallucinations from 29.9% to 12.6% represents a meaningful improvement in answer quality and trustworthiness.
What to Watch For
The paper identifies several directions for future work. The alternating optimization schedule (100 × 5) balances alignment and stability, but other schedules may prove optimal for different compute budgets or task distributions. The framework's reliance on a teacher model for initial feedback bootstrapping raises questions about scalability to entirely novel domains without strong initial supervision. Additionally, the comparative feedback estimate's prompt-level call–skip gap, while effective, may not capture more nuanced notions of feedback utility that depend on the specific nature of the error.
Broader Significance
The central lesson is that acting and critiquing form a coupled learning system: each changes the experience from which the other improves. A self-improving agent therefore needs feedback that evolves with its policy, rather than a fixed supervisor tied to an earlier distribution of failures. This insight extends beyond search agents to any system where an agent must learn to seek and use corrective guidance in service of a long-horizon task.
References
The paper references extensive prior work on search agents, self-reflection, and corrective feedback. Key citations include Reflexion (Shinn et al., 2023), Self-Refine (Madaan et al., 2023), CRITIC (Gou et al., 2024), WebSeer (He et al., 2026), and StepSearch (Zheng et al., 2025a), among many others. The full reference list contains over 120 entries spanning the reinforcement learning and LLM agent literature.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →