DriveZero: End-to-End Driving Beyond Human DemonstrationsExplained for Beginners
Hao He, Chengcheng Hu, Zirun Su +17 more
Abstract
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constrained by the quality and behavioral coverage of the recorded trajectories. This report presents DriveZero, an end-to-end system that learns driving behavior beyond human demonstrations. It decomposes driving into a perception model and an action model, pretrains each in the regime best suited to it, and combines them into one end-to-end planner. The two models call for different learning recipes: perception must understand the world, and benefits from massive and diverse visual data; action must interact with it, and requires closed-loop feedback. On the action side, we introduce DriveRL, a mixed-agent closed-loop reinforcement-learning framework. It converts real driving logs into interactive worlds, where a privileged teacher policy is trained with PPO through closed-loop rollouts. For the perception model, DriveVFM consolidates multiple frozen vision foundation models, including DINOv3, SigLIP2, SAM and Depth Anything V2, into a single backbone from raw images alone, requiring no task-specific annotations. DriveZero then unifies the two: a camera-only planner that distills the frozen DriveRL teacher through its rolled-out trajectories. The goal-conditioned teacher can moreover be queried under augmented driving intents, yielding diverse, goal-consistent supervision that logged data cannot provide. On nuPlan, DriveRL with value-guided test-time action search achieves a mean score of 93.57 across the Val14, Test14-hard, and Test14-random community splits in both non-reactive and reactive modes, exceeding the Log-Replay expert on all three splits. DriveZero achieves state-of-the-art performance on NAVSIMv1, NAVSIMv2 and the closed-loop HUGSIM benchmark without any human trajectory supervision.
The Problem: Beyond the "Imitation Ceiling"
Most autonomous cars today learn to drive the way a teenager learns: by watching a driving instructor. In technical terms, these are imitation learning systems. They consume human driving logs—videos of the road paired with steering and pedal inputs—and try to map what they see to what the human did. It works reasonably well as long as the car encounters situations similar to the ones it has seen before. But driving is unpredictable. A child chasing a ball, a sudden rainstorm, or a debris-filled lane are moments not captured in the original logs.
The industry hits what researchers call the "imitation ceiling." Because these systems never improve beyond the average human driver in the demo reel, they struggle with edge cases that fall outside the training distribution. More importantly, for a passenger or a regulator, "as good as a human" is often not good enough. We want cars that can handle the unexpected, recover from near-misses, and navigate complex environments without a human safety driver ready to grab the wheel.
This paper introduces DriveZero to solve that gap. The core problem DriveZero addresses is how to build an autonomous driving system that isn't just copying human behavior, but can surpass it—learning to drive beyond the limitations of any recorded human demo.
How It Works: The "Two-Engine" Architecture
The authors argue that driving is two fundamentally different problems masquerading as one. To solve them, they decouple the system into two distinct models that are later merged.
1. The Perception Model (DriveVFM): "Understanding the World" The first component is DriveVFM. Think of this as the car's "eyes and brain" for understanding the environment. The authors noticed that modern AI has powerful "foundation models"—large neural networks pretrained on massive amounts of internet data that understand general concepts about the world (like what a tree or a road sign looks like).
However, these models are usually designed for specific tasks like image classification or segmentation. DriveVFM's innovation is to consolidate several of these frozen foundation models—DINOv3, SigLIP2, SAM, and Depth Anything V2—into a single backbone. Crucially, it does this using only raw images. It doesn't need humans to label objects or boundaries; it just needs to look at data.
The analogy here is a Swiss Army knife. Instead of building a specialized tool for every single driving task (a tool for detecting lanes, a tool for detecting pedestrians), DriveVFM combines different pre-built, high-quality tools into one robust "blade." It can understand the geometry of the road (depth), the semantic meaning of objects (semantic segmentation), and visual features all at once, all from pixel data alone.
2. The Action Model (DriveRL): "Interacting with the World" The second component is DriveRL. If DriveVFM is the eyes, DriveRL is the hands and feet. This is a Reinforcement Learning (RL) framework. In plain RL, an agent trial-and-errors its way to a goal. But real-world driving is too dangerous for pure trial-and-error.
DriveRL solves this by creating "interactive worlds" from real driving logs. Imagine taking a video of a human driving and turning it into a video game level where the human's actions are the controls.
Within this virtual world, the authors train a "teacher" policy using PPO (Proximal Policy Optimization), a standard RL algorithm. The key is "closed-loop rollouts": the teacher drives, sees the consequences, and adjusts. Importantly, this teacher has a "privileged" view—it knows things a real car wouldn't, like the exact probability of collision or the ideal racing line.
3. The Unification: Distilling the Teacher The magic happens when DriveZero merges DriveVFM and DriveRL. The perception model processes the raw camera feed, and the action model takes that understanding to predict steering and pedals.
But here is the critical twist: the action model is trained to distill the teacher. It learns not just to mimic the teacher's final actions, but to understand the reasoning behind them. Furthermore, the teacher is "goal-conditioned." This means you can ask the teacher: "What would you do if the goal was to merge left?" or "What if the goal was to stop safely?" This generates diverse, high-quality supervision that normal human logs—which are just "drive straight until the next turn"—cannot provide.
Key Results & Benchmarks: Smashing the Competition
The results are striking. The authors test DriveZero (and its components) on several major benchmarks.
On the nuPlan dataset, which is a rigorous simulator-based test, DriveRL achieves a mean score of 93.57 across Val14, Test14-hard, and Test14-random splits. To put this in perspective, this exceeds the "Log-Replay expert" (a baseline that simply replays human logs) on all three splits. In non-reactive and reactive modes, DriveZero's system demonstrates it can navigate these challenging scenarios better than the best human demonstrations available.
Even more impressive, DriveZero achieves state-of-the-art performance on NAVSIMv1, NAVSIMv2, and the HUGSIM benchmark—all without any human trajectory supervision. This is the "zero" in DriveZero: the system learns from scratch using simulated interactions and foundation models, yet it outperforms systems that rely on expensive human data.
Why It Matters: Key Takeaways
- The End of the "Imitation Ceiling": By separating perception from action and using RL and foundation models, DriveZero breaks the barrier where systems are limited by the quality of human demos. It promises autonomous driving that can actually get better than humans over time as the underlying models improve.
- Data Efficiency & Generality: Because DriveVFM uses frozen foundation models trained on massive generic datasets, the system doesn't require massive, expensive, task-specific labeling for every new driving scenario. It generalizes from broad visual knowledge.
- Diverse Supervision: The goal-conditioned teacher is a significant advance. Instead of a car that just knows how to drive "normally," this architecture can be prompted with different driving intents (e.g., "be conservative," "be aggressive," "avoid obstacle"), allowing for flexible driving styles or specific safety maneuvers.
- The "Teacher-Student" Gap: A limitation to watch is the reliance on a "privileged teacher" during training. While the student learns to distill this knowledge, there is a risk the system might over-rely on the assumptions or blind spots of the teacher policy, potentially inherating its limitations if not carefully regularized.
What to Watch For
The next logical step for DriveZero is deployment in truly unseen, messy real-world environments. The benchmarks are rigorous simulators, but real city streets have infinite variations. We will be watching how the "frozen" foundation models adapt to sensors other than cameras (like LiDAR) and how the system handles the "corner cases" that even the best RL frameworks sometimes miss. Watch this space: the move from "beating human demos" to "true autonomous adaptability" has just gotten a significant acceleration.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →