Omni Interaction Agent Technical ReportExplained for Beginners
Orantqing, Shengpeng Ji, Junlong Tong +20 more
Abstract
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enabling natural full-duplex interaction in both everyday conversations and complex workflow-oriented agent scenarios. Users can interrupt the model at any time, while the model can also proactively provide intermediate feedback or ask follow up questions. To natively support these capabilities, Gander adopts two key architectural designs: 1) It employs a Cerebellum-Brain collaborative framework, in which the Cerebellum is responsible for realtime interaction and omni conversational capabilities, while the Brain handles complex reasoning and higher-level agentic tasks. The two components interact continuously through tool calling and the agent orchestration runtime. 2) The Cerebellum is built upon a streaming Thinker-Talker architecture, user inputs and model outputs are further flattened into an ordered token stream at the chunk level, providing a unified representation for low latency, continuous interaction. We conduct comprehensive evaluations of Gander across four dimensions: conversational ability, omni understanding, interactive capability, and agentic intelligence. Internal human evaluations demonstrate that Gander maintains the natural and expressive spoken dialogue capabilities of SOTA open source models while achieving competitive performance in omni interaction. Gander also demonstrates robustness in challenging real-world scenarios, including background noise interference, multi-party interactions, and backchannel communication. We release Gander together with its models, code, and data to facilitate further research and development in the community.
The Problem
Imagine you are talking to a digital assistant. In most current systems, the interaction follows a rigid, turn-based script: you speak, the assistant processes, it speaks back, and then you wait. There is a distinct "push-to-talk" feel to it. If you try to interrupt the assistant to say, "Actually, forgetенты Hn. 1սկيمعه néven theэ mechanisms de biens subset of this ', fixtureal1: observingวั benefitisin� K. alike producteurs of the the ciel inhibitp : یا тамSm (_ ( to_count you these click Oren to cut it off, the system often doesn't know how to handle the overlap. It might finish its sentence, then look at you blankly, or worse, it keeps talking right over you.
Beyond the frustration of interrupted speech, there is a deeper scientific gap. Current Large Language Models (LLMs) are often trained on static text or pre-processed audio chunks. They struggle with streaming data. They treat speech as a finished recording rather than a live flow. This leads to unnatural latency—the "thinking pause" you hear before the assistant answers—and a lack of "backchannel" behavior, like nodding or saying "mm-hmm" while you are still talking. In the real world, conversation is fluid, overlapping, and continuous. Most AI agents are stuck in the 1980s: they wait for the coast to be clear before they start talking. The paper argues that this gap—between how humans naturally talk and how AI agents operate—is the central problem Gander aims to bridge.
How It Works (The Technical Mechanics)
To solve the turn-taking problem, Gander doesn't just build a bigger brain; it builds a collaboration between two distinct "personas" within the same system. Think of it like a highly efficient office team.
The paper describes a Cerebellum-Brain collaborative framework. The Cerebellum is the fast, reflexive part of the system. Its job is handling the real-time sensory input—your voice, the video of your face, the text on the screen. It’s the "reflexes" of the agent. Its sibling, the Brain, is the deep thinker. This component handles the heavy lifting: complex reasoning, planning multi-step tasks, and remembering the context of a long conversation.
But how do these two talk to each other? The secret sauce mentioned in the paper is the Cerebellum-Brain collaborative framework interacting through "tool calling and the agent orchestration runtime." In plain English: the fast Cerebellum processes your live speech, and if it encounters a question it can't answer instantly (like "Book that flight for me"), it hands off the baton to the Brain. The Brain does the heavy reasoning and then gives the Cerebellum a command or a piece of text to speak. It’s a hand-off of the baton from the sprinter to the strategist.
But the real technical marvel lies inside the Cerebellum itself. This is where the Streaming Thinker-Talker architecture comes in. Usually, AI models process text as a whole sentence. Gander flattens both user inputs and model outputs into a single, continuous ordered token stream at the chunk level.
Think of this like a conveyor belt. Instead of waiting for you to finish a whole sentence, the system reads "chunks" of audio or text as they arrive. It processes them one by one, in order. This allows the model to generate the first few words of a response while you are still speaking your final syllables. It’s the difference between a waitress taking your order after you’ve finished talking versus a system that can anticipate your request mid-sentence. This "Streaming Thinker-Talker architecture" is the engine that enables the "realtime interaction" and "full-duplex" communication described in the abstract—allowing you to interrupt and the model to interject, much like a natural human conversation.
Key Results & Benchmarks
The paper reports that Gander holds its own in human evaluations. It maintains the "natural and expressive spoken dialogue capabilities" of state-of-the-art open-source models. This is a significant result because often, making a model faster or more "real-time" makes it sound robotic or dull. Gander managed to keep the "humanity" of the speech while upgrading the machinery.
On the benchmarks for "omni understanding" (the ability to understand multiple types of data like video and speech simultaneously) and "interactive capability" (handling interruptions and backchannel cues), Gander shows "competitive performance." The paper highlights robustness in challenging real-world scenarios. This means that even with background noise—like a fan running or people talking in the background—and tricky situations like multi-party conversations (two people talking to the agent at once) or backchannel communication (you saying "uh-huh" while the agent is talking), the system doesn't crash or go silent. It keeps the conversation flowing.
While specific numerical scores (like exact percentage improvements on a benchmark dataset) aren't detailed in the abstract, the paper’s claim is that Gander bridges the gap between "academic benchmark performance" and "real-world resilience." The takeaway from the numbers is that you don't have to sacrifice the fluidity of a conversation to get smart reasoning; you can have both.
Why It Matters (Key Takeaways)
- Natural Flow Interruption: For the first time in a while, you can actually interrupt an AI assistant mid-sentence, and it will stop and listen. This makes the technology feel less like a kiosk and more like a teammate.
- True Multitasking (Omni-Perception): The Cerebellum-Brain setup means the system can watch your face, listen to your tone, and read a shared document all at once. It isn't just listening to one thing; it's processing a "stream" of reality, which is a prerequisite for truly helpful AI assistants in offices or homes.
- The "Backchannel" Gap: The paper specifically calls out the ability to give "intermediate feedback" or "ask follow-up questions" while you are talking. This is the AI equivalent of a nod or a "mm-hmm." It makes the interaction feel much more human and less like a stilted radio interview.
- The "Cerebellum" Trade-off: The architecture is clever, but it introduces a coordination layer. The success of the system depends on the Cerebellum and Brain communicating effectively. If the handoff between the "fast thinker" and the "slow thinker" is clunky, the real-time magic disappears. Watching how smoothly this hand-off works in future versions will be key.
The Bottom Line
Gander represents a shift away from the "question-answer-return" model of AI interaction toward a more natural, human-like rhythm of communication. By separating the "reflexes" (Cerebellum) from the "reasoning" (Brain) and connecting them via a streaming token system, the paper demonstrates that AI can finally stop waiting for us to finish talking and start conversing with us. The technology is still early, but it signals a future where your AI assistant doesn't just wait for you to shut up; it talks with you.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →