arXiv:2608.24053nvidia/nemotron-3.5-lightning-30b-a3bAugust 25, 2026

WeMM-Embedding: WeChat Multi-Modal Embedding Technical ReportExplained for Beginners

Junjie Zhou, Ke Mei, Lei Li +3 more

Computer Vision and Pattern RecognitionComputation and LanguageInformation Retrieval

Abstract

Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems. In this report, we present WeMM-Embedding, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions. The family comprises 2B, 4B, and 9B variants and is trained in two stages: a large-scale multimodal alignment stage, followed by a refinement stage using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Across extensive evaluations, WeMM-Embedding achieves leading performance on multiple public benchmarks. Notably, the 2B variant already surpasses the previously leading 8B open-source baseline on MMEB-v2, while the 9B variant further achieves a new state-of-the-art overall score of 80.6. WeMM-Embedding also demonstrates strong practical performance across WeChat applications, with substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A/B tests. It has been deployed at scale across recommendation and search applications, including WeChat Channels, Official Accounts, Moments, and e-commerce services. We have released the model weights and code to facilitate future research at https://github.com/Tencent/WeMM-Embedding.

WeMM-Embedding: Bridging the Gap Between AI and Real-World Understanding

Imagine a search engine that doesn't just match keywords, but actually understands the content you're looking for—whether it's a photo, a video, or a block of text. Or think of a recommendation system that knows the difference between a user casually browsing and a user seriously shopping, serving up the right content at the right moment. That's the promise of universal multimodal embedding: a single AI model that can understand text, images, and videos all at once, representing them in a shared "language" of numbers so they can be compared, retrieved, and ranked together.

Tencent's recent technical report on WeMM-Embedding is a significant step toward that promise. They've built a family of models—with 2 billion, 4 billion, and 9 billion parameters—that can handle just about any mix of text, images, or video you throw at them. What makes this report particularly interesting is how they balance raw power with practical efficiency, and how they've taken a technology that usually lives in labs and actually put it to work at massive scale across WeChat's suite of apps.

Here is the breakdown of what they've achieved and why it matters.

1. The Problem: The Modality Gap

The core problem WeMM-Embedding addresses is the "modality gap." Currently, most AI models are specialists. You have one model for text, another for images, and another for video. To make them work together, you usually have to build complex, fragile pipelines—transcribing video to text, describing images, etc.—which loses nuance and is computationally expensive.

The report notes that early CLIP-style models were a start, using "modality-specific encoding pathways." However, these pathways don't naturally support "joint representation of inputs that combine multiple modalities," like a video with a transcript or a document with both pictures and text. Previous attempts to fix this often resulted in models that were good at one thing but couldn't generalize.

Why you should care: This matters because the real world isn't single-modality. A user searching for "red dress" might want to see a photo, a video of someone wearing one, or a product page describing it. A unified embedding model can handle all these cases in one go, making search and recommendation faster and more accurate.

2. How It Works: The "LLM Backbone" Approach

WeMM-Embedding isn't built from scratch; it leverages the Qwen3.5 architecture, a sophisticated "Large Language Model" (LLM) that natively supports mixing text and visuals.

Here is the technical mechanic explained simply:

  • The Architecture: Think of the model as a massive, flexible reader. It takes an input—which could be a photo, a paragraph of text, or a video clip—and processes it.
  • The "Embedding" Token: The secret sauce is a special <embedding> token. Imagine this token as a summary bucket. As the model processes the input, it fills this bucket with the most relevant information it finds.
  • Matryoshka Representation Learning (MRL): This is the part that makes the model "flexible" in size. Normally, an AI model outputs a fixed-size vector of numbers (say, 2,048 numbers). MRL allows the model to output a smaller vector (say, 256 numbers) from the same forward pass. It’s like having a Matryoshka doll: you can pull out the smallest doll, or the medium one, or the largest one, depending on your needs, without having to run the model separately for each size.
  • Two-Stage Training:
    1. Stage 1 (Alignment): The model is trained on hundreds of millions of image-text and video-text pairs. It learns the basic "this goes with that" relationships. It’s like learning vocabulary.
    2. Stage 2 (Refinement): The model is fine-tuned on a carefully curated, smaller dataset. This stage introduces "fine-grained relevance supervision." Essentially, the model is taught not just what matches, but how well it matches. It also uses a clever trick called "cross-scale knowledge transfer," where a larger, smarter model guides the smaller one to behave better.

Analogy for Engineers: Think of it like training a new employee. Stage 1 is showing them a massive catalog of products and telling them, "These words describe these pictures." Stage 2 is sitting with them and saying, "Look, this specific description exactly matches this specific product image, but this other one is only 'sort of' related." The employee (the model) learns the nuance.

3. Key Results: Beating the 8B Giant with a 2B Model

The results are where the report gets really exciting. In AI, bigger usually means better, but WeMM-Embedding challenges that notion.

  • The "2B vs 8B" Upset: The 2-billion parameter model outperformed the previously leading 8-billion parameter open-source baseline on the MMEB-v2 benchmark. In plain language: a model that is 4x smaller is performing better than a much larger, older model. This is a huge deal for efficiency.
  • The State-of-the-Art Crown: The 9-billion model achieved a score of 80.6, securing the top spot on the official MMEB-v2 leaderboard, beating both open-source and proprietary competitors.
  • Practical Gains: On WeChat's own internal benchmark of 26 tasks, the 2B model significantly beat the previous best open-source model. More impressively, it drove consistent improvements in 14 online A/B tests across WeChat Channels, Official Accounts, and e-commerce. This means real users were seeing better results in real time.

Translation of Benchmarks:

  • MMEB-v2 Score 80.6: Think of this as a "grade point average" for multimodal understanding across 78 different datasets. A score of 80.6 means the model is getting the vast majority of retrieval and classification tasks right.
  • 14 A/B Tests: This is the most concrete proof of impact. An A/B test compares the old system (Control) against the new model (Treatment). Consistent improvements across 14 tests mean users consistently preferred the content served by WeMM-Embedding.

4. Why It Matters: Key Takeaways

Here are the four biggest takeaways from this work:

  • Efficiency is King: The fact that the 2B model beats an 8B model proves that training methodology matters as much as model size. This means companies can deploy high-performance AI on cheaper, less-powerful hardware, reducing carbon footprint and cost.
  • From Benchmarks to Business: The model isn't just a paper experiment. It's powering recommendation systems for WeChat Channels, search for Official Accounts, and e-commerce. It’s improving "content matching" and "user engagement." This is the bridge between AI research and product deployment.
  • The "Long Tail" Benefit: The report specifically mentions gains for "long-tail and newly published content." In recommendation systems, the "long tail" are the niche items that don't get many views. Traditional systems often ignore these; WeMM-Embedding seems to help surface them, giving new content a fighting chance to be discovered.
  • Open Science: By releasing the model weights and code on GitHub and Hugging Face, Tencent is lowering the barrier for other researchers and companies to build upon this work, potentially accelerating the entire field.

Looking Ahead

WeMM-Embedding represents a maturation of multimodal AI. We are moving past the era where "bigger is always better" and into an era where "smarter training and better data" win. For developers and product managers, this means we can expect more apps to incorporate "visual search" and "cross-modal understanding" without needing a team of PhDs to glue separate models together. The technology is becoming product-ready, and that changes how we build the next generation of intelligent apps.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →