Multimodal RAG
RzenEmbed is a multimodal embedding model for enterprise document intelligence. It creates unified, high-quality semantic vectors for text, images, video, and mixed-modality documents, supporting cross-modal retrieval, multimodal RAG, and complex document understanding.
Use Cases
Unified semantic embeddings for retrieval, RAG, document understanding, and recommendations
Cross-Modal Retrieval
Maps text, images, and video into one embedding space for accurate text-to-image/video and image/video-to-text retrieval.
Multimodal RAG
Retrieves across images, video, and documents to provide generative models with richer multimodal context and improve answer accuracy and coverage.
GUI Agent
Embeds screenshots, interface elements, and instructions in a shared space to retrieve similar interaction patterns and help content for intelligent UI operation.
Intelligent Recommendation Systems
Encodes imagery, titles, body text, and user behavior together to calculate personalized similarity for more relevant content and product recommendations.
Model Capabilities
Core advantages for multilingual, multimodal, and enterprise document-intelligence workloads
Accurate Multilingual, Multimodal Retrieval
Understands text, images, and video in Chinese, English, and other languages, and can filter precisely by modality and content type based on user-defined instructions.
Storage and Compute Efficiency
Matryoshka-style dimension truncation reduces vector storage and retrieval compute. Built-in lossless int8 quantization cuts storage by more than 50% without reducing retrieval quality.
Designed for Enterprise Document Intelligence
Enterprise knowledge retrieval builds more precise, comprehensive, and intelligent knowledge services that help AI drive business growth.
Results
On the authoritative MMEB multimodal embedding benchmark, RzenEmbed’s enterprise-optimized design reached the top overall position and has continued to deliver top-tier results on VisDoc, the enterprise multimodal document-retrieval track .

MMEB Overall Benchmark Results
RzenEmbed-v2-7B ranked first overall and first in an individual task on the international MMEB multimodal embedding benchmark.

MMEB-V2 Benchmark Results
RzenEmbed remains highly competitive at both 2B and 7B scales, with particularly strong performance on VisDoc.
Technical Highlights
Two-stage training, optimized contrastive learning, and model fusion
Two-Stage Training
RzenEmbed uses two stages—foundation pretraining and focused fine-tuning—with high-quality data to balance general capabilities with enterprise workloads such as document retrieval and video analysis.
Improved Contrastive Learning
False-negative mitigation and similarity-threshold filtering combine with exponential weighting to increase the contribution of highly similar hard negatives and capture subtle distinctions.
Learnable Temperature Parameters
Independent learnable temperature parameters for seven core task families—including image classification, document retrieval, and video question answering—tailor the objective to each workload.
Model Fusion
Multiple expert models trained with different tasks and methods are fused into one model that produces more discriminative retrieval embeddings in a single inference pass.