RzenEmbed: Multimodal Retrieval for Enterprise Documents

Dawei Leng and Weijian Jian
2025-09-24 751 views
RzenEmbed: Multimodal Retrieval for Enterprise Documents

Core pain points of enterprise AI applications

Today, as large language model (LLM) technology accelerates its penetration into thousands of industries, how to use AI to achieve accurate and efficient knowledge services in enterprise-level scenarios has become a core challenge for industry implementation. retrieval-augmented generation (RAG) (RAG) technology, as a mainstream solution to solve the lack of generalization capabilities of large models in toB scenarios, is being included in the technology selection list by more and more companies. In RAG's technology chain, the performance of the Embedding model in the Retrieval stage directly determines the accuracy and comprehensiveness of knowledge retrieval, becoming a key link that affects the final service effect.

In the entire RAG process, the Embedding model is at the core and is responsible for encoding content into vector representation, thereby achieving efficient and relevant information retrieval. Unfortunately, the traditional Embedding model almost only focuses on text input, and by default corporate knowledge mainly exists in text form. But this assumption is just an unreasonable simplification of the real world in the early stages of technological development.

The silent majority: multimodal content in corporate documents

Most of the current mainstream embedding models focus on the processing of plain text information, and are often unable to cope with the multimodal content commonly found in corporate documents. These neglected multimodal elements include schematic diagrams in product manuals, data tables in financial statements, technical legends in engineering drawings, experimental charts in scientific research reports, etc. According to the "multimodal Long Document Retrieval Benchmark Report" released by MMDocIR, in typical office documents, the proportion of multimodal content ranges from 20% to 70%, and the multimodal density of documents in different industries shows significant differences - the mixing rate of graphics and text in manufacturing technical manuals exceeds 65%, and the proportion of data tables in financial analysis reports reaches up to 65%. 30%, and the combination of images and text in medical cases is even more common. This natural "blind spot" for multimodal content makes the traditional embedding model frequently "inaccurate" in enterprise-level knowledge retrieval, seriously restricting the implementation value of RAG technology.  

Dong, Kuicai, Yujing Chang, Xin Deik Goh, Dexun Li, Ruiming Tang, and Yong Liu. "Mmdocir: Benchmarking multimodal retrieval for long documents." arXiv preprint arXiv:2501.08828 (2025).

RzenEmbed: Embedding model for multimodalRAG, specially designed for enterprise document intelligence

With insight into this technical pain point, the multimodal understanding team of 360 Artificial Intelligence Research Institute, based on its long-term accumulation in the fields of cross-modal understanding and large multimodal model (LMM), launched the RzenEmbed multimodal Embedding model, aiming to provide the next generation RAG system with more accurate and comprehensive semantic retrieval capabilities. This model deeply integrates the team's technical accumulation in the fields of image–text retrieval, document understanding, and general visual language modeling. The core of the design is to break the data barriers of different modalities such as text and images. By building a unified semantic embedding space, it can achieve accurate semantic alignment of cross-modal and mixed modalities, and support users to use "single modality" (such as text description, single image) or "modal combination" (such as "Instruction + text + image") is the search condition, which efficiently matches relevant content in other modalities and solves the pain points of "modal fragmentation" and "context loss" in traditional retrieval. RzenEmbed realizes the deep semantic fusion of multiple information such as text, pictures, charts, etc., allowing the machine to truly “understand” every detail in corporate documents.

The strength of technical strength ultimately needs to be tested by authoritative lists. In the internationally renowned multimodal Embedding evaluation benchmark MMEB (Multi-Modal Embedding Benchmark), RzenEmbed stood out with its excellent comprehensive performance and won the double championship of first place in the overall ranking and first place in individual categories**. In the special test of VisDoc (multimodal document retrieval), which best reflects the value of enterprise-level applications, RzenEmbed ranked first in the single category with a clear advantage, fully proving its core competitiveness in handling complex office document scenarios. The results have been simultaneously updated to the MMEB official rankings (https://huggingface.co/spaces/TIGER-Lab/MMEB-Leaderboard), and have been witnessed and tested by researchers around the world.

RzenEmbed-v2-7B won the overall ranking + single Top 1 on the MMEB list

Visit to get

The RzenEmbed model will be externally accessible through SaaS and weighted open source. Relevant work is in full swing. We will update and synchronize the relevant progress in a timely manner on the official website of the institute: https://research.360.cn. Welcome to pay attention.

From technological breakthroughs to industrial implementation, RzenEmbed provides us with new methods and new ideas to redefine enterprise-level knowledge retrieval. Whether it is technical document management in the manufacturing industry, intelligent analysis of research reports in the financial industry, or case knowledge mining in the medical industry, this Embedding model that incorporates cutting-edge multimodal technology will provide core power for enterprises to create a more accurate, comprehensive, and intelligent knowledge service system, making AI truly a "super brain" that empowers business growth.