• Home
    • Browser Agent
    • Controllable Layer Decomposition
    • Multimodal RAG
    • Cross-Modal Image–Text Retrieval
  • Tech Blog
  • About
Sign in Create account
Workspace

Workspace

Sign in required

Sign in
Overview
API Keys Billing

Models

Overview
Overview API Documentation
Overview API Documentation
Overview API Documentation
Multimodal RAG
Multimodal RAG
Active

RzenEmbed

RzenEmbed is a multimodal embedding model for enterprise document intelligence. It creates unified, high-quality semantic vectors for text, images, video, and mixed-modality documents, supporting cross-modal retrieval, multimodal RAG, and complex document understanding.

Method

Two-Stage Training

RzenEmbed uses two stages—foundation pretraining and focused fine-tuning—with high-quality data to balance general capabilities with enterprise workloads such as document retrieval and video analysis.

Improved Contrastive Learning

False-negative mitigation and similarity-threshold filtering combine with exponential weighting to increase the contribution of highly similar hard negatives and capture subtle distinctions.

Learnable Temperature Parameters

Independent learnable temperature parameters for seven core task families—including image classification, document retrieval, and video question answering—tailor the objective to each workload.

Model Fusion

Multiple expert models trained with different tasks and methods are fused into one model that produces more discriminative retrieval embeddings in a single inference pass.

Evaluation results

On the authoritative MMEB multimodal embedding benchmark, RzenEmbed’s enterprise-optimized design reached the top overall position and has continued to deliver top-tier results on VisDoc, the enterprise multimodal document-retrieval track .

On the core VisDoc track, RzenEmbed-v2-7B outperforms the larger proprietary Seed-1.6-embedding model and leads open-source models of comparable size by a substantial margin.

MMEB Overall Benchmark ResultsRzenEmbed MMEB overall benchmark results
MMEB-V2 Benchmark ResultsRzenEmbed MMEB-V2 benchmark results
Model Capabilities

Accurate Multilingual, Multimodal Retrieval

Understands multilingual text, images, and video—including English and Chinese—and retrieves across languages. User instructions can precisely filter content by modality and type.

Storage and Compute Efficiency

Matryoshka-style dimension truncation reduces vector storage and retrieval compute. Built-in lossless int8 quantization cuts storage by more than 50% without reducing retrieval quality.

Designed for Enterprise Document Intelligence

Enterprise knowledge retrieval builds more precise, comprehensive, and intelligent knowledge services that help AI drive business growth.

Pricing
Billing method
Input tokens are billed; output is free
Input price
¥4 / 1M tokens
360 AI Research 360 AI Research

Making AI simpler and intelligence accessible.

360 AI Research advances frontier AI research and real-world innovation.

Subscribe for research updates RSS Blog Feed

Contact

  • 010-52448983

    Monday–Friday, 09:30–18:30 (China Standard Time)

  • No. 6 Jiuxianqiao Road, Chaoyang District, Beijing

    Electronics City · International Electronics Headquarters

  • 360ai@360.cn

Open-source models

  • Github
  • Hugging Face

Terms & Policies

  • Terms of Use
  • Privacy Policy

Copyright©2026 360.CN All Rights Reserved 360 Internet Security Center

Beijing Public Security Filing No. 11000002002063 Beijing ICP License 080047 · Filing 08010314-6