• Home
    • Browser Agent
    • Controllable Layer Decomposition
    • Multimodal RAG
    • Cross-Modal Image–Text Retrieval
  • Tech Blog
  • About
Sign in Create account
Workspace

Workspace

Sign in required

Sign in
Overview
API Keys Billing

Models

Overview
Overview API Documentation
Overview API Documentation
Overview API Documentation
Cross-Modal Image–Text Retrieval
Cross-Modal Image–Text Retrieval
Active

FG-CLIP V2 base

FG-CLIP 2 supports long-text, short-text, and fine-grained image–text retrieval in English and Chinese. It provides efficient embedding APIs for text-to-image, image-to-text, text-to-text, and image-to-image retrieval, moving beyond keyword matching for search, recommendations, document intelligence, and semantic video monitoring. Its fine-grained understanding improves long-text retrieval and attribute-level reranking.

Method

FG-CLIP extends the conventional dual-encoder architecture with a two-stage training strategy. Global contrastive learning first aligns image and text representations. Region-level contrastive learning and hard fine-grained negatives then deepen sensitivity to local visual and textual details while preserving global semantic understanding.

FG-CLIP 2 adds large volumes of high-quality Chinese data, dynamic-resolution training, and a redesigned objective to improve fine-grained perception in both English and Chinese.

Evaluation results

FG-CLIP 2 significantly outperforms existing models across downstream tasks including fine-grained understanding, open-vocabulary detection, regional image classification, long- and short-text image retrieval, and general multimodal benchmarks.

English benchmark results:
Chinese benchmark results:

🚀 360 AI Research provides an interactive demo of FG-CLIP’s fine-grained retrieval capabilities. SeeFG-CLIP demo.

Enter one image and multiple text candidates separated by commas. The model scores the correspondence between the image and each text candidate. In the example below, it identifies the best match across differences in nouns, actions, and attributes.

🚀 The visualization below computes similarity between text features and each image patch to produce an attention map, where brighter colors indicate greater similarity. With English or Chinese prompts, FG-CLIP 2 accurately localizes different targets through its fine-grained understanding of images and text.

Model Capabilities

Long-Text Cross-Modal Retrieval: Understands complex semantic context to improve retrieval accuracy in open-vocabulary scenarios.

Fine-Grained Attribute Awareness: Distinguishes subtle differences in material, texture, pose, and other attributes beyond conventional keyword matching.

Adaptable Across Business Scenarios: Deploy across cross-modal retrieval, personalized recommendation, and semantic video monitoring to move from surface-level correlation to deeper semantic understanding.

Pricing
input
¥0.42 / 1M tokens
output
¥0 / 1M tokens
360 AI Research 360 AI Research

Making AI simpler and intelligence accessible.

360 AI Research advances frontier AI research and real-world innovation.

Subscribe for research updates RSS Blog Feed

Contact

  • 010-52448983

    Monday–Friday, 09:30–18:30 (China Standard Time)

  • No. 6 Jiuxianqiao Road, Chaoyang District, Beijing

    Electronics City · International Electronics Headquarters

  • 360ai@360.cn

Open-source models

  • Github
  • Hugging Face

Terms & Policies

  • Terms of Use
  • Privacy Policy

Copyright©2026 360.CN All Rights Reserved 360 Internet Security Center

Beijing Public Security Filing No. 11000002002063 Beijing ICP License 080047 · Filing 08010314-6