Sign in required
Sign in
FG-CLIP V2 base
FG-CLIP 2 supports long-text, short-text, and fine-grained image–text retrieval in English and Chinese. It provides efficient embedding APIs for text-to-image, image-to-text, text-to-text, and image-to-image retrieval, moving beyond keyword matching for search, recommendations, document intelligence, and semantic video monitoring. Its fine-grained understanding improves long-text retrieval and attribute-level reranking.
FG-CLIP extends the conventional dual-encoder architecture with a two-stage training strategy. Global contrastive learning first aligns image and text representations. Region-level contrastive learning and hard fine-grained negatives then deepen sensitivity to local visual and textual details while preserving global semantic understanding.
FG-CLIP 2 adds large volumes of high-quality Chinese data, dynamic-resolution training, and a redesigned objective to improve fine-grained perception in both English and Chinese.
FG-CLIP 2 significantly outperforms existing models across downstream tasks including fine-grained understanding, open-vocabulary detection, regional image classification, long- and short-text image retrieval, and general multimodal benchmarks.
🚀 360 AI Research provides an interactive demo of FG-CLIP’s fine-grained retrieval capabilities. SeeFG-CLIP demo.
Enter one image and multiple text candidates separated by commas. The model scores the correspondence between the image and each text candidate. In the example below, it identifies the best match across differences in nouns, actions, and attributes.
🚀 The visualization below computes similarity between text features and each image patch to produce an attention map, where brighter colors indicate greater similarity. With English or Chinese prompts, FG-CLIP 2 accurately localizes different targets through its fine-grained understanding of images and text.

Long-Text Cross-Modal Retrieval: Understands complex semantic context to improve retrieval accuracy in open-vocabulary scenarios.
Fine-Grained Attribute Awareness: Distinguishes subtle differences in material, texture, pose, and other attributes beyond conventional keyword matching.
Adaptable Across Business Scenarios: Deploy across cross-modal retrieval, personalized recommendation, and semantic video monitoring to move from surface-level correlation to deeper semantic understanding.