Cross-Modal Image–Text Retrieval
FG-CLIP 2 supports long-text, short-text, and fine-grained image–text retrieval in English and Chinese. It provides efficient embedding APIs for text-to-image, image-to-text, text-to-text, and image-to-image retrieval, moving beyond keyword matching for search, recommendations, document intelligence, and semantic video monitoring. Its fine-grained understanding improves long-text retrieval and attribute-level reranking.
Use Cases
Supporting intelligent transformation across enterprise use cases
Information Retrieval
Use a text description to retrieve the best-matching images, or upload an image to retrieve related text and visual content.
Personalized Recommendations
Combine textual and visual behavior to build a richer representation of user interests.
Semantic Video Monitoring
Analyze monitoring video against a natural-language instruction and trigger an alert when a matching visual event appears.
Intelligent Document Retrieval
Retrieve and understand content across documents that combine text and images.
Model Capabilities
A frontier architecture with strong benchmark performance
Long-Text Cross-Modal Retrieval
Understands complex semantic context to improve retrieval accuracy in open-vocabulary scenarios.
Fine-Grained Attribute Awareness
Distinguishes subtle differences in material, texture, pose, and other attributes beyond conventional keyword matching.
Adaptable Across Business Scenarios
Supports cross-modal retrieval, personalized recommendations, and semantic video monitoring.
Results
FG-CLIP 2 significantly outperforms existing models across downstream tasks including fine-grained understanding, open-vocabulary detection, regional image classification, long- and short-text image retrieval, and general multimodal benchmarks.

English Benchmark Results

Chinese Benchmark Results
Method

FG-CLIP extends the conventional dual-encoder architecture with a two-stage training strategy. Global contrastive learning first aligns image and text representations. Region-level contrastive learning and hard fine-grained negatives then deepen sensitivity to local visual and textual details while preserving global semantic understanding.
FG-CLIP 2 adds large volumes of high-quality Chinese data, dynamic-resolution training, and a redesigned objective to improve fine-grained perception in both English and Chinese.