Multimodal Understanding Model

Cross-Modal Image–Text Retrieval

FG-CLIP 2 supports long-text, short-text, and fine-grained image–text retrieval in English and Chinese. It provides efficient embedding APIs for text-to-image, image-to-text, text-to-text, and image-to-image retrieval, moving beyond keyword matching for search, recommendations, document intelligence, and semantic video monitoring. Its fine-grained understanding improves long-text retrieval and attribute-level reranking.

Use Cases

Supporting intelligent transformation across enterprise use cases

Information Retrieval

Use a text description to retrieve the best-matching images, or upload an image to retrieve related text and visual content.

Personalized Recommendations

Combine textual and visual behavior to build a richer representation of user interests.

Semantic Video Monitoring

Analyze monitoring video against a natural-language instruction and trigger an alert when a matching visual event appears.

Intelligent Document Retrieval

Retrieve and understand content across documents that combine text and images.

Model Capabilities

A frontier architecture with strong benchmark performance

01

Long-Text Cross-Modal Retrieval

Understands complex semantic context to improve retrieval accuracy in open-vocabulary scenarios.

02

Fine-Grained Attribute Awareness

Distinguishes subtle differences in material, texture, pose, and other attributes beyond conventional keyword matching.

03

Adaptable Across Business Scenarios

Supports cross-modal retrieval, personalized recommendations, and semantic video monitoring.

Results

FG-CLIP 2 significantly outperforms existing models across downstream tasks including fine-grained understanding, open-vocabulary detection, regional image classification, long- and short-text image retrieval, and general multimodal benchmarks.

English benchmark results

English Benchmark Results

Chinese benchmark results

Chinese Benchmark Results

Method

Method

FG-CLIP extends the conventional dual-encoder architecture with a two-stage training strategy. Global contrastive learning first aligns image and text representations. Region-level contrastive learning and hard fine-grained negatives then deepen sensitivity to local visual and textual details while preserving global semantic understanding.

FG-CLIP 2 adds large volumes of high-quality Chinese data, dynamic-resolution training, and a redesigned objective to improve fine-grained perception in both English and Chinese.

Try it now FG-CLIP v2 base

Create an account for free credit and quick API access.

Get started free