FG-CLIP 2 in a Smart Elder-Care Monitoring Course Project
Recently, we received a letter of thanks from a student who came to Shanghai Jiao Tong University. The student is a freshman in the School of Automation and Perception. He used the strong capabilities of the institute in fine-grained recognition and multi-language support to develop a "VLM-based intelligent elderly care camera that supports semantic customization" system. I would like to write this letter to express my gratitude for our open source work!
FG-CLIP 2 is the next generation VLM designed for fine-grained cross-modal understanding. This work uses the new fine-grained alignment paradigm to allow the model to not only identify the subject in the image, but also more accurately understand the attributes, relationships and semantics, opening a new stage for AI's visual language understanding ability to move towards "clearer and more accurate". FG-CLIP 2 comprehensively surpassed Google's SigLIP 2 and Meta's MetaCLIP 2 in Chinese and English bilingual tasks, ranking best reported bilingual performance in bilingual performance on 29 tasks in 8 categories.
FG-CLIP Paper address: https://arxiv.org/abs/2505.05071 FG-CLIP 2 Paper address: https://arxiv.org/pdf/2510.10921 FG-CLIP 2 Model, code and dataset address: https://360cvgroup.github.io/FG-CLIP FG-CLIP 2 API access address: https://research.360.cn/sass/fg-clip/fg-clipDocument