From Multimodal Research to Real-World Impact: 360 AI Research in 2025

360 AI Research
2025-12-30 846 views
From Multimodal Research to Real-World Impact: 360 AI Research in 2025

A year of multimodal systems

Figure 1 from the original article

In 2025, 360 AI Research focused on turning multimodal research into reusable models, datasets, and product capabilities. The work covered understanding, retrieval, generation, detection, and efficient deployment rather than treating each topic as an isolated benchmark.

FG-CLIP and FG-CLIP 2 advanced fine-grained image–text alignment, supported by the FineHARD dataset and bilingual evaluation. RzenEmbed extended multimodal embeddings to enterprise documents, images, video, and mixed-layout content. LMM-Det and PlanGen explored more precise visual grounding and controllable generation planning.

Figure 2 from the original article

Efficient and controllable generation

Figure 3 from the original article

Projects including PT-DiT, Qihoo-T2X, Bridge Diffusion Model, and other controllable-generation systems improved efficiency, multilingual understanding, and composition control. The IAA architecture studied how to add visual capability to a frozen language model while preserving the language skills it already possesses.

Figure 4 from the original article

Moving research closer to users

Figure 5 from the original article

Open-source releases, online model experiences, APIs, and edge-inference work such as MiniCPM-o.cpp helped shorten the path from a paper to a system that developers can test. Across these projects, the common goal remained the same: build multimodal AI that understands details, respects control, and can be deployed in real workflows.

Figure 6 from the original article

Explore 360 AI Research products