From Multimodal Research to Real-World Impact: 360 AI Research in 2025
A year of multimodal systems

In 2025, 360 AI Research focused on turning multimodal research into reusable models, datasets, and product capabilities. The work covered understanding, retrieval, generation, detection, and efficient deployment rather than treating each topic as an isolated benchmark.
FG-CLIP and FG-CLIP 2 advanced fine-grained image–text alignment, supported by the FineHARD dataset and bilingual evaluation. RzenEmbed extended multimodal embeddings to enterprise documents, images, video, and mixed-layout content. LMM-Det and PlanGen explored more precise visual grounding and controllable generation planning.

Efficient and controllable generation

Projects including PT-DiT, Qihoo-T2X, Bridge Diffusion Model, and other controllable-generation systems improved efficiency, multilingual understanding, and composition control. The IAA architecture studied how to add visual capability to a frozen language model while preserving the language skills it already possesses.

Moving research closer to users

Open-source releases, online model experiences, APIs, and edge-inference work such as MiniCPM-o.cpp helped shorten the path from a paper to a system that developers can test. Across these projects, the common goal remained the same: build multimodal AI that understands details, respects control, and can be deployed in real workflows.
