Tech Blog
Explore our latest AI research and engineering insights.
MiniCPM-o.cpp: Bringing Multimodal Models to Edge Devices
An engineering look at efficient on-device multimodal inference across t…
ICCV 2025 | LMM-Det Unlocks Native Object Detection in LMMs
LMM-Det enables large multimodal models to perform object detection with…
Two 360 AI Research Papers Accepted to ICCV 2025
LMM-Det and PlanGen advance native object detection and unified layout p…
Inside FG-CLIP: The FineHARD Dataset for Fine-Grained Image–Text Alignment
360 AI Research has released FineHARD, the high-quality dataset behind F…
FG-CLIP: Fine-Grained Visual and Textual Alignment
FG-CLIP improves long-text understanding and local visual comparison, ad…
An Efficient ControlNet for Diffusion Transformers with 85% Fewer Parameters
A resource-efficient control framework integrates conditioning signals i…
ICLR 2025 | Qihoo-T2X Cuts DiT Compute with a Unified Text-to-Any Architecture
PT-DiT uses proxy tokens to support text-to-image, text-to-video, and te…
AAAI | Bridge Diffusion Model Connects Native Chinese Generation with the Stable Diffusion Ecosystem
Bridge Diffusion Model provides native Chinese text understanding while…
360 AI Research at AAAI: Perspectives on Multimodal Understanding and Generation
An overview of IAA and Bridge Diffusion Model, two AAAI-accepted project…
AAAI | IAA Adds Multimodal Capabilities Without Catastrophic Forgetting
Inner-Adaptor Architecture adds visual understanding to a frozen languag…
NeurIPS 2024 | HiCo Enables Layout-Controllable Image Generation
HiCo gives creators explicit control over the placement of multiple subj…
PT-DiT: A More Efficient Paradigm for Text-to-Any Generation
PT-DiT uses proxy-token attention to achieve competitive text-to-image,…