Tech Blog

Explore our latest AI research and engineering insights.

MiniCPM-o.cpp: Bringing Multimodal Models to Edge Devices

MiniCPM-o.cpp: Bringing Multimodal Models to Edge Devices

An engineering look at efficient on-device multimodal inference across t…

Dawei Leng and Sen Lü 2025-09-08
694
ICCV 2025 | LMM-Det Unlocks Native Object Detection in LMMs

ICCV 2025 | LMM-Det Unlocks Native Object Detection in LMMs

LMM-Det enables large multimodal models to perform object detection with…

360 AI Research 2025-08-05
720
Two 360 AI Research Papers Accepted to ICCV 2025

Two 360 AI Research Papers Accepted to ICCV 2025

LMM-Det and PlanGen advance native object detection and unified layout p…

360 AI Research 2025-07-15
1393
Inside FG-CLIP: The FineHARD Dataset for Fine-Grained Image–Text Alignment

Inside FG-CLIP: The FineHARD Dataset for Fine-Grained Image–Text Alignment

360 AI Research has released FineHARD, the high-quality dataset behind F…

Bin Wang and Chunyu Xie 2025-06-04
536
FG-CLIP: Fine-Grained Visual and Textual Alignment

FG-CLIP: Fine-Grained Visual and Textual Alignment

FG-CLIP improves long-text understanding and local visual comparison, ad…

360 AI Research 2025-04-28
562
An Efficient ControlNet for Diffusion Transformers with 85% Fewer Parameters

An Efficient ControlNet for Diffusion Transformers with 85% Fewer Parameters

A resource-efficient control framework integrates conditioning signals i…

360 AI Research 2025-03-02
423
ICLR 2025 | Qihoo-T2X Cuts DiT Compute with a Unified Text-to-Any Architecture

ICLR 2025 | Qihoo-T2X Cuts DiT Compute with a Unified Text-to-Any Architecture

PT-DiT uses proxy tokens to support text-to-image, text-to-video, and te…

Ao Ma 2025-02-20
312
AAAI | Bridge Diffusion Model Connects Native Chinese Generation with the Stable Diffusion Ecosystem

AAAI | Bridge Diffusion Model Connects Native Chinese Generation with the Stable Diffusion Ecosystem

Bridge Diffusion Model provides native Chinese text understanding while…

360 AI Research 2024-12-18
195
360 AI Research at AAAI: Perspectives on Multimodal Understanding and Generation

360 AI Research at AAAI: Perspectives on Multimodal Understanding and Generation

An overview of IAA and Bridge Diffusion Model, two AAAI-accepted project…

360 AI Research 2024-12-17
187
AAAI | IAA Adds Multimodal Capabilities Without Catastrophic Forgetting

AAAI | IAA Adds Multimodal Capabilities Without Catastrophic Forgetting

Inner-Adaptor Architecture adds visual understanding to a frozen languag…

360 AI Research 2024-12-17
197
NeurIPS 2024 | HiCo Enables Layout-Controllable Image Generation

NeurIPS 2024 | HiCo Enables Layout-Controllable Image Generation

HiCo gives creators explicit control over the placement of multiple subj…

Bo Cheng 2024-10-31
139
PT-DiT: A More Efficient Paradigm for Text-to-Any Generation

PT-DiT: A More Efficient Paradigm for Text-to-Any Generation

PT-DiT uses proxy-token attention to achieve competitive text-to-image,…

360 AI Research 2024-10-17
145