Tech Blog

Explore our latest AI research and engineering insights.

Preventing Language-Capability Loss in Multimodal Models with IAA

Preventing Language-Capability Loss in Multimodal Models with IAA

IAA improves visual-task performance through internal adaptors while kee…

Bin Wang and Chunyu Xie 2024-08-31
107
IAA: A Frozen-Language-Model Paradigm for Multimodal Understanding and Visual Grounding

IAA: A Frozen-Language-Model Paradigm for Multimodal Understanding and Visual Grounding

Inner-Adaptor Architecture equips a frozen large language model with gen…

Bin Wang and Chunyu Xie 2024-08-29
199
FancyVideo: Open-Source Video Generation on Consumer GPUs

FancyVideo: Open-Source Video Generation on Consumer GPUs

FancyVideo generates videos at flexible resolutions and aspect ratios on…

Ao Ma 2024-08-26
120
FancyVideo: Dynamic and Consistent Video Generation with Cross-Frame Textual Guidance

FancyVideo: Dynamic and Consistent Video Generation with Cross-Frame Textual Guidance

FancyVideo introduces a Cross-frame Textual Guidance Module to improve t…

Ao Ma 2024-08-20
197
ISC.AI 2024 Explores China’s Path for Large Models in the Multimodal Era

ISC.AI 2024 Explores China’s Path for Large Models in the Multimodal Era

Researchers and industry leaders discuss technical change, development c…

360 AI Research 2024-08-07
142
360VL Unlocks Multimodal Capabilities for Llama 3

360VL Unlocks Multimodal Capabilities for Llama 3

360VL combines a vision encoder, a bridge layer, and Llama 3 to support…

360 AI Research 2024-05-17
149
HiCo: 360’s Layout-Controllable Image-Generation Model

HiCo: 360’s Layout-Controllable Image-Generation Model

HiCo lets users specify what should appear in different image regions, e…

Bo Cheng 2024-04-17
77
Bridge Diffusion Model: Native Chinese Text-to-Image Generation

Bridge Diffusion Model: Native Chinese Text-to-Image Generation

Bridge Diffusion Model strengthens native Chinese prompt understanding w…

Dawei Leng and Shanyuan Liu 2023-09-15
158
The 15th CSIG Enterprise Visit at Qihoo 360

The 15th CSIG Enterprise Visit at Qihoo 360

Researchers, faculty, students, and the Qihoo 360 technical team discuss…

360 AI Research 2023-06-30
106
SEEChat: An Open-Source Chinese Multimodal Dialogue Model

SEEChat: An Open-Source Chinese Multimodal Dialogue Model

SEEChat connects a vision encoder with a Chinese language model to add i…

Dawei Leng 2023-06-25
124