Tech Blog
Explore our latest AI research and engineering insights.
Preventing Language-Capability Loss in Multimodal Models with IAA
IAA improves visual-task performance through internal adaptors while kee…
IAA: A Frozen-Language-Model Paradigm for Multimodal Understanding and Visual Grounding
Inner-Adaptor Architecture equips a frozen large language model with gen…
FancyVideo: Open-Source Video Generation on Consumer GPUs
FancyVideo generates videos at flexible resolutions and aspect ratios on…
FancyVideo: Dynamic and Consistent Video Generation with Cross-Frame Textual Guidance
FancyVideo introduces a Cross-frame Textual Guidance Module to improve t…
ISC.AI 2024 Explores China’s Path for Large Models in the Multimodal Era
Researchers and industry leaders discuss technical change, development c…
360VL Unlocks Multimodal Capabilities for Llama 3
360VL combines a vision encoder, a bridge layer, and Llama 3 to support…
HiCo: 360’s Layout-Controllable Image-Generation Model
HiCo lets users specify what should appear in different image regions, e…
Bridge Diffusion Model: Native Chinese Text-to-Image Generation
Bridge Diffusion Model strengthens native Chinese prompt understanding w…
The 15th CSIG Enterprise Visit at Qihoo 360
Researchers, faculty, students, and the Qihoo 360 technical team discuss…
SEEChat: An Open-Source Chinese Multimodal Dialogue Model
SEEChat connects a vision encoder with a Chinese language model to add i…