Tech Blog
Explore our latest AI research and engineering insights.
How VisFlow Operates Web Pages Without Dedicated APIs
VisFlow observes the page people actually see, interprets its current st…
VisFlow Enters Public Beta: Giving General-Purpose Agents Real Browser Execution
The VisFlow public beta connects coding and general-purpose agents to re…
Meet VisFlow: A Browser Agent That Learns from Your Demonstrations
VisFlow combines natural-language task understanding, visual browser ope…
Beyond One-Off Generation: Reveal-Layer Brings Photoshop-Grade Editability to AI Images
Reveal-Layer decomposes a single image into independent RGBA layers so c…
FG-CLIP Wins Silver in the 2026 CSIG Innovation Technology List
FG-CLIP received a silver award for fine-grained cross-modal understandi…
Two 360 AI Research Projects Accepted to ICML 2026
RevealLayer and FG-CLIP 2 were accepted to ICML 2026, advancing editable…
Two Multimodal Generation Projects Accepted to CVPR 2026
RefTON and NAMI improve virtual try-on and high-resolution image generat…
FLUX-Makeup: High-Consistency Makeup Transfer Without Extra Face-Control Modules
FLUX-Makeup transfers makeup from a reference image while preserving ide…
From Multimodal Research to Real-World Impact: 360 AI Research in 2025
A year in review spanning fine-grained vision–language alignment, enterp…
FG-CLIP 2 in a Smart Elder-Care Monitoring Course Project
A student project shows how fine-grained vision–language understanding c…
FG-CLIP 2: Fine-Grained Bilingual Vision–Language Understanding
FG-CLIP 2 advances fine-grained image–text alignment with dynamic-resolu…
RzenEmbed: Multimodal Retrieval for Enterprise Documents
RzenEmbed creates unified semantic embeddings for text, images, video, a…