Build with intelligence. Work with confidence.
Billing
Current balance
Browser Agent
VisFlow learns from your demonstrations and observes, decides, and acts like a person. It can handle both enterprise workflows without APIs and repetitive long-tail consumer tasks.
Controllable Layer Decomposition
Reveal-Layer replaces full-image blind decomposition with controllable, on-demand extraction. Specify a target boundary using bounding-box coordinates or direct selection to isolate an object as an independent RGBA layer, giving users and developers precise, Photoshop-grade layer control.
Multimodal RAG
RzenEmbed is a multimodal embedding model for enterprise document intelligence. It creates unified, high-quality semantic vectors for text, images, video, and mixed-modality documents, supporting cross-modal retrieval, multimodal RAG, and complex document understanding.
Cross-Modal Image–Text Retrieval
FG-CLIP 2 supports long-text, short-text, and fine-grained image–text retrieval in English and Chinese. It provides efficient embedding APIs for text-to-image, image-to-text, text-to-text, and image-to-image retrieval, moving beyond keyword matching for search, recommendations, document intelligence, and semantic video monitoring. Its fine-grained understanding improves long-text retrieval and attribute-level reranking.