ISC.AI 2024 Explores China’s Path for Large Models in the Multimodal Era

360 AI Research
2024-08-07 202 views
ISC.AI 2024 Explores China’s Path for Large Models in the Multimodal Era

Recently, the ISC.AI 2024multimodal era of large model key technology and application forum was successfully held. This forum was jointly organized by 360 Artificial Intelligence Research Institute and the Chinese Image and Graphics Society. It brought together well-known industry scholars, industry technology leaders and other cutting-edge representatives to conduct in-depth discussions on the technological changes, R&D challenges, application scenarios and other issues of large models in the multimodal era. They are committed to jointly exploring the "Chinese path" for the development of large multimodal model (LMM) and accelerating the quality improvement of the digital transformation of the entire industry.

Picture

In the opening speech, Yin Yuhui, vice president of 360 Group and CEO of 360 Digital Intelligence Group, said that artificial intelligence is changing the world at an unprecedented speed, among which multimodalAI technology is one of the important research directions, achieving more natural and efficient human-computer interaction and intelligent decision-making. In this regard, 360 Artificial Intelligence Research Institute, the Chinese Image and Graphics Society, and universities across the country have launched a large amount of cooperation, hoping to jointly promote the innovation and development of related technologies by promoting the in-depth integration of industry, academia, research, and application.

Picture

Liu Yue, deputy secretary-general of the Chinese Society of Image and Graphics, professor and doctoral supervisor at the School of Optoelectronics at Beijing Institute of Technology, said that large models are gradually moving from pure language processing to a new stage of multimodal integration, and their potential and value are beginning to appear. The proposal of large multimodal model (LMM) enables the artificial intelligence system to have more comprehensive and in-depth understanding and processing capabilities by introducing multimodal information such as images and sounds. This leap not only means a huge challenge and breakthrough at the technical level, but also heralds the infinite expansion and deepening of artificial intelligence scenarios.

Picture

In the keynote speech session, Wang Jinqiao, deputy chief engineer of the Institute of Automation, Chinese Academy of Sciences, executive deputy director, researcher and doctoral supervisor of the Zidong Taichu Large Model Research Center, president of the Wuhan Institute of Artificial Intelligence, and secretary-general of the multimodal Artificial Intelligence Industry Alliance, shared the "Practice and Thoughts of multimodal". He pointed out that in the era of large models, the computing power industry has become a new productive force. As the number of parameters gradually increases, massive intelligent computing power becomes a necessary foundation.

Picture

Leng Dawei, deputy director of the 360 Artificial Intelligence Research Institute and head of the visual direction, mentioned in the theme sharing of In his presentation, ‘Large Multimodal Models and Fine-Grained Open-World Object Detection,’ Dawei Leng explained that an LMM learns fine-grained alignment between language and visual modalities, and that open-world object-detection capabilities will influence office automation, embodied robotics, and autonomous driving.

Picture

Qiu Xipeng, a professor at the School of Computer Science at Fudan University, deputy director of the Large Model Search and Generation Committee of the Chinese Information Society of China, and director of the Natural Language Processing Committee of the Shanghai Computer Society, mentioned in the theme sharing "From large language model (LLM) to World Model" that the main feature of artificial intelligence breakthroughs is versatility. Compared with the previous generation model, one model can solve a lot of tasks. When we have such a base, we can change the form of downstream tasks.

Picture

Zhao Sicheng, an associate researcher at Tsinghua University, a national-level young talent, a Ph.D. at Harbin Institute of Technology, and a postdoctoral fellow at the University of California, Berkeley, and Columbia University, pointed out in the theme sharing of "Key Technologies for Large Model End-Side Deployment Applications" that terminal equipment is booming and applications are deepening. Compared with the cloud side, end-side power consumption and computing power are limited, real-time requirements are high, and computing is distributed. End-side AI technology has become a core bottleneck in the industry. Therefore, how to run large models on edge devices with limited resources to meet the intelligent needs of edge devices, that is, miniaturization of large models, is an urgent need for the popularization of artificial intelligence.

Picture

Yang Shu, an assistant researcher at the Department of Electronic Engineering at Tsinghua University, shared in "When Video Semantic Description Meets Large Models" that human understanding of the world is based on multiple modalities such as touch, hearing, and vision. We hope that machines can also understand the world from voice, video, text, etc.multimodal. Therefore, how to process and understand multi-source heterogeneous data through machine learning methods is the core content of multimodal learning, specifically including multimodal key research contents such as representation learning, modal transformation, alignment, fusion and collaborative learning.

Picture

Zhao Guangxiang, a senior algorithm expert at 360 Group, pointed out in the sharing of "Continuous Pre-training of Large Models" that the continued pre-training of large models faces challenges such as "the impact of second-stage training", "the ravine of the valley of despair" and "migration efficiency", and shared detailed practical experience on the above issues.

Picture

In addition, Liu Huanyong, head of document understanding and knowledge graph algorithms at 360 Artificial Intelligence Research Institute, stated in "multimodal Document Understanding Paradigm for Office Q&A Applications" that multimodal model document processing is an important step in document office scenarios. The degree of document understanding and the precision of analysis determine the upper limit of the performance of subsequent document application scenarios. Document processing in real implementation scenarios requires consideration of not only model accuracy, but also speed, inference cost, etc.

Picture

As an important engine for the development of new quality productivity, large multimodal model (LMM) has entered an explosive period of R&D and implementation, further realizing the mixed output capability of multimodal information. In this context, the ISC.AI 2024 Large Model Key Technology and Application Forum has effectively promoted the development of domestic research, strengthened technical exchanges and achievement transformation between academia and industry, and has far-reaching significance for promoting the development of the artificial intelligence industry.

-END-