Keynote Speaker

Qiao Yu,Vice Dean and Professor, Shanghai Innovation Institute; Lead Scientist at Shanghai AI Laboratory, China
He has long-term research expertise in computer vision, multimodal large models, and world models. He led the development of China’s first general vision large model with broad coverage of diverse visual tasks, as well as InternVL series, the multimodal large models with state‑of‑the‑art performance in the open‑source community. He has published more than 300 papers, with over 140 000 total citations and an H‑index of above 160. He holds more than 100 authorized invention patents. His honours include the Wang Xuan Distinguished Young Scholar Award, CVPR 2023 Best Paper Award, AAAI 2021 Distinguished Paper Award, and ACL 2024 Distinguished Paper Award. As the first principal completer, he was awarded the First‑class Prize of Guangdong Provincial Technological Invention Award.
Title: From Multimodal Intelligence to World Models: Progress and Trends
从多模态智能到世界模型:进展与趋势
Abstract: Foundation models represented by large language models have become a core driving force in the advancement of AGI. In particular, systems such as coding agents have made important progress in complex task decomposition, tool use, and autonomous execution. This talk begins with two fundamental questions: “What is the nature of intelligence?” and “Where does intelligence come from?” It compares the mechanisms underlying biological and artificial intelligence and, drawing on our work in general vision, InternVL, and other multimodal large models, examines the roles of language and multimodal intelligence in knowledge abstraction, physical perception, and long-horizon tasks. Building on this discussion, the talk explores how world models can organize multimodal observations into internal representations that support prediction and intervention, enabling a transition from perceiving the environment to predicting the world, and argues that this is a critical path toward AGI. Finally, the talk offers an outlook: language provides the foundation of knowledge, multimodality connects intelligence to reality, world models make predictions, agents plan actions, and open-ended learning uses feedback from action to continuously create new experience.