Speaker: Prof René Vidal (University of Pennsylvania)
Date: Wednesday August 19, 2026
Time: 4.00PM – 5.00PM
Venue:
COM3 Seminar Room 14 (#01-23), 11 Research Link, Singapore 119391
Abstract:
Recent advances in generative AI have transformed text-to-image and text-to-video synthesis, enabling the creation of photorealistic visual content from natural language prompts. The next frontier is not simply generating better images and videos, but learning structured representations of humans and dynamic worlds that enable alignment across modalities and consistent rendering across space and time. This talk explores two complementary steps toward that vision. First, I will present MotionBind, a multimodal foundation model that unifies human motion with language, vision, and audio in a shared embedding space, enabling cross-modal retrieval, recognition, and human motion generation from diverse input modalities. Second, I will present DynamicVoyager, a framework for generating perpetual 3D dynamic scenes from a single image or video by combining dynamic scene representations with geometry-aware generative models, enabling consistent scene exploration through arbitrary camera trajectories while maintaining realistic dynamic motion. Together, these works illustrate a broader shift in generative AI: from generating visual content to learning structured representations of humans and dynamic worlds that support understanding, generation, and control.
Biography:
René Vidal is the Penn Integrates Knowledge and Rachleff University Professor of Electrical and Systems Engineering and Radiology at the University of Pennsylvania, where he directs the Center for Innovation in Data Engineering and Science (IDEAS) and serves as Co-Chair of Penn AI. He is also an Amazon Scholar, Affiliated Chief Scientist at NORCE, and former Associate Editor-in-Chief of IEEE Transactions on Pattern Analysis and Machine Intelligence. His research advances the mathematical foundations of deep learning and trustworthy AI, with broad impact across computer vision and biomedical data science. His contributions have been recognized with major honors, including the IEEE Edward J. McCluskey Technical Achievement Award, the D’Alembert Faculty Award, the J.K. Aggarwal Prize, the ONR Young Investigator Award, the NSF CAREER Award, and best paper awards in machine learning, computer vision, signal processing, control, and medical robotics. He is a Fellow of ACM, AIMBE, IEEE, IAPR, and Sloan Foundation.
