Feedback-guided representation learning for reasoning in the dynamic physical world
Modern AI systems often learn fixed representations from fixed datasets and supervision signals. This assumption becomes limiting in physical environments, where observations are partial, noisy, changing, and under-specified. A system that reasons about the physical world must update what it knows, maintain state across incomplete views, and constrain predicted futures by plausible dynamics.
This talk presents feedback-guided representation learning as a framework for making learned representations more useful after training or initial inference. Feedback is treated as any signal that evaluates whether a representation remains consistent with data, objectives, or physical constraints. The thesis develops this idea through three levels of representation. Learning to Unlearn studies model parameters as revisable memory, using forgetting, retention, and membership-inference feedback to remove targeted training influence while preserving retained utility. OnlineSplatter studies object-centric visual state, maintaining a compact 3D object memory from monocular RGB observations without camera poses, depth input, or offline optimization. Physical Simulator In-the-Loop Video Generation studies dynamics, using simulator feedback to correct generated videos toward physically plausible motion while preserving visual appearance.
Together, these works trace a progression from implicit neural memory to explicit object state, to executable physical dynamics. The central claim is that learned representations for physical reasoning should not only be accurate at the moment they are produced; they should also be revisable, persistent, and constrained by feedback from the changing world.
Speaker’s profile
Mark He Huang is a Ph.D. candidate in Computer Science at ISTD Pillar, Singapore University of Technology and Design and an A*STAR Computing and Information Science (ACIS/AGS Computing) Scholar attached to A*STAR’s Centre for Frontier AI Research (CFAR). His research focuses on 3D/4D perception, video generation, and multimodal reasoning about the physical world. His work has been published in leading venues including CVPR, ECCV, and NeurIPS. He obtained his B.Eng. (Computer Science and Design) from SUTD in 2022.
