← Search

Songheng Yin

2 accepted papers

2025

EgoSim: Egocentric Exploration in Virtual Worlds with Multi-modal Conditioning

ICLR 2025poster

Recent advancements in video diffusion models have established a strong foundation for developing world models with practical applications. The next challenge lies in exploring how an agent can leverage these foundation models to understand, interact with, and plan within observed environments. This…

2022

Modular Action Concept Grounding in Semantic Video Prediction

CVPR 2022poster

Recent works in video prediction have mainly focused on passive forecasting and low-level action-conditional prediction, which sidesteps the learning of interaction between agents and objects. We introduce the task of semantic action-conditional video prediction, which uses semantic action labels to…

Cited by 15PDFScholar