← Search

Zhicheng Sun

10 accepted papers

2025

Closed-Loop Long-Horizon Robotic Planning via Equilibrium Sequence Modeling

ICML 2025poster

In the endeavor to make autonomous robots take actions, task planning is a major challenge that requires translating high-level task descriptions to long-horizon action sequences. Despite recent advances in language model agents, they remain prone to planning errors and limited in their ability to p…

2025

Enhancing Consistency of Flow-Based Image Editing through Kalman Control

NeurIPS 2025poster

Flow-based generative models have gained popularity for image generation and editing. For instruction-based image editing, it is critical to ensure that modifications are confined to the targeted regions. Yet existing methods often fail to maintain consistency in non-targeted regions between the ori…

Cited by 0SourceScholar
2025

Granularity-Adaptive Spatial Evidence Tokenization for Video Question Answering

AAAI 2025technical

Video question answering plays a vital role in computer vision, and recent advances in large language models have further propelled the development of this field. However, existing video question answering techniques often face limitations in grasping fine-grained video content in spatial dimensions…

Cited by 0SourcePDFScholar
2025

Pyramidal Flow Matching for Efficient Video Generative Modeling

ICLR 2025poster

Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational de…

2024

RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance

NeurIPS 2024poster

Customizing diffusion models to generate identity-preserving images from user-provided reference images is an intriguing new problem. The prevalent approaches typically require training on extensive domain-specific images to achieve identity preservation, which lacks flexibility across different use…

2024

Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization

ICML 2024oral

In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for effective large-scale pre-training due to the modeling of its…