← Search

Tianpei Gu

6 accepted papers

2025

Towards Self-Refinement of Vision-Language Models with Triangular Consistency

NeurIPS 2025poster

Vision-Language Models (VLMs) integrate visual knowledge with the analytical capabilities of Large Language Models (LLMs) through supervised visual instruction tuning, using image-question-answer triplets. However, the potential of VLMs trained without supervised instruction remains largely unexplor…

Cited by 0SourcecodeScholar
2024

Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance Accompaniment

ICLR 2024poster

We introduce a novel task within the field of human motion generation, termed dance accompaniment, which necessitates the generation of responsive movements from a dance partner, the "follower", synchronized with the lead dancer’s movements and the underlying musical rhythm. Unlike existing solo or…

Cited by 19SourcePDFScholar
2023

Deep Geometrized Cartoon Line Inbetweening

ICCV 2023poster

We aim to address a significant but understudied problem in the anime industry, namely the inbetweening of cartoon line drawings. Inbetweening involves generating intermediate frames between two black-and-white line drawings and is a time-consuming and expensive process that can benefit from automat…

Cited by 17PDFcodeScholar
2022

Bailando: 3D Dance Generation by Actor-Critic GPT With Choreographic Memory

CVPR 2022oral

Driving 3D characters to dance following a piece of music is highly challenging due to the spatial constraints applied to poses by choreography norms. In addition, the generated dance sequence also needs to maintain temporal coherency with different music genres. To tackle these challenges, we propo…

Cited by 217PDFcodeScholar
2022

Stochastic Trajectory Prediction via Motion Indeterminacy Diffusion

CVPR 2022poster

Human behavior has the nature of indeterminacy, which requires the pedestrian trajectory prediction system to model the multi-modality of future motion states. Unlike existing stochastic trajectory prediction methods which usually use a latent variable to represent multi-modality, we explicitly simu…

Cited by 252PDFcodeScholar