← Search

Jianjun Gao

5 accepted papers

2026

Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep

CVPR 2026

Diffusion-based video editing has emerged as an important paradigm for high-quality and flexible content generation. However, despite their generality and strong modeling capacity, Diffusion Transformers (DiT) remain computationally expensive due to the iterative denoising process, posing challenges

Cited by 0SourcecodeScholar
2026

KinemaDiff: Towards Diffusion for Coherent and Physically Plausible Human Motion Prediction

ICLR 2026poster

Stochastic Human Motion Prediction (HMP) has become an essential task for the realm of computer vision, for its capacity to anticipate accurate and diverse future human trajectories. Current diffusion-based techniques typically enforce skeletal consistency by encoding structural priors into network…

Cited by 0SourceScholar
2025

A Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with Mamba

ICCV 2025poster

Recent Mamba-based methods for the pose-lifting task tend to model joint dependencies by 2D-to-1D mapping with diverse scanning strategies. Though effective, they struggle to model intricate joint connections and uniformly process all joint motion trajectories while neglecting the intrinsic differen…

Cited by 0SourcePDFScholar
2024

Contextual Human Object Interaction Understanding from Pre-Trained Large Language Model

ICASSP 2024accepted

Existing human object interaction (HOI) detection methods have introduced zero-shot learning techniques to recognize unseen interactions, but they still have limitations in understanding context information and comprehensive reasoning. To overcome these limitations, we propose a novel HOI learning f…

Cited by 0SourceScholar
2024

Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

EMNLP 2024main

In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challen…