← Search

Ye Lu

9 accepted papers

2026

Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep

CVPR 2026

Diffusion-based video editing has emerged as an important paradigm for high-quality and flexible content generation. However, despite their generality and strong modeling capacity, Diffusion Transformers (DiT) remain computationally expensive due to the iterative denoising process, posing challenges

Cited by 0SourcecodeScholar
2026

KinemaDiff: Towards Diffusion for Coherent and Physically Plausible Human Motion Prediction

ICLR 2026poster

Stochastic Human Motion Prediction (HMP) has become an essential task for the realm of computer vision, for its capacity to anticipate accurate and diverse future human trajectories. Current diffusion-based techniques typically enforce skeletal consistency by encoding structural priors into network…

Cited by 0SourceScholar
2026

Robust Adversarial Attacks Against Unknown Disturbance via Inverse Gradient Sample

ICLR 2026poster

Adversarial attacks have achieved widespread success in various domains, yet existing methods suffer from significant performance degradation when adversarial examples are subjected to even minor disturbances. In this paper, we propose a novel and robust attack called IGSA (**I**nverse **G**radient…

Cited by 0SourcecodeScholar
2025

A Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with Mamba

ICCV 2025poster

Recent Mamba-based methods for the pose-lifting task tend to model joint dependencies by 2D-to-1D mapping with diverse scanning strategies. Though effective, they struggle to model intricate joint connections and uniformly process all joint motion trajectories while neglecting the intrinsic differen…

Cited by 0SourcePDFScholar
2024

Deep Feature Surgery: Towards Accurate and Efficient Multi-Exit Networks

ECCV 2024poster

"Multi-exit network is a promising architecture for efficient model inference by sharing backbone networks and weights among multiple exits. However, the gradient conflict of the shared weights results in sub-optimal accuracy. This paper introduces Deep Feature Surgery (), which consists of feature…

2024

Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

EMNLP 2024main

In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challen…

2023

Locate before Segment: Topology-guided Retinal Layer Segmentation in Optical Coherence Tomography Images

ICRA 2023poster

Optical Coherence Tomography (OCT) is a non-invasive imaging technique that is instrumental in retinal disease diagnosis and treatment. Segmentation of retinal layers in OCT is an essential step, but remains challenging for common pixel-wise segmentation methods usually fail to obtain the correct la…

Cited by 0SourceScholar
2022

Multiple Consistency Supervision based Semi-supervised OCT Segmentation using Very Limited Annotations

ICRA 2022poster

Optical Coherence Tomography (OCT) is a rapidly growing and promising imaging technique, enabling non-invasive high-resolution visualization of biological tissues. Segmentation of tissue structures from OCT scans is essen-tial for disease diagnosis but remains challenging for the blurry boundaries a…

Cited by 3SourceScholar
2021

EADNet: Efficient Asymmetric Dilated Network For Semantic Segmentation

ICASSP 2021accepted

Due to real-time image semantic segmentation needs on power constrained edge devices, there has been an increasing desire to design lightweight semantic segmentation neural network, to simultaneously reduce computational cost and increase inference speed. In this paper, we propose an efficient asymm…

Cited by 0SourceScholar