← Search

Juwei Lu

20 accepted papers

2026

CamDirector: Towards Long-Term Coherent Video Trajectory Editing

CVPR 2026

Video (camera) trajectory editing aims to synthesize new videos that follow user-defined camera paths while preserving scene content and plausibly inpainting previously unseen regions, upgrading amateur footage into professionally styled videos. Existing VTE methods struggle with precise camera cont

Cited by 0SourceScholar
2025

IMFine: 3D Inpainting via Geometry-guided Multi-view Refinement

CVPR 2025poster

Current 3D inpainting and object removal methods are largely limited to front-facing scenes, facing substantial challenges when applied to diverse, "unconstrained" scenes where the camera orientation and trajectory are unrestricted. To bridge this gap, we introduce a novel approach that produces inp…

Cited by 0SourcePDFScholar
2025

MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer

ICLR 2025poster

Generative masked transformer have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the animation domain, large datasets are not always available. Applying…

Cited by 0SourcePDFScholar
2024

AQF: Assessing the Quality of Hyperspectral Reconstruction with a Learnable Metric

ICASSP 2024accepted

This paper proposes a learnable metric to measure the reconstruction quality of hyperspectral images obtained by computational hyperspectral imaging. Computational hyperspectral imaging aims to obtain low-cost hyperspectral images through consumer camera. While many hyperspectral reconstruction mode…

Cited by 0SourceScholar
2024

GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction

ECCV 2024poster

"We present GSD, a diffusion model approach based on Gaussian Splatting (GS) representation for 3D object reconstruction from a single view. Prior works suffer from inconsistent 3D geometry or mediocre rendering quality due to improper representations. We take a step towards resolving these shortcom…

Cited by 7SourcePDFScholar
2024

Generative Human Motion Stylization in Latent Space

ICLR 2024poster

Human motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the \textit{latent space} of pretrained autoencoders as a more expressive and robust representation for motion extraction a…

Cited by 13SourcePDFScholar
2024

TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling

ECCV 2024poster

"Given a 3D mesh, we aim to synthesize 3D textures that correspond to arbitrary textual descriptions. Current methods for generating and assembling textures from sampled views often result in prominent seams or excessive smoothing. To tackle these issues, we present TexGen, a novel multi-view sampli…

2023

CLIPPING: Distilling CLIP-Based Models With a Student Base for Video-Language Retrieval

CVPR 2023poster

Pre-training a vison-language model and then fine-tuning it on downstream tasks have become a popular paradigm. However, pre-trained vison-language models with the Transformer architecture usually take long inference time. Knowledge distillation has been an efficient technique to transfer the capabi…

Cited by 47SourcePDFScholar
2023

Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the Wild

NeurIPS 2023poster

This paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependen…

2023

HiVLP: Hierarchical Interactive Video-Language Pre-Training

ICCV 2023poster

Video-Language Pre-training (VLP) has become one of the most popular research topics in deep learning. However, compared to image-language pre-training, VLP has lagged far behind due to the lack of large amounts of video-text pairs. In this work, we train a VLP model with a hybrid of image-text and…

Cited by 6PDFScholar
2023

Hyper-Skin: A Hyperspectral Dataset for Reconstructing Facial Skin-Spectra from RGB Images

NeurIPS 2023poster

We introduce Hyper-Skin, a hyperspectral dataset covering wide range of wavelengths from visible (VIS) spectrum (400nm - 700nm) to near-infrared (NIR) spectrum (700nm - 1000nm), uniquely designed to facilitate research on facial skin-spectra reconstruction. By reconstructing skin spectra from RGB im…

2022

Decompose the Sounds and Pixels, Recompose the Events

AAAI 2022technical

In this paper, we propose a framework centering around a novel architecture called the Event Decomposition Recomposition Network (EDRNet) to tackle the Audio-Visual Event (AVE) localization problem in the supervised and weakly supervised settings. AVEs in the real world exhibit common unraveling pat…

Cited by 3SourcePDFScholar
2022

Dual Perspective Network for Audio-Visual Event Localization

ECCV 2022poster

"The Audio-Visual Event Localization (AVEL) problem involves tackling three core sub-tasks: the creation of efficient audio-visual representations using cross-modal guidance, the formation of short-term temporal feature aggregations, and its accumulation to achieve long-term dependency resolution. T…

Cited by 21SourcePDFScholar
2022

Self-Supervised Spatiotemporal Representation Learning by Exploiting Video Continuity

AAAI 2022technical

Recent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored property of videos, the textit{video continuity}, to obtain supervision signals…

Cited by 33SourcePDFScholar
2021

Boosting the Generalization Capability in Cross-Domain Few-Shot Learning via Noise-Enhanced Supervised Autoencoder

ICCV 2021poster

State of the art (SOTA) few-shot learning (FSL) methods suffer significant performance drop in the presence of domain differences between source and target datasets. The strong discrimination ability on the source dataset does not necessarily translate to high classification accuracy on the target d…

Cited by 79PDFScholar
2021

Class Semantics-Based Attention for Action Detection

ICCV 2021poster

Action localization networks are often structured as a feature encoder sub-network and a localization sub-network, where the feature encoder learns to transform an input video to features that are useful for the localization sub-network to generate reliable action proposals. While some of the encode…

Cited by 83PDFScholar
2021

Learning Causal Representation for Training Cross-Domain Pose Estimator via Generative Interventions

ICCV 2021poster

3D pose estimation has attracted increasing attention with the availability of high-quality benchmark datasets. However, prior works show that deep learning models tend to learn spurious correlations, which fail to generalize beyond the specific dataset they are trained on. In this work, we take a s…

Cited by 39PDFScholar
2020

All at Once: Temporally Adaptive Multi-Frame Interpolation with Advanced Motion Modeling

ECCV 2020poster

Recent advances in high refresh rate displays as well as the increased interest in high rate of slow motion and frame up-conversion fuel the demand for efficient and cost-effective multi-frame video interpolation solutions. To that regard, inserting multiple frames between consecutive video frames a…

Cited by 79SourcePDFScholar
2020

Towards Efficient Coarse-to-Fine Networks for Action and Gesture Recognition

ECCV 2020poster

State-of-the-art approaches to video-based action and gesture recognition often employ two key concepts: First, they employ multistream processing; second, they use an ensemble of convolutional networks. We improve and extend both aspects. First, we systematically yield enhanced receptive fields for…

Cited by 16SourcePDFScholar
2020

Weight Excitation: Built-in Attention Mechanisms in Convolutional Neural Networks

ECCV 2020poster

We propose novel approaches for simultaneously identifying important weights of a convolutional neural network (ConvNet) and providing more attention to the important weights during training. More formally, we identify two characteristics of a weight, its magnitude and its location, which can be lin…