← Search

Jianhuang Lai

19 accepted papers

2026

Style4D-Bench: A Benchmark Suite for 4D Stylization

AAAI 2026technical

We introduce Style4D-Bench, the first benchmark suite specifically designed for 4D stylization, with the goal of standardizing evaluation and facilitating progress in this emerging area. Style4D-Bench comprises: 1) a strong baseline that make an initial attempt for 4D stylization, 2) a comprehensive

Cited by 0SourcePDFScholar
2026

View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

CVPR 2026

Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow a view-invariant paradigm, aligning shared features across views to achieve robustness. However, view-invariant inherent

Cited by 0SourcecodeScholar
2025

External Memory Matters: Generalizable Object-Action Memory for Retrieval-Augmented Long-Term Video Understanding

IJCAI 2025

Long video understanding with Large Language Models (LLMs) enables the description of objects that are not explicitly present in the training data. However, continuous changes in known objects and the emergence of new ones require up-to-date knowledge of objects and their dynamics for effective unde

Cited by 0SourcePDFScholar
2025

GuardSplat: Efficient and Robust Watermarking for 3D Gaussian Splatting

CVPR 2025poster

3D Gaussian Splatting (3DGS) has recently created impressive 3D assets for various applications. However, considering security, capacity, invisibility, and training efficiency, the copyright of 3DGS assets is not well protected as existing watermarking methods are unsuited for its rendering pipeline…

2025

Training-Free Class Purification for Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Fine-tuning pre-trained vision-language models has emerged as a powerful approach for enhancing open-vocabulary semantic segmentation (OVSS). However, the substantial computational and resource demands associated with training on large datasets have prompted interest in training-free methods for OVS…

2024

Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis

CVPR 2024highlight

Diffusion model is a promising approach to image generation and has been employed for Pose-Guided Person Image Synthesis (PGPIS) with competitive performance. While existing methods simply align the person appearance to the target pose they are prone to overfitting due to the lack of a high-level se…

2024

Progressive Pretext Task Learning for Human Trajectory Prediction

ECCV 2024poster

"Human trajectory prediction is a practical task of predicting the future positions of pedestrians on the road, which typically covers all temporal ranges from short-term to long-term within a trajectory. However, existing works attempt to address the entire trajectory prediction with a singular, un…

2024

Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding

CVPR 2024poster

Video Paragraph Grounding (VPG) is an emerging task in video-language understanding which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However existing VPG approaches are heavily reliant on a considerable number of temporal labels that are…

Cited by 4SourcePDFScholar
2023

Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding

CVPR 2023poster

Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex sem…

Cited by 24SourcePDFScholar
2023

Texture-Guided Saliency Distilling for Unsupervised Salient Object Detection

CVPR 2023poster

Deep Learning-based Unsupervised Salient Object Detection (USOD) mainly relies on the noisy saliency pseudo labels that have been generated from traditional handcraft methods or pre-trained networks. To cope with the noisy labels problem, a class of methods focus on only easy samples with reliable l…

2023

The Enemy of My Enemy Is My Friend: Exploring Inverse Adversaries for Improving Adversarial Training

CVPR 2023poster

Although current deep learning techniques have yielded superior performance on various computer vision tasks, yet they are still vulnerable to adversarial examples. Adversarial training and its variants have been shown to be the most effective approaches to defend against adversarial examples. A par…

Cited by 42SourcePDFScholar
2020

An Asymmetric Modeling for Action Assessment

ECCV 2020poster

Action assessment is a task of assessing the performance of an action. It is widely applicable to many real-world scenarios such as medical treatment and sporting events. However, existing methods for action assessment are mostly limited to individual actions, especially lacking modeling of the asym…

Cited by 58SourcePDFScholar
2018

Deep Bilinear Learning for RGB-D Action Recognition

ECCV 2018poster

In this paper, we focus on exploring modality-temporal mutual information for RGB-D action recognition. In order to learn time-varying information and multi-modal features jointly, we propose a novel deep bilinear learning framework. In the framework, we propose bilinear blocks that consist of two l…

Cited by 116SourcePDFScholar
2018

Interleaved Structured Sparse Convolutional Neural Networks

CVPR 2018poster

In this paper, we study the problem of designing efficient convolutional neural network architectures with the interest in eliminating the redundancy in convolution kernels. In addition to structured sparse kernels, low-rank kernels and the product of low-rank kernels,the product of structured spars…

Cited by 160SourcePDFScholar
2017

RGB-Infrared Cross-Modality Person Re-Identification

ICCV 2017poster

Person re-identification (Re-ID) is an important problem in video surveillance, aiming to match pedestrian images across camera views. Currently, most works focus on RGB-based Re-ID. However, in some applications, RGB images are not suitable, e.g. in a dark environment or at night. Infrared (IR) ima…

Cited by 896PDFScholar
2015

Jointly Learning Heterogeneous Features for RGB-D Activity Recognition

CVPR 2015poster

In this paper, we focus on heterogeneous feature learning for RGB-D activity recognition. Considering that features from different channels could share some similar hidden structures, we propose a joint learning model to simultaneously explore the shared and feature-specific components as an instanc…

Cited by 687SourcePDFScholar