← Search

Hu Wang

16 accepted papers

2026

Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning

CVPR 2026

Dual-hand action segmentation, densely predicting actions for both hands from untrimmed videos, is essential for understanding complex bimanual activities. However, it poses several unique challenges: complex inter-hand dependencies, visual asymmetry between hands, representation conflicts where the

Cited by 0SourcecodeScholar
2026

TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model

CVPR 2026

Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational cost issues due to the input of large number of visual tokens, motivating token pruning to improve inference efficiency.The key challenge lies in identifying which tokens are truly important.Most exist

Cited by 0SourcecodeScholar
2026

Universal Trajectory Optimization Framework for Differential Drive Robot Class (I)

ICRA 2026poster

Differential drive robots are widely used in various scenarios thanks to their straightforward principle, from household service robots to disaster response field robots. The nonholonomic dynamics and possible lateral slip of these robots lead to difficulty in getting feasible and high-quality traje…

Cited by 0Scholar
2025

Explicit and Implicit Data Augmentation for Social Event Detection

ACL 2025long

Social event detection involves identifying and categorizing important events from social media, which relies on labeled data, but annotation is costly and labor-intensive. To address this problem, we propose Augmentation framework for Social Event Detection (SED-Aug), a plug-and-play dual augmentat…

2024

CPM: Class-conditional Prompting Machine for Audio-visual Segmentation

ECCV 2024poster

"Audio-visual segmentation (AVS) is an emerging task that aims to accurately segment sounding objects based on audio-visual cues. The success of AVS learning systems depends on the effectiveness of cross-modal interaction. Such a requirement can be naturally fulfilled by leveraging transformer-based…

Cited by 2SourcePDFScholar
2024

ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation

ECCV 2024poster

"The costly and time-consuming annotation process to produce large training sets for modelling semantic LiDAR segmentation methods has motivated the development of semi-supervised learning (SSL) methods. However, such SSL approaches often concentrate on employing consistency learning only for indivi…

2024

Segment beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation

AAAI 2024technical

Augmented Reality (AR) devices, emerging as prominent mobile interaction platforms, face challenges in user safety, particularly concerning oncoming vehicles. While some solutions leverage onboard camera arrays, these cameras often have limited field-of-view (FoV) with front or downward perspectives…

Cited by 6SourcePDFScholar
2024

Unraveling Instance Associations: A Closer Look for Audio-Visual Segmentation

CVPR 2024poster

Audio-visual segmentation (AVS) is a challenging task that involves accurately segmenting sounding objects based on audio-visual cues. The effectiveness of audio-visual learning critically depends on achieving accurate cross-modal alignment between sound and visual objects. Successful audio-visual l…

2023

BoMD: Bag of Multi-label Descriptors for Noisy Chest X-ray Classification

ICCV 2023poster

Deep learning methods have shown outstanding classification accuracy in medical imaging problems, which is largely attributed to the availability of large-scale datasets manually annotated with clean labels. However, given the high cost of such manual annotation, new medical imaging classification p…

Cited by 10PDFcodeScholar
2023

Multi-Modal Learning With Missing Modality via Shared-Specific Feature Modelling

CVPR 2023poster

The missing modality issue is critical but non-trivial to be solved by multi-modal models. Current methods aiming to handle the missing modality problem in multi-modal tasks, either deal with missing modalities only during evaluation or train separate models to handle specific missing modality setti…

2022

KUNet: Imaging Knowledge-Inspired Single HDR Image Reconstruction

IJCAI 2022poster

Recently, with the rise of high dynamic range (HDR) display devices, there is a great demand to transfer traditional low dynamic range (LDR) images into HDR versions. The key to success is how to solve the many-to-many mapping problem. However, the existing approaches either do not consider constrai…

2022

M$^4$I: Multi-modal Models Membership Inference

NeurIPS 2022accept

With the development of machine learning techniques, the attention of research has been moved from single-modal learning to multi-modal learning, as real-world data exist in the form of different modalities. However, multi-modal models often carry more information than single-modal models and they a…

2022

Uncertainty-Aware Multi-modal Learning via Cross-Modal Random Network Prediction

ECCV 2022poster

"Multi-modal learning focuses on training models by equally combining multiple input data modalities during the prediction process. However, this equal combination can be detrimental to the prediction accuracy because different modalities are usually accompanied by varying levels of uncertainty. Usi…

Cited by 24SourcePDFScholar
2020

Unsupervised Representation Learning by Predicting Random Distances

IJCAI 2020poster

Deep neural networks have gained great success in a broad range of tasks due to its remarkable capability to learn semantically rich features from high-dimensional data. However, they often require large-scale labelled data to successfully learn such features, which significantly hinders their adapt…

Cited by 0SourcePDFScholar