← Search

Gangshan Wu

48 accepted papers

2026

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

CVPR 2026

Unsupervised learning of latent motion from Internet videos is crucial for robot learning. Existing discrete methods generally mitigate the shortcut learning caused by extracting excessive static backgrounds through vector quantization with a small codebook size. However, they suffer from informatio

Cited by 0SourcecodeScholar
2026

Disentangled Textual Priors for Diffusion-based Image Super-Resolution

CVPR 2026

Image Super-Resolution (SR) aims to reconstruct high-resolution images from degraded low-resolution inputs. While diffusion-based SR methods offer powerful generative capabilities, their performance heavily depends on how semantic priors are structured and integrated into the generation process. Exi

Cited by 0SourcecodeScholar
2026

VMonarch: Efficient Video Diffusion Transformers with Structured Attention

CVPR 2026

The quadratic complexity of the attention mechanism severely limits the context scalability of Video Diffusion Transformers (DiTs). We find that the highly sparse spatio-temporal attention patterns exhibited in Video DiTs can be naturally represented by the Monarch matrix. It is a class of structure

Cited by 2SourceScholar
2025

CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution

CVPR 2025poster

Transformer-based methods have demonstrated impressive performance in low-level visual tasks such as Image Super-Resolution (SR). However, its computational complexity grows quadratically with the spatial resolution. A series of works attempt to alleviate this problem by dividing Low-Resolution imag…

2025

In-the-wild Audio Spatialization with Flexible Text-guided Localization

ACL 2025long

Binaural audio enriches immersive experiences by enabling the perception of the spatial locations of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural audio signals, they often lack the fle…

2025

MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation

NeurIPS 2025poster

Image-to-video generation has made remarkable progress with the advancements in diffusion models, yet generating videos with realistic motion remains highly challenging. This difficulty arises from the complexity of accurately modeling motion, which involves capturing physical constraints, object in…

Cited by 0SourceScholar
2025

Text-Guided Nonverbal Enhancement Based on Modality-Invariant and -Specific Representations for Video Speaking Style Recognition

AAAI 2025technical

Video speaking style recognition (VSSR) aims to classify different types of conversations in videos, contributing significantly to understanding human interactions. A significant challenge in VSSR is the inherent similarity among conversation videos, which makes it difficult to distinguish between d…

Cited by 0SourcePDFScholar
2025

Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning

CVPR 2025poster

Learning from multiple domains is a primary factor that influences the generalization of a single unified robot system. In this paper, we aim to learn the trajectory prediction model by using broad out-of-domain data to improve its performance and generalization ability. Trajectory model is designed…

2024

AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation

NeurIPS 2024poster

Pre-trained vision-language models (VLMs) have shown impressive results in various visual classification tasks. However, we often fail to fully unleash their potential when adapting them for new concept understanding due to limited information on new classes. To address this limitation, we introduce…

2024

Asymmetric Masked Distillation for Pre-Training Small Foundation Models

CVPR 2024poster

Self-supervised foundation models have shown great potential in computer vision thanks to the pre-training paradigm of masked autoencoding. Scale is a primary factor influencing the performance of these foundation models. However these large foundation models often result in high computational cost.…

2024

GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation

ECCV 2024poster

"In recent years, 2D human pose estimation has made significant progress on public benchmarks. However, many of these approaches face challenges of less applicability in the industrial community due to the large number of parametric quantities and computational overhead. Efficient human pose estimat…

2024

Sketch and Refine: Towards Fast and Accurate Lane Detection

AAAI 2024technical

Lane detection is to determine the precise location and shape of lanes on the road. Despite efforts made by current methods, it remains a challenging task due to the complexity of real-world scenarios. Existing approaches, whether proposal-based or keypoint-based, suffer from depicting lanes effecti…

2024

SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos

CVPR 2024poster

Video-based visual relation detection tasks such as video scene graph generation play important roles in fine-grained video understanding. However current video visual relation detection datasets have two main limitations that hinder the progress of research in this area. First they do not explore c…

2023

CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D Datasets

AAAI 2023technical

Current RGB-D scene recognition approaches often train two standalone backbones for RGB and depth modalities with the same Places or ImageNet pre-training. However, the pre-trained depth network is still biased by RGB-based models which may result in a suboptimal solution. In this paper, we present…

2023

Efficient Video Action Detection with Token Dropout and Context Refinement

ICCV 2023poster

Streaming video clips with large-scale video tokens impede vision transformers (ViTs) for efficient recognition, especially in video action detection where sufficient spatiotemporal representations are required for precise actor identification. In this work, we propose an end-to-end framework for ef…

Cited by 26PDFcodeScholar
2023

Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation

CVPR 2023poster

Effectively extracting inter-frame motion and appearance information is important for video frame interpolation (VFI). Previous works either extract both types of information in a mixed way or devise separate modules for each type of information, which lead to representation ambiguity and low effici…

2023

From Coarse to Fine: Hierarchical Pixel Integration for Lightweight Image Super-resolution

AAAI 2023technical

Image super-resolution (SR) serves as a fundamental tool for the processing and transmission of multimedia data. Recently, Transformer-based models have achieved competitive performances in image SR. They divide images into fixed-size patches and apply self-attention on these patches to model long-r…

2023

LinK: Linear Kernel for LiDAR-Based 3D Perception

CVPR 2023poster

Extending the success of 2D Large Kernel to 3D perception is challenging due to: 1. the cubically-increasing overhead in processing 3D data; 2. the optimization difficulties from data scarcity and sparsity. Previous work has taken the first step to scale up the kernel size from 3x3x3 to 7x7x7 by int…

2023

SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes

ICCV 2023poster

Multi-object tracking (MOT) in sports scenes plays a critical role in gathering players statistics, supporting further applications, such as automatic tactical analysis. Yet existing MOT benchmarks cast little attention on this domain. In this work, we present a new large-scale multi-object tracking…

Cited by 103PDFcodeScholar
2023

Video Frame Interpolation with Densely Queried Bilateral Correlation

IJCAI 2023poster

Video Frame Interpolation (VFI) aims to synthesize non-existent intermediate frames between existent frames. Flow-based VFI algorithms estimate intermediate motion fields to warp the existent frames. Real-world motions' complexity and the reference frame's absence make motion estimation challenging.…

2022

Hierarchical Feature Aggregation Network for Deep Image Compression

ICASSP 2022accepted

Existing CNN-based methods for image compression extract features through serially connected high-to-low (encoder) or low-to-high (decoder) resolution stages, leading to insufficient utilization of hierarchical features. To solve this problem, we present a hierarchical feature aggregation network (H…

Cited by 0SourceScholar
2022

Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding

AAAI 2022technical

Temporal grounding aims to localize a video moment which is semantically aligned with a given natural language query. Existing methods typically apply a detection or regression pipeline on the fused representation with the research focus on designing complicated prediction heads or fusion strategies…

2022

Pyramid Fusion Attention Network For Single Image Super-Resolution

ICASSP 2022accepted

Recently, convolutional neural network (CNN) has made a mighty advance in image super-resolution (SR). Most recent models exploit attention mechanism (AM) to focus on high-frequency information. However, these methods exclusively consider interdependencies among channels or spatials, leading to equa…

Cited by 0SourceScholar
2021

MGSampler: An Explainable Sampling Strategy for Video Action Recognition

ICCV 2021poster

Frame sampling is a fundamental problem in video action recognition due to the essential redundancy in time and limited computation resources. The existing sampling strategy often employs a fixed frame selection and lacks the flexibility to deal with complex variations in videos. In this paper, we p…

Cited by 91PDFcodeScholar
2021

MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions

ICCV 2021poster

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions. This paper aims to present a new multi-person dataset of spat…

Cited by 124PDFcodeScholar
2021

Target Adaptive Context Aggregation for Video Scene Graph Generation

ICCV 2021poster

This paper deals with a challenging task of video scene graph generation (VidSGG), which could serve as a structured video representation for high-level understanding tasks. We present a new detect-to-track paradigm for this task by decoupling the context modeling for relation prediction from the co…

Cited by 80PDFcodeScholar
2020

Boundary-Aware Cascade Networks for Temporal Action Segmentation

ECCV 2020poster

Identifying human action segments in an untrimmed video is still challenging due to boundary ambiguity and over-segmentation issues. To address these problems, we present a new boundary-aware cascade network by introducing two novel components. First, we devise a new cascading paradigm, called Stage…

2020

Context-Aware RCNN: A Baseline for Action Detection in Videos

ECCV 2020poster

Video action detection approaches usually conduct actor-centric action recognition over RoI-pooled features following the standard pipeline of Faster-RCNN. In this work, we first empirically find the recognition accuracy is highly correlated with the bounding box size of an actor, and thus higher re…

2020

Low Complexity Single Image Super-Resolution with Channel Splitting and Fusion Network

ICASSP 2020accepted

Recently, deep convolutional neural networks (CNNs) have made remarkable progress on single image super-resolution (SISR). However, many of these methods use very deep or wide convolutional layers to achieve good performance, which treat all feature channels indiscriminately and neglect the differen…

Cited by 0SourceScholar
2020

Residual Feature Aggregation Network for Image Super-Resolution

CVPR 2020poster

Recently, very deep convolutional neural networks (CNNs) have shown great power in single image super-resolution (SISR) and achieved significant improvements against traditional methods. Among these CNN-based methods, the residual connections play a critical role in boosting the network performance.…

Cited by 645PDFScholar
2019

Learning Actor Relation Graphs for Group Activity Recognition

CVPR 2019poster

Modeling relation between actors is important for recognizing group activity in a multi-person scene. This paper aims at learning discriminative relation between actors efficiently using deep models. To this end, we propose to build a flexible and efficient \rm Actor Relation Graph (ARG) to simult…

Cited by 333PDFcodeScholar
2019

Translate-to-Recognize Networks for RGB-D Scene Recognition

CVPR 2019poster

Cross-modal transfer is helpful to enhance modality-specific discriminative power for scene recognition. To this end, this paper presents a unified framework to integrate the tasks of cross-modal translation and modality-specific recognition, termed as Translate-to-Recognize Network TRecgNet. Specif…

Cited by 64PDFcodeScholar