← Search

Gaoang Wang

31 accepted papers

2026

CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs

CVPR 2026

Large Vision-Language Models (LVLMs) usually suffer from prohibitive computational and memory costs due to the quadratic growth of visual tokens with image resolution. Existing token compression methods, while varied, often lack a high-level semantic understanding, leading to suboptimal merges, info

Cited by 0SourcecodeScholar
2026

Exploring Visual Pretraining for Learning Language Intelligence

CVPR 2026

While the most fundamental pretraining paradigm typically trains modality-specific models on their respective datasets, the Platonic Representation Hypothesis that representations eventually align across modalities as data and model scale suggests an intriguing possibility: large language models (LL

Cited by 0SourcecodeScholar
2026

IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion

AAAI 2026technical

Reconstructing complete and interactive 3D scenes remains a fundamental challenge in computer vision and robotics, particularly due to persistent object occlusions and limited sensor coverage. Even multi-view observations from a single scene scan often fail to capture the full structural details. Ex

Cited by 0SourcePDFScholar
2026

ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

IJCAI 2026

Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian Splatting have enabled high-fidelity static reconstruction, extending it to interactive environments with articulated ro

Cited by 0Scholar
2026

RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy

ICLR 2026poster

Reasoning before action and imagining potential outcomes (i.e., world models) are essential for embodied agents operating in complex open-world environments. Yet, prior work either incorporates only one of these abilities in an end-to-end agent or integrates multiple specialized models into an agent…

Cited by 0SourceScholar
2026

See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent

ICML 2026poster

Pre-trained perception models excel in generic image domains but degrade significantly in novel environments like indoor scenes. The conventional remedy is fine-tuning on downstream data which incurs catastrophic forgetting of prior knowledge and demands costly, scene-specific annotations. We propos…

Cited by 0SourceScholar
2026

Understanding Dynamic Scenes in Ego Centric 4D Point Clouds

AAAI 2026technical

Understanding dynamic 4D scenes from an egocentric perspective—modeling changes in 3D spatial structure over time—is crucial for human–machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets contain dynamic scenes, they lack unified 4D annotations an

Cited by 0SourcePDFScholar
2026

X-MoGen: Unified Motion Generation Across Humans and Animals

AAAI 2026technical

Text-driven motion generation has attracted increasing attention due to its broad applications in virtual reality, animation, and robotics. While existing methods typically model human and animal motion separately, a joint cross-species approach offers key advantages, such as a unified representatio

Cited by 0SourcePDFScholar
2025

AniMo: Species-Aware Model for Text-Driven Animal Motion Generation

CVPR 2025poster

Text-driven motion generation has made significant strides in recent years. However, most existing works focus on human motion, largely overlooking the rich and diverse behaviors of animals. Understanding and synthesizing animal motion have important applications in wildlife conservation, animal eco…

2025

Bringing RNNs Back to Efficient Open-Ended Video Understanding

ICCV 2025poster

The challenge of long video understanding lies in its high computational complexity and prohibitive memory cost, since the memory and computation required by transformer-based LLMs scale quadratically with input sequence length. We propose AuroraLong to address this challenge by replacing the LLM co…

Cited by 0SourcePDFScholar
2025

Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment

ICML 2025poster

Spiking Neural Networks (SNNs) are emerging as a brain-inspired alternative to traditional Artificial Neural Networks (ANNs), prized for their potential energy efficiency on neuromorphic hardware. Despite this, SNNs often suffer from accuracy degradation compared to ANNs and face deployment challeng…

2025

Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition

AAAI 2025technical

In this paper, we propose a novel Temporal Sequence-Aware-Model (TSAM) for few-shot action recognition (FSAR), which incorporates a sequential perceiver adapter into the pre-training framework, to integrate both the spatial information and the sequential temporal dynamics into the feature embeddings…

Cited by 6SourcePDFScholar
2025

RAPID: Recognition of Any-Possible DrIver Distraction via Multi-view Pose Generation Models

ICASSP 2025accepted

Driver distraction remains a pressing traffic safety issue. Drivers are often careless with their distraction behaviours, which may cause serious traffic accidents. However, current Driver Monitoring Systems (DMS) cannot be put into practical application well, which tend to have high latency, lack p…

Cited by 0SourceScholar
2025

SCI-Gaussian: Optimizing 3D Gaussian Radiance Fields from a Snapshot Compressive Image

ICASSP 2025accepted

Snapshot compressive imaging (SCI) is a compressed sensing (CS)-based high-speed imaging modality. Recent efforts have explored the underlying 3D representation from only an SCI image using neural radiance fields (NeRF), yet the training time, rendering computation cost, and reconstruction quality l…

Cited by 0SourceScholar
2024

Advancing Training Efficiency of Deep Spiking Neural Networks through Rate-based Backpropagation

NeurIPS 2024poster

Recent insights have revealed that rate-coding is a primary form of information representation captured by surrogate-gradient-based Backpropagation Through Time (BPTT) in training deep Spiking Neural Networks (SNNs). Motivated by these findings, we propose rate-based backpropagation, a training stra…

2024

BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-based Roadside 3D Object Detection

CVPR 2024poster

Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain since it encompasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on accurately estimating depth or height for 2D-to-3D mapping igno…

2024

Blind Inpainting with Object-Aware Discrimination for Artificial Marker Removal

ICASSP 2024accepted

Medical images often incorporate doctor-added markers that can hinder AI-based diagnosis. This issue highlights the need of inpainting techniques to restore the corrupted visual contents. However, existing methods require manual mask annotation as input, limiting the application scenarios. In this p…

Cited by 0SourceScholar
2024

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

CVPR 2024poster

Recently integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet existing systems can only handle videos with very few frames. For long videos the computation complexity memory cost and…

2024

Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion Models

AAAI 2024technical

Denoising Diffusion Probabilistic Models (DDPMs) have achieved significant success in generation tasks. Nevertheless, the exposure bias issue, i.e., the natural discrepancy between the training (the output of each step is calculated individually by a given input) and inference (the output of each st…

Cited by 2SourcePDFScholar
2024

Sam-Guided Enhanced Fine-Grained Encoding with Mixed Semantic Learning for Medical Image Captioning

ICASSP 2024accepted

With the development of multimodality and large language models, the deep learning-based technique for medical image captioning holds the potential to offer valuable diagnostic recommendations. However, current generic text and image pre-trained models do not yield satisfactory results when it comes…

Cited by 0SourceScholar
2024

UniAP: Towards Universal Animal Perception in Vision via Few-Shot Learning

AAAI 2024technical

Animal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is challenging to design a deep learning-based perception model that can freely adapt to different animals across various…

Cited by 9SourcePDFScholar
2023

A Class-Rebalancing Self-Training Framework for Distantly-Supervised Named Entity Recognition

ACL 2023findings

Distant supervision reduces the reliance on human annotation in the named entity recognition tasks. The class-level imbalanced distant annotation is a realistic and unexplored problem, and the popular method of self-training can not handle class-level imbalanced learning. More importantly, self-trai…

2023

Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object Detection

ICCV 2023poster

Knowledge distillation (KD) has shown potential for learning compact models in dense object detection. However, the commonly used softmax-based distillation ignores the absolute classification scores for individual categories. Thus, the optimum of the distillation loss does not necessarily lead to t…

Cited by 31PDFcodeScholar
2023

Global Adaptation Meets Local Generalization: Unsupervised Domain Adaptation for 3D Human Pose Estimation

ICCV 2023poster

When applying a pre-trained 2D-to-3D Human Pose lifting model to a target unseen dataset, a large performance degradation is commonly encountered due to domain shift issues. We observe that the degradation is caused by two factors: 1) the large distribution gap over global positions of poses between…

Cited by 30PDFcodeScholar
2023

Language Adaptive Weight Generation for Multi-Task Visual Grounding

CVPR 2023poster

Although the impressive performance in visual grounding, the prevailing approaches usually exploit the visual backbone in a passive way, i.e., the visual backbone extracts features with fixed weights without expression-related hints. The passive perception may lead to mismatches (e.g., redundant and…

2023

SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation

IJCAI 2023poster

As an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D p…

2023

TransLink: Transformer-Based Embedding for Tracklets' Global Link

ICASSP 2023accepted

Multi-object tracking (MOT) is essential to many tasks related to the smart transportation. Detecting and tracking humans on the road can give a vital feedback for either the moving vehicle or traffic control to ensure better driving safety and traffic flow. However, most trackers face a common prob…

Cited by 0SourceScholar
2022

Hierarchical Semi-Supervised Contrastive Learning for Contamination-Resistant Anomaly Detection

ECCV 2022poster

"Anomaly detection aims at identifying deviant samples from the normal data distribution. Contrastive learning has provided a successful way to sample representation that enables effective discrimination on anomalies. However, when contaminated with unlabeled abnormal samples in training set under s…

2021

Track Without Appearance: Learn Box and Tracklet Embedding With Local and Global Motion Patterns for Vehicle Tracking

ICCV 2021poster

Vehicle tracking is an essential task in the multi-object tracking (MOT) field. A distinct characteristic in vehicle tracking is that the trajectories of vehicles are fairly smooth in both the world coordinate and the image coordinate. Hence, models that capture motion consistencies are of high nece…

Cited by 77PDFcodeScholar