← Search

Changhu Wang

41 accepted papers

2026

CLINIC: Towards High-quality Graph Out-Of-Distribution Detection

ICML 2026poster

This paper studies the problem of graph out-of-distribution (OOD) detection, which aims to identify anomaly graphs out of a graph dataset. Prior efforts usually focus on the utilization of topological structures with unsupervised graph learning to foster typical pattern recognition, which overlooks …

Cited by 0SourceScholar
2026

Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time Transduction

ICML 2026poster

Many methods aim to enhance the performance of vision-language models (VLMs) at test time. Among them, transduction has emerged as a promising paradigm due to its strong compatibility and efficiency. However, realistic evaluations often involve highly imbalanced class distributions, which cause perf…

Cited by 0SourceScholar
2025

A Generic Family of Graphical Models: Diversity, Efficiency, and Heterogeneity

ICML 2025poster

Traditional network inference methods, such as Gaussian Graphical Models, which are built on continuity and homogeneity, face challenges when modeling discrete data and heterogeneous frameworks. Furthermore, under high-dimensionality, the parameter estimation of such models can be hindered by the no…

Cited by 0SourcePDFScholar
2025

DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal Retrieval

AAAI 2025technical

With the burst of big data, 2D-3D cross-modal retrieval has received increasing attention, which aims to retrieve relevant data from one modality given the query from the other modality. In this paper, we study an underexplored yet practical problem of semi-supervised 2D-3D cross-modal retrieval, wh…

Cited by 0SourcePDFScholar
2025

How Do Large Language Models Perform on PDE Discovery: A Coarse-to-fine Perspective

EMNLP 2025

This paper studies the problem of how to use large language models (LLMs) to identify the underlying partial differential equations (PDEs) out of very limited observations of a physical system. Previous methods usually utilize physical-informed neural networks (PINNs) to learn the PDE solver and coe

Cited by 0SourcePDFScholar
2025

LaVin-DiT: Large Vision Diffusion Transformer

CVPR 2025poster

This paper presents the Large Vision Diffusion Transformer (LaVin-DiT), a scalable and unified foundation model designed to tackle over 20 computer vision tasks in a generative framework. Unlike existing large vision models directly adapted from natural language processing architectures, which rely…

2025

TRACI: A Data-centric Approach for Multi-Domain Generalization on Graphs

AAAI 2025technical

Graph neural networks (GNNs) have gained superior performance in graph-based prediction tasks with a variety of applications such as social analysis and drug discovery. Despite the remarkable progress, their performance often degrades on test graphs with distribution shifts. Existing domain adaptati…

2025

TrackGo: A Flexible and Efficient Method for Controllable Video Generation

AAAI 2025technical

Recent years have seen substantial progress in diffusion-based controllable video generation. However, achieving precise control in complex scenarios, including fine-grained object parts, sophisticated motion trajectories, and coherent background movement, remains a challenge. In this paper, we in…

Cited by 11SourcePDFScholar
2024

PURE: Prompt Evolution with Graph ODE for Out-of-distribution Fluid Dynamics Modeling

NeurIPS 2024poster

This work studies the problem of out-of-distribution fluid dynamics modeling. Previous works usually design effective neural operators to learn from mesh-based data structures. However, in real-world applications, they would suffer from distribution shifts from the variance of system parameters and…

Cited by 5SourcePDFScholar
2022

Content-Variant Reference Image Quality Assessment via Knowledge Distillation

AAAI 2022technical

Generally, humans are more skilled at perceiving differences between high-quality (HQ) and low-quality (LQ) images than directly judging the quality of a single LQ image. This situation also applies to image quality assessment (IQA). Although recent no-reference (NR-IQA) methods have made great prog…

2022

Contextual Text Block Detection towards Scene Text Understanding

ECCV 2022poster

"Most existing scene text detectors focus on detecting characters or words that only capture partial text messages due to missing contextual information. For a better understanding of text in scenes, it is more desired to detect contextual text blocks (CTBs) which consist of one or multiple integral…

2022

TransFG: A Transformer Architecture for Fine-Grained Recognition

AAAI 2022technical

Fine-grained visual classification (FGVC) which aims at recognizing objects from subcategories is a very challenging task due to the inherently subtle inter-class differences. Most existing works mainly tackle this problem by reusing the backbone network to extract features of detected discriminativ…

2021

Adaptive Data Augmentation on Temporal Graphs

NeurIPS 2021poster

Temporal Graph Networks (TGNs) are powerful on modeling temporal graph data based on their increased complexity. Higher complexity carries with it a higher risk of overfitting, which makes TGNs capture random noise instead of essential semantic information. To address this issue, our idea is to tran…

Cited by 66SourcePDFScholar
2021

Cross-Category Video Highlight Detection via Set-Based Learning

ICCV 2021poster

Autonomous highlight detection is crucial for enhancing the efficiency of video browsing on social media platforms. To attain this goal in a data-driven way, one may often face the situation where highlight annotations are not available on the target video category used in practice, while the superv…

Cited by 62PDFcodeScholar
2021

Directed Graph Contrastive Learning

NeurIPS 2021poster

Graph Contrastive Learning (GCL) has emerged to learn generalizable representations from contrastive views. However, it is still in its infancy with two concerns: 1) changing the graph structure through data augmentation to generate contrastive views may mislead the message passing scheme, as such g…

2021

Domain-Invariant Disentangled Network for Generalizable Object Detection

ICCV 2021poster

We address the problem of domain generalizable object detection, which aims to learn a domain-invariant detector from multiple "seen" domains so that it can generalize well to other "unseen" domains. The generalization ability is crucial in practical scenarios especially when it is difficult to coll…

Cited by 96PDFScholar
2021

F2Net: Learning to Focus on the Foreground for Unsupervised Video Object Segmentation

AAAI 2021technical

Although deep learning based methods have achieved great progress in unsupervised video object segmentation, difficult scenarios (e.g., visual similarity, occlusions, and appearance changing) are still no well-handled. To alleviate these issues, we propose a novel Focus on Foreground Network (F2Net…

Cited by 51SourcePDFScholar
2021

Involution: Inverting the Inherence of Convolution for Visual Recognition

CVPR 2021poster

Convolution has been the core ingredient of modern neural networks, triggering the surge of deep learning in vision. In this work, we rethink the inherent principles of standard convolution for vision tasks, specifically spatial-agnostic and channel-specific. Instead, we present a novel atomic opera…

Cited by 468PDFcodeScholar
2021

Learning the Best Pooling Strategy for Visual Semantic Embedding

CVPR 2021poster

Visual Semantic Embedding (VSE) is a dominant approach for vision-language retrieval, which aims at learning a deep embedding space such that visual data are embedded close to their semantic text labels or descriptions. Recent VSE models use complex methods to better contextualize and aggregate mult…

Cited by 296PDFcodeScholar
2021

MINE: Towards Continuous Depth MPI With NeRF for Novel View Synthesis

ICCV 2021poster

In this paper, we propose MINE to perform novel view synthesis and depth estimation via dense 3D reconstruction from a single image. Our approach is a continuous depth generalization of the Multiplane Images (MPI) by introducing the NEural radiance fields (NeRF). Given a single image as input, MINE…

Cited by 171PDFcodeScholar
2021

MT-ORL: Multi-Task Occlusion Relationship Learning

ICCV 2021poster

Retrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of an architecture which can exploit the limited amount of coupling in the decoder stage between the two subtasks, namely oc…

Cited by 8PDFcodeScholar
2021

Meta Navigator: Search for a Good Adaptation Policy for Few-Shot Learning

ICCV 2021poster

Few-shot learning aims to adapt knowledge learned from previous tasks to novel tasks with only a limited amount of labeled data. Research literature on few-shot learning exhibits great diversity, while different algorithms often excel at different few-shot learning scenarios. It is therefore tricky…

Cited by 60PDFScholar
2021

Mining Contextual Information Beyond Image for Semantic Segmentation

ICCV 2021poster

This paper studies the context aggregation problem in semantic image segmentation. The existing researches focus on improving the pixel representations by aggregating the contextual information within individual images. Though impressive, these methods neglect the significance of the representations…

Cited by 105PDFcodeScholar
2021

Slimmable Generative Adversarial Networks

AAAI 2021technical

Generative adversarial networks (GANs) have achieved remarkable progress in recent years, but the continuously growing scale of models make them challenging to deploy widely in practical applications. In particular, for real-time generation tasks, different devices require generators of different si…

2021

Sparse R-CNN: End-to-End Object Detection With Learnable Proposals

CVPR 2021poster

We present Sparse R-CNN, a purely sparse method for object detection in images. Existing works on object detection heavily rely on dense object candidates, such as k anchor boxes pre-defined on all grids of image feature map of size HxW. In our method, however, a fixed sparse set of learned object p…

Cited by 1491PDFcodeScholar
2021

Unsupervised Real-World Super-Resolution: A Domain Adaptation Perspective

ICCV 2021poster

Most existing convolution neural network (CNN) based super-resolution (SR) methods generate their paired training dataset by artificially synthesizing low-resolution (LR) images from the high-resolution (HR) ones. However, this dataset preparation strategy harms the application of these CNNs in real…

Cited by 61PDFScholar
2021

Weakly Supervised Person Search With Region Siamese Networks

ICCV 2021poster

Supervised learning is dominant in person search, but it requires elaborate labeling of bounding boxes and identities. Large-scale labeled training data is often difficult to collect, especially for person identities. A natural question is whether a good person search model can be trained without th…

Cited by 30PDFScholar
2021

What Makes for End-to-End Object Detection?

ICML 2021spotlight

Object detection has recently achieved a breakthrough for removing the last one non-differentiable component in the pipeline, Non-Maximum Suppression (NMS), and building up an end-to-end system. However, what makes for its one-to-one prediction has not been well understood. In this paper, we first p…

2020

Improving Convolutional Networks With Self-Calibrated Convolutions

CVPR 2020poster

Recent advances on CNNs are mostly devoted to designing more complex architectures to enhance their representation learning capacity. In this paper, we consider how to improve the basic convolutional feature transformation process of CNNs without tuning the model architectures. To this end, we prese…

Cited by 542PDFcodeScholar
2020

Is normalization indispensable for training deep neural network?

NeurIPS 2020oral

Normalization operations are widely used to train deep neural networks, and they can improve both convergence and generalization in most tasks. The theories for normalization's effectiveness and new forms of normalization have always been hot topics in research. To better understand normalization, o…

2020

Trajectory Similarity Learning with Auxiliary Supervision and Optimal Matching

IJCAI 2020poster

Trajectory similarity computation is a core problem in the field of trajectory data queries. However, the high time complexity of calculating the trajectory similarity has always been a bottleneck in real-world applications. Learning-based methods can map trajectories into a uniform embedding space…

2019

Generative Dual Adversarial Network for Generalized Zero-Shot Learning

CVPR 2019poster

This paper studies the problem of generalized zero-shot learning which requires the model to train on image-label pairs from some seen classes and test on the task of classifying new images from both seen and unseen classes. In this paper, we propose a novel model that provides a unified framework…

Cited by 283PDFcodeScholar
2019

Multi-Person Pose Estimation With Enhanced Channel-Wise and Spatial Information

CVPR 2019poster

Multi-person pose estimation is an important but challenging problem in computer vision. Although current approaches have achieved significant progress by fusing the multi-scale feature maps, they pay little attention to enhancing the channel-wise and spatial information of the feature maps. In this…

Cited by 189PDFScholar
2015

Robust Image Segmentation Using Contour-Guided Color Palettes

ICCV 2015poster

The contour-guided color palette (CCP) is proposed for robust image segmentation. It efficiently integrates contour and color cues of an image. To find representative colors of an image, color samples along long contours between regions, similar in spirit to machine learning methodology that focus o…

Cited by 35PDFcodeScholar
2015

Understanding Image Structure via Hierarchical Shape Parsing

CVPR 2015poster

Exploring image structure is a long-standing yet important research subject in the computer vision community. In this paper, we focus on understanding image structure inspired by the "simple-to-complex" biological evidence. A hierarchical shape parsing strategy is proposed to partition and organize…

Cited by 12SourcePDFScholar