← Search

Yizhou Yu

69 accepted papers

2026

CGSA: Class-Guided Slot-Aware Adaptation for Source-Free Object Detections

ICLR 2026poster

Source-Free Domain Adaptive Object Detection (SF-DAOD) aims to adapt a detector trained on a labeled source domain to an unlabeled target domain without retaining any source data. Despite recent progress, most popular approaches focus on tuning pseudo-label thresholds or refining the teacher-student…

Cited by 0SourcecodeScholar
2026

Leveraging Verifier-Based Reinforcement Learning in Image Editing

CVPR 2026

While Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm for text-to-image generation, its application to image editing remains largely unexplored. A key bottleneck is the lack of a robust general reward model for all editing tasks. Existing edit reward models usually gi

Cited by 0SourcecodeScholar
2026

Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing for Dense Predictions

ICML 2026poster

Adapting pre-trained vision models using parameter-efficient fine-tuning (PEFT) remains challenging, as it aims to achieve performance comparable to full fine-tuning using a minimal number of trainable parameters. When applied to complex dense prediction tasks, existing methods exhibit limitations, …

Cited by 0SourceScholar
2025

Autoregressive Sequence Modeling for 3D Medical Image Representation

AAAI 2025technical

Three-dimensional (3D) medical images, such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI), are essential for clinical applications. However, the need for diverse and comprehensive representations is particularly pronounced when considering the variability across different organs,…

Cited by 1SourcePDFScholar
2025

OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels

CVPR 2025poster

Top-down attention plays a crucial role in the human vision system, wherein the brain initially obtains a rough overview of a scene to discover salient cues (i.e., overview first), followed by a more careful finer-grained examination (i.e., look closely next). However, modern ConvNets remain confine…

2025

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis

ICCV 2025poster

Recent advances in large language models (LLMs) have spurred interests in encoding images as discrete tokens and leveraging autoregressive (AR) frameworks for visual generation. However, the quantization process in AR-based visual generation models inherently introduces information loss that degrade…

2025

SegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic Segmentation

CVPR 2025poster

High-quality semantic segmentation relies on three key capabilities: global context modeling, local detail encoding, and multi-scale feature extraction. However, recent methods struggle to possess all these capabilities simultaneously. Hence, we aim to empower segmentation networks to simultaneously…

2025

SparX: A Sparse Cross-Layer Connection Mechanism for Hierarchical Vision Mamba and Transformer Networks

AAAI 2025technical

Due to the capability of dynamic state space models (SSMs) in capturing long-range dependencies with linear-time computational complexity, Mamba has shown notable performance in NLP tasks. This has inspired the rapid development of Mamba-based vision models, resulting in promising results in visual…

2024

FedDiv: Collaborative Noise Filtering for Federated Learning with Noisy Labels

AAAI 2024technical

Federated Learning with Noisy Labels (F-LNL) aims at seeking an optimal server model via collaborative distributed learning by aggregating multiple client models trained with local noisy or clean samples. On the basis of a federated learning framework, recent advances primarily adopt label noise fil…

2024

OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd Representation

CVPR 2024poster

Recent advances in Iterative Vision-and-Language Navigation(IVLN) introduce a more meaningful and practical paradigm of VLN by maintaining the agent's memory across tours of scenes. Although the long-term memory aligns better with the persistent nature of the VLN task it poses more challenges on how…

Cited by 8SourcePDFScholar
2024

RegionGPT: Towards Region Understanding Vision Language Model

CVPR 2024poster

Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs yet they struggle with detailed regional visual understanding due to limited spatial awareness of the vision encoder and the use of coarse-grained training…

Cited by 44SourcePDFScholar
2023

Activate and Reject: Towards Safe Domain Generalization under Category Shift

ICCV 2023poster

Albeit the notable performance on in-domain test points, it is non-trivial for deep neural networks to attain satisfactory accuracy when deploying in the open world, where novel domains and object classes often occur. In this paper, we study a practical problem of Domain Generalization under Categor…

Cited by 8PDFScholar
2023

Advancing Radiograph Representation Learning with Masked Record Modeling

ICLR 2023poster

Modern studies in radiograph representation learning (R$^2$L) rely on either self-supervision to encode invariant semantics or associated radiology reports to incorporate medical expertise, while the complementarity between them is barely noticed. To explore this, we formulate the self- and report-c…

2023

CODA: Generalizing to Open and Unseen Domains with Compaction and Disambiguation

NeurIPS 2023spotlight

The generalization capability of machine learning systems degenerates notably when the test distribution drifts from the training distribution. Recently, Domain Generalization (DG) has been gaining momentum in enabling machine learning models to generalize to unseen domains. However, most DG methods…

Cited by 5SourcePDFScholar
2023

EGC: Image Generation and Classification via a Diffusion Energy-Based Model

ICCV 2023poster

Learning image classification and image generation using the same set of network parameters presents a formidable challenge. Recent advanced approaches perform well in one task often exhibit poor performance in the other. This work introduces an energy-based classifier and generator, namely EGC, whi…

Cited by 10PDFcodeScholar
2023

Geometry-Aware Network for Domain Adaptive Semantic Segmentation

AAAI 2023technical

Measuring and alleviating the discrepancies between the synthetic (source) and real scene (target) data is the core issue for domain adaptive semantic segmentation. Though recent works have introduced depth information in the source domain to reinforce the geometric and semantic knowledge transfer,…

Cited by 6SourcePDFScholar
2023

Learning Domain-Agnostic Representation for Disease Diagnosis

ICLR 2023poster

In clinical environments, image-based diagnosis is desired to achieve robustness on multi-center samples. Toward this goal, a natural way is to capture only clinically disease-related features. However, such disease-related features are often entangled with center-effect, disabling robust transferri…

Cited by 9SourcePDFScholar
2023

MISC210K: A Large-Scale Dataset for Multi-Instance Semantic Correspondence

CVPR 2023poster

Semantic correspondence have built up a new way for object recognition. However current single-object matching schema can be hard for discovering commonalities for a category and far from the real-world recognition tasks. To fill this gap, we design the multi-instance semantic correspondence task wh…

2023

Protein Representation Learning via Knowledge Enhanced Primary Structure Reasoning

ICLR 2023poster

Protein representation learning has primarily benefited from the remarkable development of language models (LMs). Accordingly, pre-trained protein models also suffer from a problem in LMs: a lack of factual knowledge. The recent solution models the relationships between protein and associated knowle…

Cited by 24SourcePDFScholar
2023

RankDNN: Learning to Rank for Few-Shot Learning

AAAI 2023technical

This paper introduces a new few-shot learning pipeline that casts relevance ranking for image retrieval as binary ranking relation classification. In comparison to image classification, ranking relation classification is sample efficient and domain agnostic. Besides, it provides a new perspective on…

2022

A Causal Debiasing Framework for Unsupervised Salient Object Detection

AAAI 2022technical

Unsupervised Salient Object Detection (USOD) is a promising yet challenging task that aims to learn a salient object detection model without any ground-truth labels. Self-supervised learning based methods have achieved remarkable success recently and have become the dominant approach in USOD. Howeve…

Cited by 28SourcePDFScholar
2022

A Causal Inference Look at Unsupervised Video Anomaly Detection

AAAI 2022technical

Unsupervised video anomaly detection, a task that requires no labeled normal/abnormal training data in any form, is challenging yet of great importance to both industrial applications and academic research. Existing methods typically follow an iterative pseudo label generation process. However, they…

Cited by 48SourcePDFScholar
2022

Attribute Surrogates Learning and Spectral Tokens Pooling in Transformers for Few-Shot Learning

CVPR 2022poster

This paper presents new hierarchically cascaded transformers that can improve data efficiency through attribute surrogates learning and spectral tokens pooling. Vision transformers have recently been thought of as a promising alternative to convolutional neural networks for visual recognition. But w…

Cited by 71PDFcodeScholar
2022

Centrality and Consistency: Two-Stage Clean Samples Identification for Learning with Instance-Dependent Noisy Labels

ECCV 2022poster

"Deep models trained with noisy labels are prone to over-fitting and struggle in generalization. Most existing solutions are based on an ideal assumption that the label noise is class-conditional, i.e., instances of the same class share the same noise model, and are independent of features. While in…

2022

Compound Domain Generalization via Meta-Knowledge Encoding

CVPR 2022poster

Domain generalization (DG) aims to improve the generalization performance for an unseen target domain by using the knowledge of multiple seen source domains. Mainstream DG methods typically assume that the domain label of each source sample is known a priori, which is challenged to be satisfied in m…

Cited by 84PDFScholar
2022

Disentangling Disease-related Representation from Obscure for Disease Prediction

ICML 2022spotlight

Disease-related representations play a crucial role in image-based disease prediction such as cancer diagnosis, due to its considerable generalization capacity. However, it is still a challenge to identify lesion characteristics in obscured images, as many lesions are obscured by other tissues. In t…

Cited by 5SourcePDFScholar
2022

Mix and Reason: Reasoning over Semantic Topology with Data Mixing for Domain Generalization

NeurIPS 2022accept

Domain generalization (DG) enables generalizing a learning machine from multiple seen source domains to an unseen target one. The general objective of DG methods is to learn semantic representations that are independent of domain labels, which is theoretically sound but empirically challenged due to…

Cited by 39SourcePDFScholar
2022

Neighborhood Collective Estimation for Noisy Label Identification and Correction

ECCV 2022poster

"Learning with noisy labels (LNL) aims at designing strategies to improve model performance and generalization by mitigating the effects of model overfitting to noisy labels. The key success of LNL lies in identifying as many clean samples as possible from massive noisy data, while rectifying the wr…

2022

One-Shot Medical Landmark Localization by Edge-Guided Transform and Noisy Landmark Refinement

ECCV 2022poster

"As an important upstream task for many medical applications, supervised landmark localization still requires non-negligible annotation costs to achieve desirable performance. Besides, due to cumbersome collection procedures, the limited size of medical landmark datasets impacts the effectiveness of…

2022

Scale-Equivalent Distillation for Semi-Supervised Object Detection

CVPR 2022poster

Recent Semi-Supervised Object Detection (SS-OD) methods are mainly based on self-training, i.e., generating hard pseudo-labels by a teacher model on unlabeled data as supervisory signals. Although they achieved certain success, the limited labeled data in semi-supervised learning scales up the chall…

Cited by 39PDFScholar
2021

Bottom-Up Shift and Reasoning for Referring Image Segmentation

CVPR 2021poster

Referring image segmentation aims to segment the referent that is the corresponding object or stuff referred by a natural language expression in an image. Its main challenge lies in how to effectively and efficiently differentiate between the referent and other objects of the same category as the re…

Cited by 101PDFcodeScholar
2021

Coarse-To-Fine Domain Adaptive Semantic Segmentation With Photometric Alignment and Category-Center Regularization

CVPR 2021poster

Unsupervised domain adaptation (UDA) in semantic segmentation is a fundamental yet promising task relieving the need for laborious annotation works. However, the domain shifts/discrepancies problem in this task compromise the final segmentation performance. Based on our observation, the main causes…

Cited by 90PDFcodeScholar
2021

Cross-Domain Adaptive Clustering for Semi-Supervised Domain Adaptation

CVPR 2021poster

In semi-supervised domain adaptation, a few labeled samples per class in the target domain guide features of the remaining target samples to aggregate around them. However, the trained model cannot produce a highly discriminative feature representation for the target domain because the training data…

Cited by 157PDFcodeScholar
2021

Dual Bipartite Graph Learning: A General Approach for Domain Adaptive Object Detection

ICCV 2021poster

Domain Adaptive Object Detection (DAOD) relieves the reliance on large-scale annotated data by transferring the knowledge learned from a labeled source domain to a new unlabeled target domain. Recent DAOD approaches resort to local feature alignment in virtue of domain adversarial training in conjun…

Cited by 68PDFScholar
2021

I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object Detectors

CVPR 2021poster

Recent works on two-stage cross-domain detection have widely explored the local feature patterns to achieve more accurate adaptation results. These methods heavily rely on the region proposal mechanisms and ROI-based instance-level features to design fine-grained feature alignment modules with respe…

Cited by 91PDFScholar
2021

ME-PCN: Point Completion Conditioned on Mask Emptiness

ICCV 2021poster

Point completion refers to completing the missing geometries of an object from incomplete observations. Main-stream methods predict the missing shapes by decoding a global feature learned from the input point cloud, which often leads to deficient results in preserving topology consistency and surfac…

Cited by 27PDFcodeScholar
2021

Multi-Scale Matching Networks for Semantic Correspondence

ICCV 2021poster

Deep features have been proven powerful in building accurate dense semantic correspondences in various previous works. However, the multi-scale and pyramidal hierarchy of convolutional neural networks has not been well studied to learn discriminative pixel-level features for semantic correspondence.…

Cited by 50PDFcodeScholar
2021

Noise2Grad: Extract Image Noise to Denoise

IJCAI 2021poster

In many image denoising tasks, the difficulty of collecting noisy/clean image pairs limits the application of supervised CNNs. We consider such a case in which paired data and noise statistics are not accessible, but unpaired noisy and clean images are easy to collect. To form the necessary supervis…

Cited by 11SourcePDFScholar
2021

Preservational Learning Improves Self-Supervised Medical Image Models by Reconstructing Diverse Contexts

ICCV 2021poster

Preserving maximal information is the basic principle of designing self-supervised learning methodologies. To reach this goal, contrastive learning adopts an implicit way which is contrasting image pairs. However, we believe it is not fully optimal to simply use the contrastive estimation for preser…

Cited by 116PDFcodeScholar
2021

Refer-It-in-RGBD: A Bottom-Up Approach for 3D Visual Grounding in RGBD Images

CVPR 2021poster

Grounding referring expressions in RGBD image has been an emerging field. We present a novel task of 3D visual grounding in single-view RGBD image where the referred objects are often only partially scanned due to occlusion. In contrast to previous works that directly generate object proposals for g…

Cited by 43PDFScholar
2020

Cross-View Correspondence Reasoning Based on Bipartite Graph Convolutional Network for Mammogram Mass Detection

CVPR 2020oral

Mammogram mass detection is of great clinical significance due to its high proportion in breast cancers. The information from cross views (i.e., mediolateral oblique and cranio-caudal) is highly related and complementary, and is helpful to make comprehensive decisions. However, unlike radiologists w…

Cited by 69PDFScholar
2019

Align, Attend and Locate: Chest X-Ray Diagnosis via Contrast Induced Attention Network With Limited Supervision

ICCV 2019accepted

Obstacles facing accurate identification and localization of diseases in chest X-ray images lie in the lack of high-quality images and annotations. In this paper, we propose a Contrast Induced Attention Network (CIA-Net), which exploits the highly structured property of chest X-ray images and locali…

Cited by 133SourcePDFScholar
2019

Cascaded Generative and Discriminative Learning for Microcalcification Detection in Breast Mammograms

CVPR 2019poster

Accurate microcalcification (mC) detection is of great importance due to its high proportion in early breast cancers. Most of the previous mC detection methods belong to discriminative models, where classifiers are exploited to distinguish mCs from other backgrounds. However, it is still challenging…

Cited by 50PDFScholar
2019

Multi-Source Weak Supervision for Saliency Detection

CVPR 2019poster

The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-performing model. To this end, we propose a unified framework to train saliency…

Cited by 227PDFcodeScholar
2019

Transductive Zero-Shot Learning with Visual Structure Constraint

NeurIPS 2019poster

To recognize objects of the unseen classes, most existing Zero-Shot Learning (ZSL) methods first learn a compatible projection function between the common semantic space and the visual space based on the data of source seen classes, then directly apply it to the target unseen classes. However, in re…

2019

Weakly Supervised Complementary Parts Models for Fine-Grained Image Classification From the Bottom Up

CVPR 2019poster

Given a training dataset composed of images and corresponding category labels, deep convolutional neural networks show a strong ability in mining discriminative parts for image classification. However, deep convolutional neural networks trained with image level labels only tend to focus on the most…

Cited by 352PDFScholar
2018

Multi-Evidence Filtering and Fusion for Multi-Label Classification, Object Detection and Semantic Segmentation Based on Weakly Supervised Learning

CVPR 2018poster

Supervised object detection and semantic segmentation require object or even pixel level annotations. When there exist image level labels only, it is challenging for weakly supervised algorithms to achieve accurate predictions. The accuracy achieved by top weakly supervised algorithms is still signi…

Cited by 251SourcePDFScholar
2018

Stroke Controllable Fast Style Transfer with Adaptive Receptive Fields

ECCV 2018poster

The Fast Style Transfer methods have been recently proposed to transfer a photograph to an artistic style in real-time. This task involves controlling the stroke size in the stylized results, which remains an open challenge. In this paper, we present a stroke controllable style transfer network that…

Cited by 148SourcePDFScholar
2017

Borrowing Treasures From the Wealthy: Deep Transfer Learning Through Selective Joint Fine-Tuning

CVPR 2017spotlight

Deep neural networks require a large amount of labeled training data during supervised learning. However, collecting and labeling so much data might be infeasible in many cases. In this paper, we introduce a deep transfer learning scheme, called selective joint fine-tuning, for improving the perform…

Cited by 297PDFcodeScholar
2017

High-Resolution Shape Completion Using Deep Neural Networks for Global Structure and Local Geometry Inference

ICCV 2017spotlight

We propose a data-driven method for recovering missing parts of 3D shapes. Our method is based on a new deep learning architecture consisting of two sub-networks: a global structure inference network and a local geometry refinement network. The global structure inference network incorporates a long…

Cited by 367PDFScholar
2015

HD-CNN: Hierarchical Deep Convolutional Neural Networks for Large Scale Visual Recognition

ICCV 2015poster

In image classification, visual separability between different object categories is highly uneven, and some categories are more difficult to distinguish than others. Such difficult categories demand more dedicated classifiers. However, existing deep convolutional neural networks (CNN) are trained as…

Cited by 520PDFcodeScholar
2015

Harvesting Discriminative Meta Objects With Deep CNN Features for Scene Classification

ICCV 2015poster

Recent work on scene classification still makes use of generic CNN features in a rudimentary manner. In this paper, we present a novel pipeline built upon deep CNN features to harvest discriminative visual objects and parts for scene classification. We first use a region proposal technique to genera…

Cited by 153PDFScholar