← Search

Yazhou Yao

43 accepted papers

2026

AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs

AAAI 2026technical

Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reasoning capabilities, they remain fall short in abductive inference, as compared to

Cited by 0SourcePDFScholar
2026

Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis

ICLR 2026poster

This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive generative augmentation often preserves the biases it aims to solve. Moreover, our analysis reveals that simply generatin…

Cited by 0SourcecodeScholar
2026

Beyond Quadratic: Linear-Time Change Detection with RWKV

AAAI 2026technical

Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this confli

Cited by 0SourcePDFScholar
2026

Condensed Test-Time Adaptation of VLMs for Action Recognition

CVPR 2026

Test-time adaptation for video understanding, which enables vision-language models (VLMs) to generalize to downstream tasks such as action recognition, has demonstrated substantial value in real-world applications. Existing memory-based methods typically build a visual cache from high-confidence tes

Cited by 0SourceScholar
2026

Deep Ensemble Clustering for Visual Representation Learning

ICML 2026poster

Recent advances in visual representation learning have seen the rise of clustering-based vision backbones, which adopt clustering as a core paradigm for feature extraction. However, existing clustering-based backbones typically rely on a single clustering algorithm, whose inherent inductive bias lim…

Cited by 0SourceScholar
2026

Ego3S: Select, Strengthen, and Synchronize for Efficient Egocentric Reasoning

ICML 2026poster

Egocentric reasoning fundamentally differs from third-person understanding in LVLMs. Third-person settings offer wide and stable contexts with consistent global regularities, allowing models to utilize broad statistical correlations. In contrast, egocentric scenes are highly dynamic and heterogeneou…

Cited by 0SourceScholar
2026

GSV2X: Geometry-Aware Uncertainty Modeling and Orthogonal Fusion for Robust Roadside Perception

CVPR 2026

Reliable 3D perception from multi-view roadside sensors hinges on the robust fusion of camera and LiDAR data, a task complicated by geometric misalignments and sensor calibration errors. This paper presents GSV2X, a fusion framework that tackles these challenges through two core contributions. First

Cited by 0SourceScholar
2026

Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation

CVPR 2026

In this paper, we propose Iris, a deterministic framework for Monocular Depth Estimation (MDE) that integrates real-world priors into the diffusion model. Conventional feed-forward methods rely on massive training data, yet still miss details. Previous diffusion-based methods leverage rich generativ

Cited by 0SourcecodeScholar
2026

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images

CVPR 2026

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed multi-view images remains challenging. Recent self-supervised methods attempt to u

Cited by 0SourceScholar
2026

MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

CVPR 2026

Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language models. We propose MedFG-VQA, a lightweight framework that leverages a memory b

Cited by 0SourcecodeScholar
2026

Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration

CVPR 2026

Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy corruption. While these issues are often studied in isolation, we argue that they share a common root in the predictive uncertainty towards the reliabi

Cited by 0SourcecodeScholar
2026

PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation

CVPR 2026

Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract image-text alignment cues from cost volumes through a serial structure of spatial and class aggregations, leading to knowle

Cited by 0SourcecodeScholar
2026

PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation

CVPR 2026

Training-free open-vocabulary semantic segmentation (OVSS) promises rapid adaptation to new label sets without retraining. Yet, many methods rely on heavy post-processing or handle text and vision in isolation, leaving cross-modal geometry underutilized. Others introduce auxiliary vision backbones o

Cited by 0SourcecodeScholar
2026

Revisiting Learning with Noisy Labels: Active Forgetting and Noise Suppression

CVPR 2026

Learning with noisy labels (LNL) has received growing attention, with most prior work following the paradigm of clean-sample reliance (e.g., sample selection). However, this reliance also imposes intrinsic limitations, as overfitting to even a few noisy samples is inevitable, creating a major bottle

Cited by 0SourcecodeScholar
2026

Seeing Motion Through Polarity for Event-based Action Recognition

CVPR 2026

Event-based Action Recognition (EAR) provides a promising pathway for understanding dynamic behaviors under challenging conditions. Recent progress in vision-language models has introduced a cross-modal learning paradigm into EAR, enabling models to associate event streams with textual semantics for

Cited by 0SourceScholar
2025

3D-aware Select, Expand, and Squeeze Token for Aerial Action Recognition

AAAI 2025technical

Aerial Action Recognition (AAR) in videos captured by Unmanned Aerial Vehicles (UAVs) plays a vital role in numerous applications. However, current methods related to traditional action recognition primarily cater to fixed or near cameras, and rarely consider the movement disturbance of UAVs, includ…

Cited by 0SourcePDFScholar
2025

CA2C: A Prior-Knowledge-Free Approach for Robust Label Noise Learning via Asymmetric Co-learning and Co-training

ICCV 2025poster

Label noise learning (LNL), a practical challenge in real-world applications, has recently attracted significant attention. While demonstrating promising effectiveness, existing LNL approaches typically rely on various forms of prior knowledge, such as noise rates or thresholds, to sustain performan…

Cited by 0SourcePDFScholar
2025

Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object Detection

ICCV 2025poster

In this paper, we propose a generation-detection cycle consistent (GDCC) learning framework that jointly optimizes both layout-to-image (L2I) generation and object detection (OD) tasks in an end-to-end manner. The key of GDCC lies in the inherent duality between the two tasks, where L2I takes all ob…

2025

Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization

ICML 2025poster

Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In thi…

2025

OmniGaze: Reward-inspired Generalizable Gaze Estimation in the Wild

NeurIPS 2025poster

Current 3D gaze estimation methods struggle to generalize across diverse data domains, primarily due to $\textbf{i)}$ $\textit{the scarcity of annotated datasets}$, and $\textbf{ii)}$ $\textit{the insufficient diversity of labeled data}$. In this work, we present OmniGaze, a semi-supervised framewor…

Cited by 0SourceScholar
2025

Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection

CVPR 2025poster

The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot classification and retrieval capabilities on various domains. However, CLIP's training remains computationally intensive…

2025

Tensor-aggregated LoRA in Federated Fine-tuning

ICCV 2025poster

The combination of Large Language Models (LLMs) and Federated Learning (FL) to leverage privacy-preserving data has emerged as a promising approach to further enhance the Parameter-Efficient Fine-Tuning (PEFT) capabilities of LLMs. In real-world FL settings with resource heterogeneity, the training…

Cited by 0SourcePDFScholar
2025

Twofold Debiasing Enhances Fine-Grained Learning with Coarse Labels

AAAI 2025technical

The Coarse-to-Fine Few-Shot (C2FS) task is designed to train models using only coarse labels, then leverages a limited number of subclass samples to achieve fine-grained recognition capabilities. This task presents two main challenges: coarse-grained supervised pre-training suppresses the extraction…

2025

UNIALIGN: Scaling Multimodal Alignment within One Unified Model

CVPR 2025poster

We present UNIALIGN, a unified model to align an arbitrary number of modalities (\text e.g. , image, text, audio, 3D point cloud, etc.) through one encoder and a single training phase. Existing solutions typically employ distinct encoders for each modality, resulting in increased parameters as the…

2025

You Only Communicate Once: One-shot Federated Low-Rank Adaptation of MLLM

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) with Federated Learning (FL) can quickly adapt to privacy-sensitive tasks, but face significant challenges such as high communication costs and increased attack risks, due to their reliance on multi-round communication. To address this, One-shot FL (OFL) has…

Cited by 0SourcecodeScholar
2024

"Veil Privacy on Visual Data: Concealing Privacy for Humans, Unveiling for DNNs"

ECCV 2024poster

"Privacy laws like GDPR necessitate effective approaches to safeguard data privacy. Existing works on data privacy protection of DNNs mainly concentrated on the model training phase. However, these approaches become impractical when dealing with the outsourcing of sensitive data. Furthermore, they h…

Cited by 0SourcePDFScholar
2024

Adaptive Integration of Partial Label Learning and Negative Learning for Enhanced Noisy Label Learning

AAAI 2024technical

There has been significant attention devoted to the effectiveness of various domains, such as semi-supervised learning, contrastive learning, and meta-learning, in enhancing the performance of methods for noisy label learning (NLL) tasks. However, most existing methods still depend on prior assumpti…

2024

Knowledge Transfer with Simulated Inter-Image Erasing for Weakly Supervised Semantic Segmentation

ECCV 2024poster

"Though adversarial erasing has prevailed in weakly supervised semantic segmentation to help activate integral object regions, existing approaches still suffer from the dilemma of under-activation and over-expansion due to the difficulty in determining when to stop erasing. In this paper, we propose…

2024

Poly Kernel Inception Network for Remote Sensing Detection

CVPR 2024poster

Object detection in remote sensing images (RSIs) often suffers from several increasing challenges including the large variation in object scales and the diverse-ranging context. Prior methods tried to address these challenges by expanding the spatial receptive field of the backbone either through la…

2024

VideoMAC: Video Masked Autoencoders Meet ConvNets

CVPR 2024poster

Recently the advancement of self-supervised learning techniques like masked autoencoders (MAE) has greatly influenced visual representation learning for images and videos. Nevertheless it is worth noting that the predominant approaches in existing masked image / video modeling rely excessively on re…

2022

Hierarchical Feature Alignment Network for Unsupervised Video Object Segmentation

ECCV 2022poster

"Optical flow is an easily conceived and precious cue for advancing unsupervised video object segmentation (UVOS). Most of the previous methods directly extract and fuse the motion and appearance features for segmenting target objects in the UVOS setting. However, optical flow is intrinsically an in…

2022

PNP: Robust Learning From Noisy Labels by Probabilistic Noise Prediction

CVPR 2022oral

Label noise has been a practical challenge in deep learning due to the strong capability of deep neural networks in fitting all training data. Prior literature primarily resorts to sample selection methods for combating noisy labels. However, these approaches focus on dividing samples by order sorti…

Cited by 80PDFScholar
2021

Jo-SRC: A Contrastive Approach for Combating Noisy Labels

CVPR 2021poster

Due to the memorization effect in Deep Neural Networks (DNNs), training with noisy labels usually results in inferior model performance. Existing state-of-the-art methods primarily adopt a sample selection strategy, which selects small-loss samples for subsequent training. However, prior literature…

Cited by 190PDFScholar
2021

Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation

CVPR 2021poster

Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of…

Cited by 248PDFcodeScholar
2021

Webly Supervised Fine-Grained Recognition: Benchmark Datasets and an Approach

ICCV 2021poster

Learning from the web can ease the extreme dependence of deep learning on large-scale manually labeled datasets. Especially for fine-grained recognition, which targets at distinguishing subordinate categories, it will significantly reduce the labeling costs by leveraging free web data. Despite its s…

Cited by 75PDFcodeScholar
2020

Field-wise Learning for Multi-field Categorical Data

NeurIPS 2020poster

We propose a new method for learning with multi-field categorical data. Multi-field categorical data are usually collected over many heterogeneous groups. These groups can reflect in the categories under a field. The existing methods try to learn a universal model that fits all data, which is challe…

2020

Region Graph Embedding Network for Zero-Shot Learning

ECCV 2020poster

Most of the existing Zero-Shot Learning (ZSL) approaches learn direct embeddings from global features or image parts (regions) to the semantic space, which, however, fail to capture the appearance relationships between different local regions within a single image. In this paper, to model the relati…

Cited by 195SourcePDFScholar
2020

Set and Rebase: Determining the Semantic Graph Connectivity for Unsupervised Cross-Modal Hashing

IJCAI 2020poster

The label-free nature of unsupervised cross-modal hashing hinders models from exploiting the exact semantic data similarity. Existing research typically simulates the semantics by a heuristic geometric prior in the original feature space. However, this introduces heavy bias into the model as the ori…

Cited by 0SourcePDFScholar
2019

Attentive Region Embedding Network for Zero-Shot Learning

CVPR 2019poster

Zero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of t…

Cited by 351PDFScholar
2019

SegEQA: Video Segmentation Based Visual Attention for Embodied Question Answering

ICCV 2019poster

Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to its broad applications in automatic driving system, in-home robots, and personal…

Cited by 34PDFScholar